Apparatus and method for implementing processing information based grouping
By embedding processing information packets into audio signal data packets, the problem of transmitting spatial and priority information of audio signals in immersive voice and audio services is solved, enabling high-quality audio transmission and rendering under low bit rate and high fault tolerance conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NOKIA TECHNOLOGIES OY
- Filing Date
- 2024-08-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to effectively process and transmit spatial and priority information of audio signals in immersive voice and audio services, leading to a decline in audio quality under low bit rate and high fault-tolerant transmission conditions.
By embedding processing information packets, including size, type, orientation, scene information, and priority parameters, into audio signal data packets, the system assists in audio signal processing at the renderer and utilizes real-time delivery protocols for transmission and reception.
It enables the effective transmission and processing of spatial and priority information of immersive speech and audio signals under low bit rate and high fault tolerance conditions, thereby improving audio quality and rendering effects.
Smart Images

Figure CN122122878A_ABST
Abstract
Description
Technical Field
[0001] This application relates to apparatus and methods for implementing packet processing based on information processing, but is not limited to implementing packet processing based on audio information in the Immersive Voice and Audio Services (IVAS) Real-Time Transport Protocol (RTP) payload. Background Technology
[0002] Immersive audio codecs are being implemented to support a variety of operating points, from low bit rate operation to transparency. An example of such a codec is the Immersive Voice and Audio Services (IVAS) codec, designed for use over communication networks such as 3GPP 4G / 5G networks. Such immersive services include, for example, immersive voice and audio use in applications such as Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), as well as spatial voice communications including remote conferencing. This audio codec is expected to handle the encoding, decoding, and rendering of speech, music, and general audio. Furthermore, it is expected to support channel-based and scene-based audio inputs, including spatial information about the sound field and sound sources. The codec is also expected to operate with low latency to enable dialogue services and support high fault tolerance under various transmission conditions.
[0003] The input signal is presented to the IVAS encoder in one of the supported formats (and in certain allowed combinations of formats). Similarly, the decoder is expected to output audio in supported formats. A pass-through mode has been proposed, in which the audio can be provided in its original format after transmission (encoding / decoding).
[0004] Additionally, RTP (Real-Time Transport Protocol) is designed for end-to-end, real-time delivery of streaming media and provides facilities for jitter compensation and detection of packet loss and out-of-order delivery. RTP allows data to be delivered to multiple destinations via IP multicast or to a specific destination via IP unicast. Most RTP implementations are built on top of the User Datagram Protocol (UDP). Other transport protocols can also be utilized. RTP is used in conjunction with other protocols such as H.323 and the Real-Time Streaming Protocol (RTSP).
[0005] The RTP specification describes two protocols: RTP and RTCP. RTP is used for the transmission of multimedia data, and its companion protocol RTCP is used to periodically send control information and Quality of Service (QoS) parameters.
[0006] RTP sessions are typically initiated between a client and a server, or between a client and another client (or in a multi-party topology), using signaling protocols such as H.323, Session Initiation Protocol (SIP), or RTSP. These protocols typically use Session Description Protocols (SDPs) (such as those defined by RFC 8866) to specify the parameters used for the session. Summary of the Invention
[0007] According to a first aspect, an apparatus is provided, comprising components configured to: acquire at least one audio signal; generate at least one audio signal data packet; acquire processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generate at least one processing information packet including the processing information; and associate the at least one processing information packet with the at least one audio signal data packet in a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0008] The component can also be configured to determine the size of at least one processing information, wherein the component configured to generate at least one group of processing information including the processing information can be configured to generate at least one group of processing information including a size field, the size field indicating the size of at least one processing information.
[0009] The component can also be configured to: determine the type of processing information, wherein the component configured to generate at least one group of processing information including the processing information can be configured to generate at least one group of processing information including a type field, the type field indicating the type of at least one group of processing information.
[0010] The component can also be configured to: determine the size and type of the processing information, wherein the component configured to generate at least one group of processing information including the processing information can be configured to generate at least one group of processing information including a size field and a type field, wherein the size field indicates the size of at least one group of processing information; and the type field indicates the type of at least one group of processing information.
[0011] The types of information processed may include one of the following: a directional indicator parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within at least one audio signal.
[0012] A component configured to generate at least one processing information group including processing information can be configured to generate at least one processing information group including at least one of the following: using an indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid.
[0013] The processing information may include one of the following: a azimuth parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within at least one audio signal.
[0014] The component can also be configured to: generate at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information; and append at least one processing information packet header to at least one processing information packet in a manner that at least one processing information packet is identified in another device as including at least one processing data packet including processing information rather than audio data.
[0015] The component configured to attach at least one processing information group header to at least one processing information group can be configured to: attach at least one processing information group header to a first block and attach at least one processing information group to a separate second block; and attach at least one processing information group header immediately before the associated processing information group of at least one processing information group.
[0016] A component configured to associate at least one processing information packet with at least one audio signal data packet can be configured to attach at least one processing information packet with at least one audio signal data packet.
[0017] A component configured to associate at least one processing information packet with at least one audio signal data packet may be configured to perform one of the following: appending each of the at least one processing information packet to the at least one audio signal data packet; and appending the at least one processing information packet immediately before the associated at least one processing information packet.
[0018] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0019] This component can also be configured to send at least one audio signal data packet and at least one processing information packet as a real-time transmission protocol.
[0020] This component can also be configured to send at least one processing information packet as part of a Real-Time Transport Protocol header extension.
[0021] This component can also be configured to send at least one processed data packet as separate packets according to a real-time delivery protocol.
[0022] This component can be configured to determine whether to send at least one processing information packet as part of the Real-Time Transport Protocol header extension or as a separate packet according to the Real-Time Transport Protocol based on one of the following: session negotiation preferences; or the protocol indicated by out-of-band signaling.
[0023] The component can also be configured to send a signal to at least one additional device, including at least one packet of processing information.
[0024] The component can also be configured to send at least one supported processing information packet type to at least one additional device via signaling.
[0025] The component can also be configured to negotiate at least one supported processing information packet type with at least one additional device for at least one session.
[0026] This component can also be configured to send at least one processing information packet as part of the session description.
[0027] According to a second aspect, an apparatus is provided, comprising components configured to: receive an audio packet stream, the audio packet stream comprising: at least one audio signal data packet, the at least one audio signal data packet comprising at least one audio signal; at least one processing information packet, the at least one processing information packet comprising processing information, the processing information comprising at least one parameter associated with the at least one audio signal; and process the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0028] The at least one group of processing information may include a size field that indicates the size of the at least one group of processing information.
[0029] The at least one processing information group may include a type field indicating the type of processing information within the at least one processing information group, and wherein a component configured to process at least one audio signal based on the processing information may be configured to identify the type of processing information within the at least one processing information group based on the type field.
[0030] The at least one processing information group may include: a size field indicating the size of the at least one processing information group; and a type field indicating the type of processing information within the at least one processing information group, wherein a component configured to process the at least one audio signal based on the processing information may be configured to identify the type of processing information within the at least one processing information group.
[0031] The type of processing information may include one of the following: a directional indicator parameter associated with the at least one audio signal; a scene information parameter associated with the at least one audio signal; an audio signal or stream priority parameter indicating the priority of audio signals within the at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within the at least one audio signal.
[0032] The at least one processing information packet may also include at least one of the following: a usage indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid with respect to the processing of at least one subsequent audio signal data packet.
[0033] The processing information may include one of the following: a azimuth parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating the priority audio object within the at least one audio signal.
[0034] The audio data packet stream may also include at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information.
[0035] At least one processing information packet header and at least one processing information packet may be arranged in one of the following: at least one processing information packet header in a first block and at least one processing information packet in a separate second block; and at least one processing information packet header immediately preceding the associated processing information packet of at least one processing information packet.
[0036] At least one processing information packet and at least one audio signal data packet may be arranged in one of the following ways: each of the at least one processing information packet precedes the at least one audio signal data packet; and the at least one processing information packet immediately precedes the associated at least one audio signal data packet.
[0037] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0038] This component can be configured to receive the audio packet stream as a real-time transport protocol transmission.
[0039] The component can also be configured to receive the at least one processing information packet as part of a real-time transport protocol header extension.
[0040] The component can also be configured to receive at least one processing information packet as a separate packet according to a real-time transmission protocol.
[0041] This component can be configured to determine whether to receive at least one processing information packet as part of a RealTime Transport protocol header extension or as a separate packet according to the RealTime Transport protocol based on one of the following: session negotiation preferences; or the protocol indicated by out-of-band signaling.
[0042] The component can also be configured to receive a signal from at least one additional device indicating that it includes at least one packet of processing information.
[0043] The component can also be configured to receive a signal from at least one additional device that identifies at least one supported processing information packet type.
[0044] This component can also be configured to retain at least one of the supporting processing information groups of the identifier.
[0045] The component can also be configured to negotiate at least one supported processing information packet type with at least one additional device for at least one session.
[0046] This component can also be configured to receive at least one processing information packet as part of a session description.
[0047] According to a third aspect, a method for an apparatus is provided, comprising: acquiring at least one audio signal; generating at least one audio signal data packet; acquiring processing information, the processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generating at least one processing information packet including the processing information; and associating the at least one processing information packet with the at least one audio signal data packet in a manner in which the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0048] The method may further include determining the size of at least one processing information, wherein generating at least one processing information group including the processing information may include: generating at least one processing information group including a size field, the size field indicating the size of at least one processing information.
[0049] The method may further include: determining the type of processing information, wherein generating at least one processing information group including the processing information may include generating at least one processing information group including a type field that indicates the type of at least one processing information.
[0050] The method may further include: determining the size and type of the processing information, wherein generating at least one processing information group including the processing information may include generating at least one processing information group including a size field and a type field, the size field indicating the size of at least one processing information; the type field indicating the type of at least one processing information.
[0051] The types of information processed may include one of the following: a directional indicator parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within at least one audio signal.
[0052] Generating at least one processing information group that includes processing information may include generating at least one processing information group that includes at least one of the following: using an indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid.
[0053] The processing information may include one of the following: a azimuth parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within at least one audio signal.
[0054] The method may further include: generating at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information; and appending the at least one processing information packet header to the at least one processing information packet in such a manner that the at least one processing information packet is identified in another device as including at least one processing data packet including processing information rather than audio data.
[0055] Attaching the header of the at least one processing information group to the at least one processing information group may include: attaching the header of the at least one processing information group to a first block and attaching the at least one processing information group to a separate second block; and attaching the header of the at least one processing information group immediately before the associated processing information group of the at least one processing information group.
[0056] Associating at least one processing information packet with at least one audio signal data packet may include attaching at least one processing information packet to at least one audio signal data packet.
[0057] Associating at least one processing information packet with at least one audio signal data packet may include one of the following: appending each of the at least one processing information packet to the preceding at least one audio signal data packet; or appending the at least one processing information packet immediately before the associated at least one processing information packet.
[0058] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0059] The method may also include transmitting at least one audio signal data packet and at least one audio processing data packet as a real-time transport protocol.
[0060] Sending at least one processing information packet may include sending at least one processing information packet as part of the Real-Time Transport Protocol header extension.
[0061] Sending at least one processing information packet may include sending at least one processing information packet as separate packets according to the real-time transport protocol.
[0062] The method may also include determining whether to send at least one processing information packet as part of a RealTime Transport Protocol header extension or as a separate packet according to the RealTime Transport Protocol based on one of the following: session negotiation preference; or protocol indicated by out-of-band signaling.
[0063] The method may also include signaling at least one processing information packet to at least one additional device.
[0064] The method may also include signaling the at least one supported processing information packet type to at least one additional device.
[0065] The method may also include negotiating at least one supported processing information packet type with at least one additional device for at least one session.
[0066] The method can also be configured to send at least one processing information packet as part of the session description.
[0067] According to a fourth aspect, a method for an apparatus is provided, comprising: receiving an audio packet stream, the audio packet stream comprising: at least one audio signal data packet, the at least one audio signal data packet comprising at least one audio signal; at least one processing information packet, the at least one processing information packet comprising processing information, the processing information comprising at least one parameter associated with the at least one audio signal; and processing the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0068] The at least one group of processing information may include a size field that indicates the size of the at least one group of processing information.
[0069] The at least one processing information group may include a type field indicating the type of processing information within the at least one processing information group, and wherein processing the at least one audio signal based on the processing information may include identifying the type of processing information within the at least one processing information group based on the type field.
[0070] The at least one processing information group may include: a size field indicating the size of the at least one processing information group; and a type field indicating the type of processing information within the at least one processing information group, wherein processing the at least one audio signal based on the processing information may include identifying the type of processing information within the at least one processing information group.
[0071] The type of processing information may include one of the following: a directional indicator parameter associated with the at least one audio signal; a scene information parameter associated with the at least one audio signal; an audio signal or stream priority parameter indicating the priority of audio signals within the at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within the at least one audio signal.
[0072] The at least one processing information packet may also include at least one of the following: a usage indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid with respect to the processing of at least one subsequent audio signal data packet.
[0073] The processing information may include one of the following: a azimuth parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating the priority audio object within the at least one audio signal.
[0074] The audio data packet stream may also include at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information.
[0075] At least one processing information packet header and at least one processing information packet may be arranged in one of the following: at least one processing information packet header in a first block and at least one processing information packet in a separate second block; and at least one processing information packet header immediately preceding the associated processing information packet of at least one processing information packet.
[0076] At least one processing information packet and at least one audio signal data packet may be arranged in one of the following ways: each of the at least one processing information packet precedes the at least one audio signal data packet; and the at least one processing information packet immediately precedes the associated at least one audio signal data packet.
[0077] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0078] Audio packet streams can be received and transmitted as a real-time delivery protocol.
[0079] Receiving an audio packet stream that includes at least one processing information packet may include receiving the at least one processing information packet as part of a Real-Time Transport Protocol header extension.
[0080] Receiving at least one audio packet stream, including at least one processing information packet, may include receiving at least one processing information packet as separate packets according to a real-time transport protocol.
[0081] The method may also include determining whether to receive an audio packet stream comprising at least one processing information packet as part of a RealTime Transport Protocol header extension or as a separate packet according to the RealTime Transport Protocol based on one of the following: session negotiation preference; or protocol indicated by out-of-band signaling.
[0082] The method may also include receiving a signal from at least one additional device indicating that it includes at least one processing information packet.
[0083] The method may also include receiving a signal from at least one additional device that identifies at least one supported processing information packet type.
[0084] The method may also include at least one of the supporting processing information groups that retain the identifier.
[0085] The method may also include negotiating at least one supported processing information packet type with at least one additional device for at least one session.
[0086] The method may also include receiving at least one processing information packet as part of a session description.
[0087] According to a fifth aspect, an apparatus is provided, comprising at least one processor and at least one memory, the at least one memory including computer program code, the at least one memory and the computer program code being configured together with the at least one processor such that the apparatus at least: acquires at least one audio signal; generates at least one audio signal data packet; acquires processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generates at least one processing information packet including the processing information; and associates the at least one processing information packet with the at least one audio signal data packet in a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0088] The apparatus can also be configured to determine the size of at least one processing information, wherein the apparatus configured to generate at least one processing information group including the processing information can be configured to generate at least one processing information group including a size field, the size field indicating the size of at least one processing information.
[0089] The apparatus can also be configured to: determine the type of processing information, wherein the apparatus configured to generate at least one group of processing information including the processing information can be configured to generate at least one group of processing information including a type field, the type field indicating the type of at least one group of processing information.
[0090] The apparatus can also be configured to: determine the size and type of the processed information, wherein the apparatus configured to generate at least one group of processed information including the processed information can be configured to generate at least one group of processed information including a size field and a type field, the size field indicating the size of at least one group of processed information; the type field indicating the type of at least one group of processed information.
[0091] The type of processing information may include one of the following: a directional indicator parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal in at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object in at least one audio signal.
[0092] The means for generating at least one processing information packet including processing information can be made to generate at least one processing information packet including at least one of the following: using an indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid.
[0093] The processing information may include one of the following: a azimuth parameter associated with the at least one audio signal; a scene information parameter associated with the at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within the at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within the at least one audio signal.
[0094] The apparatus may also be configured to: generate at least one processing information packet header, the at least one processing information packet header including an indicator that identifies at least one processing information packet as including processing information; and append at least one processing information packet header to at least one processing information packet in a manner that at least one processing information packet is identified in another apparatus as including at least one processing data packet including processing information rather than audio data.
[0095] The means of attaching the header of the at least one processing information group to the at least one processing information group can be made to: attach the header of the at least one processing information group to a first block and attach the at least one processing information group to a separate second block; and attach the header of the at least one processing information group immediately before the associated processing information group of the at least one processing information group.
[0096] The device that is configured to associate at least one processing information packet with at least one audio signal data packet can be configured to attach at least one processing information packet with at least one audio signal data packet.
[0097] The means for associating at least one processing information packet with at least one audio signal data packet can be made to perform one of the following: appending each of the at least one processing information packet to the at least one audio signal data packet before; or appending at least one processing information packet immediately before the associated at least one processing information packet.
[0098] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0099] The device can also be configured to transmit at least one audio signal data packet and at least one audio processing data packet as a real-time transmission protocol.
[0100] The device can be configured to send at least one processing information packet as part of a real-time transport protocol header extension.
[0101] The device can be configured to send at least one processing information packet as separate packets according to a real-time transmission protocol.
[0102] The device can also be configured to send at least one processing information packet as part of a real-time transport protocol header extension.
[0103] The device can also be configured to send at least one processing information packet as separate packets according to a real-time transmission protocol.
[0104] The apparatus can be configured to determine whether to send at least one processing information packet as part of a Real-Time Transport Protocol (RTP) header extension or as a separate packet according to RTP, based on one of the following: session negotiation preferences; or protocols indicated by out-of-band signaling.
[0105] The device can be made to send a signal including at least one processing information packet to at least one other device.
[0106] The device can be configured to send the at least one supported processing information packet type to at least one other device by signaling.
[0107] The device can be configured to negotiate at least one supported processing information packet type with at least one other device for at least one session.
[0108] The apparatus can be configured to send at least one processing information packet as part of a session description. According to a sixth aspect, an apparatus is provided comprising at least one processor and at least one memory, the at least one memory including computer program code, the at least one memory and the computer program code being configured, together with the at least one processor, such that the apparatus at least: receives an audio packet stream, the audio packet stream including: at least one audio signal data packet, the at least one audio signal data packet including at least one audio signal; at least one processing information packet, the at least one processing information packet including processing information, the processing information including at least one parameter associated with the at least one audio signal; and processes the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0109] The at least one group of processing information may include a size field that indicates the size of the at least one group of processing information.
[0110] The at least one processing information group may include a type field indicating the type of processing information within the at least one processing information group, and wherein the device that is configured to process the at least one audio signal based on the processing information may be configured to identify the type of processing information within the at least one processing information group based on the type field.
[0111] The at least one processing information group may include: a size field indicating the size of the at least one processing information group; and a type field indicating the type of processing information within the at least one processing information group, wherein the device that is configured to process the at least one audio signal based on the processing information may be configured to identify the type of processing information within the at least one processing information group.
[0112] The type of processing information may include one of the following: a directional indicator parameter associated with the at least one audio signal; a scene information parameter associated with the at least one audio signal; an audio signal or stream priority parameter indicating the priority of audio signals within the at least one audio signal; an active speaker indicator; and a priority parameter indicating a priority audio object within the at least one audio signal.
[0113] The at least one processing information packet may also include at least one of the following: a usage indicator configured to describe how the processing information should be used; and a validity indicator configured to describe how long the processing information is valid with respect to the processing of at least one subsequent audio signal data packet.
[0114] The processing information may include one of the following: a azimuth parameter associated with at least one audio signal; a scene information parameter associated with at least one audio signal; an audio signal or stream priority parameter indicating the priority of an audio signal within at least one audio signal; an active speaker indicator; and a priority parameter indicating the priority audio object within the at least one audio signal.
[0115] The audio data packet stream may also include at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information.
[0116] At least one processing information packet header and at least one processing information packet may be arranged in one of the following: at least one processing information packet header in a first block and at least one processing information packet in a separate second block; and at least one processing information packet header immediately preceding the associated processing information packet of at least one processing information packet.
[0117] At least one processing information packet and at least one audio signal data packet may be arranged in one of the following ways: each of the at least one processing information packet precedes the at least one audio signal data packet; and the at least one processing information packet immediately precedes the associated at least one audio signal data packet.
[0118] The at least one audio signal data packet may be an immersive voice and audio service data packet.
[0119] This device can enable the received audio packet stream to be transmitted as a real-time transport protocol.
[0120] The device can also be configured to receive the at least one processing information packet as part of a real-time transport protocol header extension.
[0121] The device can be configured to receive at least one processing information packet as a separate packet according to a real-time transmission protocol.
[0122] The apparatus can be configured to determine whether to receive an audio packet stream comprising at least one processing information packet as part of a RealTime Transport protocol header extension or as a separate packet according to the RealTime Transport protocol based on one of the following: session negotiation preference; or protocol indicated by out-of-band signaling.
[0123] The device can also be configured to receive a signal from at least one other device indicating that it includes at least one packet of processing information.
[0124] The device can also be configured to receive a signal from at least one additional device that identifies at least one supported processing information packet type.
[0125] The device can also be configured to retain at least one of the supporting processing information groups of the identifier.
[0126] The device can also be configured to negotiate at least one supported processing information packet type with at least one other device for at least one session.
[0127] The device can also be configured to receive at least one processing information packet as part of a session description.
[0128] According to a seventh aspect, an apparatus is provided, comprising: an acquisition circuit system configured to acquire at least one audio signal; a generation circuit system configured to generate at least one audio signal data packet; an acquisition circuit system configured to acquire processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; a generation circuit system configured to generate at least one processing information packet including the processing information; and an additional circuit system configured to associate the at least one processing information packet with the at least one audio signal data packet in such a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0129] According to an eighth aspect, an apparatus is provided, comprising: a receiving circuit system configured to receive an audio packet stream, the audio packet stream including: at least one audio signal data packet, the at least one audio signal data packet including at least one audio signal; at least one processing information packet, the at least one processing information packet including processing information including at least one parameter associated with the at least one audio signal; and a processing circuit system configured to process the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0130] According to a ninth aspect, a computer program is provided, including instructions [or a computer-readable medium including program instructions], for causing a device to perform at least the following: acquiring at least one audio signal; generating at least one audio signal data packet; acquiring processing information, the processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generating at least one processing information packet including the processing information; and associating the at least one processing information packet with the at least one audio signal data packet in a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0131] According to a tenth aspect, a computer program is provided, including instructions [or a computer-readable medium including program instructions], for causing an apparatus to perform at least the following: receiving an audio packet stream, the audio packet stream including: at least one audio signal data packet, the at least one audio signal data packet including at least one audio signal; at least one processing information packet, the at least one processing information packet including processing information, the processing information including at least one parameter associated with the at least one audio signal; and processing the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0132] According to an eleventh aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing a device to perform at least the following: acquiring at least one audio signal; generating at least one audio signal data packet; acquiring processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generating at least one processing information packet including the processing information; and associating the at least one processing information packet with the at least one audio signal data packet in a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0133] According to a twelfth aspect, a non-transitory computer-readable medium is provided, comprising program instructions for causing an apparatus to perform at least the following: receiving an audio packet stream, the audio packet stream comprising: at least one audio signal data packet, the at least one audio signal data packet comprising at least one audio signal; at least one processing information packet, the at least one processing information packet comprising processing information, the processing information comprising at least one parameter associated with the at least one audio signal; and processing the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0134] According to a thirteenth aspect, an apparatus is provided, comprising: components for acquiring at least one audio signal; components for generating at least one audio signal data packet; components for acquiring processing information, the processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; components for generating at least one processing information packet including the processing information; and components for associating the at least one processing information packet with the at least one audio signal data packet in a manner in which the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0135] According to a fourteenth aspect, an apparatus is provided, comprising: a component for receiving an audio packet stream, the audio packet stream including: at least one audio signal data packet, the at least one audio signal data packet including at least one audio signal; at least one processing information packet, the at least one processing information packet including processing information including at least one parameter associated with the at least one audio signal; and a component for processing the at least one audio signal based on the processing information such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0136] According to a fifteenth aspect, a computer-readable medium is provided, including program instructions for causing an apparatus to perform at least the following: acquiring at least one audio signal; generating at least one audio signal data packet; acquiring processing information including at least one parameter associated with the at least one audio signal, and the processing information being used to assist processing of the at least one audio signal at a renderer; generating at least one processing information packet including the processing information; and associating the at least one processing information packet with the at least one audio signal data packet in a manner that the at least one processing information packet is configured to assist processing of at least one subsequent audio signal data packet.
[0137] According to a sixteenth aspect, a computer-readable medium is provided, including program instructions for causing an apparatus to perform at least the following: receiving an audio packet stream, the audio packet stream including: at least one audio signal data packet, the at least one audio signal data packet including at least one audio signal; at least one processing information packet, the at least one processing information packet including processing information including at least one parameter associated with the at least one audio signal; and processing the at least one audio signal based on the processing information, such that the at least one processing information packet is configured to assist in the processing of at least one subsequent audio signal data packet.
[0138] An apparatus comprising components for performing actions as described above.
[0139] An apparatus configured to perform the actions described above.
[0140] A computer program comprising program instructions for causing a computer to perform the methods described above.
[0141] A computer program product stored on a medium can cause a device to perform the methods described herein.
[0142] Electronic devices may include devices as described herein.
[0143] Chipsets may include devices as described herein.
[0144] The embodiments of this application are intended to solve problems associated with the prior art. Attached Figure Description
[0145] To better understand this application, reference will now be made to the accompanying drawings, in which:
[0146] Figure 1 An example server and a peer-to-peer remote conferencing system in which embodiments can be implemented are illustrated schematically;
[0147] Figure 2 An example packet format with an Enhanced Voice Service (EVS) Header-Full payload structure is illustrated schematically;
[0148] Figure 3 This schematically illustrates some implementations, for example Figure 1 The example encoder configuration shown;
[0149] Figure 4 The illustration shows, according to some embodiments, as follows Figure 3 The flowchart shown is a method for operating the encoder;
[0150] Figure 5An example Header-FullIVAS payload structure according to some embodiments is schematically shown;
[0151] Figure 6a and 6b An example of the content table (ToC) byte structure in an EVS Header-Full frame and Header-FullIVAS frame according to some embodiments is illustrated.
[0152] Figure 7 An example payload structure of a processing information (PI) frame with and without a Header-Full IVAS frame is illustrated schematically according to some embodiments.
[0153] Figure 8 An example payload structure of multiple Header-Full IVAS frames and PI frames in the same RTP packet according to some embodiments is illustrated schematically.
[0154] Figure 9 Examples of different PI frame structures according to some embodiments are illustrated schematically;
[0155] Figure 10a and 10b An example structure of a "type" identifier in a PI frame according to some embodiments is illustrated schematically;
[0156] Figure 11 An example of an orientation frame structure according to some embodiments is illustrated schematically;
[0157] Figure 12 An example ISM audio object priority frame structure according to some embodiments is illustrated schematically;
[0158] Figure 13 An example of a priority frame for the IVAS input format is illustrated schematically according to some embodiments;
[0159] Figure 14 An example of an active speaker label structure according to some embodiments is illustrated schematically;
[0160] Figure 15 An example of an ISM tag according to some embodiments is illustrated schematically;
[0161] Figure 16 The illustration shows, for example, some embodiments. Figure 1 The example decoder configuration shown;
[0162] Figure 17 An illustration according to some embodiments is shown as follows Figure 16The flowchart shown illustrates the method for operating the decoder; and
[0163] Figure 18 An example device suitable for implementing the illustrated apparatus is shown. Detailed Implementation
[0164] The following will describe in further detail the suitable devices and possible mechanisms for providing efficient IVAS audio.
[0165] Example systems in which embodiments may be implemented are in Figure 1 It is shown in the middle.
[0166] For example, Figure 1 An example remote conferencing system in which some embodiments can be implemented is shown. This example shows two locations or rooms, room A 100 and room B 102. Room A 100 includes a “speaker” or user, speaker 1 103. Room B 102 includes a “speaker” or user, speaker RX 141.
[0167] In the following example, room A contains a suitable remote conferencing device (or more generally, telecommunications device 110), configured to spatially capture and encode an audio environment, and further configured to present a spatial audio signal to the room. Each of the other rooms may contain a suitable remote conferencing device (or more generally, telecommunications device, such as device 120 in room B), configured to present a spatial audio signal to the room, and further configured to at least capture and encode mono audio, and optionally configured to spatially capture and encode an audio environment.
[0168] In the following examples, each room is provided with components for spatially capturing, encoding, receiving, and presenting spatial audio signals to a suitable listener. It should be understood that other embodiments may exist in which the system includes some means configured to capture and encode audio signals only (in other words, the means is a "transmitting" means only), and other means configured to receive and render audio signals only (in other words, the means is a "receiving" means only). In these embodiments, the system in which the embodiments may be implemented may include means with different audio signal capture / rendering capabilities.
[0169] The remote conferencing devices (for each site or room) 110, 120 can be configured to call remote conferences controlled by and implemented through the server 111.
[0170] In some embodiments, the communication or teleconference system includes a (point-to-point) communication system (rather than) Figure 1The server-based system shown in the diagram has several embodiments that can be implemented within it. Thus, for example, two or more UEs can be configured to interact directly with each other (e.g., to enable immersive audio calls between users). In such a scenario, one UE can be configured to transmit the spatial environment as a stream and capture speech as an audio object or audio source using a proximity microphone (e.g., a lavalier microphone). The sending UE can be configured to encode the spatial environment audio signal into a MASA format stream and the proximity microphone audio signal into an object format stream. These two audio streams can then be transmitted as separate IVAS streams. Additionally, the sending UE can be configured to encode processing information during encoding to transmit the PI frame along with the IVAS frame to the receiving UE.
[0171] The remote conferencing device can be configured to spatially capture and encode the audio environment, and further configured to render spatial audio signals into the room. For simplicity, only the communication or signaling path from room A 100 to room B 102 is shown in this example, but full-duplex or multipoint communication systems including multiple signaling paths can be implemented using the methods described herein without significant innovative input.
[0172] The remote conferencing devices (for each site or room) 110, 120 are also configured to communicate with each other to enable remote conferencing functionality.
[0173] like Figure 1 As shown, devices 110, 120, and server 111 can include suitable encoder and decoder functions. For example, device 110 is shown to include an (IVAS) encoder 101, server 111 is shown to include an (IVAS) decoder and encoder 121, and device 120 is shown to include an (IVAS) decoder 131. In this way, object 120 (representing the audio signal of user or speaker 1111) can be encoded by encoder 101, which generates a bitstream 106 to be passed to server 111. Server 111 can then decode (optionally mix with other objects and otherwise process the audio signal) and then encode it to generate a bitstream 108 to be passed to device 120. Device 120 can then decode the audio signal and present it to the user or speaker "Speaker RX" 141.
[0174] Although this example illustrates a teleconference application, encoder / decoder functionality can be applied to any suitable media stream.
[0175] Furthermore, the IVAS decoder / renderer for each of the remote conferencing devices 102 can be configured to process multiple input streams, which may each originate from a different encoder.
[0176] As mentioned earlier, RTP is designed for end-to-end real-time delivery of streaming media and provides facilities for jitter compensation, packet loss, and out-of-order delivery detection. RTP is also designed to carry a variety of multimedia formats, allowing the delivery of new formats without modifying the RTP standard. Therefore, the information required by the application-specific aspects of the protocol is not included in the general RTP header. For a class of applications (e.g., audio, video), an RTP profile can be defined. For a media format (e.g., a specific video encoding format), an associated RTP payload format can be defined. Therefore, each instance of RTP in a particular application may require both a profile and a payload format specification.
[0177] This profile is configured to define the codecs used to encode payload data, as well as the mapping to the payload format code in the protocol field "Payload Type (PT)" of the RTP header.
[0178] For example, an RTP profile for audio and video conferencing with minimal control is defined in RFC 3551. This profile defines the set of static payload type assignments and the dynamic mechanism for mapping between payload formats and PT values using the Session Description Protocol (SDP). This latter mechanism is used in newer video codecs, such as the RTP payload format for H.264 video defined in RFC 6184, or the RTP payload format for High Efficiency Video Coding (HEVC) defined in RFC 7798.
[0179] An RTP session can be established for each multimedia stream. Audio and video streams can be implemented using separate RTP sessions, allowing the receiver to selectively receive components of a particular stream. Furthermore, the RTP specification can be configured to recommend port numbers for RTP, and additionally recommends using the next odd-numbered port number for the associated RTCP session. In applications using multiplexing protocols, a single port can be used for both RTP and RTCP.
[0180] Each RTP stream can include RTP packets, and each RTP packet can include an RTP header and a payload pair.
[0181] Enhanced Voice Service (EVS) is a mono voice codec standardized in 3GPP and described in the TS 26.445 specification document. This codec can operate in two modes: EVS master mode and EVS AMR-WB IO (Adaptive Multi-Rate Wideband Interoperability).
[0182] The IVAS codec can be considered an extension of the EVS codec, therefore the IVAS and EVS codecs can share some similarities in design and implementation. While the RTP payload structure has not yet been explicitly defined for the IVAS codec, it is expected to be similar to the RTP payload structure in EVS.
[0183] The EVS RTP payload format is described in Annex A of 3GPP TS 26.445. In EVS, the RTP payload format is divided into two different implementations: the Compact format and the Header-Full format. In the EVS Compact payload format, for EVS Master mode, the RTP packet consists of only a single EVS voice frame. For EVS AMR-WB IO mode, the compact RTP packet also includes a 3-bit Codec Mode Request (CMR) field before the voice frame. In the EVS Compact format, the different modes and bit rates used for voice frames are identified by the size of the RTP payload. For example, as shown in Table A.1 of Annex A of TS 26.445, a 328-bit RTP packet is assigned to EVS Master mode with a bit rate of 16.4 kbps.
[0184] In the EVS Header-Full format, the RTP payload consists of multiple voice frames accompanied by optional CMR bytes and a Table of Contents (ToC) byte. The CMR byte is used to request a change in the bit rate or coding mode that the receiver wants to receive. This request is sent as a CMR byte as part of the Header-Full EVS packet. In the EVS AMR-WB IOCompact format, the CMR function also appears as a 3-bit signaling at the beginning of the packet.
[0185] The ToC byte in the EVS Header-Full format describes the mode and bit rate used for the accompanying EVS-encoded speech frame. The RTP payload structure for the EVS Header-Full format is as follows: Figure 2 This is further shown in the middle.
[0186] For example, Figure 2 Payload structures 201, 203, 205, 207, and 208 are shown, all of which are in Header-Full format with a single ToC frame, multiple ToC frames, and a single ToC frame.
[0187] Each of the payload structures includes a ToC / CMR byte, wherein for each of these bytes, the ToC / CMR byte begins with a first bit 211, 221, 231, which can be used to distinguish the bytes between ToC (where the bit value is 0) and CMR (where the bit value is 1).
[0188] In payload structure 201, where there is a single voice frame in the payload, the voice data 215 is preceded by a single ToC byte 213 (with optional zero padding 217 at the end of the payload).
[0189] In the example, the payload structure of a single speech frame can have an optional CMR byte, as shown in payload structure 205. Payload structure 205 differs from payload structure 201 in that it has a CMR byte 232 (with its associated indicator bit 231 having a value of 1) preceding the ToC byte 213.
[0190] In some example structures, payload structures 203 and 207 can contain two voice frames. In this example, the additional ToC byte 223 (with its associated indicator bit 221 having a value of 0) and ToC byte 213 are located at the beginning of the payload, followed by voice data frames 215 and 225, in the corresponding order of the ToC bytes.
[0191] Furthermore, the multi-voice frame payload structure can have optional CMR bytes, as shown in payload structure 207. Payload structure 207 differs from payload structure 203 in that a CMR byte 232 (with its associated indicator bit 231 value of 1) is present before ToC bytes 213, 223.
[0192] The IVAS codec supports stereo and spatial / immersive formats at various bit rates. However, at the lowest IVAS bit rates (down to 13.2 kbps), the encoding quality of the audio signal is degraded due to the insufficient bit rate. Therefore, it typically has very limited capabilities to provide audio processing, rendering-related, or even non-audio-related auxiliary information within the codec payload.
[0193] The following embodiments therefore describe apparatus and methods for enabling the transmission and consumption of information related to the transmitted audio bitstream (other than the encoded IVAS bitstream). These examples are intended to precisely represent one or more associated encoded audio frames and their associated processing information, but allow for any subsequent extensions based on user- or application-defined requirements.
[0194] Therefore, in some embodiments, additional information can be sent as part of the IVAS RTP packet / stream to facilitate the rendering or consumption of the IVAS audio bitstream data. For example, this information can include information such as the sender's location, priority information for ISM objects, or other non-audio information. In some implementations, processing information can be signaled as part of the RTP header extension, using a one- or two-byte header determined by the amount of data carried as part of the processing information packet.
[0195] Therefore, the concept discussed in further detail in the embodiments herein is that the apparatus and methods for transmitting (audio) processing information (PI) include suitable components for providing audio processing information as RTP payload packets in addition to immersive audio-coded data packets (IVAS frames), and further transmitting them in the same RTP stream carrying the IVAS frames. In these embodiments, the described RTP payload structure is designed to enable in-band delivery of PI packets while achieving a compact header size for the delivery of immersive audio-coded data (IVAS frames) for precise bitrate allocation networks, and enabling general extension support for over-the-top (OTT) delivery networks without precise bitrate allocation constraints. The compact header size for IVAS is maintained by encapsulating the processing information as separate RTP packets or carrying it as separate packets in an RTP header extension. Therefore, there is no need to extend the IVAS header required for all IVAS frames.
[0196] In some embodiments, this can be implemented with regard to a suitable encoder configured to receive audio processing information (and, in some embodiments, update frequency information and the type of audio processing information). The audio processing information can be employed by a suitable decoder or renderer device to assist in the processing of encoded and grouped audio signals (e.g., audio signals within IVAS audio data packets). This assistance can be as simple as being able to tag sources and / or identify the order of audio sources or audio objects. Furthermore, this assistance can enable or control modifications to the processing or rendering of the audio signals. Then, as discussed in further detail herein, for RTP packets corresponding to update frequencies, the audio processing information is grouped with a ToC header identifying the packet type. Additionally, a PI type header including a predetermined number of bytes can then be appended. In some PI types, the payload size follows the PI type header. The PI payload can then be appended to the PI payload packet.
[0197] In some embodiments, the inclusion of processing information packets in the RTP stream is implicitly negotiated or signaled. For example, the IVAS codec can be configured to implicitly indicate the application or use of processing information packets. In some embodiments, specific processing information packets are explicitly indicated as part of session negotiation. Furthermore, whether processing information is packetized as part of an RTP header extension or transmitted as separate packets can be determined by specific session negotiation preferences or via any appropriate out-of-band signaling.
[0198] In other embodiments, the sending UE or SDP proposal creator can be configured to instruct processing information packets to support packetization. In some examples, the receiver can be configured to select which processing information packets are supported. The supported information packets can then be retained.
[0199] Similarly, in some embodiments, session negotiation can include deciding whether information packets are sent as part of an RTP HE or delivered as separate packets. Depending on implementation preferences, the sending UE and the receiving UE can negotiate a mutually acceptable method.
[0200] In another embodiment, processing information packets can be carried as part of the session description. This is more suitable for scenarios where processing information is expected to be less dynamic or where certain values need to be used as default starting values.
[0201] For example, session parameters can be attributes used for IVAS media streams. a=pi-frame-included; / / Indicates that pi frames will be used a=pi-frame-included:mode; / / The mode can be 0, which is equal to a separate RTP PI frame or a PI frame appended to an IVAS frame; a mode value of 1 indicates that the PI frame data is carried as part of the RTP header extension. a=pi-frame-included:mode: / /
[0202] Regarding the decoder, in some embodiments, RTP packets are received, and then the payload type is determined based on the ToC byte.
[0203] The embodiments described herein can be employed by transmitting (audio) processing information (PI) frames as separate payload frames, except that the encoded audio frames are part of an IVAS RTP packet stream. Frames can be identified or distinguished from encoded audio frames using a Content Table (ToC) byte located before or before the frame. In some embodiments, PI frames can be transmitted as separate IVAS RTP packets (i.e., PI frames are not in the same packet as IVAS audio frames), or PI frames can be in the same IVAS RTP payload packet along with other IVAS frames. In some embodiments, PI frames can be identified in a manner similar to that described above, wherein the identification is provided using a ToC byte. Furthermore, as previously stated, PI frames can be transmitted as part of an RTP header extension for IVAS frames.
[0204] As mentioned above, IVAS is expected to have two different formats for RTP packets, similar to EVS: Compact and Header-Full formats.
[0205] In the EVS Compact format, RTP packets consist only of encoded EVS voice frames. Similar behavior is defined in the IVAS Compact format, where RTP packets consist only of encoded IVAS frames. In some embodiments, payload differentiation is based on a unique packet size defined by a bit rate.
[0206] In the EVS Header-Full format, voice frames are accompanied by a Table of Contents (ToC) byte, which describes the mode and bit rate used for the frame. Additionally, packets can include a Codec Mode Request (CMR) byte, which is used to modify the encoding on the receiver / decoder side. In the IVAS Header-Full format, the modified ToC byte is included in the RTP packet, as described below.
[0207] Additionally, in EVS, the CMR byte can be used to request a change in the bit rate or encoding mode that the receiver wants to receive. This request can be sent as a CMR byte as part of a Header-Full EVS packet. In the EVS AMR-WB IO Compact format, the CMR function can also appear as a 3-bit signaling at the beginning of a packet, as described in the EVS description [TS 26.445, Annex A].
[0208] In the following embodiments, the ToC byte can be modified in the IVAS application to indicate the availability of any additional (audio) processing information (PI) within the data stream, and further identify the type of additional (audio) processing information (PI) within the data stream.
[0209] According to some embodiments, (IVAS) encoders (such as example encoder 101) are in Figure 3 It is shown in the middle.
[0210] In some embodiments, the encoder is configured to receive or otherwise acquire audio processing information 300, update frequency 302, and audio processing type information 304. Additionally, the encoder includes an RTP generator 301 configured to receive the audio processing information 300, update frequency 302, and audio processing type information 304, and generate RTP packets including a PI payload.
[0211] In some embodiments, the RTP generator 301 includes an audio processing information packetizer (where the ToC header identifies the packet type) 303.
[0212] In some embodiments, the RTP generator 301 further includes a PI type header appender 305, which is configured to append PI type information to the packet.
[0213] In some embodiments, the RTP generator 301 further includes a PI payload attacher 307. The PI payload attacher 307 is configured to attach a PI payload to the generated packet before sending it.
[0214] Figure 4 It shows Figure 3 The example RTP generator shown operates as illustrated.
[0215] The first operation is to acquire or receive audio processing information, update the frequency, and select one of the types of audio processing information, such as... Figure 4 As shown in step 401.
[0216] Then, for RTP packets corresponding to the update frequency, packets are generated, such as... Figure 4 As shown in step 403.
[0217] The grouping operation in step 403 can be, for example, divided into a first operation of grouping audio processing information using a ToC header that identifies the type of grouping, as shown in step 405.
[0218] The packet generation in step 403 may also include appending a PI type header containing a predetermined number of bytes. In some PI types, the payload size follows the PI type, as shown in step 407.
[0219] The group generation operation in step 403 also includes attaching the PI payload to the PI payload group, as shown in step 409.
[0220] A Header-Full IVAS frame includes one or more Content Table (ToC) indicators and their associated IVAS speech frames. Figure 5 An example grouping structure for the complete header IVAS payload format is shown.
[0221] The first example packet structure, as shown in 501, represents the payload structure when a single IVAS frame is encoded in Header-Full format. In this example, the IVAS audio frame 505 is preceded by a single ToC byte 503.
[0222] In addition, example packet structures 511 and 521 are shown, which represent alternative versions of the payload structure when two IVAS frames are encoded in Header-Full format.
[0223] Therefore, the first multi-IVAS frame packet structure 511 includes a structure in which all ToC bytes 513, 514 are located at the beginning of the payload and before IVAS frames 515, 516.
[0224] The additional multi-IVAS frame packet structure 521 includes ToC bytes 523, 527 that can be located before each corresponding or associated IVAS frame 525 and 529.
[0225] Example packet structures 501 and 511 are similar to the packet structures for the EVS codec described in TS 26.445, where all EVS ToC bytes precede the EVS speech frame. The ToC bytes differ in EVS and IVAS, but both describe the content of the speech frame within the packet and have similar high-level functionality in RTP packets.
[0226] Examples of the ToC byte structure of the EVS and IVAS Header-Full frame ToC byte structure according to some embodiments are shown below. Figure 6a and Figure 6b It is shown in the middle.
[0227] about Figure 6a According to some embodiments, the EVS Header-Full frame ToC byte structure 601 is shown. The header type (H) identifier bit 603 is used to distinguish the CMR byte from the ToC byte, the (F) bit 605 is used to indicate whether there is another ToC byte after this frame, and the frame type (FT) index 607 is used to indicate the EVS bit rate and mode of the frame.
[0228] Similarly, Figure 6bThe ToC byte structure 651 for IVAS frames is shown. Compared to the EVS ToC byte structure, the (H) bit is omitted because the CMR byte is not expected to be part of the IVAS packet. The ToC indicator describes the data in the IVAS speech frame, which can be used to deduce the content and size of each frame.
[0229] Therefore, in some embodiments, the bits in the IVAS ToC byte structure 651 can have the following structure: F (1 bit) 653: If set to 1, this bit indicates that the corresponding frame in this payload is followed by another voice or PI frame, meaning that this entry is followed by another ToC byte. If set to 0, this bit indicates that this frame is the last frame in this payload, and that there are no other header entries following this entry. Extra (2-3 bits) 655 (only when Extra is 2 bits) or 655 and 657 (when Extra is 3 bits): These bits can be reserved for future use. If 4 bits are reserved for the FT indicator, the additional 3 bits can be used in part to identify other frame content besides bit rate / frame size (e.g., SPEECH_LOST, NO_DATA, SID), as shown below. FT (4-5 bits) 659 (only when FT is 4 bits) or 659 and 657 (when FT is 5 bits): Frame type index. These bits can be configured to indicate the bit rate or other frame content indication used for the frame. The receiver can determine the size of the received frame directly from the bit rate or from a predetermined frame size (e.g., for SPEECH_LOST and NO_DATA frames) based on the content indication. The FT bits can indicate the bit rate of, for example, IVAS voice frames, SPEECH_LOST frames, NO_DATA frames, or comfort noise (SID) frames.
[0230] Examples of FT bit values (for 5 bits) are presented in the table below, which will result in additional bits. At this point in codec development, it has not been decided whether the three types (SPEECH_LOST, NO_DATA, SID) will be supported in the final codec. SPEECH_LOST, NO_DATA, and SID frames are part of the EVS codec, and at least the SID and NO_DATA frame types are likely to be included in the IVAS specification. However, it should be understood that the final IVAS codec specification may have different frame types than those described here.
[0231] In some embodiments, 4 bits are reserved for the FT portion. The frame type bit value and content indication when using 4 bits are presented below. In these embodiments, the FT portion identifies the bit rate, where the FT bits are bit values other than the defined values (e.g., 1111), and can be used to identify other aspects, such as SPEECH_LOST, NO_DATA, and SID frames, using a combination of the FT indicator OTHER (the defined value) and additional bits.
[0232] As shown above, when 5 bits are reserved for the FT indicator, 15 bits are available for future FT bit allocation (bit allocations 10001–11111). In some embodiments, when 4 bits are reserved for the FT indicator, and FT value 1111 is used, 5 bits are available for future use. FT value 1110, when combined with additional bits, provides an additional 8 bits for future use. These available bit allocations can be used in the future to indicate the content of a frame to be defined, such as a PI frame as described below.
[0233] These bit allocation values are merely examples, and it should be understood that in some embodiments, the bit allocation values can be configured in other ways.
[0234] In some embodiments, the name of the supplementary information frame can be different from the name of the audio processing information frame.
[0235] In some embodiments, additional (non-audio) data can be added to streaming IVAS RTP packets to enable better rendering or consumption of IVAS audio bitstream data. For example, in some embodiments, any external location from the transmitting UE side can be included and encoded into (audio) processing information (PI) frames. In another example, in ISM mode, the importance and / or content type of the ISM audio object can be included in the PI frame.
[0236] Figure 7 An example of how PI frames can be integrated into IVAS RTP packets is shown.
[0237] For example, option 701 is shown, in which the PI frame 705 is sent as a separate IVAS RTP packet. In other words, in some embodiments, the PI frame is sent within the same RTP packet without any IVAS voice (or any encoded audio) frame. This approach can be adopted if the bit rate limit for a single RTP packet is very tight. For example, in cases where adding the PI frame to the same packet as the IVAS voice frame would increase the size of the RTP packet too much (e.g., the resulting frame is larger than the maximum transmission unit (MTU) size used by the network).
[0238] Another option 711 is that the PI frame is integrated into the same RTP packet as the IVAS voice frame. In this example, all ToC bytes 713, 714 in the packet structure 711 are located at the beginning of the payload, which includes the PI frame 715 and the IVAS frame 716.
[0239] In some embodiments, the ToC bytes associated with the payload frame can be positioned such that the ToC bytes associated with the payload frame are adjacent to each other. Thus, as shown by structure 721, ToC byte 723 and the associated PI frame 725 are positioned together, and ToC byte 727 and the associated IVAS frame 729 are positioned together.
[0240] In some embodiments, when the PI frame does not contain critical data for rendering, a completely separate structure including the PI frame and IVAS frame payloads can be implemented. Such groups can be configured to be sent with lower priority.
[0241] For example, in cases where important information affecting rendering and / or that should be closely time-aligned with audio data (e.g., IVAS frames) is being sent, a grouping structure of combined PI and IVAS frame payloads can be used instead, as they guarantee the simultaneous delivery of audio frame (IVAS frame) and PI frame data.
[0242] about Figure 8 An example of how to integrate multiple IVAS voice frames or PI frames in the same RTP packet structure is shown.
[0243] For example, in some embodiments, such as those shown in example structures 801, 811, and 821, the ToC byte is positioned before the data frame.
[0244] For example, structure 801 is a structure in which ToC bytes 802, 803, and 804 precede the data frames, and these data frames are arranged such that PI frame 805 precedes all IVAS speech frames 806 and 807. In this case, the data in PI frame 805 can be applied to the first IVAS speech frame 806, or to all IVAS frames 806 and 807 in the RTP packet. For example, if the PI data contains orientation information, this orientation can be applied to the first IVAS frame, or to all IVAS frames. The interpretation of which IVAS frames the PI data should be applied to can depend on the content of the PI frame or the content specified in the codec specification.
[0245] In another example structure 811, ToC bytes 812, 813, and 814 precede the data frames, which are arranged such that the PI frame 816 lies between two IVAS speech frames 815 and 817. In this case, the PI data can be interpreted as being applied only to the latter IVAS speech frame 817.
[0246] Another example structure 821 is one in which ToC bytes 822, 823, and 824 precede the data frames, which are arranged such that the RTP packet includes two PI frames 825 and 826 preceding the IVAS frame 827. In some embodiments, these two PI frames include different types of data. For example, the first PI frame 825 may include orientation information, while the second PI frame 826 may include priority information for ISM audio objects.
[0247] In some embodiments, such as those shown in example structures 831, 841, and 851, the structure causes the ToC bytes to be interleaved between data frames and located in adjacent frames to which they are associated.
[0248] For example, in structure 831, ToC byte 832 is associated with and precedes PI frame 833, ToC byte 834 is associated with and precedes IVAS frame 835, and ToC byte 836 is associated with and precedes IVAS frame 837. In this case, the data in PI frame 833 can be applied to the first IVAS voice frame 835, or to all IVAS frames 835 and 837 in the RTP data packet (as described above with respect to example structure 801).
[0249] Another example structure 841 is one in which ToC byte 842 is associated with and precedes IVAS frame 843, ToC byte 844 is associated with and precedes PI frame 845, and ToC byte 846 is associated with and precedes IVAS frame 847. In this case, the data in PI frame 845 can only be applied to the subsequent IVAS voice frame 847.
[0250] Another example structure 851 is one in which ToC byte 852 is associated with and precedes PI frame 853, ToC byte 854 is associated with and precedes PI frame 855, and ToC byte 856 is associated with and precedes IVAS frame 857. In this example, the two PI frames contain different types of data.
[0251] PI frames are identified with the help of ToC bytes and can follow rules such as those regarding... Figure 6b The same structure is shown. As previously described and presented in the table above, in some embodiments (with 5 bits), the FT bits have 15 unused bit allocation slots, which can be defined. One (or more) of these bit values can be used to identify the PI frame; for example, a defined bit allocation (such as 10010) can be used as the PI frame identifier.
[0252] In some embodiments, when 4 bits are reserved for the FT indicator, one of the additional bits given above and the free available combination of the FT can be used to identify the PI frame, for example, the bit value (011+1111).
[0253] In one embodiment, and regarding Figure 9 As shown, PI frames can be structured in various ways.
[0254] For example, the example PI frame structure 901 is a structure in which the frame includes a size identifier or size field 902 at the beginning of the frame, followed by a frame type identifier or type field 903, and then PI data 904. In some embodiments, the size identifier information (or size field) is encoded using N bits, where N is a fixed integer (and can be predefined in the IVAS specification). The frame type identifier or type field can be configured to indicate the type of data carried by the PI frame. For example, the type can be orientation data, priority of an audio object in ISM mode, scene information, or any suitable non-audio information. The type identifier or field can also include usage and validity information, as explained later. In some embodiments, the type identifier or field can be encoded using a predefined number of bits.
[0255] Another example of the PI frame structure 911 is a structure in which the frame includes a frame type identifier 913, followed by PI data 914. In this example, frame size information is not encoded into the PI frame, but the size of the PI frame is predefined (e.g., in the IVAS specification). Therefore, the size of the PI frame is known to the receiver when reading RTP packets, ensuring that the correct number of bits is read for that frame.
[0256] Another example PI frame structure 921 is one in which the frame includes a size identifier 922 at the beginning of the frame, followed by PI data 924. The size identifier information is also encoded using N bits, where N is a fixed integer. In this example structure, the type of the PI frame cannot be modified within the frame. However, in some embodiments, different types can be defined using the FT bits in the ToC bytes. For example, when 5 bits are allocated for the FT indicator, or when 4 bits are allocated for the FT indicator in the case of a combination of additional bits and the FT indicator. Therefore, instead of having a single PI_FRAME frame content indicator, it is possible to have different FT bit allocations, or additional bits and FT combinations reserved for different types of PI frames.
[0257] Example FT bit usage (with 5 FT bits) is presented in the table below, where the PI_FRAME type can be used for any general PI data, the ORIENTATION_DATA type can be used for orientation-based data, the SCENE_INFO type can be used for scene information data, and the ISM_PRIORITY type can be used to indicate the frame of a priority audio object in an ISM stream.
[0258] In some embodiments, the combination of the additional 3 bits for the above frame type + FT indicator (4 bits) can be presented in the table below.
[0259] In some embodiments, the PI frame structure 931 includes only PI data 934. In these embodiments, the size of the PI frame is fixed, and the data type is predefined. In these embodiments, only the PI frame structure reduces the total number of bits required for the PI frame, but requires the use of freely available FT bits or combinations of extra bits + FT bits in the ToC bytes. However, this significantly limits possible size / type variations because they are predefined and fixed (e.g., by specification) or limited by the number of free FT bits or combinations of extra bits + FT bits.
[0260] Furthermore, an example structure 941 is shown, which is similar to structure 901, where the type indicator or field 943 precedes the size indicator or field 942. This approach allows for different numbers of bits to be reserved for the size field for different PI frame types.
[0261] For example, some PI types may have longer data segments, such as label names, which may require more bits to be reserved for the size field. Additionally, structure 941 allows for PI frames of arbitrary size (e.g., in the future, if new PI frame types are developed).
[0262] about Figure 10a and Figure 10b Examples of some embodiments of the structure based on the type indicator (or field or identifier) as described above are shown.
[0263] about Figure 10a An example structure of the type indicator 1001 (or field or identifier) in the PI frame is presented. For example, a type field can include the following:
[0264] Format 1003 – The format field describes the format and semantics of the PI frame payload. In other words, the format describes what type of data is carried in the PI frame, such as orientation data.
[0265] Use field 1005 – This field describes how the PI data should be used. For example, using field 1005 can describe how the current PI data should be applied to the next IVAS frame or all IVAS frames in an RTP packet. Alternatively, this field can describe how the PI data should not be considered in the rendering, i.e., how this data is more general information to the receiver.
[0266] Validity 1007 – The validity field describes how long a PI frame is valid at the receiver. For example, this field can describe how many processing frames the receiver should apply the PI data described in the frame, and how long to stop applying the data if no new PI data is received within X frames. Validity can also be indicated in other time formats besides the number of processing frames, such as in milliseconds. Furthermore, some audio frame values may be “held” while others may be only “transient.” This allows the renderer to perform rendering accordingly.
[0267] about Figure 10b Another example structure 1051 for a type indicator, wherein the type indicator also includes size information 1053.
[0268] In some embodiments, the maximum size of a PI frame can be limited to minimize the additional bits introduced by the PI frame. For example, the frame can be limited to a maximum size of K bits, which is common to all IVAS bit rates. Alternatively, the maximum size of the PI frame can be linked to the bit rate used by the IVAS voice frame, for example, by making the maximum number of bits used for the PI frame a certain percentage share of the number of bits used for the voice frame. Furthermore, the use of PI frames can be limited to use only at the highest IVAS bit rate to reduce the risk of adding excessive payload to the transmitted packets.
[0269] about Figure 11 The first example application of PI frames is shown. Figure 11 The example shown in (and the subsequent diagrams) illustrates that the size of the group (nBits 1100) is outside the type indicator (in other words, the type indicator is a similar to...). Figure 10a (Example in the example). However, in some embodiments, the size can be included in the type indicator, such as Figure 10b The examples in the document illustrate how other fields (format, usage, validity, and PI data) can be used in a PI frame.
[0270] Figure 11 An example of how a PI frame can be used to transmit azimuth data is shown. In the first example 1101, the PI frame includes a size (nbits) indicator 1103 and a type indicator 1105.
[0271] The type indicator or field can indicate the orientation format (ORIENTATION) 1121 and that the orientation data should be applied to all IVAS frames (ALL_FRAMES) 1123. The valid time period for orientation is set to infinity (INF) 1125. In other words, the specified orientation should be applied to all incoming frames until another orientation data frame arrives.
[0272] First example 1101 illustrates that orientation data is sent as a spherical index 1107. In some embodiments, the spherical index 1107 indicates a predefined orientation within a spherical grid. The accuracy of the encoded orientation depends on the resolution of the spherical grid. In other words, the number of orientations the grid has.
[0273] The second example 1111 encodes the direction directly as yaw angle 1116, pitch angle 1117, and roll angle 1118.
[0274] In some other embodiments, only one or some of the azimuth angles can be transmitted. These combinations can be identified by an identifier preceding the angle value. For example, these combinations can be identified by a 2-bit identifier, such as (00 – yaw only), (01 – pitch only), (10 – yaw and pitch), (11 – yaw, pitch, and roll).
[0275] In some embodiments, identifiers with a higher number of bits can be used to include more angle combinations.
[0276] about Figure 12 This demonstrates an example of how PI frames can be used to send priority for audio objects in the IVAS ISM format.
[0277] The number of bits (nBits) used for PI frames 1200 includes a size (nbits) indicator 1203 and a type indicator 1205.
[0278] In this embodiment, the type indicator 1205 includes a type identifier, and the ISM_Priority 1221 value identifies the PI frame as a priority identifier.
[0279] The type indicator 1205 also includes a use field or portion with a THIS_FRAME 1223 value, which indicates that the priority value to be transmitted should only be applied to the current IVAS frame, i.e., IVAS frames following this PI frame.
[0280] Furthermore, in this example, the effective time period is shown as zero frame with a value of 0_FRAMES 1225, which indicates that the priority value is only valid for the current frame.
[0281] Since the type usage field is shown as THIS_FRAME, the priority value will only be valid for the current frame, regardless of the value in the validity field.
[0282] Priority values (ISM_impN) 1206, 1207, 1208, and 1209 indicate the priority of audio objects in the IVAS ISM stream (by definition, it can have up to 4 objects).
[0283] about Figure 13 This demonstrates how PI frames can be used to indicate priority among multiple IVAS input formats.
[0284] The number of bits (nBits) 1300 used for the PI frame includes a size (nbits) indicator 1303, a type indicator 1305, and PI data in the form of an IVAS_format_priority value 1309.
[0285] In this example, the type indicator 1305 includes a type indicator IVAS_IF_PRIORITY value of 1321. Additionally, the type indicator includes a use field or portion with an ALL_FRAMES value of 1323, which indicates that the priority value transmitted should be applied to all frames following this PI frame.
[0286] Furthermore, in this example, the effective time period is shown as 50 frames with a value of 50_FRAMES 1325, which indicates that the priority value is only valid for IVAS formats present in the current frame. Therefore, the effective time period is set to 50 frames, meaning that if a new priority value is not set during the 50 processing frames, the priority value for the additional IVAS input format will remain unchanged during that time period.
[0287] Figure 14 and Figure 15 Further examples are shown of how PI frames can be used to send more general data that is not directly linked to rendering. For example, such data could be, for instance, tags for active speakers (such as those provided by...). Figure 14 (as shown) or tags for audio objects (such as those provided by...) Figure 15 (As shown). Other examples of more general data or information could be the ideal speaker setup for playback, or information related to how the transmitted audio was captured. For example, this capture information could be the constellation of the microphone array.
[0288] For example, Figure 14 Includes the size nBits 1403 indicator and the type indicator 1405.
[0289] In this embodiment, the type indicator 1405 includes a type identifier, and the SPEAKER_LABEL 1421 value indicates that the PI frame is a speaker tag identifier.
[0290] The type 1405 indicator also includes a use field or portion with a THIS_FRAME value of 1423, which indicates that the priority value to be transmitted should only be applied to the current IVAS frame, i.e., IVAS frames following this PI frame.
[0291] Furthermore, in this example, the effective time period is shown as the zero frame with a value of 0_FRAMES 1425, which indicates that the priority value is only valid for the current frame.
[0292] For example, Figure 15 Includes the size nBits 1503 indicator and the type indicator 1505.
[0293] In this embodiment, the type indicator 1505 includes a type identifier, and the ISM_LABEL 1521 value indicates that the PI frame is an ISM tag identifier.
[0294] The Type 1505 indicator also includes a use field or portion with a THIS_FRAME value of 1523, which indicates that the priority value sent should only be applied to the current IVAS frame, i.e., IVAS frames following this PI frame.
[0295] Furthermore, in this example, the effective time period is shown as the zero frame with a value of 0_FRAMES 1525, which indicates that the priority value is only valid for the current frame.
[0296] According to some embodiments of the (IVAS) decoder, such as example decoder 131, in Figure 16 It is shown in the middle.
[0297] In some embodiments, the decoder is configured to receive or otherwise acquire RTP packets 1600. Additionally, the decoder includes an RTP extractor 1601 configured to receive RTP packets 1600 and extract PI payloads from these RTP packets, and to enable the decoder to render audio signals based on the RTP packet (IVAS) payloads.
[0298] In some embodiments, the RTP extractor 1601 includes a payload type determiner 1603. The payload type determiner 1603 is configured to determine the PI information type from the RTP packets, thereby enabling the determination of PI information and the application of PI information based on the determined PI information type.
[0299] Figure 17 It shows Figure 16 The example RTP extractor operation is shown below.
[0300] The first operation is to acquire or receive one of the RTP packets, such as... Figure 17 As shown in step 1701.
[0301] Then, for the received RTP packets, the payload type is determined as follows: Figure 17 As shown in step 1703.
[0302] Then, the payload is applied based on a defined payload type, such as... Figure 17 As shown in step 1705.
[0303] about Figure 18 An example electronic device is shown. This device can be any suitable electronic device or apparatus. For example, in some embodiments, device 1800 can be a mobile device, user equipment, tablet computer, computer, audio playback device, etc.
[0304] In some embodiments, device 1800 includes at least one processor or central processing unit 1807. Processor 1807 can be configured to execute various program codes, such as those described herein.
[0305] In one embodiment, device 1800 includes memory 1811. In some embodiments, at least one processor 1807 is coupled to memory 1811. Memory 1811 can be any suitable storage component. In some embodiments, memory 1811 includes program code segments for storing program code implementable at processor 1807. Furthermore, in some embodiments, memory 1811 can also include stored data segments for storing data, such as data that has been processed or will be processed according to the embodiments described herein. The implemented program code stored in the program code segments and the data stored in the stored data segments can be retrieved by processor 1807 when needed via memory-processor coupling.
[0306] In one embodiment, device 1800 includes a user interface 1805. In some embodiments, user interface 1805 may be coupled to processor 1807. In some embodiments, processor 1807 may control the operation of user interface 1805 and receive input from user interface 1805. In some embodiments, user interface 1805 enables a user to input commands to device 1800, such as via a keyboard. In some embodiments, user interface 1805 enables a user to obtain information from device 1800. For example, user interface 1805 may include a display configured to display information from device 1800 to a user. In some embodiments, user interface 1805 may include a touchscreen or touch interface that enables information to be input to device 1800 and also displays information to the user of device 1800.
[0307] In some embodiments, device 1800 includes an input / output port 1809. In some embodiments, input / output port 1809 includes a transceiver. In these embodiments, the transceiver can be coupled to processor 1807 and configured to enable communication with other devices or electronic devices, such as via a wireless communication network. In some embodiments, the transceiver, or any suitable transceiver or transmitter and / or receiver component, can be configured to communicate with other electronic devices or devices via a wired or wired connection.
[0308] The transceiver can communicate with other devices via any suitable known communication protocol. For example, in some embodiments, the transceiver can use a suitable Universal Mobile Telecommunications System (UMTS) protocol, a Wireless Local Area Network (WLAN) protocol (such as, for example, IEEE 802.X), a suitable short-range radio frequency communication protocol (such as Bluetooth), or an Infrared Data Communication Channel (IRDA).
[0309] The transceiver input / output port 1809 can be configured to receive signals and, in some embodiments, acquire focus parameters as described herein.
[0310] In some embodiments, device 1800 may be employed to generate suitable audio signals using processor 1807 that executes appropriate code. Input / output port 1809 may be coupled to any suitable audio output device, such as a multi-channel speaker system and / or headphones (which may be head-tracking headphones or non-tracking headphones) or similar.
[0311] Generally, various embodiments of the present invention can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software that may be executable by a controller, microprocessor, or other computing device, but the invention is not limited thereto. While various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be well understood that, by way of non-limiting example, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0312] Embodiments of the present invention can be implemented by computer software executable by a data processor of a mobile device (such as in a processor entity), or by hardware, or by a combination of software and hardware. Furthermore, it should be noted in this regard that any block in the logical flow shown in the figures can represent a program step, or an interconnected logic circuit, block, or function, or a combination of program steps and logic circuits, blocks, and functions. Software can be stored on memory blocks implemented within memory chips or processors, magnetic media such as hard disks or floppy disks, and physical media such as optical media, for example, DVDs and their data variants, CDs.
[0313] The memory can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor can be of any type suitable for the local technical environment and, as a non-limiting example, can include one or more of the following: general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), gate-level circuits, and processors based on multi-core processor architectures.
[0314] Embodiments of the present invention can be practiced in various components, such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Complex and powerful software tools can be used to transform logic-level designs into semiconductor circuit designs ready to be etched and formed on a semiconductor substrate.
[0315] Programs such as those provided by Synopsys, based in Mountain View, California, and Cadence Design, based in San Jose, California, use established design rules and pre-stored design module libraries to automatically route and position components on semiconductor chips. Once the design for the semiconductor circuit is complete, the resulting standardized electronic format (such as Opus, GDSII, etc.) can be sent to a semiconductor manufacturing plant or "wafer fab" for fabrication.
[0316] The foregoing description, through exemplary and non-limiting examples, has provided a comprehensive and detailed description of exemplary embodiments of the invention. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and appended claims, given the foregoing description. Nevertheless, all such and similar modifications will still fall within the scope of the invention as defined by the appended claims. 3GPP Third Generation Partnership Project AMR-WB IO Adaptive Multi-Rate Broadband Interoperability PI Processing Information (Audio) EVS Enhanced Voice Service ISM has independent streams with metadata (i.e., based on the audio type of the object). IVAS Immersive Voice and Audio Services kbps RTP (Real-Time Transport Protocol) SDP Session Description Protocol
Claims
1. An apparatus comprising components configured as follows: Acquire at least one audio signal; Generate at least one audio signal data packet; Obtain processing information, the processing information including at least one parameter associated with the at least one audio signal, and the processing information is used to assist the processing of the at least one audio signal at the renderer; Generate at least one processing information group including the processing information; and The at least one processing information packet is associated with the at least one audio signal data packet in such a way that the at least one processing information packet is configured to assist the processing of at least one subsequent audio signal data packet.
2. The apparatus of claim 1, wherein the component is further configured to determine the size of the at least one processing information, wherein the component configured to generate at least one processing information group including the processing information is configured to generate the at least one processing information group including a size field, the size field indicating the size of the at least one processing information.
3. The apparatus of claim 1, wherein the component is further configured to: The type of the processing information is determined, wherein the component configured to generate at least one processing information group including the processing information is configured to generate the at least one processing information group including a type field, the type field indicating the type of the at least one processing information.
4. The apparatus of claim 1, wherein the component is further configured to: The size and type of the processing information are determined, wherein the component configured to generate at least one processing information group including the processing information is configured to generate the at least one processing information group including a size field and a type field, wherein the size field indicates the size of the at least one processing information; and the type field indicates the type of the at least one processing information.
5. The apparatus according to any one of claims 3 or 4, wherein the type of the processed information includes one of the following: A directional indicator parameter associated with the at least one audio signal; Scene information parameters associated with the at least one audio signal; An audio signal or stream priority parameter that indicates the priority of the audio signal within the at least one audio signal; Active speaker indicator; as well as Priority parameters of the priority audio object in the at least one audio signal.
6. The apparatus according to any one of claims 2 to 5, wherein the component configured to generate at least one processing information packet including the processing information is configured to generate at least one processing information packet including at least one of the following: Indicators are used, configured to describe how the processed information should be used; and A validity indicator is configured to describe how long the processed information is valid.
7. The apparatus according to any one of claims 1 to 6, wherein the processing information includes one of the following: Orientation parameters associated with the at least one audio signal; Scene information parameters associated with the at least one audio signal; An audio signal or stream priority parameter that indicates the priority of the audio signal within the at least one audio signal; Active speaker indicator; as well as Priority parameters of the priority audio object in the at least one audio signal.
8. The apparatus according to any one of claims 1 to 7, wherein the component is further configured to: Generate at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information; and The header of the at least one processing information packet is appended to the at least one processing information packet in such a way that the at least one processing information packet is identified in another device as including at least one processing data packet containing processing information rather than audio data.
9. The apparatus of claim 8, wherein the component configured to attach the at least one processing information packet header to the at least one processing information packet is configured to: The header of the at least one processed information group is appended to the first block, and the at least one processed information group is appended to a separate second block; and The header of the at least one processing information group is immediately appended before the associated processing information group of the at least one processing information group.
10. The apparatus according to any one of claims 1 to 9, wherein the component configured to associate the at least one processing information packet with the at least one audio signal data packet is configured to perform one of the following: Append each of the at least one processing information group to the preceding at least one audio signal data group; or The at least one processing information group is immediately appended to the at least one associated processing information group.
11. The apparatus according to any one of claims 1 to 10, wherein the at least one audio signal data packet is an immersive voice and audio service data packet.
12. The apparatus according to any one of claims 1 to 11, wherein the component is further configured to transmit the at least one audio signal data packet and the at least one audio processing data packet as a real-time transport protocol.
13. An apparatus comprising a component configured as follows: Receive an audio packet stream, the audio packet stream comprising: At least one audio signal data packet, wherein the at least one audio signal data packet includes at least one audio signal; At least one processing information group, the at least one processing information group including processing information, the processing information including at least one parameter associated with the at least one audio signal; as well as The at least one audio signal is processed based on the processing information, such that the at least one processing information group is configured to assist the processing of at least one subsequent audio signal data group.
14. The apparatus of claim 13, wherein the at least one processing information group includes a size field indicating the size of the at least one processing information.
15. The apparatus of claim 13, wherein the at least one processing information group includes a type field indicating the type of processing information within the at least one processing information group, and wherein the component configured to process the at least one audio signal based on the processing information is configured to identify the type of processing information within the at least one processing information group based on the type field.
16. The apparatus of claim 13, wherein the at least one processed information packet comprises: A size field, wherein the size field indicates the size of the at least one group of processed information; as well as A type field indicating the type of the processing information within the at least one processing information group, wherein the component configured to process the at least one audio signal based on the processing information is configured to identify the type of the processing information within the at least one processing information group.
17. The apparatus according to any one of claims 15 or 16, wherein the type of the processed information includes one of the following: A directional indicator parameter associated with the at least one audio signal; Scene information parameters associated with the at least one audio signal; An audio signal or stream priority parameter that indicates the priority of the audio signal within the at least one audio signal; Active speaker indicator; as well as Priority parameters of the priority audio object in the at least one audio signal.
18. The apparatus according to any one of claims 14 to 17, wherein the at least one processed information packet further comprises at least one of the following: Indicators are used, configured to describe how the processed information should be used; and A validity indicator is configured to describe how long the processing information is valid with respect to the processing of the at least one subsequent audio signal data packet.
19. The apparatus according to any one of claims 13 to 18, wherein the processing information includes one of the following: Orientation parameters associated with the at least one audio signal; Scene information parameters associated with the at least one audio signal; An audio signal or stream priority parameter that indicates the priority of the audio signal within the at least one audio signal; Active speaker indicator; as well as A priority parameter indicating the priority audio object in the at least one audio signal.
20. The apparatus of any one of claims 13 to 19, wherein the audio data packet stream further comprises at least one processing information packet header, the at least one processing information packet header including an indicator that the at least one processing information packet includes processing information; and 21. The apparatus of claim 8, wherein the at least one processing information packet header and the at least one processing information packet are arranged in one of the following: The at least one processed information packet header is in the first block, and the at least one processed information packet is in a separate second block; and The header of the at least one processing information group immediately precedes the associated processing information group of the at least one processing information group.
22. The apparatus according to any one of claims 13 to 21, wherein the at least one processing information packet and the at least one audio signal data packet are arranged in one of the following: Each of the at least one processed information packet precedes the at least one audio signal data packet; and The at least one processing information packet immediately precedes the at least one associated audio signal data packet.
23. The apparatus according to any one of claims 13 to 22, wherein the at least one audio signal data packet is an immersive voice and audio service data packet.
24. The apparatus according to any one of claims 13 to 23, wherein the component is configured to receive the audio packet stream as a real-time transport protocol transmission.
25. A method for an apparatus, the method comprising: Acquire at least one audio signal; Generate at least one audio signal data packet; Obtain processing information, the processing information including at least one parameter associated with the at least one audio signal, and the processing information is used to assist the processing of the at least one audio signal at the renderer; Generate at least one processing information group that includes the processing information; as well as The at least one processing information packet is associated with the at least one audio signal data packet in such a way that the at least one processing information packet is configured to assist the processing of at least one subsequent audio signal data packet.
26. A method for an apparatus, the method comprising: Receive an audio packet stream, the audio packet stream comprising: At least one audio signal data packet, wherein the at least one audio signal data packet includes at least one audio signal; At least one processing information packet, the at least one processing information packet including processing information, the processing information including at least one parameter associated with the at least one audio signal; and The at least one audio signal is processed based on the processing information, such that the at least one processing information group is configured to assist the processing of at least one subsequent audio signal data group.