Signaling of audio effect metadata in bitstream

By carrying audio effect metadata in the bitstream, parsing and applying sound field manipulation effects, the problems of high bit rate encoding in surround sound formats and difficult user manipulation are solved, and sound field reproduction with free manipulation by users and post-effect modification by creators is realized.

CN114631332BActive Publication Date: 2025-10-10QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080073035.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-04
Filing Date
2020-10-29
Publication Date
2025-10-10
Estimated Expiration
2040-10-29

AI Technical Summary

Technical Problem

Existing technologies require high bit rates in surround sound format encoding, and users cannot freely manipulate the sound field, which hinders users from experiencing other aspects of the original sound field and makes it difficult for content creators to change the effects after production.

Method used

By carrying audio effect metadata in the bitstream and parsing effect identifiers and parameter values, it allows users and creators to manipulate the sound field, including effects such as focus, zoom, rotation, and pan, and supports user interaction and restriction mechanisms.

Benefits of technology

It enables users to freely manipulate the sound field experience, supports content creators to change effects after production, ensures that the sound field reproduction meets the creator's intentions, and allows users to choose original or effect-processed audio versions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631332B_ABST
    Figure CN114631332B_ABST
Patent Text Reader

Abstract

Methods, systems, computer-readable media, and apparatuses for manipulating a soundfield are presented. Some configurations include receiving a bitstream including metadata and a soundfield description; parsing the metadata to obtain an effect identifier and at least one effect parameter value; and applying an effect identified by the effect identifier to the soundfield description. The applying can include applying the identified effect to the soundfield description using the at least one effect parameter value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Greek Provisional Patent Application No. 20190100493, filed on November 4, 2019, entitled “SIGNALLING OF AUDIO EFFECT METADATA IN ABITSTREAM,” the entire contents of which are hereby incorporated by reference.

[0003] Aspects of the present disclosure relate to audio signal processing. Background Art

[0004] The development of surround sound has made many entertainment output formats available today. The range of surround sound formats on the market includes the popular 5.1 home theater system format, which has been most successful in entering the living room beyond stereo. This format includes the following six channels: left front (L), right front (R), center or front center (C), left rear or left surround (Ls), right rear or right surround (Rs), and low frequency effects (LFE). Other examples of surround sound formats include the increasingly popular 7.1 format and the futuristic 22.2 format developed by NHK (Nippon HosoKyokai or Japan Broadcasting Corporation) for use in, for example, ultra-high-definition television standards. Surround sound formats may be required to encode audio in two dimensions (2D) and / or three dimensions (3D). However, these 2D and / or 3D surround sound formats require high bit rates to correctly encode the audio in 2D and / or 3D.

[0005] In addition to channel-based formats, new audio formats for enhanced reproduction are becoming available, such as object-based and scene-based (e.g., High-Order Ambisonics or HOA) codecs. Audio objects encapsulate individual pulse code modulation (PCM) audio streams, along with their three-dimensional (3D) position coordinates and other spatial information (e.g., object coherence) encoded as metadata. The PCM stream is typically encoded using, for example, a transform-based scheme (e.g., MPEG Layer 3 (MP3), AAC, MDCT-based coding). Metadata can also be encoded for transmission. On the decoding and rendering side, the metadata is combined with the PCM data to recreate the 3D sound field.

[0006] Scene-based audio is often encoded using an Ambisonics format such as B-format. The channels of a B-format signal correspond to the spherical harmonics of the sound field, rather than the speaker feeds. A first-order B-format signal has up to four channels (one omnidirectional channel W and three directional channels X, Y, Z); a second-order B-format signal has up to nine channels (four first-order channels and five additional channels R, S, T, U, V); and a third-order B-format signal has up to 16 channels (nine second-order channels and seven additional channels K, L, M, N, O, P, Q).

[0007] Advanced audio codecs (e.g., object-based codecs or scene-based codecs) can be used to represent the sound field (i.e., the distribution of air pressure in space and time) over an area to support multi-directional and immersive reproduction. Incorporating head-related transfer functions (HRTFs) into the rendering process can be used to enhance these qualities for headphones. Summary of the Invention

[0008] According to a general configuration, a method for manipulating a sound field includes receiving a bitstream including metadata and a sound field description; parsing the metadata to obtain an effect identifier and at least one effect parameter value; and applying the effect identified by the effect identifier to the sound field description. Applying may include applying the identified effect to the sound field description using the at least one effect parameter value. Also disclosed is a computer-readable storage medium including code that, when executed by at least one processor, causes the at least one processor to perform the method.

[0009] According to a general configuration, an apparatus for manipulating a sound field includes: a decoder configured to receive a bitstream including metadata and a sound field description and parse the metadata to obtain an effect identifier and at least one effect parameter value; and a renderer configured to apply an effect identified by the effect identifier to the sound field description. The renderer is configured to use the at least one effect parameter value to apply the identified effect to the sound field description. Also disclosed is an apparatus including a memory configured to store computer-executable instructions and a processor coupled to the memory and configured to execute the computer-executable instructions to perform these parsing and rendering operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The various aspects of the present disclosure are illustrated by way of example.In the drawings, like reference numerals denote similar elements.

[0011] Figure 1 An example of user orientation for manipulating a sound field is shown.

[0012] Figure 2A Describes the sequence of audio content generation and reproduction.

[0013] Figure 2B A sequence of audio content generation and reproduction according to a general configuration is depicted.

[0014] Figure 3A A flow chart of method M100 according to a general configuration is shown.

[0015] Figure 3B An example of two metadata fields related to audio effects is shown.

[0016] Figure 3C An example of three metadata fields related to audio effects is shown.

[0017] Figure 3D An example of a value table for the effect identifier metadata field is shown.

[0018] Figure 4A An example of a sound field comprising three sound sources is shown.

[0019] Figure 4B Shows the Figure 4A The result of focusing the sound field.

[0020] Figure 5A An example of rotating the sound field relative to a reference direction is shown.

[0021] Figure 5B An example of replacing the reference direction of the sound field with a different direction is shown.

[0022] Figure 6A Examples of desired panning of the sound field and user position are shown.

[0023] Figure 6B Shows the desired translation applied to Figure 6A The result of the sound field.

[0024] Figure 7A An example of three metadata fields related to audio effects is shown.

[0025] Figure 7B An example of four metadata fields related to audio effects is shown.

[0026] Figure 7C A block diagram of an implementation M200 of method M100 is shown.

[0027] Figure 8A An example of a user wearing a user tracking device is shown.

[0028] Figure 8B Motion (eg, of a user) in six degrees of freedom (6DOF) is illustrated.

[0029] Figure 9A An example of a restriction flag metadata field associated with a plurality of effect identifiers is shown.

[0030] Figure 9B An example of a plurality of restriction flag metadata fields, each associated with a corresponding effect identifier, is shown.

[0031] Figure 9C An example of a restriction flag metadata field associated with a duration metadata field is shown.

[0032] Figure 9D An example of audio effect metadata encoded within an extension payload is shown.

[0033] Figure 10 An example of different levels of scaling and / or zeroing for different hotspots is shown.

[0034] Figure 11A An example of a soundfield including five sound sources around a user’s position is shown.

[0035] Figure 11B An example of a result of performing an angle compression operation on a soundfield of Figure 11A is shown.

[0036] Figure 12A A block diagram of a system according to a general configuration is shown.

[0037] Figure 12B A block diagram of a device A100 according to a general configuration is shown.

[0038] Figure 12C A block diagram of an implementation A200 of a device A100 is shown.

[0039] Figure 13A A block diagram of a device F100 according to a general configuration is shown.

[0040] Figure 13B A block diagram of an implementation F200 of a device F100 is shown.

[0041] Figure 14 An example of a scene space is shown.

[0042] Figure 15 An example 400 of a VR device is shown.

[0043] Figure 16 is a diagram depicting an example of an implementation 800 of a wearable device.

[0044] Figure 17 A block diagram of a system 900 that can be implemented within a device is shown. DETAILED DESCRIPTION

[0045] The sound field as described herein can be two-dimensional (2D) or three-dimensional (3D). The one or more arrays for capturing the sound field may include a linear transducer array. Additionally or alternatively, the one or more arrays may include a spherical transducer array. One or more arrays may also be positioned within the scene space, and such arrays may include arrays with fixed positions and / or arrays with positions that may change during the event (e.g., mounted on a person, wire, or drone). For example, one or more arrays within the scene space may be mounted on a person participating in the event, such as an athlete and / or official (e.g., a referee) at a sporting event, a performer and / or an orchestra conductor at a music event, and the like.

[0046] Multiple distributed transducer arrays (e.g., microphones) can be used to record the sound field in order to obtain a sound field in a large scene space (e.g., Figure 14 Spatial audio can be captured on a scene space (e.g., a baseball stadium, football field, cricket field, etc.) as shown. For example, capture can be performed using one or more arrays of sound sensing transducers (e.g., microphones) located outside the scene space (e.g., along the periphery of the scene space). These arrays can be positioned (e.g., oriented and / or distributed) so that certain areas of the sound field are sampled more densely or less densely than other areas (e.g., depending on the importance of the area of ​​interest). Such positioning can change over time (e.g., corresponding to changes in the focus of interest). The arrangement can vary depending on the size / type of venue or with maximum coverage and reduced blind spots. The generated sound field can include audio that has been captured from another source (e.g., a commentator in a broadcast booth) and added to the sound field of the scene space.

[0047] Audio formats that provide more accurate sound field modeling (e.g., object- and scene-based codecs) can also allow for spatial manipulation of the sound field. For example, a user may prefer to change the reproduced sound field in any one or more of the following ways: making sounds from a particular direction louder or softer than sounds from other directions; hearing sounds from a particular direction more clearly than sounds from other directions; hearing sounds from only one direction and / or muting sounds from a particular direction; rotating the sound field; moving a sound source within the sound field; or moving the user's position within the sound field. For example, user selections or modifications as described herein can be performed using a mobile device (e.g., a smartphone), a tablet device, or any other interactive device or devices.

[0048] It can be similar to selecting a region of interest in an image or video (e.g. Figure 1Such user interaction or pointing (e.g., soundfield rotation, zooming into the audio scene) can be performed in a manner similar to that shown in FIG. 1 1 (e.g., in a manner similar to that shown in FIG. 1 1 ). The user can indicate a desired audio operation on the touchscreen, e.g., by performing a spread ("reverse pinch" or "pinch out") or touch-and-hold gesture to indicate a desired zoom, a touch-and-drag gesture to indicate a desired rotation, etc. The user can indicate a desired audio manipulation by gesture (e.g., for optical and / or sound detection): by moving her finger or hand in a desired direction to indicate a zoom, by performing a grab-and-move gesture to indicate a desired rotation, etc. The user can indicate a desired audio manipulation by changing the position and / or orientation of a handheld device (e.g., a smartphone or other device equipped with an inertial measurement unit (IMU) (e.g., including one or more accelerometers, gyroscopes, and / or magnetometers)) that is capable of recording such changes.

[0049] Although audio manipulation (e.g., zooming, focusing) is described above as a purely consumer-side process, content creators can wish to be able to apply such effects during the production of media content that includes a soundfield. Examples of such produced content can include recordings of live events, such as sports or musical performances, as well as recordings of scripted events (e.g., movies or plays). The content can be audiovisual (e.g., a video or a movie), or purely audio (e.g., a recording of a concert), and can include one or both of recorded (i.e., captured) audio and generated (e.g., synthesized, meaning synthesized rather than captured) audio. Content creators can want to manipulate recorded and / or generated soundfields for a variety of reasons, such as for dramatic effect, to provide emphasis, to direct the listener's attention, to improve intelligibility, etc. The product of such processing is audio content (e.g., a file or bitstream) with pre-applied audio effects (like Figure 2A as shown in FIG. 1 1 ).

[0050] While producing audio content in this form can ensure that the soundfield can be reproduced as intended by the content creator, such production can also prevent the user from experiencing other aspects of the originally recorded soundfield. For example, the result of a user's attempt to zoom in on a certain region of the soundfield can not be optimal, as audio information for that region can no longer be available in the generated content. Producing audio content in this manner can also prevent the consumer from being able to reverse the creator's manipulations, and can even prevent the content creator from being able to modify the produced content in a desired manner. For example, the content creator can be dissatisfied with the audio processing, and can want to change the effects after the fact. As the audio information needed to support such changes can be lost during production, being able to change the effects after production can require that the original soundfield be stored separately as a backup (e.g., the creator can be required to maintain a separate soundfield archive prior to applying the effects).

[0051] Systems, methods, apparatuses, and devices as disclosed herein can be implemented to send intended audio operations as metadata. For example, captured audio content can be stored in its original format (i.e., without intended audio effects), and the intended audio effects behavior of the creator can be stored as metadata in the bitstream. A consumer of the content can decide whether she wants to listen to the original audio or to the audio with the intended creator's audio effects (as shown in FIG. 6, for example). If the consumer selects the creator's audio effects version, then the audio rendering will process the audio based on the signaled audio effects behavior metadata. If the consumer selects the original version, then the consumer can also be allowed to freely apply audio effects to the original audio stream. Figure 2B

[0052] Some illustrative configurations will now be described with reference to the drawings, which form a part of the specification. Although specific configurations are described in which one or more aspects of the disclosure can be implemented, other configurations can be used and various modifications can be made without departing from the scope of the disclosure or the spirit of the appended claims.

[0053] ​Unless expressly limited by its context, the term "signal" is used herein to indicate any of its ordinary meanings, including the state of a memory location (or a group of memory locations) represented on a line, bus, or other transmission medium. Unless expressly limited by its context, the term "generate" is used herein to indicate any of its ordinary meanings, such as calculating or otherwise producing. Unless expressly limited by its context, the term "calculate" is used herein to indicate any of its ordinary meanings, such as calculating, evaluating, estimating, and / or selecting from a plurality of values. Unless expressly limited by its context, the term "obtain" is used herein to indicate any of its ordinary meanings, such as calculating, deriving, receiving (e.g., from an external device), and / or retrieving (e.g., from a memory array element). Unless expressly limited by its context, the term "select" is used herein to indicate any of its ordinary meanings, such as identifying, indicating, applying, and / or using at least one of a group of two or more, but not all. Unless expressly limited by its context, the term "determine" is used herein to indicate any of its ordinary meanings, such as deciding, establishing, concluding, calculating, selecting, and / or evaluating. When the term "comprising" is used in this specification and claims, it does not exclude other elements or operations. The term "based on" (such as "A is based on B") is used to mean any of its ordinary meanings, including the following: (i) "derived from" (for example, "B is a precursor of A"), (ii) "based at least on" (for example, "A is at least based on B"), and, where appropriate in the particular context, (iii) "equal to" (for example, "A is equal to B"). Similarly, the word "responsive to" is used to mean any of its ordinary meanings, including "at least responsive to". Unless otherwise stated, the terms "at least one of A, B, and C", "one or more of A, B, and C", "at least one of A, B, and C", and "one or more of A, B, C" mean "A and / or B and / or C". Unless otherwise stated, the terms "each of A, B, and C" and "each of A, B, and C" mean "A and B and C".

[0054] Unless otherwise stated, any disclosure of the operation of an apparatus having particular features is also expressly intended to disclose a method having similar features (and vice versa), and any disclosure of the operation of an apparatus according to a particular configuration is also expressly intended to disclose a method according to a similar configuration (and vice versa). The term "configuration" may be used to refer to a method, apparatus, and / or system indicated by its specific context. Unless the specific context indicates otherwise, the terms "method," "process," "procedure," and "technique" may be used generically and interchangeably. A "task" with multiple subtasks is also a method. Unless the specific context indicates otherwise, the terms "apparatus" and "device" may also be used generically and interchangeably. The terms "element" and "module" are generally used to represent a portion of a larger configuration. Unless expressly limited by its context, the term "system" is used herein to mean any of its ordinary meanings, including "a group of elements that interact to serve a common purpose."

[0055] Unless initially introduced by a definite article, ordinal terms (e.g., "first," "second," "third," etc.) used to modify a claim element do not, by themselves, denote any priority or order of the claim element relative to another claim element, but rather simply distinguish the claim element from another claim element having the same name (but using the ordinal term). Unless expressly limited by its context, each of the terms "plurality" and "set" is used herein to denote an integer quantity greater than one.

[0056] Figure 3A A flowchart of a method M100 for manipulating a sound field according to a general configuration including tasks T100, T200, and T300 is shown. Task T100 receives a bitstream including metadata (e.g., one or more metadata streams) and a sound field description (e.g., one or more audio streams). For example, the bitstream can include separate audio and metadata streams formatted to conform to International Telecommunication Union Recommendation (ITU-R) BS2076-1 (Audio Definition Model, June 2017).

[0057] The sound field description can be based on predetermined regions of interest within the sound field, for example, including different audio streams for different regions (e.g., object-based schemes for some regions and HOA schemes for other regions). For example, it may be desirable to encode regions with a high degree of wavefield concentration using object-based or HOA schemes, and to encode regions with a low degree of wavefield concentration (e.g., ambiance, crowd noise, applause) using HOA or plane wave expansion.

[0058] Object-based schemes may reduce sound sources to point sources and may not preserve directional patterns (e.g., directional changes in sound emitted by, for example, a shouting musician or a trumpet player). HOA schemes (more generally, coding schemes based on a hierarchical set of basis function coefficients) are generally more efficient than object-based schemes when encoding a large number of sound sources (more objects can be represented by smaller HOA coefficients compared to object-based schemes). Benefits of using HOA schemes may include being able to evaluate and / or represent the sound field for different listener positions without the need to detect and track individual objects. Rendering of HOA-encoded audio streams is generally flexible and independent of the speaker configuration. HOA encoding is also generally efficient under free-field conditions, so that translation of the user's virtual listening position can be performed within the active area close to the nearest source.

[0059] Task T200 parses the metadata to obtain an effect identifier and at least one effect parameter value. Task T300 applies the effect identified by the effect identifier to the sound field description. The information signaled in the metadata stream may include the type of audio effect to be applied to the sound field: for example, any one or more of focus, zoom, zeroing, rotation, and panning. For each effect to be applied, the metadata may be implemented to include a corresponding effect identifier ID10 identifying the effect (for example, a different value corresponding to each of zoom, zeroing, focus, rotation, and panning; a mode indicator for indicating a desired mode, such as conference or meeting mode, etc.). Figure 3D An example of a value table for effect identifier ID10 is shown that assigns unique identifier values ​​to each of a plurality of different audio effects and also provides signaling of one or more special configurations or modes (e.g., a conference or meeting mode as described below; a transition mode, such as a fade in or fade out; a mode for mixing one or more sound sources and / or mixing one or more other sound sources; a mode for enabling or disabling reverb and / or equalization, etc.).

[0060] For each identified effect, the metadata may include a corresponding set of effect parameter values ​​PM10 (e.g., Figure 3B ). For example, such parameters may include: an indication of a region of interest for the associated audio effect (e.g., the spatial direction and the size and / or width of the region); one or more values ​​of effect-specific parameters (e.g., the intensity of the focus effect); etc. Examples of these parameters are discussed in more detail below with reference to specific effects.

[0061] It may be desirable to allocate more bits of the metadata stream to carry parameter values ​​for one effect than for another. In one example, the number of bits allocated for the parameter value of each effect is a fixed value for the encoding scheme. In another example, the number of bits allocated for the parameter value of each identified effect is indicated within the metadata stream (e.g., as Figure 3C ).

[0062] A focusing effect can be defined as an enhanced directionality for a particular source or region. Parameters defining how the desired focusing effect is applied may include: the direction of the focused region or source, the strength of the focusing effect, and / or the width of the focused region. The direction may be indicated in three dimensions, for example, as an azimuth and elevation angle corresponding to the center of the region or source. In one example, the focusing effect is applied during rendering by decoding the focused source or region at a higher HOA order (more generally, by adding one or more levels of a hierarchical set of basis function coefficients) and / or by decoding other sources or regions at a lower HOA order. Figure 4A An example of a focused sound field on the source SS10 is shown, Figure 4B An example of the same sound field after applying a focusing effect is shown (it should be noted that the sound sources shown in the sound field figures herein may indicate, for example, audio objects in an object-based representation or virtual sources in a scene-based representation). In this example, the focusing effect is applied by increasing the directionality of source SS10 and increasing the diffusivity of other sources SS20 and SS30.

[0063] A zoom effect can be applied to increase the sound level of the sound field in a desired direction. Parameters defining how the desired zoom effect is applied may include: the direction of the area to be boosted. The direction may be indicated in three dimensions, for example, as an azimuth and elevation angle corresponding to the center of the area. Other parameters defining the zoom effect that may be included in the metadata may include: one or both of: the strength of the level boost and the size (e.g., width) of the area to be boosted. For a zoom effect implemented using a beamformer, defining the parameters may include: selecting a beamformer type (e.g., FIR or IIR); selecting a set of beamformer weights (e.g., one or more series of tap weights); time-frequency masking values; and the like.

[0064] A nulling effect may be applied to reduce the sound level of the sound field in a desired direction. Parameters defining how the desired nulling effect is applied may be similar to parameters defining how the desired scaling effect is applied.

[0065] A rotation effect can be applied by rotating the sound field to a desired orientation. The parameters defining the desired rotation of the sound field may indicate the direction to rotate to a defined reference orientation (e.g., Figure 5AAlternatively, the desired rotation may be indicated as a rotation of the reference direction to a different specified direction within the sound field (e.g., as shown in Figure 5B (as shown in the equivalent).

[0066] A panning effect may be applied to pan a sound source to a new position within the sound field. Parameters defining the desired panning may include direction and distance (or, rotation angle relative to the user's position). Figure 6A An example of a sound field with three sound sources SS10, SS20, SS30 and a desired pan TR10 of source SS20 is shown; Figure 6B The sound field after applying a panning TR10 is shown.

[0067] Each sound field modification indicated in the metadata may be linked to a specific moment in the sound field stream (e.g. by a timestamp included in the metadata, e.g. Figure 7A and 7B ). For implementations that indicate more than one sound field modification at a shared timestamp, the metadata may also include information identifying the temporal priority between the modifications (e.g., "apply the indicated rotation effect to the sound field, then apply the indicated focus effect to the rotated sound field").

[0068] As described above, it may be desirable to enable the user to select an original version of the sound field or a version modified by the audio effects metadata, and / or to modify the sound field in a manner that differs partially or entirely from the effects indicated in the effects metadata. The user may actively indicate such commands: for example, on a touch screen, by gestures, by voice commands, etc. Alternatively or in addition, the user commands may be generated by passive user interaction via a device that tracks the user's movements and / or orientation (e.g., a user tracking device that may include an inertial measurement unit (IMU)). Figure 8A One example of such a device, UT 10, is shown that also includes a display and headphones. An IMU may include one or more accelerometers, gyroscopes, and / or magnetometers to indicate and quantify motion and / or direction.

[0069] Figure 7C A flowchart of an implementation scheme M200 of method M100 is shown, which includes an implementation scheme T350 of task T400 and task T300. Task T400 receives at least one user command (e.g., through active and / or passive user interaction). Based on at least one of (A) at least one effect parameter value or (B) at least one user command, task T350 applies the effect identified by the effect identifier to the sound field description. Method M200 can be performed, for example, by an implementation of a user tracking device UT10 that receives audio and metadata streams and generates corresponding audio to the user via headphones.

[0070] To support an immersive VR experience, it may be necessary to adjust the provided audio environment in response to changes in the listener's virtual position. For example, it may be necessary to support six degrees of freedom (6DOF) virtual movement. Figure 8A and 8B As shown in , 6DOF includes three rotational motions of 3DOF and three translational motions: forward / backward (surge), up / down (heave), and left / right (sway). Examples of 6DOF applications include: remote users virtually participating in spectator activities such as sporting events (e.g., baseball games). For users wearing devices such as user tracking device UT10, it may be desirable to perform sound field rotation based on passive user commands generated by the device UT10 (e.g., indicating the user's current forward viewing direction as the desired reference direction of the sound field), rather than based on the rotation effect indicated by the content creator in the metadata stream as described above.

[0071] It may be desirable to allow content creators to limit the extent to which effects described in metadata can be changed downstream. For example, it may be desirable to impose spatial restrictions that allow users to apply effects only in specific areas and / or prevent users from applying effects in specific areas. Such restrictions may apply to all signaled effects or a specific set of effects, or the restrictions may apply to only a single effect. In one example, a spatial restriction allows a user to apply a zoom effect only in a specific area. In another example, a spatial restriction prevents a user from applying a zoom effect in another specific area (e.g., a confidential and / or private area). In another example, it may be desirable to impose temporal restrictions that allow users to apply effects only during specific intervals and / or prevent users from applying effects during specific intervals. Again, such restrictions may apply to all signaled effects or a specific set of effects, or the restrictions may apply to only a single effect.

[0072] To support such restrictions, the metadata can include a flag to indicate the desired restriction. For example, a restriction flag can indicate whether one or more (perhaps all) of the effects indicated in the metadata can be overridden by user interaction. Additionally or alternatively, the restriction flag can indicate whether the user is allowed or prohibited from making changes to the sound field. This disabling can apply to all effects, or one or more effects can be specifically enabled or disabled. Restrictions can apply to the entire file or bitstream, or can be associated with a specific time period within a file or bitstream. In another example, an effect identifier can be implemented to use different values ​​to distinguish between a restricted version of an effect (e.g., which cannot be removed or overwritten) and an unrestricted version of the same effect (which can be applied or ignored based on the consumer's choice).

[0073] Figure 9A An example of a metadata stream is shown where restriction flag RF10 is applied to two identified effects. Figure 9BAn example of a metadata stream is shown where a separate restriction flag is applied to each of two different effects. Figure 9C An example is shown in which the restriction flag is accompanied in the metadata stream by a restriction duration RD10 indicating the duration for which the restriction is valid.

[0074] An audio file or stream may include one or more versions of effects metadata, and different versions of such effects metadata may be provided for the same audio content (e.g., as user suggestions from the content creator). For example, different versions of effects metadata may provide different areas of focus for different viewers. In one example, different versions of effects metadata may describe effects that zoom in on different people (e.g., actors, athletes) in a video. Content creators may tag audio sources and / or directions of interest (e.g., Figure 10 ), and the corresponding video stream can be configured to support user selection of the desired metadata stream (obtained by selecting the corresponding feature in the video stream). In another example, different versions of user-generated metadata can be shared via social media (e.g., for live events with many different audience perspectives, such as arena-scale music events). For example, different versions of effects metadata can describe different changes to the same sound field to correspond to different video streams. Different versions of audio effects metadata bitstreams can be downloaded or streamed separately, perhaps from a different source than the sound field itself.

[0075] The effects metadata may be created under human guidance (e.g., by a content creator) and / or automatically based on one or more design criteria. For example, in a teleconferencing application, it may be desirable to automatically select the single loudest sound source or audio from multiple talking sources and reduce the importance of other audio components of the sound field (e.g., discard or reduce the volume). The corresponding effects metadata stream may include a flag indicating "conference mode." In one example, if Figure 3C As shown in , one or more possible values ​​of the effect identifier field of the metadata (e.g., effect identifier ID10) are assigned to indicate the selection of the mode. Parameters defining how the conference mode is applied may include: the number of sources to be amplified (e.g., the number of people at the conference table, the number of people to be speaking, etc.). The number of sources may be selected by a live user, a content creator, and / or automatically. For example, face, motion, and / or person detection may be performed on one or more corresponding video streams to identify a direction of interest and / or support suppression of noise arriving from other directions.

[0076] Other parameters defining how conferencing mode is applied may include metadata for enhancing the extraction of sources from the sound field (e.g., beamformer weights, time-frequency masking values, etc.). The metadata may also include one or more parameter values ​​indicating a desired rotation of the sound field. The sound field may be rotated based on the location of the loudest sound sources: for example, automatic rotation of a remote user's video and audio is supported so that the loudest speakers are located in front of the remote user. In another example, the metadata may indicate automatic rotation of the sound field to allow for a two-person discussion in front of the remote user. In another example, the parameter values ​​may indicate a compression (or other remapping) of the angular range of the recorded sound field (e.g., as Figure 11A ), so that the remote participant can perceive the other participants as being in front of her rather than behind her (e.g., as shown in Figure 11B ).

[0077] The audio effects metadata stream as described herein may be carried in the same transport as the corresponding audio stream (or streams), or may be received in a separate transport, or even from a different source (e.g., as described above). In one example, the effects metadata stream is stored or transmitted in a dedicated extension payload (e.g., in a Figure 9D ) in the afx_data field shown, which is an existing feature in the Advanced Audio Coding (AAC) codec (e.g., as defined in ISO / IEC 14496-3:2009) and newer codecs. The data in this extended payload can be processed by devices that understand this type of extended payload (e.g., decoders and renderers) and can be ignored by other devices. In another example, an audio effects metadata stream as described herein can be standardized for an audio or audiovisual codec. For example, such an approach can be implemented as an amendment in an audio group that is part of a standardized representation of immersive environments (e.g., MPEG-H (e.g., as described in Advanced Television Systems Committee (ATSC) Doc. A / 342-3:2017) and / or MPEG-I (e.g., as described in ISO / IEC 23090)). In another example, an audio effects metadata stream as described herein can be implemented according to the Coding Independent Code Point (CICP) specification. Other use cases for the audio effects metadata stream as described herein include encoding within the IVAS (Immersive Voice and Audio Services) codec (e.g., as part of a 3GPP implementation).

[0078] Although described with respect to AAC, the techniques may be performed using any type of psychoacoustic audio coding that allows for extended payloads and / or extended packets (e.g., filler elements or other information containers that include an identifier followed by filler data) or otherwise allows for backward compatibility, as described in more detail below. Examples of other psychoacoustic audio codecs include: Audio Codec 3 (AC-3), Apple Lossless Audio Codec (ALAC), MPEG-4 Audio Lossless Stream (ALS), Enhanced AC-3, Free Lossless Audio Codec (FLAC), Monkey's Audio, MPEG-1 Audio Layer II (MP2), MPEG-1 Audio Layer III (MP3) Opus, and Windows Media Audio (WMA).

[0079] Figure 12A A block diagram of a system for processing a bitstream including audio data and audio effect metadata as described herein is shown. The system includes an audio decoding stage that is configured to parse the audio effect metadata (e.g., received in an extension payload) and provide the metadata to an audio rendering stage. The audio rendering stage is configured to use the audio effect metadata to apply the audio effects desired by the creator. The audio rendering stage can also be configured to receive user interactions to manipulate the audio effects and take these user commands into account (if permitted).

[0080] Figure 12B A block diagram of an apparatus A100 according to a general configuration including a decoder DC10 and a sound field renderer SR10 is shown. The decoder DC10 is configured to receive a bitstream BS10 including metadata MD10 and a sound field description SD10 (e.g., as described herein with respect to task T100), and to parse the metadata MD10 to obtain an effect identifier and at least one effect parameter value (e.g., as described herein with respect to task T200). The renderer SR10 is configured to apply an effect identified by the effect identifier to the sound field description SD10 (e.g., as described herein with respect to task T300) to generate a modified sound field MS10. For example, the renderer SR10 can be configured to apply the identified effect to the sound field description SD10 using the at least one effect parameter value.

[0081] Renderer SR10 can be configured to apply a focusing effect to the sound field, for example, by rendering selected areas of the sound field at a higher resolution than other areas and / or by rendering other areas to have a higher diffuseness. In one example, an apparatus or device performing task T300 (e.g., renderer SR10) is configured to implement the focusing effect by requesting additional information (e.g., higher-order HOA coefficient values) of the focused source or area from a server via a wired and / or wireless connection (e.g., Wi-Fi and / or LTE).

[0082] The renderer SR10 may be configured to apply a scaling effect to the sound field, for example, by applying a beamformer (e.g., according to parameter values ​​carried in corresponding fields of the metadata). The renderer SR10 may be configured to apply a rotation or translation effect to the sound field, for example, by applying a corresponding matrix transformation to a set of HOA coefficients (or more generally, to a hierarchical set of basis function coefficients) and / or by moving audio objects within the sound field accordingly.

[0083] Figure 12C A block diagram of an embodiment A200 of an apparatus A100 including a command processor CP10 is shown. The processor CP10 is configured to receive metadata MD10 and at least one user command UC10 as described herein, and to generate at least one effect command EC10 based on the at least one user command UC10 and at least one effect parameter value (e.g., in accordance with one or more restriction flags in the metadata). The renderer SR10 is configured to apply the identified effect to the sound field description SD10 using the at least one effect command EC10 to generate a modified sound field MS10.

[0084] Figure 13A A block diagram of an apparatus for manipulating a sound field F100 according to a general configuration is shown. Apparatus F100 includes a unit MF100 for receiving a bitstream comprising metadata (e.g., one or more metadata streams) and a sound field description (e.g., one or more audio streams) (e.g., as described herein with respect to task T100). For example, unit MF100 for receiving includes a transceiver, a modem, a decoder DC10, one or more other circuits or devices configured to receive a bitstream BS10, or a combination thereof. Apparatus F100 also includes a unit MF200 for parsing metadata to obtain an effect identifier and at least one effect parameter value (e.g., as described herein with respect to task T200). For example, unit MF200 for parsing includes a decoder DC10, one or more other circuits or devices configured to parse metadata MD10, or a combination thereof. Apparatus F100 also includes a unit MF300 for applying an effect identified by an effect identifier to a sound field description (e.g., as described herein with respect to task T300). For example, the unit MF300 may be configured to apply the identified effect by applying a matrix transformation to the sound field description using at least one effect parameter value. In some examples, the unit MF300 for applying the effect comprises a renderer SR10, a processor CP10, one or more other circuits or devices configured to apply the effect to the sound field description SD10, or a combination thereof.

[0085] Figure 13BA block diagram of an embodiment F200 of apparatus F100 is shown, comprising a unit MF400 for receiving at least one user command (e.g., through active and / or passive user interaction) (e.g., as described herein with respect to task T400). For example, the unit MF400 for receiving at least one user command comprises a processor CP10, one or more other circuits or devices configured to receive at least one user command UC10, or a combination thereof. The apparatus F200 further comprises a unit MF350 (an implementation of the unit MF300) for applying an effect identified by an effect identifier to a sound field description based on at least one of (A) at least one effect parameter value or (B) at least one user command. In one example, the unit MF350 comprises a unit for combining at least one effect parameter value with the user command to obtain at least one correction parameter. In another example, parsing the metadata comprises parsing the metadata to obtain a second effect identifier, and the unit MF350 comprises a unit for determining not to apply the effect identified by the second effect identifier to the sound field description. In some examples, the unit MF350 for applying effects includes a renderer SR10, a processor CP10, one or more other circuits or devices configured to apply effects to the sound field description SD10, or a combination thereof. The apparatus F200 can be embodied, for example, by an implementation of a user tracking device UT10 that receives audio and metadata streams and generates corresponding audio to the user via headphones.

[0086] Hardware for virtual reality (VR) may include one or more screens that present a visual scene to a user, one or more sound transducers (e.g., a speaker array or a head-mounted transducer array) that provide a corresponding audio environment, and one or more sensors for determining the user's position, orientation, and / or movement. Figure 8A The user tracking device UT10 shown in FIG is an example of a VR headset. To support an immersive experience, such a headset can detect the orientation of the user's head in three degrees of freedom (3DOF): rotation of the head around the up-down axis (yaw), tilt of the head in the front-to-back plane (pitch), and tilt of the head in the left-to-right plane (roll), and adjust the provided audio environment accordingly.

[0087] Computer-mediated reality systems are being developed to allow a computing device to augment or add, remove or subtract, replace or change, or generally modify an existing reality of a user's experience. As a few examples, computer-mediated reality systems can include virtual reality (VR) systems, augmented reality (AR) systems, and mixed reality (MR) systems. The perceptual success of a computer-mediated reality system is often tied to the ability of such a system to provide a realistic immersive experience in terms of video and audio, such that the video and audio experience align in a manner that a user finds natural and expected. While the human visual system is more sensitive than the human auditory system (e.g., in terms of the perceived localization of various objects within a scene), ensuring adequate auditory experience is an increasingly important factor in ensuring a realistic immersive experience, especially as video experiences improve, enabling better localization of video objects, thereby enabling a user to better identify the source of audio content.

[0088] In VR technology, a virtual information can be presented to a user using a head-mounted display such that the user can visually experience an artificial world on a screen in front of their eyes. In AR technology, the real world is augmented by visual objects that can be superimposed (e.g., overlaid) on physical objects in the real world. This augmentation can insert new visual objects and / or mask visual objects in the real-world environment. In MR technology, the line between real or synthetic / virtual and the user's visual experience becomes indistinguishable. The technology as described herein can be used with a VR device 400 as shown in FIG. 4 to improve the experience of a user 402 of the device via earphones 404 of the device. Figure 15

[0089] Video, audio, and other sensory data can play an important role in a VR experience. To participate in a VR experience, a user 402 can wear a VR device 400 (which can also be referred to as a VR headset 400) or other wearable electronic device. A VR client device (e.g., the VR headset 400) can track head movements of the user 402 and adjust video data displayed via the VR headset 400 to account for the head movements, thereby providing an immersive experience in which the user 402 can experience a virtual world displayed in the video data in visual three-dimensions.

[0090] While VR (and other forms of AR and / or MR) can allow a user 402 to visually reside in a virtual world, a VR headset 400 can often lack the ability to place the user in an audible virtual world. In other words, a VR system (which can include a computer responsible for rendering video data and audio data (for ease of illustration, referred to herein as a VR client device)) can provide a realistic immersive experience in terms of video, but can lack a realistic immersive experience in terms of audio. This can be due to a variety of factors, including the fact that the human auditory system is less sensitive than the human visual system (e.g., in terms of the perceived localization of various objects within a scene), and the fact that the human auditory system is less sensitive to changes in audio than the human visual system is to changes in video. Figure 15 ​), and the VR headset 400) may not be able to support full 3D immersion in hearing (in some cases, realistically reflected in the virtual scene displayed to the user via the VR headset 400).

[0091] While full three-dimensional audible rendering still presents some challenges, the techniques in this disclosure bring one step closer to that goal. The audio aspects of AR, MR, and / or VR can be divided into three separate immersive categories. The first category provides the lowest level of immersion and is called three degrees of freedom (3DOF). 3DOF refers to audio rendering that takes into account the movement of the head in three degrees of freedom (yaw, pitch, and roll), allowing the user to look around freely in any direction. However, 3DOF cannot account for translational (and directional) head movement where the head is not centered about the optical and acoustic center of the sound field.

[0092] The second category, called 3DOF plus (or "3DOF+"), provides three degrees of freedom (yaw, pitch, and roll) in addition to the limited spatial translational (and directional) motion caused by the movement of the head away from the optical and acoustic centers within the sound field. 3DOF+ can provide support for perceptual effects such as motion parallax, which can enhance immersion.

[0093] The third category, called six degrees of freedom (6DOF), renders audio data in a way that takes into account the three degrees of freedom of head movement (yaw, pitch, and roll), but also takes into account the translation of the person in space (x, y, and z translation). Spatial translation can be induced, for example, by sensors tracking the person's position in the physical world, by input controllers, and / or by a rendering program that simulates the user's transportation within the virtual space.

[0094] The audio aspect of VR may not be as immersive as the video aspect, potentially reducing the overall immersion of the user experience. However, with advances in processors and wireless connectivity, 6DOF rendering may be achieved using wearable AR, MR, and / or VR devices. In addition, it may be possible in the future to take into account the movement of vehicles with AR, MR, and / or VR device capabilities and provide an immersive audio experience. In addition, ordinary technicians will recognize that mobile devices (e.g., mobile phones, smartphones, tablets) can also implement VR, AR, and / or MR technologies.

[0095] According to the techniques described in this disclosure, various ways of adjusting audio data (whether in audio channel format, audio object format, and / or audio scene-based format) can enable 6DOF audio rendering. 6DOF rendering provides a more immersive listening experience by rendering audio data in a manner that accounts for three degrees of freedom of head motion (yaw, pitch, and roll) as well as translational motion (e.g., in a spatial three-dimensional coordinate system x, y, z). In implementation, where head motion may not be centered about the optical and acoustic center, adjustments can be made to provide 6DOF rendering without being limited to a spatial two-dimensional coordinate system. As disclosed herein, the following figures and description enable 6DOF audio rendering.

[0096] Figure 16 is a diagram depicting an example of an embodiment 800 of a wearable device that can operate in accordance with various aspects of the technology described in this disclosure. In various examples, the wearable device 800 can represent a VR headset (e.g., the VR headset 400 described above), an AR headset, an MR headset, or an extended reality (XR) headset. Augmented reality "AR" can refer to computer-rendered images or data that are overlaid on the real world where a user is actually located. Mixed reality "MR" can refer to computer-rendered images or data that are locked to a specific location in the real world, or can refer to a variant of VR in which some computer-rendered 3D elements and some photographed real elements are combined into an immersive experience that simulates the user's physical presence in the environment. Extended reality "XR" can refer to a collective term for VR, AR, and MR.

[0097] Wearable device 800 may represent other types of devices, such as watches (including so-called “smart watches”), glasses (including so-called “smart glasses”), headphones (including so-called “wireless headphones” and “smart headphones”), smart clothing, smart jewelry, etc. Regardless of whether it represents a VR device, a watch, glasses, and / or headphones, wearable device 800 may communicate with a computing device that supports wearable device 800 via a wired connection or a wireless connection.

[0098] In some cases, the computing device supporting wearable device 800 can be integrated within wearable device 800, and thus, wearable device 800 can be considered the same device as the computing device supporting wearable device 800. In other instances, wearable device 800 can communicate with a separate computing device capable of supporting wearable device 800. In this regard, the term "supporting" should not be construed as requiring a separate, dedicated device, but rather, one or more processors configured to perform various aspects of the techniques described in the present disclosure can be integrated within wearable device 800, or integrated within a computing device separate from wearable device 800.

[0099] For example, when the wearable device 800 represents the VR device 400, a separate dedicated computing device (e.g., a personal computer including one or more processors) can render audio and video content, and the wearable device 800 can determine the translational head movement on which the dedicated computing device can render the audio content (as a speaker feed) based on the translational head movement in accordance with various aspects of the techniques described in this disclosure. As another example, when the wearable device 800 represents smart glasses, the wearable device 800 can include a processor (e.g., one or more processors) that determines the translational head movement (by connecting within one or more sensors of the wearable device 800) and renders the speaker feed based on the determined translational head movement.

[0100] As shown, the wearable device 800 includes a rear-facing camera, one or more directional speakers, one or more tracking and / or recording cameras, and one or more light-emitting diode (LED) lights. In some examples, the LED lights may be referred to as "super bright" LED lights. In addition, the wearable device 800 includes one or more eye-tracking cameras, a high-sensitivity audio microphone, and optical / projection hardware. The optical / projection hardware of the wearable device 800 may include durable translucent display technology and hardware.

[0101] The wearable device 800 also includes connection hardware, which may represent one or more network interfaces that support multi-mode connections, such as 4G communication, 5G communication, and the like. The wearable device 800 also includes an ambient light sensor and a bone conduction transducer. In some cases, the wearable device 800 may also include one or more passive and / or active cameras with fisheye lenses and / or telephoto lenses. According to various techniques of the present disclosure, the steering angle of the wearable device 800 may be used to select an audio representation of the sound field (e.g., one of the mixed order surround (MOA) representations) for output via the directional speakers (headphones 404) of the wearable device 800. It should be understood that the wearable device 800 can take on a variety of different form factors.

[0102] Despite Figure 16 Although not shown in the example of FIG, the wearable device 800 may include an orientation / translation sensor unit, such as a combination of micro-electromechanical systems (MEMS) for sensing, or any other type of sensor capable of providing information supporting head and / or body tracking. In one example, the orientation / translation sensor unit may represent a MEMS for sensing translational motion, similar to those used in cellular telephones (e.g., so-called "smartphones").

[0103] Although described with respect to a specific example of a wearable device, a person of ordinary skill in the art will understand that Figure 15 and Figure 16The related descriptions can apply to other examples of wearable devices. For example, other wearable devices (e.g., smart glasses) can include sensors through which translational head motion can be obtained. As another example, other wearable devices (e.g., smart watches) can include sensors through which translational motion is obtained. Thus, the techniques described in this disclosure should not be limited to a particular type of wearable device, but rather any wearable device can be configured to perform the techniques described in this disclosure.

[0104] Figure 17 A block diagram of a system 900 that can be implemented within a device (e.g., wearable device 400 or 800) is shown. The system 900 includes a processor 420 (e.g., one or more processors) that can be configured to perform the methods M100 or M200 as described herein. The system 900 also includes a memory 120 coupled to the processor 420, a sensor 110 (e.g., an ambient light sensor, a direction and / or tracking sensor of the device 800), a vision sensor 130 (e.g., a night vision sensor, a tracking and recording camera, an eye tracking camera, and a rear-facing camera of the device 800), a display device 100 (e.g., an optical piece / projector of the device 800), an audio capture device 112 (e.g., a high-sensitivity microphone of the device 800), a speaker 470 (e.g., an earpiece 404 of the device 400, a directional speaker of the device 800), a transceiver 480, and an antenna 490. In particular aspects, the system 900 includes a modem, in addition to or instead of the transceiver 480. For example, the modem, the transceiver 480, or both are configured to receive a signal representing a bitstream BS10 and provide the bitstream BS10 to a decoder DC10.

[0105] The various elements of embodiments of apparatuses or systems (e.g., apparatuses A100, A200, F100, and / or F200) as disclosed herein can be embodied using any combination of hardware and software and / or firmware thought appropriate for the intended application. For example, these elements can be manufactured as electronic and / or optical devices, e.g., residing on the same chip or between two or more chips in a chipset. One example of such a device is a fixed or programmable array of logic elements (e.g., transistors or logic gates), and any of these elements can be implemented as one or more such arrays. Any two or more, or even all, of these elements can be implemented within the same array or arrays. Such one or more arrays can be implemented within one or more chips (e.g., within a chipset including two or more chips).

[0106] A processor or other device for processing as disclosed herein can be manufactured as one or more electronic and / or optical devices, which are, for example, located on the same chip or between two or more chips in a chipset. An example of such a device is a fixed or programmable array of logic elements (e.g., transistors or logic gates), and any of these elements can be implemented as one or more such arrays. Such one or more arrays can be implemented within one or more chips (e.g., within a chipset comprising two or more chips). Examples of such arrays include fixed or programmable arrays of logic elements, such as microprocessors, embedded processors, IP cores, DSPs (digital signal processors), FPGAs (field programmable gate arrays), ASSPs (application-specific standard products), and ASICs (application-specific integrated circuits). A processor or other unit for processing as disclosed herein can also be embodied as one or more computers (e.g., machines comprising one or more arrays that are programmed to execute one or more sets of instructions or sequences of instructions) or other processors. It is possible that a processor as described herein may be used to perform tasks or execute other instruction sets that are not directly related to the process of implementing method M100 or M200 (or another method disclosed with reference to the operation of the apparatus or system described herein), such as tasks related to another operation of a device or system in which the processor is embedded (e.g., a voice communication device such as a smartphone or smart speaker). Portions of the methods disclosed herein may also be executed under the control of one or more other processors.

[0107] Each task of the method disclosed herein (e.g., method M100 and / or M200) can be directly embodied in hardware, in a software module executed by a processor, or in a combination of the two. In a typical application of the implementation of the method disclosed herein, a logic element (e.g., logic gate) array is configured to perform one, multiple, or even all tasks of the method. One or more (possibly all) tasks can also be implemented as code (e.g., one or more groups of instructions), embodied in a computer program product (e.g., one or more data storage media such as a disk, flash memory or other non-volatile memory card, semiconductor memory chip, etc.), which can be read and / or executed by a machine (e.g., a computer) comprising a logic element array (e.g., a processor, microprocessor, microcontroller or other finite state machine). The tasks of implementing the method disclosed herein can also be performed by more than one such array or machine. In these or other embodiments, these tasks can be performed in a device for wireless communication such as a cellular phone or other device with such communication capabilities. Such a device can be configured to communicate with a circuit-switched and / or packet-switched network (e.g., using one or more protocols such as VoIP). For example, such a device may include RF circuitry configured to receive and / or transmit encoded frames.

[0108] In one or more exemplary aspects, the operations described herein can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, the operations can be stored on a computer-readable medium or transmitted as one or more instructions or code. The term "computer-readable medium" includes both computer-readable storage media and communication (e.g., transmission) media. By way of example, and not limitation, a computer-readable storage medium can include an array of memory elements, such as semiconductor memory (which can include, but is not limited to, dynamic or static RAM, ROM, EEPROM, and / or flash RAM), or ferroelectric, magnetoresistive, elliptical, polymer, or phase-change memory; CD-ROM or other optical disk storage devices; and / or magnetic disk storage or other magnetic storage devices. Such storage media can store information in the form of instructions or data structures that can be accessed by a computer. Communication media can include any medium that can be used to carry desired program code in the form of instructions or data structures and that can be accessed by a computer, including any medium that facilitates the transfer of a computer program from one location to another. Furthermore, any connection can be appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, wireless and / or microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, wireless and / or microwave is included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc TM (Blu-ray Disc Association, Universal City, CA), where magnetic disks typically reproduce data magnetically, while optical disks use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0109] In one example, a non-transitory computer-readable storage medium includes code that, when executed by at least one processor, causes the at least one processor to perform a method for characterizing a portion of a sound field as described herein. Further examples of such storage media include media further including code that, when executed by at least one processor, causes the at least one processor to perform the following operations: receive a bitstream including metadata and a sound field description (e.g., as described herein with reference to task T100); parse the metadata to obtain an effect identifier and at least one effect parameter value (e.g., as described herein with reference to task T200); and apply the effect identified by the effect identifier to the sound field description (e.g., as described herein with reference to task T300). The applying may include applying the identified effect to the sound field description using the at least one effect parameter value.

[0110] Implementation examples are described in the following numbered clauses:

[0111] Clause 1. A method of manipulating a sound field, the method comprising: receiving a bitstream comprising metadata and a sound field description; parsing the metadata to obtain an effect identifier and at least one effect parameter value; and applying an effect identified by the effect identifier to the sound field description.

[0112] Clause 2. The method of clause 1, wherein the parsing the metadata comprises parsing the metadata to obtain a timestamp corresponding to the effect identifier, and wherein the applying the identified effect comprises applying the identified effect to the portion of the sound field description corresponding to the timestamp using the at least one effect parameter value.

[0113] Clause 3. The method of Clause 1, wherein applying the identified effect comprises combining the at least one effect parameter value with a user command to obtain at least one modified parameter value.

[0114] Clause 4. The method of any of clauses 1 to 3, wherein applying the identified effect comprises rotating the sound field to a desired direction.

[0115] Clause 5. The method of any of Clauses 1 to 3, wherein the at least one effect parameter value comprises an indicated direction, and wherein applying the identified effect comprises rotating the sound field to the indicated direction using the at least one effect parameter value.

[0116] Clause 6. A method according to any one of clauses 1 to 3, wherein the at least one effect parameter value comprises an indicated direction, and wherein the effect identified by applying comprises increasing the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions using the at least one effect parameter value.

[0117] Clause 7. A method according to any one of clauses 1 to 3, wherein the at least one effect parameter value comprises an indicated direction, and wherein the effect identified by applying comprises reducing the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions using the at least one effect parameter value.

[0118] Clause 8. The method of any of clauses 1 to 3, wherein the at least one effect parameter value indicates a position within the sound field, and wherein applying the identified effect comprises translating a sound source to the indicated position using the at least one effect parameter value.

[0119] Clause 9. A method according to any one of clauses 1 to 3, wherein the at least one effect parameter value includes an indication of direction, and wherein applying the identified effect comprises: using the at least one effect parameter value to increase the directionality of at least one of a sound source of the sound field or the region of the sound field relative to another sound source of the sound field or the region of the sound field.

[0120] Clause 10. The method of any one of clauses 1 to 3, wherein applying the identified effect comprises applying a matrix transformation to the sound field description.

[0121] Clause 11. The method of Clause 10, wherein the matrix transformation comprises at least one of a rotation of the sound field and a translation of the sound field.

[0122] Clause 12. The method of any one of clauses 1 to 3, wherein the sound field description comprises a hierarchical set of basis function coefficients.

[0123] Clause 13. The method of any of clauses 1 to 3, wherein the sound field description comprises a plurality of audio objects.

[0124] Clause 14. The method of any of clauses 1 to 3, wherein the parsing the metadata comprises parsing the metadata to obtain a second effect identifier, and wherein the method comprises determining not to apply the effect identified by the second effect identifier to the sound field description.

[0125] Clause 15. An apparatus for manipulating a sound field, the apparatus comprising: a decoder configured to receive a bitstream comprising metadata and a sound field description, and to parse the metadata to obtain an effect identifier and at least one effect parameter value; and a renderer configured to apply an effect identified by the effect identifier to the sound field description.

[0126] Clause 16. The apparatus of Clause 15, further comprising a modem configured to: receive a signal representing the bitstream; and provide the bitstream to the decoder.

[0127] Clause 17. An apparatus for manipulating a sound field, the apparatus comprising: a memory configured to store a bitstream comprising metadata and a sound field description; and a processor coupled to the memory configured to: parse the metadata to obtain an effect identifier and at least one effect parameter value; and apply an effect identified by the effect identifier to the sound field description.

[0128] Clause 18. The apparatus of clause 17, wherein the processor is configured to parse the metadata to obtain a timestamp corresponding to the effect identifier, and to apply the identified effect by using the at least one effect parameter value to apply the identified effect to the portion of the sound field description corresponding to the timestamp.

[0129] Clause 19. The device of Clause 17, wherein the processor is configured to combine the at least one effect parameter value with a user command to obtain at least one modified parameter.

[0130] Clause 20. Apparatus according to any of clauses 17 to 19, wherein the at least one effect parameter value comprises an indicated direction, and wherein the processor is configured to apply the identified effect by rotating the sound field to the indicated direction using the at least one effect parameter value.

[0131] Clause 21. An apparatus as described in any of clauses 17 to 19, wherein the at least one effect parameter value includes an indicated direction, and wherein the processor is configured to apply the identified effect by using the at least one effect parameter value to increase the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions.

[0132] Clause 22. An apparatus as described in any of clauses 17 to 19, wherein the at least one effect parameter value includes an indicated direction, and wherein the processor is configured to apply the identified effect using the at least one effect parameter value to reduce the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions.

[0133] Clause 23. Apparatus according to any of clauses 17 to 19, wherein the at least one effect parameter value indicates a position within the sound field, and wherein the processor is configured to apply the identified effect by translating the sound source to the indicated position using the at least one effect parameter value.

[0134] Clause 24. An apparatus as described in any of clauses 17 to 19, wherein the at least one effect parameter value includes an indication of a direction, and wherein the processor is configured to apply the identified effect using the at least one effect parameter value to increase the directionality of at least one of a sound source of the sound field or the region of the sound field relative to another sound source of the sound field or the region of the sound field.

[0135] Clause 25. Apparatus according to any of clauses 17 to 19, wherein the processor is configured to apply the identified effect by applying a matrix transformation to the sound field description using the at least one effect parameter value.

[0136] Clause 26. The apparatus of Clause 25, wherein the matrix transformation comprises at least one of a rotation of the sound field and a translation of the sound field.

[0137] Clause 27. Apparatus according to any of clauses 17 to 19, wherein the sound field description comprises a hierarchical set of basis function coefficients.

[0138] Clause 28. Apparatus according to any of clauses 17 to 19, wherein the sound field description comprises a plurality of audio objects.

[0139] Clause 29. The apparatus of any of clauses 17 to 19, wherein the processor is configured to parse the metadata to obtain a second effect identifier and to determine not to apply the effect identified by the second effect identifier to the sound field description.

[0140] Clause 30. The apparatus of any of clauses 17 to 19, wherein the apparatus comprises an application specific integrated circuit, the application specific integrated circuit comprising the processor.

[0141] Clause 31. An apparatus for manipulating a sound field, the apparatus comprising: a receiving unit for receiving a bitstream comprising metadata and a sound field description; a parsing unit for parsing the metadata to obtain an effect identifier and at least one effect parameter value; and an applying unit for applying an effect identified by the effect identifier to the sound field description.

[0142] Clause 32. The apparatus of clause 31, wherein at least one of the means for receiving, the means for parsing, or the means for applying is integrated into at least one of a mobile phone, a tablet device, a wearable electronic device, a camera device, a virtual reality headset, an augmented reality headset, or a vehicle.

[0143] Those skilled in the art will also appreciate that the various exemplary logic blocks, configurations, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or a combination thereof. The various exemplary components, blocks, configurations, modules, circuits, and steps described above have been generally described around their functions. Whether such functions are implemented as hardware or as processor-executable instructions depends on the specific application and the design constraints imposed on the entire system. A skilled person may implement the described functions in a flexible manner for each specific application, but such implementation decisions should not be interpreted as departing from the scope of protection of this disclosure.

[0144] In conjunction with the steps of the method or algorithm described in the embodiments disclosed herein, it is possible to directly embody hardware, a software module executed by a processor, or a combination thereof. The software module may be located in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a compact disc read-only memory (CD-ROM), or any other form of non-temporary storage medium known in the art. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative embodiment, a storage medium may also be a component of a processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and the storage medium may reside in a computing device or a user terminal as discrete components.

[0145] The above description focuses on the present disclosure to enable any person skilled in the art to implement or use the disclosed embodiments. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein but is intended to be consistent with the broadest scope of the principles and novel features defined by the appended claims.

Claims

1. A method for manipulating a sound field, the method comprising: receiving a bitstream comprising metadata and a sound field description, wherein the sound field description comprises a hierarchical set of basis function coefficients, the metadata comprising a first effect identifier specifying a type of a first effect to be applied to the sound field description and at least one effect parameter value defining how the effect is to be applied, comprising a second effect identifier specifying a type of a second effect to be applied to the sound field description and comprising a restriction flag indicating whether the second effect is disabled or enabled; parsing the metadata to obtain the first effect identifier, the second effect identifier, the at least one effect parameter value, and the restriction flag; determining not to apply the second effect identified by the second effect identifier to the sound field description based on the restriction flag indicating that the second effect is disabled; and The first effect identified by the first effect identifier is applied to the sound field description.

2. The method according to claim 1, wherein The parsing the metadata comprises parsing the metadata to obtain a timestamp corresponding to the first effect identifier, and wherein the applying the first effect comprises applying the first effect to the portion of the sound field description corresponding to the timestamp using the at least one effect parameter value.

3. The method according to claim 1, wherein Said applying the first effect comprises combining the at least one effect parameter value with a user command to obtain at least one modified parameter value.

4. The method according to claim 1, wherein The applying the first effect includes rotating the sound field to a desired direction.

5. The method according to claim 1, wherein The at least one effect parameter value comprises an indicated direction, and wherein said applying the first effect comprises rotating the sound field to the indicated direction using the at least one effect parameter value.

6. The method according to claim 1, wherein The at least one effect parameter value comprises an indicated direction, and wherein applying the first effect comprises increasing the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions using the at least one effect parameter value.

7. The method according to claim 1, wherein The at least one effect parameter value comprises an indicated direction, and wherein applying the first effect comprises using the at least one effect parameter value to reduce a sound level of the sound field in the indicated direction relative to a sound level of the sound field in other directions.

8. The method according to claim 1, wherein The at least one effect parameter value indicates a position within the sound field, and wherein applying the first effect comprises translating a sound source to the indicated position using the at least one effect parameter value.

9. The method according to claim 1, wherein The at least one effect parameter value includes an indication direction, and wherein applying the first effect includes: using the at least one effect parameter value to increase the directionality of at least one of the sound source of the sound field or the area of ​​the sound field relative to another sound source of the sound field or the area of ​​the sound field.

10. The method according to claim 1, wherein The applying the first effect comprises applying a matrix transformation to the sound field description.

11. The method according to claim 10, wherein: The matrix transformation includes at least one of a rotation of the sound field and a translation of the sound field.

12. The method according to claim 1, wherein The sound field description also includes a plurality of audio objects.

13. A device for manipulating a sound field, the device comprising: a decoder configured to receive a bitstream comprising metadata and a sound field description, wherein the sound field description comprises a hierarchical set of basis function coefficients, the metadata comprising a first effect identifier specifying a type of a first effect to be applied to the sound field description and at least one effect parameter value defining how the effect is to be applied, comprising a second effect identifier specifying a type of a second effect to be applied to the sound field description and comprising a restriction flag indicating whether the second effect is disabled or enabled, and to parse the metadata to obtain the first effect identifier, the second effect identifier, the at least one effect parameter value and the restriction flag, and to determine not to apply the second effect identified by the second effect identifier to the sound field description based on the restriction flag indicating that the second effect is disabled; A renderer is configured to apply the first effect identified by the first effect identifier to the sound field description.

14. The apparatus of claim 13, further comprising a modem configured to: receiving a signal representing the bitstream; and The bitstream is provided to the decoder.

15. A device for manipulating a sound field, the device comprising: a memory configured to store a bitstream comprising metadata and a sound field description, wherein the sound field description comprises a hierarchical set of basis function coefficients, the metadata comprising a first effect identifier specifying a type of a first effect to be applied to the sound field description and at least one effect parameter value defining how the effect is to be applied, comprising a second effect identifier specifying a type of a second effect to be applied to the sound field description and comprising a restriction flag indicating whether the second effect is disabled or enabled; and a processor coupled to the memory, configured to: parsing the metadata to obtain the first effect identifier, the second effect identifier, the at least one effect parameter value, and the restriction flag; determining not to apply the second effect identified by the second effect identifier to the sound field description based on the restriction flag indicating that the second effect is disabled; and The first effect identified by the first effect identifier is applied to the sound field description.

16. The apparatus according to claim 15, wherein The processor is configured to parse the metadata to obtain a timestamp corresponding to the first effect identifier and to apply the first effect by using the at least one effect parameter value to apply the first effect to the portion of the sound field description corresponding to the timestamp.

17. The apparatus according to claim 15, wherein The processor is configured to combine the at least one effect parameter value with a user command to obtain at least one modified parameter.

18. The apparatus according to claim 15, wherein The at least one effect parameter value comprises an indicated direction, and wherein the processor is configured to apply the first effect by rotating the sound field to the indicated direction using the at least one effect parameter value.

19. The apparatus according to claim 15, wherein The at least one effect parameter value comprises an indicated direction, and wherein the processor is configured to apply the first effect by increasing the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions using the at least one effect parameter value.

20. The apparatus of claim 15, wherein The at least one effect parameter value comprises an indicated direction, and wherein the processor is configured to apply the first effect by using the at least one effect parameter value to reduce the sound level of the sound field in the indicated direction relative to the sound level of the sound field in other directions.

21. The apparatus of claim 15, wherein: The at least one effect parameter value indicates a position within the sound field, and wherein the processor is configured to apply the first effect by translating a sound source to the indicated position using the at least one effect parameter value.

22. The apparatus of claim 15, wherein: The at least one effect parameter value includes an indication direction, and wherein the processor is configured to apply the first effect by using the at least one effect parameter value to increase the directionality of at least one of a sound source of the sound field or the area of ​​the sound field relative to another sound source of the sound field or the area of ​​the sound field.

23. The apparatus of claim 15, wherein: The processor is configured to apply the first effect by applying a matrix transformation to the sound field description using the at least one effect parameter value.

24. The apparatus according to claim 23, wherein The matrix transformation includes at least one of a rotation of the sound field and a translation of the sound field.

25. The apparatus of claim 15, wherein: The sound field description also includes a plurality of audio objects.

26. The apparatus of claim 15, wherein: The apparatus includes an application specific integrated circuit including the processor.

Citation Information

Patent Citations

  • System and method for adaptive audio signal generation, coding and rendering

    CN105792086A

  • Object-based audio loudness management

    US20150245153A1

  • Method to align an immersive video and an immersive sound field

    US20170372748A1

  • Metadata transcoding

    US20170373857A1