Incorporating jitter into spatial audio objects
By introducing positional jitter patterns to audio objects, the perceptual accuracy and flexibility of audio playback are improved, addressing inaccuracies in existing adaptive audio systems.
Patent Information
- Application Number
- PCT/US2025/015749
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2025-02-13
- Publication Date
- 2025-08-21
AI Technical Summary
Existing adaptive audio systems lack the ability to introduce jitter or movement into the spatial position of audio objects, leading to inaccuracies in perceptual accuracy and limited flexibility in defining and controlling the playback experience.
Introduce positional jitter patterns that define a periodically varying movement of audio objects, separately from their static or dynamic positions, to enhance perceptual accuracy and flexibility in audio playback.
Enhances perceptual accuracy of audio object positions by mimicking real-world sound perception and provides additional flexibility in controlling the playback experience through varied jitter patterns.
Smart Images

Figure US2025015749_21082025_PF_FP_ABST
Abstract
Description
INCORPORATING JITTER INTO SPATIAL AUDIO OBJECTSTECHNICAL FIELD
[0001] This application relates generally to techniques for introducing jitter (e.g., slight movements) into the spatial position of audio objects for spatial audio ecosystems.SUMMARY
[0002] Various aspects of the present disclosure relate to devices, systems, and methods to introduce jitter (e.g., slight movements) into the spatial position of audio objects for spatial audio ecosystems in order to improve the perceptual accuracy of audio object locations (e.g., naturalness) and / or provide additional flexibility to define and / or control the playback experience.
[0003] In one aspect of the present disclosure, there is provided a method of rendering audio. The method may include receiving audio reproduction data including audio objects and associated metadata for each audio object. Each of the audio objects may include an audio object position. The method may also include receiving reproduction environment data including an indication of a number of reproduction speakers in a reproduction environment and an indication of a location of each reproduction speaker within the reproduction environment. The method may also include rendering the audio objects into one or more speaker feed signals based, at least in part, on the associated metadata. Each speaker feed signal may correspond to at least one of the reproduction speakers within the reproduction environment. Rendering the audio objects may include rendering a first audio object (i) according to a first audio object position of the first audio object, and (ii) in accordance with a positional jitter pattern that defines a periodically varying movement of the first audio object. The periodically varying movement may be separately defined with respect to any movement defined by the first audio object position of the first audio object.
[0004] In addition to any combination of features described above, the associated metadata of the first audio object may include audio object position data indicative of the audio object position of the first audio object, and positional jitter data indicative of the positional jitter pattern of the first audio object. In addition to any combination of features described above, the positional jitter data may indicate whether the first audio object is to be rendered in accordance with the positional jitter pattern. In addition to any combination of features described above, the positional jitter data may indicate at least one of a group consisting of a type of the positional jitter pattern, a velocity of positional jitter associated with the positional jitter pattern, a frequency of the positional jitterpattern, an amplitude of the positional jitter pattern, and combinations thereof. In addition to any combination of features described above, the positional jitter data may be sidecar data that is updated only at an onset and at an ending of the rendering of the first audio object.
[0005] In addition to any combination of features described above, the positional jitter pattern may include a sinusoidal pattern such that the first audio object position of the first audio object oscillates back and forth along a positional path between a first positional point and a second positional point to cover spatial positions between the first positional point and the second positional point.
[0006] In addition to any combination of features described above, the positional jitter pattern may include a square wave such that the first audio object position of the first audio object periodically alternates between being located at a first positional point for a first amount of time and being located at a second positional point for a second amount of time.
[0007] In addition to any combination of features described above, the positional jitter pattern may include a noise pattern modeled according to a type of noise.
[0008] In addition to any combination of features described above, the positional jitter pattern may be set to be a first pattern in response to determining that accelerometer data indicative of movement of a head of a listener indicates that the head of the listener is moving. The positional jitter pattern may also be set to be a second pattern different than the first pattern in response to determining that the accelerometer data indicates that the head of the listener is stationary.
[0009] In addition to any combination of features described above, the positional jitter pattern may include a positional path derived from accelerometer data indicative of human movements during listening.
[0010] In addition to any combination of features described above, the first audio object may include a static audio object such that the first audio object remains stationary within the reproduction environment except for movement in accordance with the positional jitter pattern.
[0011] In addition to any combination of features described above, the first audio object may include a dynamic audio object such that the first audio object spatially moves within the reproduction environment according to movement defined by the first audio object position. The movement defined by the first audio object position may be separately defined with respect to the periodically varying movement defined by the positional jitter pattern.
[0012] In addition to any combination of features described above, rendering the audio objects may include rendering a second audio object (i) according to a second audio object position of the second audio object, and (ii) in accordance with a second positional jitter pattern that defines a second periodically varying movement of the second audio object. The second periodically varying movement may be separately defined with respect to any movement defined by the second audio object position of the second audio object. The second positional jitter pattern may be different than the positional jitter pattern according to which the first audio object is rendered.
[0013] In addition to any combination of features described above, rendering the audio objects may include rendering a second audio object (i) according to a second audio object position of the second audio object, and (ii) without the positional jitter pattern or any other positional jitter pattern.
[0014] In another aspect of the present disclosure, there is provided a computing apparatus that may include at least one electronic processor. The computing apparatus may also include a memory storing instructions, which when executed by the at least one electronic processor, cause the computing apparatus to perform the method and / or any combination of features described above.
[0015] In another aspect of the present disclosure, there is provided a non-transitory computer- readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method and / or any combination of features described above.
[0016] In another aspect of the present disclosure, there is provided a method of authoring audio. The method may include generating metadata including positional jitter metadata for each audio object included in audio reproduction data including a plurality of audio objects. The positional jitter metadata may be indicative of a positional jitter pattern of a respective audio object. The positional jitter pattern may define a periodically varying movement of the respective audio object. The periodically varying movement may be separately defined with respect to a spatial location or any spatial movement defined by an audio object position of the respective audio object.
[0017] In addition to any combination of features described above, generating the metadata may include generating audio object position data indicative of the audio object position of the respective audio object.
[0018] Other aspects of the embodiments will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG. 1 illustrates an example speaker placement in a surround system (e.g., 5.1.4 surround) that provides height speakers for playback of height channels, according to some example embodiments.
[0020] FIG. 2 illustrates a combination of channel and object-based data to produce an adaptive audio mix, according to some example embodiments.
[0021] FIG. 3 is a block diagram that provides examples of components of an authoring and / or rendering apparatus, according to some example embodiments.
[0022] FIG. 4A is a block diagram that represents some components that may be used for audio content creation, according to some example embodiments.
[0023] FIG. 4B is a block diagram that represents some components that may be used for audio playback in a reproduction environment, according to some example embodiments.
[0024] FIG. 5 illustrates a workflow diagram for adaptive audio content creation and rendering workflow, according to some example embodiments.
[0025] FIGS. 6A and 6B illustrate flowcharts of a method and sub-method, respectively, of rendering audio where at least one audio object is rendered to include positional jitter, according to some example embodiments.
[0026] FIG. 7 illustrates a workflow diagram for adaptive audio content creation and rendering workflow in a situation involving streaming or other applications where the audio data is compressed before it is rendered, according to some example embodiments.DETAILED DESCRIPTION
[0027] Aspects of the one or more embodiments described herein may be implemented in an audio or audio-visual system that processes source audio information in a mixing, rendering, and playback system that includes one or more computers or processing devices executing software instructions. Any of the described embodiments may be used alone or together with one another in any combination.
[0028] The introduction of digital cinema and the development of true three-dimensional ("3D") or virtual 3D content has created new standards for sound, such as the incorporation of multiple channels of audio to allow for greater creativity for content creators and a more enveloping and realistic auditory experience for audiences. Expanding beyond traditional speaker feeds andchannel-based audio as a means for distributing spatial audio is critical, and there has been considerable interest in a model-based audio description that allows the listener to select a desired playback configuration with the audio rendered specifically for their chosen configuration. The spatial presentation of sound utilizes audio objects, which arc audio signals with associated parametric source descriptions of apparent source position (e.g., 3D coordinates), apparent source width, and other parameters. Further advancements include a next generation spatial audio (also referred to as "adaptive audio") format that comprises a mix of audio objects and traditional channel-based speaker feeds along with positional metadata for the audio objects. In a spatial audio decoder, the channels are sent directly to their associated speakers or down-mixed to an existing speaker set, and audio objects are rendered by the decoder in a flexible (adaptive) manner. The parametric source description associated with each object, such as a positional trajectory in 3D space, is taken as an input along with the number and position of speakers connected to the decoder. The Tenderer then utilizes certain algorithms, such as a panning law, to distribute the audio associated with each object across the attached set of speakers. The authored spatial intent of each object is thus optimally presented over the specific speaker configuration that is present in the listening room. Similar methods may be used for wearable devices (e.g., headphones, earbuds, other head-mounted devices, etc.) such that the authored spatial intent of each object is optimally presented over the specific speaker configuration of the wearable device. For example, audio (e.g., audio objects) may be binaurally rendered on a wearable device. In some instances, audio output by the speakers of a wearable device may be output in accordance with a head-related transfer function (HRTF) (e.g., a personalized HRTF (pHRTF)) of the wearer of the wearable device that takes into account the anatomy of the head, ears, and / or shoulders of the wearer to allow the audio output by the wearable device to seem more natural / realistic.
[0029] The advent of advanced object-based audio has significantly increased the complexity of the rendering process and the nature of the audio content transmitted to various different arrays of speakers. For example, cinema sound tracks may comprise many different sound elements corresponding to images on the screen, dialog, noises, and sound effects that emanate from different places on the screen and combine with background music and ambient effects to create the overall auditory experience. Accurate playback requires that sounds be reproduced in a way that corresponds as closely as possible to what is shown on screen with respect to sound source position, intensity, movement, and depth. Although advanced 3D audio systems (such as the Dolby® Atmos™ system) have largely been designed and deployed for cinema applications, consumer levelsystems are also being developed to bring a similar audio experience to home, mobile, automotive, entertainment venue, and extended reality (XR / virtual reality (VR) / augmented reality (AR)) environments. Accordingly, the disclosed techniques may be applied to numerous different arrays of speakers in a room and / or may be applied to different arrays of speakers included in wearable devices (e.g., binaural headphones, earbuds, and / or the like). Additionally, while audio content rendered as disclosed herein is often associated with corresponding video content on a playback device, the disclosed techniques may also be used when the audio content is not associated with video / imagery. For example, audio object positions and jitter may be rendered based on aesthetic decisions and / or based on intended wellness effects such as calming or concentration.
[0030] In some instances, the techniques disclosed herein are implemented as part of an audio system that is configured to work with a sound format and processing system that may be referred to as a "spatial audio system" or "adaptive audio system." Such a system is based on an audio format and rendering technology to allow enhanced audience immersion, greater artistic control, and system flexibility and scalability. An overall adaptive audio system generally comprises an audio encoding, distribution, and decoding system configured to generate one or more bitstreams containing both conventional channel-based audio elements and audio object coding elements. Such a combined approach provides greater coding efficiency and rendering flexibility compared to either channel-based or object-based approaches taken separately.
[0031] An example implementation of an adaptive audio system and associated audio format is the Dolby® Atmos™ platform. Such a system incorporates a height (up / down) dimension that may be implemented as a 5.1.4 surround system, or similar surround sound configuration. FIG. 1 illustrates the speaker placement in a present surround system (e.g., 5.1.4 surround) that provides height speakers for playback of height channels. The speaker configuration of the 5.1.4 system 100 is composed of five speakers 102 in the floor plane, four speakers 104 in the height plane, and a subwoofer 106. In general, these speakers may be used to produce sound that is designed to emanate from any position more or less accurately within the room. Predefined speaker configurations, such as those shown in FIG. 1, can naturally limit the ability to accurately represent the position of a given sound source. For example, a sound source cannot be panned further left than the left speaker itself. This applies to every speaker, therefore forming a one-dimensional (e.g., left-right), two-dimensional (e.g., left-right, front-back), or three-dimensional (e.g., left-right, front- back, up-down) geometric shape, in which the downmix is constrained. Various different speaker configurations and types may be used in such a speaker configuration. For example, certainenhanced audio systems may use speakers in a 5.1, 7.1, 5.1.2, 5.1.4, 7.1.4, 9.1, 11.1, 13.1, 19.4, or other configuration. The speaker types may include full range direct speakers, speaker arrays, surround speakers, height speakers, subwoofers, tweeters, and other types of speakers. While FIG.1 illustrates a surround sound speaker system provided in a room, an adaptive audio signal (c.g., a Dolby® Atmos™ signal) additionally or alternatively may be binaurally rendered to a wearable device (e.g., headphones, earbuds, etc.) of a listener.
[0032] Audio objects can be considered groups of sound elements that may be perceived to emanate from a particular physical location or locations in the listening environment (e.g., a movie theatre, a room in a listener’s home where a surround sound speaker system is used, a room / area where a listener is using a wearable device, and / or the like). Such objects can be static (that is, stationary) or dynamic (that is, moving). Audio objects are controlled by metadata that defines the position of the sound at a given point in time (i.e., audio object position data indicative of an audio object position of an audio object), along with other functions. The audio object position may include a plurality of different spatial points that define spatial movement of a dynamic audio object configured to move in 3D space. In some instances, the audio object position includes a single spatial point that defines a stationary spatial location of a static audio object. When objects are played back, they are rendered according to the positional metadata (e.g., the audio object position data of the audio object position) using the speakers that are present, rather than necessarily being output to a predefined physical channel. A track in a session can be an audio object, and standard panning data is analogous to positional metadata. In this way, in a video viewing environment, audio content placed on the screen might pan in effectively the same way as with channel-based content, but audio content placed in the surrounds can be rendered to an individual speaker if desired (e.g., referred to as a “snap” parameter that is used to avoid timbral side effects from interpolation across multiple speaker positions). While the use of audio objects provides the desired control for discrete effects, other aspects of a soundtrack may work effectively in a channel-based environment. For example, many ambient effects or reverberation actually benefit from being fed to arrays of speakers. Although these could be treated as objects with sufficient width to fill an array, it is beneficial to retain some channel-based functionality.
[0033] The adaptive audio system is configured to support audio beds in addition to audio objects, where beds are effectively channel-based sub-mixes or stems. These can be delivered for final playback (rendering) either individually, or combined into a single bed, depending on the intent of the content creator. These beds can be created in different channel-based configurationssuch as 5.1, 7.1, and 9.1, arrays that include overhead speakers such as shown in FIG. 1, and / or wearable devices with binaural speakers. FIG. 2 illustrates the combination of channel and objectbased data to produce an adaptive audio mix, according to some example embodiments. As shown in process 200, the channel-based data 202, which, for example, may be 5.1 or 7.1 surround sound data provided in the form of pulse-code modulated (PCM) data is combined with audio object data 204 to produce an adaptive audio mix 208. The audio object data 204 is produced by combining the elements of the original channel-based data with associated metadata that specifies certain parameters pertaining to the location of the audio objects. As shown conceptually in FIG. 2, the authoring tools provide the ability to create audio programs that contain a combination of speaker channel groups and object channels simultaneously. For example, an audio program could contain one or more speaker channels optionally organized into groups (or tracks, e.g., a stereo or 5.1 track), descriptive metadata for one or more speaker channels, one or more object channels, and descriptive metadata for one or more object channels.
[0034] FIG. 3 is a block diagram that provides examples of components of an authoring and / or rendering apparatus. In this example, the device 300 includes an interface system 305. The interface system 305 may include a network interface, such as a wireless network interface. Alternatively, or additionally, the interface system 305 may include a universal serial bus (USB) interface or another such interface.
[0035] The device 300 (i.e., a computing apparatus) includes a logic system 310. The logic system 310 may include an electronic processor, such as a general purpose single- or multi-chip processor. The logic system 310 may include a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, or discrete hardware components, or combinations thereof. The logic system 310 may be configured to control the other components of the device 300. Although no interfaces between the components of the device 300 are shown in FIG. 3, the logic system 310 may be configured with interfaces for communication with the other components. The other components may or may not be configured for communication with one another, as appropriate.
[0036] The logic system 310 may be configured to perform audio authoring and / or rendering functionality, including but not limited to the types of audio authoring and / or rendering functionality described herein. In some such implementations, the logic system 310 may be configured to operate (at least in part) according to software stored one or more non-transitorymedia. The non-transitory media may include memory associated with the logic system 310, such as random access memory (RAM) and / or read-only memory (ROM). The non-transitory media may include memory of the memory system 315. The memory system 315 may include one or more suitable types of non-transitory storage media, such as flash memory, a hard drive, etc. The memory system 315 may include a non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform any one or combination of the methods and functionality described herein.
[0037] The speakers 320 may include one or more speakers 320 that form an array of speakers 320. For example, the number of speakers may be two speakers in a binaural wearable device. As another example, the number of speakers may be many more than two speakers in a surround sound speaker system in a room of a home and / or in a movie theatre.
[0038] The display system 330 may include one or more suitable types of display, depending on the manifestation of the device 300. For example, the display system 330 may include a liquid crystal display, a plasma display, a bistable display, etc.
[0039] The user input system 335 may include one or more devices configured to accept input from a user. In some implementations, the user input system 335 may include a touch screen that overlays a display of the display system 330. The user input system 335 may include a mouse, a track ball, a gesture detection system, a joystick, one or more GUIs and / or menus presented on the display system 330, buttons, a keyboard, switches, etc. In some implementations, the user input system 335 may include the microphone 325: a user may provide voice commands for the device 300 via the microphone 325. The logic system may be configured for speech recognition and for controlling at least some operations of the device 300 according to such voice commands.
[0040] The power system 340 may include one or more suitable energy storage devices, such as a nickel-cadmium battery or a lithium-ion battery. The power system 340 may be configured to receive power from an electrical outlet.
[0041] FIG. 4A is a block diagram that represents some components that may be used for audio content creation. The system 400 may, for example, be used for audio content creation in mixing studios and / or dubbing stages. In this example, the system 400 includes an audio and metadata authoring tool 405 and a rendering tool 410. Each of the authoring tool 405 and the rendering tool 410 may be a separate instance of the device 300 shown in FIG. 3, and may have similar or different combinations of the components shown in FIG. 3. In this implementation, the audio and metadataauthoring tool 405 and the rendering tool 410 include audio connect interfaces 407 and 412, respectively, which may be configured for communication via Audio Engineering Society / European Broadcasting Union (AES / EBU), Multichannel Audio Digital Interface (MADI), analog, etc. The audio and metadata authoring tool 405 and the rendering tool 410 include network interfaces 409 and 417, respectively, which may be configured to send and receive metadata via Transmission Control Protocol / Intemet Protocol (TCP / IP) or any other suitable protocol. The interface 420 is configured to output audio data to speakers.
[0042] The system 400 may, for example, include an existing authoring system, such as a Pro Tools™ system, running a metadata creation tool (i.e., a panner as described herein) as a plugin. The panner could also run on a standalone system (e.g. a personal computer (PC) or a mixing console) connected to the rendering tool 410, or could run on the same physical device as the rendering tool 410. In the latter case, the panner and Tenderer could use a local connection e.g., through shared memory. The panner graphical user interface (GUI) could also be remoted on a tablet device, a laptop, etc. The rendering tool 410 may comprise a rendering system that includes a sound processor that is configured for executing rendering software. The rendering system may include, for example, a personal computer, a laptop, etc., that includes interfaces for audio input / output and an appropriate logic system.
[0043] FIG. 4B is a block diagram that represents some components that may be used for audio playback in a reproduction environment (e.g., a movie theater). The system 450 includes a cinema server 455 and a rendering system 460 in this example. Each of the cinema server 455 and the rendering system 460 may be a separate instance of the device 300 shown in FIG. 3, and may have similar or different combinations of the components shown in FIG. 3. The cinema server 455 and the rendering system 460 include network interfaces 457 and 462, respectively, which may be configured to send and receive audio objects via TCP / IP or any other suitable protocol. The interface 464 is configured to output audio data to speakers. While the server 455 is referred to as a cinema server 455 and an example of the reproduction environment is provided as a movie theatre, FIG. 4B may additionally or alternatively be representative of other types of servers and reproduction environments. For example, the reproduction environment may be a room in a personal home that includes a surround sound speaker system, and the server 455 may be a local server at the home or a server from which a device at the home downloads audio content including audio objects. As another example, the reproduction environment may be a room / area where alistener is using a wearable device, and the server 455 may be a local server at the home or another server from which the wearable device downloads audio content including audio objects.
[0044] While existing adaptive audio systems allow content creators to generate spatial audio scenes that that include one or more beds and one or more audio objects, where the audio objects each include a static or dynamic audio object position within a rendering / reproduction environment, existing adaptive audio systems do not introduce jitter / movement into the spatial position of audio objects. Introducing jitter / movement (e.g., periodically varying movement) of one or more audio objects to define a spatial movement of the audio object that is separately defined from the static or dynamic audio position of the audio object increases perceptual accuracy of audio object positions as perceived by a listener and / or may provide additional benefits as disclosed herein.
[0045] For example, even when a head of a listener is seemingly being head still, the head is always in some sort of slight motion due to minor head tremor. This slight motion affects the listener’ s perception of sound emanating from real world objects and emanating from an audio system. Because audio objects output by existing adaptive audio systems are output at a precise static or dynamic audio object position within the rendering / reproduction environment, these audio objects may be perceived differently than a real-world sound emanating from an object at the same spatial position as the audio object position. To address this technological problem, introducing jitter / movement to the audio objects output by an adaptive audio system increases the perceptual accuracy (e.g., naturalness) of the audio object compared to a real-world sound emanating from an object at the same spatial position as the audio object position. Additionally, introducing jitter / movement control to audio objects adds additional flexibility to define and / or control the playback experience (e.g., emphasize an audio object by controlling it to jitter or jitter more than other audio objects). In this way, jitter / movement control provides a technological solution to a technological problem with existing adaptive audio systems by providing a tool to ensure the intended audio object position is conveyed accurately in various playback scenarios and by providing a tool for enabling new forms of artistic expression. For example, different jitter patterns may be defined for different playback environments / systems (e.g., to improve perceptual accuracy / naturalness), and / or jitter may be used for other purposes such as attention steering (i.e., triggering application of jitter to an audio object to emphasize or de-emphasize the audio object relative to the spatial audio scene). Additional uses for various specific jitter patterns are also envisioned.
[0046] FIG. 5 illustrates a workflow diagram for adaptive audio content creation and rendering workflow 500 according to some example embodiments. At block 505, audio is separated into beds and audio objects in a digital audio workstation (DAW) such as the system 400 and 450 explained previously herein that may include a Pro Tools™ system. At block 510, the beds and the audio objects are mixed and automated in the DAW. Spatial audio mixing and / or automation may be performed, for example, using Dolby® Atmos™ plugins such as a native Pro Tools™ panner, a Dolby® Atmos™ Music Panner, or the like. As indicated in FIG. 5, blocks 505 and 510 are content preparation steps / tasks that may be performed, for example, by an audio and metadata authoring tool such as the audio and metadata authoring tool 405 of FIG. 4A.
[0047] Once mixing and automation is completed, at block 515, the audio is sent through a renderer (e.g., a Dolby® Atmos™ renderer) and rendered into a compatible adaptive audio format (e.g., a Dolby® Atmos™ compatible format such as Audio Definition Model Broadcast Wave Format (ADM BWF), MP4, etc.). As indicated in FIG. 5, block 515 is a content rendering step / task that may be performed, for example, by a rendering tool / system such as the rendering tool / system 410 and / or 460 of FIGS. 4 A and 4B.
[0048] Blocks 505, 510, and 515 may be included in existing adaptive audio systems such as the Dolby® Atmos™ system. Additional details with respect to the actions performed at blocks 505, 510, and 515 to author and render 3D audio are disclosed in U.S. Patent Application No.14 / 126,901, now U.S. Patent No. 9,204,236, which is hereby incorporated by reference in its entirety and appended herein.
[0049] As indicated in FIG. 5 and in accordance with the improvements to existing adaptive audio systems mentioned previously herein, positional jitter of an audio object (e.g., periodically varying movement of the audio object position of one or more audio objects) may be enabled in the content rendering stage of the workflow 500 (e.g., between blocks 510 and 515 as shown in FIG. 5), for example, by the rendering tool / system 410 and / or 460 of FIGS. 4A and 4B. In some instances, at block 520, a positional jitter pattern is defined for one or more audio objects as explained in greater detail below. For example, positional jitter metadata may be defined by a user (e.g., using the authoring tool 405 during content preparation and rendering as shown in FIG. 7) and / or may be defaulted to a predetermined positional jitter pattern (e.g., using the rendering tool / system 410 and / or 460 during content rendering as shown in FIG. 5) as explained in greater detail below. At block 525, the positional jitter is enabled for one or more of the audio objects. The workflow 500 then proceeds to block 515 to render the audio in accordance the positional jitter pattern for eachaudio object as well as in accordance with other parameters of each audio object such as each audio object’s audio object position that may be separately defined with respect to any movement defined by the positional jitter pattern for the audio object.
[0050] FIGS. 6 A and 6B illustrate flowcharts of a method 600 and sub-method 630, respectively, of rendering audio where at least one audio object is rendered to include positional jitter according to some example embodiments. The method 600 and sub-method 630 are described as being performed by an electronic processor (i.e., one or more electronic processors) of the device 300 of FIG. 3 that may represent one or more of the devices 405, 410, 455, and / or 460 of the systems 400 and 450 of FIGS. 4A and 4B. The electronic processor that performs the methods and sub-methods described herein may include any one or a combination of electronic processors located within a single device or distributed among various devices and / or systems (e.g., the device(s) 300 and / or the systems 400 and / or 450). Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other clement configured in a certain manner, for example, to make multiple determinations, the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations. To reiterate, those electronic processors and processing may be distributed. While a particular order of processing steps, message receptions, and / or message transmissions is indicated in FIGS. 6A and 6B as an example, timing and ordering of such steps, receptions, and transmissions may vary where appropriate without negating the purpose and advantages of the examples set forth in detail throughout the remainder of this disclosure.
[0051] At block 605, the logic system 310 (e.g., an electronic processor of the logic system 310) of the device 300 receives, via the interface system 305, audio reproduction data including audio objects and associated metadata for each audio object. In some instances, each of the audio objects includes an audio object position. As indicated previously herein, the audio object position may indicate that the audio object is a static audio object such that the audio object remains stationary within a reproduction environment (e.g., perceived as a stationary object by a listener).Alternatively, the audio object position may indicate that the audio object is a dynamic audio object such that the audio object spatially moves within the reproduction environment according to movement defined by the audio object position (e.g., perceived as a moving object by the listener).
[0052] In some instances, the associated metadata for each audio object includes audio object position data indicative of the audio object position of the audio object. For example, the audioobject position data may indicate whether the audio object is a static audio object or a dynamic audio object. For a static audio object, the audio object position data may indicate a 3D spatial (i.e., positional) position / location of the static audio object in the reproduction environment. For a dynamic audio object, the audio object position data may indicate a 3D spatial path in the reproduction environment along which the dynamic audio object may move. Additionally, the audio object position data may include timing information associated with the position and / or movement of the audio object (e.g., how long the audio object is configured to be located in a certain 3D spatial position, a speed or acceleration at which a dynamic audio object is configured to move along the 3D spatial path, etc.).
[0053] At block 610, the logic system 310 receives, via the interface system 305, reproduction environment data including an indication of a number of reproduction speakers in a reproduction environment and an indication of a location of each reproduction speaker within the reproduction environment. As indicated previously herein, the reproduction environment may include a movie theatre, a room in a listener’s home where a surround sound speaker system is used, a room / area where a listener is using a wearable device, and / or the like. The number of reproduction speakers may vary, for example, from two speakers in a binaural wearable device to many more than two speakers in a surround sound speaker system in a room of a home and / or in a movie theatre.
[0054] At block 615, the logic system 310 renders the audio objects into one or more speaker feed signals based, at least in part, on the associated metadata. In some instances, each speaker feed signal corresponds to at least one of the reproduction speakers within the reproduction environment. As indicated in FIG. 6A, after completion of block 615, the method 600 may proceed back to block 605 to repeat to continue receiving audio data and rendering audio content based on the audio data.
[0055] In some instances, at block 615, the logic system 310 renders the audio objects according to sub-method 630 shown in FIG. 6B. While blocks 635 and 640 of the sub-method 630 are shown as separate blocks, in some instances, the blocks 635 and 640 are executed contemporaneously with each other (i.e., in conjunction with each other) to render audio objects in the reproduction environment.
[0056] At block 635, the logic system 310 renders a first audio object according to a first audio object position of the first audio object. In some instances, the first audio object may be rendered according to its audio object position data (included in associated metadata) that is indicative of the audio object position of the first audio object. For example, a static audio object may be renderedsuch that the static audio object is perceived to be located at a stationary spatial position within the reproduction environment. On the other hand, a dynamic audio object may be rendered such that the dynamic audio object is perceived to be moving along a predetermined spatial path within the reproduction environment (e.g., horizontally across the reproduction environment when the audio object is associated with, for example, a vehicle driving horizontally across the display system 330).[00571 At block 640, the logic system 310 renders the first audio object in accordance with a positional jitter pattern that defines a periodically varying movement of the first audio object. The periodically varying movement of the positional jitter pattern is separately defined with respect to any movement (and / or location) defined by the first audio object position of the first audio object. In other words, the positional jitter pattern is separately defined with respect to the first audio object position. Accordingly, a static audio object that is rendered according to a positional jitter pattern remains stationary within the reproduction environment except for movement in accordance with the positional jitter pattern. Despite the movement defined by the positional jitter pattern, the static audio object is nevertheless considered to be a static audio object because its audio object position (e.g., as separately defined by audio object position data) is stationary. In other words, the movement caused by the positional jitter pattern may not be perceived by the listener as movement of the audio object within the reproduction environment. Turning to a dynamic audio object, the dynamic audio object may be configured to spatially move within the reproduction environment according to movement defined by the first audio object position (e.g., along a predetermined spatial path within the reproduction environment that can be perceived by the listener as movement of the audio object within the reproduction environment). However, the movement defined by the first audio object position is separately defined with respect to the periodically varying movement defined by the positional jitter pattern. Thus, the positional jitter movement of the dynamic audio object may not, itself, be perceived by the listener as movement of the audio object within the reproduction environment.
[0058] In general, movement of the first audio object in accordance with the positional jitter pattern involves movement within a smaller area / volume than movement of a dynamic audio object defined by the audio object position. For example, a dynamic audio object associated with a bouncing ball on the display system 330 may include an audio object position that includes periodically varying movement (i.e., a first periodically varying movement) in an up-down direction between two 3D spatial positions. In instances in which the dynamic audio object associated with the bouncing ball is also rendered according to a positional jitter pattern that includes its ownperiodically varying movement (i.e., a second periodically varying movement), a distance between 3D spatial positions of the second periodically varying movement of the positional jitter pattern may be approximately one, two, three, or more orders of magnitude smaller than a distance between the 3D spatial positions of the first periodically varying movement defined by the audio object position. Accordingly, the second periodically varying movement of the positional jitter pattern may be less perceptible to the listener than the first periodically varying movement defined by the audio object position.
[0059] In some instances, movement of an audio object in accordance with a positional jitter pattern is shaped in any one or combination of different manners. A few examples of types of the positional jitter pattern (i.e., positional jitter pattern shapes) are explained immediately below. The positional jitter patterns may be implemented on one or more two-dimensional planes within the reproduction environment to create a modulation in a three-dimensional space.
[0060] As one example, the positional jitter pattern includes a sinusoidal pattern (i.e., a sine wave) such that the first audio object position of the first audio object oscillates back and forth along a positional path between a first positional point (i.e., 3D spatial point / location) and a second positional point to cover spatial positions (i.e., the 3D spatial path) between the first positional point and the second positional point. The sinusoidal pattern may include a modulation in three- dimensional space (e.g., sin_x(wt) * sin_y(wt) * sin_y(wt)). In some instances, this type of motion may take place according to a velocity (e.g., how quickly the audio object moves between the first positional point and the second positional point) that is user-defined or that is predetermined by a default velocity setting. Additionally, an amplitude of a waveshape of the sine wave (e.g., how far the movement occurs in each direction in, for example, a two-dimensional plane in which the sine wave is defined) may be user-defined or may be predetermined by a default amplitude setting.
[0061] As another example, the positional jitter pattern includes a square wave such that the first audio object position of the first audio object periodically alternates between being located at a first positional point for a first amount of time and being located at a second positional point for a second amount of time (e.g., without covering spatial position between the first positional point and the second positional point). For example, such movement in three-dimensional space may cause the first audio object position to switch between different comers of a three-dimensional cube. The first amount of time may be the same or different than the second amount of time. Additionally, one or both of the first amount of time and the second amount of time may change while the first audio object is being rendered (e.g., the first amount of time and the second amount of time may havetheir values swapped with each other halfway through a time period during which the first audio object is rendered). In some instances, the first amount of time and the second amount of time may be user-defined or predetermined by a default timing setting. Additionally, an amplitude of the square wave that indicates a distance between the two positional points (e.g., how far away the first positional point is from the second positional point) may be user-defined or may be predetermined by a default amplitude setting.
[0062] As another example, the positional jitter pattern may include a custom pattern that is determined by a content creator / user. In some instances, details of the custom pattern that may be selected include any one or a combination of a waveshape of the positional jitter pattern, a velocity, acceleration, and / or frequency of positional jitter associated with the positional jitter pattern, an amplitude of the positional jitter pattern (e.g., a maximum distance that the audio object is configured to jitter between positional points or positional paths, one or more 2D or 3D boundaries within which positional jitter is configured to occur, etc.), and the like.
[0063] As another example, the positional jitter pattern includes a noise pattern modeled according to a type of noise. There are many different noise patterns that may be modeled and applied including, pink noise, white noise, brown noise, and others to induce randomness into the positional jitter pattern of motion. Additionally, noise shapes may be combined with other positional jitter patterns to alter the shape of the other positional jitter patterns. The noise pattern may be implemented across one, two, or three of the spatial axes in three-dimensional space.
[0064] As another example, the positional jitter pattern includes a positional path derived from accelerometer data indicative of human movements during listening. In some instances, accelerometer data obtained by a head-mounted device being worn by a listener may be processed and shaped into object motion to create a positional jitter pattern derived from natural human movements. In some instances, the accelerometer data may be obtained from a test listener to estimate a head tremor signal. Such accelerometer data may then be used to create a positional jitter pattern for other listeners. In other instances, the accelerometer data may be obtained in real-time from a listener who is currently consuming audio such that the positional jitter pattern used for one or more audio objects matches the physiological tremor of the listener (e.g., a head tremor) in realtime. For example, the rendering tool / system 410, 460 may receive accelerometer data from a headmounted device. In some instances, data corresponding to the physiological tremor of the listener may be obtained / received in additional or alternative ways besides using an accelerometer. For example, other sensors and / or cameras physically in contact with the listener or distanced from thelistener but configured to observe the listener may be used, and may provide data related to movement of the listener’s head to the rendering tool / system 410, 460. In some instances, the positional jitter pattern of an audio object is set to be a first pattern in response to determining that accelerometer data indicative of movement of a head of a listener indicates that the head of the listener is moving. Additionally, the positional jitter pattern may be set to be a second pattern different than the first pattern in response to determining that the accelerometer data indicates that the head of the listener is stationary. In other words, jitter of audio objects may be controlled to be different depending on whether a listener is moving versus stationary or moving more than a threshold amount versus moving less than a threshold amount. In some instances, jitter of audio objects may be controlled to present or absent depending on whether the listener is moving versus stationary or moving more than a threshold amount versus moving less than a threshold amount.
[0065] The above-noted positional jitter patterns / shapes are merely examples. Other patterns / shapes may be used in other instances (e.g., a gaussian pattern).
[0066] In the present disclosure, a positional jitter pattern is described as defining a periodically varying movement of an audio object. In this context, the term “periodically” is intended to include a broad definition meaning “from time to time” or “occasionally.” This broad definition may also encompass narrower definitions of “periodically” such as “at regularly occurring intervals.” In other words, the disclosed “periodically varying movement” in accordance with the positional jitter pattern may occur at regular intervals (e.g., at regularly occurring intervals in accordance with a repeated pattern / waveform such as a sine wave, a square wave, and / or the like) or at irregular intervals (e.g., randomly, quasi-randomly, in accordance with random noise patterns, and / or the like). In some instances, the periodically varying movement in accordance with the positional jitter pattern includes a constant or varying type of the positional jitter pattern, a constant or varying velocity of positional jitter associated with the positional jitter pattern, a constant or varying frequency of the positional jitter pattern, a constant or varying amplitude of the positional jitter pattern, and combinations thereof.
[0067] While the periodically varying movement of the positional jitter pattern is separately defined with respect to any movement defined by the audio object position of an audio object, the periodically varying movement of the positional jitter pattern may utilize the audio object position of the audio object to render the periodically varying movement. For example, a spatial (i.e., positional) point / location defined by the audio object position may be used as one of the positional points between which the audio object jitters / moves in accordance with positional jitter pattern. Asanother example, the spatial point / location defined by the audio object position may be used as a center point of the positional jitter pattern to define two or more points or a 2D or 3D virtual boundary within which the audio object jitter s / moves in accordance with the positional jitter pattern.
[0068] As indicated previously herein, the associated metadata of an audio object includes audio object position data indicative of the audio object position of the audio object. With the modification to render the audio object in accordance with a positional jitter pattern as explained herein, the associated metadata may also include positional jitter data indicative of the positional jitter pattern of the audio object. In some instances, the positional jitter data indicates whether an audio object is to be rendered in accordance with the positional jitter pattern (see “jitter tag” column in Table 1 below). For example, positional jitter for each audio object may be enabled or disabled as indicated based on positional jitter data in metadata. In some instances, the positional jitter data indicates at least one of a group consisting of a type of the positional jitter pattern (e.g., a shapc / pattcrn) (sec “pattern” column in Table 1 below), a velocity of positional jitter associated with the positional jitter pattern, a frequency of the positional jitter pattern, an amplitude of the positional jitter pattern (see “amplitude” column in Table 1 below), and combinations thereof.
[0069] This positional jitter data included in the metadata is a new type of metadata that is not included in existing adaptive audio systems and may be referred to as side data, sidecar data, or supplemental enhancement information (SEI) messaging if working in Moving Picture Experts Groups (MPEG) systems. In some instances, the sidecar positional jitter data is updated only at an onset and at an ending of the rendering of an audio object, which means that the sidecar positional jitter data has a low payload, for example, compared to audio object position data of a dynamic audio object.
[0070] For example, FIG. 7 illustrates a workflow diagram for adaptive audio content creation and rendering workflow 700 in a situation involving streaming or other applications where the audio data is compressed before it is rendered, according to some example embodiments. The workflow 700 is similar to the workflow 500 of FIG. 5 except that the workflow 700 involves a compression and decompression stage due to the situation in which the audio is being rendered (e.g., streaming).
[0071] At block 705 (which is similar to block 505 of FIG. 5), audio is separated into beds and audio objects in a DAW. At block 710 (which is similar to block 510 of FIG. 5), the beds and the audio objects are mixed and automated in the DAW. At block 715 (which is similar to block 520 of FIG. 5), a positional jitter pattern is defined for one or more audio objects as explained herein.
[0072] At block 720, the positional jitter patterns for each audio object are consolidated into a positional jitter metadata package. For example, Table 1 below shows positional jitter data included in sidecar metadata for an audio object. In some instances, the positional jitter data may include additional or fewer parameters than those shown in Table 1. For example, Table 1 may also include a “velocity” column, a “frequency” column, and / or a “time” column that indicates when jitter should occur (in instances when jitter does not occur for the entirety of the rendering of the audio object). In some instances, the “time” column may indicate when jitter should occur with reference to a time since the audio object has been rendered in order to reduce data payload compared to indicating a time when the jitter should occur relative to an overall audio scene. The consolidated positional jitter metadata package may include numerous instances of positional jitter data that are each configured to control a certain audio object. Alternatively, a single instance of positional jitter data may be used to control some or all audio objects being rendered in some applications.Table 1:
[0073] By performing block 715 and 720, the logic system 310 (e.g., an electronic processor of the logic system 310) of the device 300 embodied as an authoring tool such as the authoring tool 405 generates metadata including positional jitter metadata (e.g., based on user input, based on default positional jitter values, etc.) for each audio object included in audio reproduction data including a plurality of audio objects. As explained previously herein, the positional jitter metadata is indicative of a positional jitter pattern of a respective audio object. In some instances, the positional jitter pattern defines a periodically varying movement of the respective audio object. The periodically varying movement defined by the positional jitter pattern may be separately definedwith respect to a spatial location or any spatial movement defined by an audio object position of the respective audio object. In some instances, during performance of blocks 715 and 720, the logic system 310 generates the metadata to additionally include audio object position data indicative of the audio object position of the respective audio object. This audio object position data separately defines a spatial location and / or spatial movement of the respective audio object within the reproduction environment. In other words, the spatial location and / or spatial movement of the respective audio object defined by the audio object position data is separately defined with respect to the positional jitter pattern of the respective audio object. Accordingly, while much of the present disclosure discloses rendering of audio objects according to their respective audio object position and their respective positional jitter pattern that is separately defined with respect to their respective audio object position, the present disclosure also discloses authoring of the respective audio object position and / or the respective positional jitter pattern for each audio object.
[0074] As indicated in FIG. 7, blocks 705-720 arc content preparation steps / tasks that may be performed, for example, by an audio and metadata authoring tool such as the audio and metadata authoring tool 405 of FIG. 4A. Once the content preparation and rendering blocks 705-720 are completed, at block 725, conventional Dolby® Atmos™ or binaural stream compression and decompression may occur with respect to the sound sources. Separately but contemporaneously, at block 730, optional compression and decompression of the positional jitter metadata (i.e., sidecar data) may occur.
[0075] At block 735, the decompressed audio data is separated into beds and audio objects. At block 740, positional jitter metadata is associated with one or more of the audio objects. At block 745 (which is similar to block 515 of FIG. 5), the audio data (including its associated positional jitter metadata) is prepared for speaker waveforms or binaural waveforms and audio is rendered into a compatible adaptive audio format (e.g., a Dolby® Atmos™ compatible format such as ADM BWF, MP4, etc.). As indicated in FIG. 7, blocks 735, 740, and 745 are content playback steps / tasks that may be performed, for example, by a rendering tool / system such as the rendering tool / system 460 of FIG. 4B.
[0076] When rendering an audio object to include positional jitter, positional jitter data of the audio object may be updated only at an onset and at an ending of the rendering of the first audio object because the time period in between the onset and the ending of the rendering of the audio object may involve repetitive jitter motion along the same spatial path (or along the same spatial path relative to the audio object position of the audio object for dynamic audio objects).Accordingly, the payload of the positional jitter metadata may be low since different positional jitter metadata may not be used at different times during the rendering of the audio object. In other words, because the positional jitter data is stored in sidecar metadata that can be applied after decompression of the audio data as indicated in FIG. 7, increased levels of entropy (i.c., randomness) in rendered audio objects may be achieved without having to compress and decompress high volumes of data that may otherwise be associated with the movement achieved by the positional jitter metadata.
[0077] As indicated throughout the disclosure, positional jitter metadata may be defined using the authoring tool 405 during content preparation and rendering, for example, as indicated in FIG. 7. The positional jitter metadata may then be provided to (i) a rendering tool / system 410 and / or 460 for playback as indicated in FIG. 7 and / or (ii) to a server 455 to later be provided to a rendering tool / system 410 and / or 460 for playback. Additionally or alternatively, positional jitter metadata may be defined using the rendering tool / system 410 and / or 460 during content rendering, for example, as indicated in FIG. 5. For example, the positional jitter metadata may include user- selected and / or default positional jitter patterns that set at the rendering tool / system 410 and / or 460 and that are applied to audio objects during content rendering.
[0078] As indicated throughout the disclosure, different audio objects in an audio scene may be rendered according to the same or different positional jitter patterns. For example, a first audio object may be rendered (i) according to a first audio object position of the first audio object, and (ii) in accordance with a positional jitter pattern that defines a periodically varying movement of the first audio object. Continuing this example, a second audio object may be rendered (i) according to a second audio object position of the second audio object, and (ii) in accordance with a second positional jitter pattern that defines a second periodically varying movement of the second audio object. Further continuing this example, the second positional jitter pattern is different than the positional jitter pattern according to which the first audio object is rendered. Additionally, as explained previously herein, the periodically varying movement of the positional jitter pattern of the first audio object is separately defined with respect to any movement defined by the first audio object position of the first audio object. Furthermore, the second periodically varying movement of the second positional jitter pattern of the second audio object is separately defined with respect to any movement defined by the second audio object position of the second audio object.
[0079] In some instances, one or more audio objects in an audio scene may be rendered according to one or more positional jitter patterns while one or more other audio objects in the audioscene are not rendered according to a positional jitter pattern. For example, as a continuation of the example immediately above, a third audio object may be rendered (i) according to a third audio object position of the third audio object, and (ii) without the positional jitter pattern or any other positional jitter pattern. In some instances, positional jitter of an audio object may be enabled and disabled over time. For example, an audio object may initially be rendered without positional jitter but may be rendered with positional jitter after a certain period of time. Continuing this example, the positional jitter of the audio object may be disabled after another certain period of time to again have the audio object rendered without positional jitter. Along similar lines and as indicated previously herein, because audio object may have their own positional jitter data, different audio objects may be jittered with different jitter onset / enable and jitter offset / disable times such that different audio objects in an audio scene may be rendered according to include or not include positional jitter at different times.
[0080] In some instances, different positional jitter characteristics may be applied to one or more audio objects in different listening / playback / reproduction environments to enhance user experience (e.g., realness / naturalness of audio objects perceived by the user / listener).
[0081] While the disclosure primarily describes positional jitter / movement with respect to audio objects, in some instances, positional jitter / movement may additionally or alternatively be applied to audio beds in a similar manner.
[0082] It is to be understood that the embodiments are not limited in its application to the details of the configuration and arrangement of components set forth herein or illustrated in the accompanying drawings. The embodiments are capable of being practiced or of being carried out in various ways. Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted,” “connected,” “supported,” and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
[0083] In addition, it should be understood that embodiments may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that,in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on non-transitory computer-readable medium) executable by one or more electronic processors, such as a microprocessor and / or application specific integrated circuits (“ASICs”). As such, it should be noted that a plurality of hardware and software based devices, as well as a plurality of different structural components, may be utilized to implement the embodiments. For example, “servers” and “computing devices” described in the specification can include one or more electronic processors, one or more computer-readable medium modules, one or more input / output interfaces, and various connections (e.g., a system bus) connecting the various components.
[0084] Various features and advantages are set forth in the following claims.
Claims
CLAIMSWhat is claimed is:
1. A method of rendering audio, the method comprising: receiving audio reproduction data including audio objects and associated metadata for each audio object, wherein each of the audio objects includes an audio object position; receiving reproduction environment data including an indication of a number of reproduction speakers in a reproduction environment and an indication of a location of each reproduction speaker within the reproduction environment; and rendering the audio objects into one or more speaker feed signals based, at least in part, on the associated metadata, wherein each speaker feed signal corresponds to at least one of the reproduction speakers within the reproduction environment; wherein rendering the audio objects includes rendering a first audio object:(i) according to a first audio object position of the first audio object, and(ii) in accordance with a positional jitter pattern that defines a periodically varying movement of the first audio object, wherein the periodically varying movement is separately defined with respect to any movement defined by the first audio object position of the first audio object.
2. The method of claim 1, wherein the associated metadata of the first audio object includes: audio object position data indicative of the audio object position of the first audio object; and positional jitter data indicative of the positional jitter pattern of the first audio object.
3. The method of claim 2, wherein the positional jitter data indicates whether the first audio object is to be rendered in accordance with the positional jitter pattern.
4. The method of claim 2 or claim 3, wherein the positional jitter data indicates at least one of a group consisting of a type of the positional jitter pattern, a velocity of positional jitter associated with the positional jitter pattern, a frequency of the positional jitter pattern, an amplitude of the positional jitter pattern, and combinations thereof.
5. The method of any one of claims 2-4, wherein the positional jitter data is sidecar data that is updated only at an onset and at an ending of the rendering of the first audio object.
6. The method of any one of the preceding claims, wherein the positional jitter pattern includes a sinusoidal pattern such that the first audio object position of the first audio object oscillates back and forth along a positional path between a first positional point and a second positional point to cover spatial positions between the first positional point and the second positional point.
7. The method of any one of claims 1-5, wherein the positional jitter pattern includes a square wave such that the first audio object position of the first audio object periodically alternates between being located at a first positional point for a first amount of time and being located at a second positional point for a second amount of time.
8. The method of any one of the preceding claims, wherein the positional jitter pattern is set to be a first pattern in response to determining that accelerometer data indicative of movement of a head of a listener indicates that the head of the listener is moving; wherein the positional jitter pattern is set to be a second pattern different than the first pattern in response to determining that the accelerometer data indicates that the head of the listener is stationary; and wherein the first pattern includes a positional path derived from the accelerometer data indicative of human movements during listening.
9. The method of any one of the preceding claims, wherein the first audio object includes a static audio object such that the first audio object remains stationary within the reproduction environment except for movement in accordance with the positional jitter pattern.
10. The method of any one of the preceding claims, wherein the first audio object includes a dynamic audio object such that the first audio object spatially moves within the reproduction environment according to movement defined by the first audio object position, wherein the movement defined by the first audio object position is separately defined with respect to the periodically varying movement defined by the positional jitter pattern.
11. The method of any one of the preceding claims, wherein rendering the audio objects includes rendering a second audio object:(i) according to a second audio object position of the second audio object; and(ii) in accordance with a second positional jitter pattern that defines a second periodically varying movement of the second audio object, wherein the second periodically varying movement is separately defined with respect to any movement defined by the second audio object position of thesecond audio object, wherein the second positional jitter pattern is different than the positional jitter pattern according to which the first audio object is rendered.
12. A computing apparatus, comprising: at least one electronic processor; and memory storing instructions, which when executed by the at least one electronic processor, cause the computing apparatus to perform the method of any one of claims 1-11.
13. A non-transitory computer-readable storage medium storing instructions which, when executed by a computing apparatus, cause the computing apparatus to perform the method of any one of claims 1-11.
14. A method of authoring audio, the method comprising: generating metadata including positional jitter metadata for each audio object included in audio reproduction data including a plurality of audio objects; wherein the positional jitter metadata is indicative of a positional jitter pattern of a respective audio object; and wherein the positional jitter pattern defines a periodically varying movement of the respective audio object, wherein the periodically varying movement is separately defined with respect to a spatial location or any spatial movement defined by an audio object position of the respective audio object.
15. The method of claim 14, wherein generating the metadata includes generating audio object position data indicative of the audio object position of the respective audio object.
Citation Information
Patent Citations
System and Tools for Enhanced 3D Audio Authoring and Rendering
US20140119581A1
System and tools for enhanced 3D audio authoring and rendering
US9204236B2
Method and apparatus for the reproduction of multi-channel audio signals
EP1021062A2
Rendering virtual audio sources using loudspeaker map deformation
EP3145220A1
Method and System for Spatialization of Sound by Dynamic Movement of the Source
US20100183159A1