Providing object-based multi-sensory experience

By processing object-based sensory data and metadata, the flexibility of delivering multi-sensory content across different facilities is addressed, enabling the creation and rendering of multi-sensory experiences across devices and enhancing the flexibility and scalability of lighting fixtures and actuators.

CN121752989APending Publication Date: 2026-03-27DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, the delivery of multi-sensory content is limited by the customized nature of actuators, making it impossible to flexibly extend light experiences and tactile content across different facilities. Furthermore, existing systems only extend screen visual effects through algorithms and cannot achieve cross-device conversion of creative intent.

Method used

By introducing object-based sensory data and metadata, the system receives and processes the sensory data to generate actuator control signals, enabling the creation and rendering of multi-sensory experiences across devices, including effects such as lighting, touch, and airflow.

Benefits of technology

It enables flexible expansion of multi-sensory experiences in different playback environments, allows creative intent to be consistently presented on different actuators, and enhances the flexibility and scalability of lighting fixtures and actuators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121752989A_ABST
    Figure CN121752989A_ABST
Patent Text Reader

Abstract

Some disclosed examples relate to receiving, by a control system, a content bitstream comprising encoded object-based sensory data, the encoded object-based sensory data comprising one or more sensory objects and corresponding sensory metadata, the encoded object-based sensory data corresponding to a sensory effect, comprising lighting, haptic, airflow, one or more position actuators, or a combination thereof, to be provided by one or more sensory actuators in the environment. Some disclosed examples involve extracting, by a control system, object-based sensory metadata from a content bitstream, and providing, by the control system, the object-based sensory metadata to a sensory renderer. In some examples, the content bitstream may also include one or more encoded audio objects and / or encoded video data synchronized with the encoded object-based sensory metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 669,232, filed July 10, 2024; U.S. Provisional Application No. 63 / 514,107, filed July 17, 2023; U.S. Provisional Application No. 63 / 514,096, filed July 17, 2023; and U.S. Provisional Application No. 63 / 514,094, filed July 17, 2023, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure relates to providing multi-sensory experiences, some of which include light-based experiences. Background Technology

[0003] Unless otherwise indicated herein, the methods described in this section are not prior art to the claims of this application and are not acknowledged as prior art by virtue of their inclusion in this section.

[0004] Media content delivery typically focuses on audio and video experiences. The delivery of multi-sensory content has been limited due to the customized nature of actuation. For example, lighting is widely used as an artistic and functional expression of concerts. However, each facility is specifically designed for a particular set of lighting fixtures. Delivering lighting design outside of that set of fixtures targeted by the system design is generally not feasible. Other systems attempting to provide a broader light experience merely extend the screen visuals through algorithms, but are not specifically created. Haptic content is designed for specific haptic devices. If another device (such as a game controller, mobile phone, or even a haptic device from a different brand) is used, the creative intent of the content cannot be translated to a different actuator. Summary of the Invention

[0005] At least some aspects of this disclosure can be implemented via methods such as audio processing methods. In some cases, these methods can be implemented at least in part by a control system such as those disclosed herein. Some methods may involve receiving a content bitstream comprising encoded object-based sensory data by the control system. The encoded object-based sensory data may include one or more sensory objects and corresponding sensory metadata. The encoded object-based sensory data may correspond to sensory effects provided by one or more sensory actuators in the environment. These sensory effects may include lighting, tactile sensation, airflow, one or more position actuators, or combinations thereof. Some methods may involve extracting the object-based sensory data from the content bitstream by the control system and providing the object-based sensory data to a sensory renderer by the control system.

[0006] In some examples, the object-based sensory metadata may include sensory spatial metadata, which at least indicates the spatial location used to render the object-based sensory data within the environment, the region used to render the object-based sensory data within the environment, or a combination thereof. According to some examples, the object-based sensory data may not correspond to a specific sensory actuator in the environment. In some examples, the object-based sensory data may include abstract sensory reproduction information that allows the sensory renderer to reproduce one or more created sensory effects via one or more sensory actuator types, via various numbers of sensory actuators, and from various sensory actuator locations in the environment.

[0007] Some methods may involve the sensory renderer receiving object-based sensory data and environment descriptor data corresponding to the positioning of one or more sensory actuators in the environment. Some methods may involve the sensory renderer receiving actuator descriptor data corresponding to the characteristics of these sensory actuators in the environment. Some methods may involve the sensory renderer providing one or more actuator control signals for controlling these sensory actuators in the environment to produce one or more sensory effects indicated by the object-based sensory data. Some methods may involve the one or more sensory actuators in the environment providing the one or more sensory effects.

[0008] According to some examples, the content bitstream may also include one or more encoded audio objects synchronized with the encoded object-based sensory data. These audio objects may include one or more audio signals and corresponding audio object metadata. Some such methods may involve the control system extracting the audio objects from the content bitstream and providing these audio objects to an audio renderer. In some examples, the audio object metadata may include at least audio object spatial metadata indicating the spatial location of the audio objects used to render the one or more audio signals within the environment. Some methods may involve the audio renderer receiving the one or more audio objects; the audio renderer receiving amplifier data corresponding to one or more amplifiers in the environment; and the audio renderer providing one or more amplifier control signals to control the one or more amplifiers in the environment to play audio corresponding to the one or more audio objects and synchronized with the one or more sensory effects. Some methods may involve the one or more amplifiers in the environment playing audio corresponding to the one or more audio objects.

[0009] According to some examples, the content bitstream may include encoded video data synchronized with the encoded audio object and the encoded object-based sensory data. Some methods may involve the control system extracting video data from the content bitstream and the control system providing the video data to a video renderer. Some methods may involve the video renderer receiving the video data and the video renderer providing one or more video control signals to control one or more display devices in the environment to render one or more images corresponding to the one or more video control signals and synchronized with the one or more audio objects and the one or more sensory effects. Some methods may involve the one or more display devices in the environment rendering the one or more images corresponding to the one or more video control signals.

[0010] In some examples, the environment can be a virtual environment. According to other examples, the environment can be a physical, real-world environment. For example, the environment could be a room environment or a vehicle environment.

[0011] Some or all of the operations, functions, and / or methods described herein can be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory computer-readable media. Such non-transitory media may include one or more memory devices such as those described herein, including but not limited to one or more random access memory (RAM) devices, read-only memory (ROM) devices, etc. Therefore, some innovative aspects of the subject matter described in this disclosure can be implemented in one or more computer-readable non-transitory media on which software is stored.

[0012] At least some aspects of this disclosure can be implemented via apparatus. For example, one or more devices may be able to perform at least partially the methods disclosed herein. In some embodiments, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. The control system may be configured to perform some or all of the disclosed methods.

[0013] Details of one or more embodiments of the subject matter described in this specification are set forth in the following figures and description. Other features, aspects, and advantages will become apparent from the description, figures, and claims. Note that the relative dimensions in the following figures may not be drawn to scale. Attached Figure Description

[0014] The disclosed embodiments will now be described by way of example only with reference to the accompanying drawings.

[0015] Figure 1 This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure.

[0016] Figure 2 Example elements of the endpoint are shown.

[0017] Figure 3 An example of an actuator element is shown.

[0018] Figure 4 Example components of a system for creating and playing multi-sensory (MS) experiences are shown.

[0019] Figure 5 Example components of another system for creating and playing MS experiences are shown.

[0020] Figure 6 It shows that it can be made by Figure 5 An example of a graphical user interface (GUI) presented on a display device for a scene creation tool.

[0021] Figure 7A It shows that it can be made by Figure 5 Another example of a graphical user interface (GUI) presented on a display device for a scene creation tool.

[0022] Figure 7B It is a flowchart outlining an example of a method that can be performed by an apparatus or system such as the apparatus or system disclosed herein.

[0023] Figure 8A , Figure 8B and Figure 8C Three examples of projecting the lighting of the viewing environment onto a two-dimensional (2D) plane are shown.

[0024] Figure 9A Example elements of the light and shadow renderer are shown.

[0025] Figure 9B It is a flowchart outlining an example of a method that can be performed by an apparatus or system such as the apparatus or system disclosed herein.

[0026] Figure 10 This shows another example of a GUI that can be rendered by a display device of a scene creation tool.

[0027] Figure 11 It is a flowchart outlining an example of a method that can be performed by an apparatus or system such as the apparatus or system disclosed herein.

[0028] Figure 12 Example components of a system for creating and playing multi-sensory (MS) experiences are shown.

[0029] Figure 13 It is a flowchart outlining an example of a method that can be performed by an apparatus or system such as the apparatus or system disclosed herein. Detailed Implementation

[0030] This document describes techniques related to providing multi-sensory media content. In the following description, numerous examples and specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that this disclosure, as defined by the claims, may include some or all of these examples, either alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.

[0031] The following description details various methods, processes, and procedures. While specific steps may be described in a particular order, this order is primarily for convenience and clarity. A particular step may be performed more than once, may occur before or after other steps (even if these steps are described in a different order), and may occur in parallel with other steps. A second step is only necessary if the first step must be completed before the second step can begin. This will be specifically indicated when it is unclear from the context.

[0032] In this document, the terms “and,” “or,” and “and / or” are used. These terms should be understood to have inclusive meanings. For example, “A and B” can at least mean: “both A and B,” or “at least both A and B.” As another example, “A or B” can at least mean: “at least A,” “at least B,” “both A and B,” or “at least both A and B.” As yet another example, “A and / or B” can at least mean: “A and B,” or “A or B.” When XOR is intended to be used, it will be specifically indicated (e.g., “either A or B,” or “at most one of A and B”).

[0033] This document describes the various processing functions associated with structures such as blocks, elements, components, and circuits. Typically, these structures can be implemented by one or more processors controlled by one or more computer programs.

[0034] As mentioned above, media content delivery typically focuses on audio and video experiences. Due to the personalized nature of actuation, the delivery of multi-sensory (MS) content has always been limited.

[0035] This application describes methods for expanding the creative palette of content creators, allowing for the creation and delivery of spatial MS experiences at scale. Some of these methods involve introducing new layers of abstraction to allow the delivery of created MS experiences to different endpoints using different types of lighting fixtures or actuators. As used herein, the term "endpoint" is synonymous with "playback environment" or simply "environment," meaning an environment that includes one or more actuators that can be used to deliver the MS experience. Such endpoints can include rooms (such as the living room in a home), cars, movie theaters, nightclubs, or other venues. Some of the disclosed methods involve creating, delivering, and / or rendering object-based sensory data, which can include sensory objects and corresponding sensory metadata. This abstraction enables creative intent to be implemented in an object-based format without prior knowledge of specific controller actuations, thereby achieving greater flexibility and scalability across endpoints for lighting fixtures and actuators. In this document, the MS experience delivered via object-based sensory data may be referred to as a "flexibly extended MS experience."

[0036] acronym MS - Multisensory MSIE – MS Immersive Experience AR - Augmented Reality VR - Virtual Reality PC — Personal Computer Figure 1 This is a block diagram illustrating examples of components of an apparatus capable of implementing various aspects of this disclosure. (Compared to other methods provided herein...) Figure 1 Sample, Figure 1 The types and quantities of elements shown are provided by way of example only. Other embodiments may include more, fewer, and / or different types and quantities of elements. According to some examples, device 101 may be or may include a device configured to perform at least some of the methods disclosed herein, such as a smart audio device, laptop computer, cellular phone, tablet device, smart home hub, etc. In some such embodiments, device 101 may be or may include a server configured to perform at least some of the methods disclosed herein.

[0037] In this example, apparatus 101 includes at least an interface system 105 and a control system 110. In some embodiments, the control system 110 may be configured to perform at least partially the methods disclosed herein. In some embodiments, the control system 110 may be configured to receive a content bitstream via the interface system 105, the content bitstream including encoded object-based sensory metadata. This encoded object-based sensory metadata may correspond to sensory effects such as lighting, touch, airflow, one or more position actuators, or combinations thereof, provided by multiple sensory actuators in the environment. In some embodiments, the control system 110 may be configured to extract object-based sensory metadata from the content bitstream and to provide the object-based sensory metadata to a sensory renderer.

[0038] According to some examples, object-based sensory metadata may include sensory spatial metadata, which at least indicates the spatial location used to render the object-based sensory metadata within the environment, the region used to render the object-based sensory metadata within the environment, or a combination thereof. In some implementations, the object-based sensory metadata does not correspond to any specific sensory actuator in the environment. In some examples, object-based sensory metadata may include abstract sensory reproduction information, thereby allowing a sensory renderer to reproduce created sensory effects via various sensory actuator types, via various numbers of sensory actuators, and from various sensory actuator locations in the environment; these sensory effects may also be referred to herein as intended sensory effects.

[0039] In some examples, the content bitstream may also include encoded audio objects synchronized with encoded object-based sensory metadata. The audio object may include an audio signal and corresponding audio object metadata. In some such implementations, the control system 110 may be configured to extract audio objects from the content bitstream and provide the audio objects to an audio renderer. According to some examples, the audio object may include an audio signal and corresponding audio object metadata. The audio object metadata may at least include audio object spatial metadata indicating the spatial location of the audio object used to render the audio signal within the environment.

[0040] Interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more Universal Serial Bus (USB) interfaces). According to some implementations, interface system 105 may include one or more wireless interfaces. Interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system, and / or a gesture sensor system. In some examples, interface system 105 may include a control system 110 and a memory system (such as...) Figure 1 One or more interfaces between the optional memory system 115 shown. However, in some cases, the control system 110 may include a memory system.

[0041] For example, the control system 110 may include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.

[0042] In some implementations, the control system 110 may reside in more than one device. For example, a portion of the control system 110 may reside in a device within the environment (such as a laptop computer, tablet computer, smart audio device, etc.), and another portion of the control system 110 may reside in a device outside the environment (such as a server). In other examples, a portion of the control system 110 may reside in a device within the environment, and another portion of the control system 110 may reside in one or more other devices within the environment.

[0043] Some or all of the methods described herein can be executed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non-transitory media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. One or more non-transitory media may, for example, reside on... Figure 1 In the optional memory system 115 and / or control system 110 shown. Therefore, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media on which software is stored. For example, the software may include instructions for controlling at least one device to process audio data. For example, the software may be provided by, for example, Figure 1 The control system 110 and other control system components perform the operation.

[0044] In some examples, device 101 may include Figure 1The optional microphone system 120 shown is included. The optional microphone system 120 may include one or more microphones. In some embodiments, the one or more microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc.

[0045] According to some embodiments, device 101 may include Figure 1 The optional actuator system 125 is shown herein. The optional actuator system 125 may include one or more loudspeakers, one or more haptic devices, one or more luminaires (also referred to herein as lighting fixtures), one or more fans or other airflow devices, one or more display devices (including, but not limited to, one or more televisions), one or more position actuators, one or more other types of devices for providing an MS experience, or combinations thereof. As used herein, the term “light fixture” generally refers to various types of light sources, including individual light sources, light source groups, light strips, etc. A “light fixture” can be movable, so in this context, the word “fixture” does not mean that the fixture must be in a fixed position in space. As used herein, the term “position actuator” generally refers to a device configured to change the position or orientation of a person or object, such as a motion simulator seat. A loudspeaker may sometimes be referred to herein as a “speaker.” In some embodiments, the optional actuator system 125 may include a display system comprising one or more displays, such as one or more light-emitting diode (LED) displays, one or more organic light-emitting diode (OLED) displays, etc. In some examples where device 101 includes a display system, optional sensor system 130 may include a touch sensor system and / or gesture sensor system proximate to one or more displays of the display system. According to some such embodiments, control system 110 may be configured to control the display system to render a graphical user interface (GUI), such as a GUI associated with implementing one of the methods disclosed herein.

[0046] In some embodiments, device 101 may include Figure 1 The optional sensor system 130 shown may include a touch sensor system, a gesture sensor system, one or more cameras, etc.

[0047] This application describes a method for creating flexible, expandable multi-sensory (MS) immersive experiences (MSIEs) and delivering them to different playback environments (which may also be referred to herein as endpoints). Such endpoints may include rooms (such as living rooms in a home), cars, cinemas, nightclubs or other venues, AR / VR headsets, PCs, mobile devices, etc.

[0048] Figure 2 Example elements of an endpoint are shown. In this example, the endpoint is a living room 1001, which contains multiple actuators 008, some furniture 1010, and a person 1000 (also referred to herein as a user) who will flexibly expand the MS experience. Actuators 008 are devices capable of altering the environment 1001 in which the user 1000 is located. Actuators 008 may include one or more televisions or other display devices, one or more lighting fixtures (also referred to herein as lamps), one or more loudspeakers, etc.

[0049] The number, arrangement, and capability of the actuators 008 in space 1001 can vary significantly between different endpoint types. For example, the number, arrangement, and capability of actuators 008 in a car typically differ from those in a living room, nightclub, etc. In many embodiments, the number, arrangement, and / or capability of actuators 008 can also vary significantly between different instances of the same type (e.g., between a small living room with 2 actuators 008 and a large living room with 16 actuators 008). This disclosure describes various methods for creating flexible, expandable MSIEs and extending them to these heterogeneous endpoints.

[0050] Figure 3 An example of actuator elements is shown. In this example, the actuator is an illuminator 1100, which includes a network module 1101, a control module 1102, and a light emitter 1103. According to this example, the light emitter 1103 includes one or more light-emitting devices, such as light-emitting diodes, configured to emit light into the environment where the illuminator 1100 resides. In this example, the network module 1101 is configured to provide network connectivity to one or more other devices in space, such as devices that send commands to control the illuminator 1100 to emit light. According to this example, the control module 1102 is configured to receive signals via the network module 1101 and control the light emitter 1103 accordingly.

[0051] Other examples of actuators may include network module 1101 and control module 1102, but may include other types of actuation elements. Some such actuators may include one or more loudspeakers, one or more haptic devices, one or more fans or other airflow devices, one or more position actuators, one or more display devices, etc.

[0052] Figure 4 Example components of a system for creating and playing multisensory (MS) experiences are shown. (This is in contrast to other systems provided in this article.) Figure 1 Sample, Figure 4The types and quantities of elements shown are provided by way of example only. Other implementations may include more, fewer, and / or different types and quantities of elements. According to some examples, system 300 may be or may include one or more devices configured to perform at least some of the methods disclosed herein. In some examples, system 300 may include devices configured to perform at least some of the methods disclosed herein. Figure 1 One or more instances of the control system 110.

[0053] According to the examples in this disclosure, the method for creating and delivering an object-based MS Immersive Experience (MSIE) involves applying a set of techniques for creating, delivering, and rendering object-based sensory data (which may include sensory objects and corresponding sensory metadata) to actuator 008. Some examples are described in the following paragraphs.

[0054] Object-based representation: In various disclosed implementations, multi-sensory (MS) effects are represented using content that can be referred to herein as “sensory objects.” According to some such implementations, characteristics such as layer type and priority can be assigned to, associated with, and attached to each sensory object, thereby enabling the content creator’s intent to be represented in the rendered experience. Detailed examples of sensory object characteristics are described below.

[0055] In this example, system 300 includes a content creation tool 000 configured to design multi-sensory (MS) immersive content and to output object-based sensory data 005, individually or in combination with corresponding audio data 011 and / or video data 012, depending on a specific implementation. The object-based sensory data 005 may include timestamp information and information indicating the type, characteristics, etc., of the sensory object. In this example, the object-based sensory data 005 is not “channel-based” data corresponding to one or more specific sensory actuators in the playback environment, but rather generalized to multiple playback environments with multiple actuator types, multiple actuators, etc. In some examples, the object-based sensory data 005 may include object-based light data, object-based tactile data, object-based airflow data, or object-based position actuator data, object-based olfactory data, object-based smoke data, object-based data for one or more other types of sensor effects, or combinations thereof. According to some examples, the object-based sensory data 005 may include sensory objects and corresponding sensory metadata. For example, if object-based sensory data 005 includes object-based light data, then the object-based light data may include light object location metadata, light object color metadata, light object size metadata, light object intensity metadata, light object shape metadata, light object diffusion metadata, light object gradation metadata, light object priority metadata, light object layer metadata, or a combination thereof. Although in this example, content creation tool 000 is shown as providing a stream of object-based sensory data 005 to experience player 002, in alternative examples, content creation tool 000 may generate object-based sensory data 005 that is stored for later use. An example of a graphical user interface for a light object-based content creation tool is described below.

[0056] MS Object Renderer: Various disclosed embodiments provide a renderer configured to render MS effects to actuators in a playback environment. According to this example, system 300 includes an MS renderer 001 configured to render object-based sensory data 005 into actuator control signals 310, at least in part based on environment and actuator data 004. In this example, MS renderer 001 is configured to output the actuator control signals 310 to MS controllers 003 configured to control actuators 008. In some examples, MS renderer 001 may be configured to receive light objects and object-based lighting metadata indicating a desired lighting environment, as well as lighting information about the local lighting environment. The lighting information is a generic type of environment and actuator data 004 and may include one or more characteristics of one or more controllable light sources in the local lighting environment. In some examples, MS renderer 001 may be configured to determine the approximate drive level of each of the one or more controllable light sources in relation to the desired lighting environment. Some alternative examples may include separate renderers for each type of actuator 008, such as one renderer for a luminaire, another for a haptic device, another for an airflow device, and so on. According to some examples, the MS renderer 001 (or one of the MS controllers 003) may be configured to output drive levels to at least one of the controllable light sources. In some implementations, the MS renderer 001 may be configured to adapt to changing conditions. Some examples of implementations of the MS renderer 001 are described in more detail below.

[0057] The environment and actuator data 004 may include content referred to herein as a “room descriptor” that describes the actuator’s location (e.g., according to an x, y, z coordinate system or a spherical coordinate system). In some examples, the environment and actuator data 004 may indicate actuator orientation and / or placement characteristics (e.g., oriented and north-facing, omnidirectional, occlusion information, etc.). According to some examples, the environment and actuator data 004 may indicate actuator orientation and / or placement characteristics according to a 3 × 3 matrix, in which three elements (e.g., elements in the first row) represent spatial location (x, y, z), three other elements (e.g., elements in the second row) represent orientation (roll, pitch, yaw), and three other elements (e.g., elements in the third row) indicate scale or size (sx, sy, sz). In some examples, the environment and actuator data 004 may include a device descriptor describing actuator characteristics associated with the MS renderer 001, such as the color gamut and intensity range of a luminaire, the airflow velocity range and (multiple) directions of an airflow device, etc.

[0058] In this example, system 300 includes an experience player 002 configured to receive object-based sensory data 005', audio data 011', and video data 012', provide the object-based sensory data 005' to an MS renderer 001, provide the audio data 011' to an audio renderer 006, and provide the video data 012' to a video renderer 007. In this example, the reference numerals for the object-based sensory data 005', audio data 011', and video data 012' received by the experience player 002 include an apostrophe (') to indicate that the data may be encoded in some cases. Similarly, the object-based sensory data 005', audio data 011', and video data 012' output by the experience player 002 do not include an apostrophe to indicate that the data may have been decoded by the experience player 002 in some cases. According to some examples, the experience player 002 may be a media player, game engine, or a component integrated into a television, DVD player, soundbar, set-top box, or service provider media device (such as Chromecast, Apple TV, or Amazon FireTV). In some examples, the experience player 002 may be configured to receive encoded object-based sensory data 005, as well as encoded audio data 011 and / or encoded video data 012. In some such examples, the encoded object-based sensory data 005' may be received as part of the same bitstream as the encoded audio data 011' and / or encoded video data 012'. Some examples are described in more detail below. According to some examples, the experience player 002 may be configured to extract object-based sensory data 005' from the content bitstream and provide decoded object-based sensory data 005 to an MS renderer 001, decoded audio data 011 to an audio renderer 006, and decoded video data 012 to a video renderer 007. In some examples, the timestamp information in the object-based sensory data 005' may be used, for example, by the experience player 102, MS renderer 001, audio renderer 106, video renderer 107, or all of them, to synchronize the effects associated with the object-based sensory data 005' with the audio data 111' and / or the video data 112', which may also include timestamp information.

[0059] According to this example, system 300 includes MS controllers 003 configured to communicate with various actuator types using an application programming interface (API) or one or more similar interfaces. Generally, each actuator will require a specific type of control signal to produce the desired output from the renderer. According to this example, MS controller 003 is configured to map the output from MS renderer 001 to the control signals for each actuator. For example, a Philips Hue™ bulb receives control information in a specific format to turn on a lamp with a digital representation of a specific saturation, brightness, and hue, as well as a desired drive level.

[0060] In some examples, the room descriptor can also describe the size and orientation of the playback environment itself to establish a relative or absolute coordinate system to which all objects are positioned. For example, in a living room, the display screen can be considered as the front, or in some cases, as the center of the front, and the floor and ceiling can be considered as vertical boundaries. In some such examples, the room descriptor can also indicate the boundaries corresponding to the left, right, front, and back walls relative to the front position. According to some examples, the room descriptor can also be provided in the form of a matrix (such as a 3 × 3 matrix). This room descriptor information is used to describe the physical dimensions of the playback environment (e.g., expressed in physical distance units such as meters). In some such examples, sensory object positioning, sensory object size, and sensory object orientation can be described in units relative to the room size, such as in the range of -1 to 1. In some cases, the room descriptor can also describe the preferred viewing position based on the matrix.

[0061] The type, number, and arrangement of actuators 008 will generally vary depending on the specific implementation. In some examples, actuators 008 may include lamps and / or light strips (also referred to herein as “illuminators”), vibration motors, airflow generators, position actuators, or combinations thereof.

[0062] Similarly, the type, quantity, and arrangement of the display device 010 and the loudspeaker 009 will generally vary depending on the specific implementation. Figure 4 In the example shown, audio data 011 and video data 012 are rendered by audio renderer 006 and video renderer 007 to loudspeaker 009 and display device 010, respectively.

[0063] As described above, according to some embodiments, system 300 may include methods configured to perform at least some of the methods disclosed herein. Figure 1The control system 110 may be one or more instances. In some such examples, one instance of the control system 110 may implement a content creation tool 000, and another instance of the control system 110 may implement an experience player 002. In some examples, one instance of the control system 110 may implement an audio renderer 006, a video renderer 007, a multi-sensory renderer 001, or a combination thereof. According to some examples, an instance of the control system 110 configured to implement the experience player 002 may also be configured to implement an audio renderer 006, a video renderer 007, a multi-sensory renderer 001, or a combination thereof.

[0064] Multi-sensory rendering synchronization Object-based MS rendering involves flexibly rendering different modalities to endpoints / playback environments. Endpoints have different capabilities depending on various factors, including but not limited to the following: • Number of actuators • Modalities of these actuators (e.g., lighting and airflow control devices and tactile devices); • The types of these actuators (e.g., white smart lights versus RGB smart lights, or haptic vests versus haptic cushions) and • The positioning / layout of these actuators.

[0065] To render object-based sensory content to any endpoint, some processing of the object signals (e.g., intensity, color, pattern, etc.) is typically required. The processing of the signal path for each modality should not alter the relative phase of any feature within the object signal. For example, suppose lightning is presented in both tactile and visual modalities. The signal processing chain corresponding to the actuator control signals should not cause any type of sensory object signal (tactile or visual) to introduce a time delay sufficient to alter the perceptual synchronicity between the two modalities. The required level of synchronization may depend on various factors, such as whether the experience is interactive and what other modalities are involved. Depending on the specific context, the maximum time difference can range, for example, from approximately 10 ms to 100 ms.

[0066] Touch Rendering of object-based haptic content Object-based haptic content conveys the sensory aspects of a scene through abstract sensory representations rather than channel-based schemes. For example, instead of simply defining haptic content as a single-channel time-correlated amplitude signal, which is then played from a specific haptic actuator (such as a vibrating haptic motor) on a vest worn by the user, object-based haptic content can be defined by the sensation it is intended to convey. More specifically, in one example, there could be a haptic object representing the sensory effect of impact. Associated with this object are: • Spatial positioning of tactile objects; • Spatial direction / vector of tactile effect; • The intensity of the tactile effect; • Tactile spatial and temporal frequency data; and • Time-dependent amplitude signal.

[0067] Based on some examples, this type of haptic object can be automatically created in interactive experiences such as video games, for example, in a racing game when another car crashes into the player's car from behind. In this example, the MS renderer will determine how to render the spatial modality of this effect to a set of haptic actuators in the endpoints. In some examples, the renderer does this based on information about: • Various types of tactile devices are available, such as tactile vests and tactile gloves, tactile cushions and tactile controllers; • The location of each haptic device relative to (multiple) users (some haptic devices may not be coupled to (multiple) users, for example, a vibrator mounted on the floor or a seat); • Each haptic device offers a type of actuation, such as kinematic and vibratory haptic feedback; • Start-up and stop delay for each haptic device (in other words, the speed at which each haptic device can be turned on and off); • The dynamic response of each haptic device (how much the amplitude can change); • The time-frequency response of each haptic device (the time-frequency response that a haptic device can provide); • Spatial distribution of addressable actuators within each haptic device: For example, a haptic vest could have dozens of addressable haptic actuators distributed across the user's torso; and • The time response of any haptic sensor used to render closed-loop haptic effects (e.g., active force feedback kinetic haptic devices).

[0068] These properties of the haptic modality of the endpoints inform the renderer how best to render a particular haptic effect. Consider again the car crash effect example. In this example, the player wears a haptic vest, haptic armbands, and haptic gloves. According to this example, the haptic shockwave effect is spatially located at the point where the car hits the player. The shockwave vector is determined by the relative velocity of the player's car and the car that hits the player. The spatial and temporal spectrum of the shockwave effect is created based on the type of materials the virtual car is expected to use, as well as other virtual world characteristics. The renderer then renders the shockwave using a set of haptic devices at the endpoints, based on the shockwave vector and the physical location of the haptic devices relative to the user.

[0069] The signals sent to each specific actuator are preferably provided in such a way that the sensory effect is consistent across all available (potentially heterogeneous) actuators. For example, due to the lack of capability of other actuators, the renderer may not render very high frequencies only to one of the haptic actuators (e.g., a haptic armband). Otherwise, when a shockwave moves through the player's body, the haptic effect perceived by the user will decrease as the wave moves through the vest, into the armband, and finally into the gloves because the haptic vest and gloves worn by the user are not capable of rendering such high frequencies.

[0070] Some types of abstract tactile effects include: • Shockwave effect, as described above; • Barrier effects, such as haptic effects used to represent spatial constraints in virtual worlds (e.g., in video games). If an active or resistive kinetic actuator is present on the input device (e.g., force feedback on a steering wheel or joystick), this effect can be rendered by applying resistance to the user's input. If no such actuator is available at the endpoint, in some examples, vibratory haptic feedback consistent with the collision of an in-game avatar with a barrier can be rendered; • Presence, such as indicating the presence of a large object (like a train) approaching a scene. This type of haptic effect can be rendered using the low-frequency rumble of certain haptic device actuators. It can also be rendered using pressure applied by an air bladder to provide tactile spatial feedback. • User interface feedback, such as a click from a virtual button. For example, this type of haptic effect can be rendered to the nearest actuator on the user's body that performed the click, such as a haptic glove worn by the user. Alternatively or additionally, this type of haptic effect can also be rendered to a vibrator coupled to the chair the user is sitting on. This type of haptic effect can be defined, for example, using a time-dependent amplitude signal. However, this signal can be modified (modulated, frequency-shifted, etc.) to best suit the haptic device(s) that will provide the haptic effect; • Motion Sensation. These haptic effects are designed to make the user perceive some form of motion. These haptic effects can be rendered by actuators on an actual moving user (e.g., a moving platform / seat). In some examples, the actuators can provide auxiliary modalities (e.g., via video) to enhance the motion being rendered; and • Trigger Sequence. These haptic effects are primarily characterized by their time-dependent amplitude signals. This signal can be rendered across multiple actuators and, in doing so, can be amplified. This amplification can include splitting the signal across multiple actuators in time or frequency. Some examples may involve amplifying the signal itself so that the sum of the haptic actuator outputs does not match the original signal.

[0071] Spatial effects and non-spatial effects Spatial effects are spatial effects constructed in a way that conveys certain spatial information about the multi-sensory scene being rendered. For example, if the playback environment is a room, then a shockwave moving through the room will be rendered differently to each haptic device, depending on the position and size of one or more haptic objects being rendered at a particular time, based on the location of each haptic device within the room.

[0072] In some examples, non-spatial effects can be targeted at specific parts of a user's body, regardless of the user's position or orientation. One example is a haptic device providing enhanced vibrations to a user's back to indicate immediate danger. Another example is a haptic device providing strong vibrations to indicate injury to a specific area of ​​the body.

[0073] Some effects can be non-narrative effects. These effects are often associated with user interface feedback, such as tactile feedback used to indicate that a user has completed a level or clicked a button on a menu item. Non-narrative effects can be spatial or non-spatial.

[0074] Types of tactile devices Receiving information about the different types of haptic devices available at the endpoints allows the renderer to determine which types of sensory effects and rendering strategies it can use. For example, local haptic device data instructing the user to wear both a haptic glove and a vibrating haptic vest (or at least local haptic device data indicating the presence of the haptic glove and vibrating haptic vest in the playback environment) allows the renderer to render a consistent recoil effect on both devices when the user fires a gun in the virtual world. The actual actuator control signals sent to the haptic devices may differ from cases where only a single device is available. For example, if the user is only wearing the vest, the actuator control signals used to actuate the vest may differ in terms of the actuator control signal's activation timing, maximum amplitude, frequency, and decay time, or combinations thereof.

[0075] Equipment positioning Understanding the location of haptic devices at endpoints allows renderers to consistently render spatial effects. For example, knowing the location of vibrating motors in a lounge allows the renderer to send actuator control signals to each vibrating motor in the lounge in a way that conveys spatial effects, such as shock waves propagating within the room. Additionally, although the location of wearable haptic devices is implied by their type (e.g., gloves on a user's hand), the renderer can also use knowledge of the location of these wearable haptic devices to convey both spatial and non-spatial effects.

[0076] Types of actuation provided by tactile devices Haptic devices can provide a range of different actuations, and thus provide a perceived sensation. These are generally divided into two basic categories: 1. Vibrational tactile sensation, such as vibration; or 2. Kinesthetic feedback, such as resistance feedback or propulsion feedback.

[0077] Actuation of any type can be static or dynamic, with dynamic effects changing in real time based on some sensor inputs. Examples include touchscreens that use vibration haptic actuators to render textures and position sensors that measure the position of the user's (multiple) fingers(s).

[0078] Furthermore, the physical construction of these actuators varies considerably and affects many other properties of the device. An example of this is the significant differences in activation delay or time-frequency response across the following types of haptic devices: •Eccentric rotating mass; • Linear resonant actuator; • Piezoelectric actuators; and • Linear magnetic ram.

[0079] The renderer should be configured to take into account the startup latency of a specific haptic device type when rendering signals that will be actuated by a haptic device in an endpoint.

[0080] Start-up and shutdown delay of tactile devices The start-up delay of a haptic device refers to the delay between the time when the actuator control signal is sent to the device and the time when the device physically responds. The stop delay refers to the delay between the time when the actuator control signal is sent to bring the device's output to zero and the time when the device stops actuating.

[0081] Time-frequency response Time-frequency response refers to the frequency range in which a haptic device can be actuated in a steady state, with the signal amplitude as a function of time.

[0082] Spatial frequency response Spatial frequency response refers to the frequency range in which the signal amplitude is a function of the spacing between the actuators of a tactile device. Devices with closely spaced actuators have a higher spatial frequency response.

[0083] Dynamic range Dynamic range refers to the difference between the minimum and maximum amplitude of a physical actuation.

[0084] Characteristics of sensors in closed-loop haptic devices Some dynamic effects use sensors to update actuation signals based on certain observed states. The sampling frequency of time and space, as well as noise characteristics, will limit the ability to update the control loop of the actuator providing the dynamic effect.

[0085] airflow Another modality that some multi-sensory immersive experiences (MSIEs) can use is airflow. Airflow can be rendered consistently with one or more other modalities, such as audio, video, lighting effects, and / or haptics. Unlike dedicated (e.g., channel-based) setups designed solely for 4D experiences in theaters (which may include "wind effects"), some airflow effects can also be provided at other endpoints that typically include airflow, such as a car or living room. Unlike channel-based systems, airflow sensory effects can be represented as airflow objects, which may include the following characteristics: • Spatial positioning; • The direction of the expected airflow effect; • Intensity / airflow velocity; and / or • Air temperature.

[0086] Some examples of airflow objects can be used to represent the movement of a bird flying by. To render the airflow actuator at the endpoint, information about the following can be provided to MS Renderer 001: • Types of airflow equipment, such as fans, air conditioners, and heaters; • The position of each airflow device relative to the user's location or the user's expected location; • The capabilities of airflow devices, such as their ability to control direction, airflow, and temperature; • The control level for each actuator, such as airflow speed and temperature range; and • The response time of each actuator, for example, how long it takes to reach a selected speed.

[0087] Examples of airflow usage at different endpoints In vehicles such as cars, object-based metadata can be used to create experiences such as: • During scenes in horror movies or games, mimic the feeling of "chills down your spine" by using airflow down a chair; • Simulate the movement of a bird flying by; and / or • Create a gentle breeze in the sea view.

[0088] Within the small, enclosed space of a typical vehicle, temperature changes can likely be achieved over a relatively short period compared to temperature variations in a larger environment, such as a living room. In one example, MS Render 001 could raise the air temperature when a player enters a "lava level" or other hot area during gameplay. Some examples could include additional elements, such as confetti in the vents, to celebrate an event, such as a goal scored by a user's favorite football team.

[0089] In one example, airflow in a living space or other room can be synchronized with the breathing rhythm of guided meditation. In another example, airflow can be synchronized with the intensity of exercise, increasing airflow or decreasing temperature as intensity increases. In some examples, spatial control during rendering may be relatively limited. For instance, many existing airflow actuators are optimized for heating and / or air conditioning, rather than for providing sensory actuation that offers spatial diversity.

[0090] A combination of light, airflow, and touch Car example The following examples are described with reference to automobiles, but are applicable to other vehicles such as trucks, vans, etc. In some examples, a user interface may be present on the steering wheel or on a touchscreen near or within the dashboard. According to some examples, the following actuators may be present in a car: 1. Individually addressable lights are spatially distributed throughout the vehicle in the following manner: o on the dashboard; o is below the footrest space; o on the door; and o is located within the central control console.

[0091] 2. Individually controllable air conditioning / heating vents are distributed throughout the vehicle as follows: o is in the front dashboard; o is below the footrest space; o is located in the center console, facing the rear seats; o is on the side pillar; o in the seat; and o. Guide windshield (for defogging).

[0092] 3. A individually controllable seat with vibrating tactile feedback; and 4. Individually controllable floor mats with vibrating tactile feedback.

[0093] In this example, the modes supported by these actuators include the following: • Lights with individually addressable LEDs throughout the car, as well as indicator lights on the dashboard and steering wheel; • Airflow via controlled air conditioning vents; •Touch, including: o Steering wheel: Tactile vibration feedback; o Dashboard touchscreen: haptic feedback and texture rendering; and o Seat: Touch, vibration, and movement.

[0094] In one example, a live music stream is rendered to four users seated in the front row. In this example, MS Renderer 001 attempts to optimize the experience for multiple viewing positions. During construction, before the artists take the stage and after the previous performances have concluded, the content includes: • Interlude music; •Low-intensity lighting; and • Represents the tactile content of a crowd colliding.

[0095] In addition to the rendered audio and video streams, light content also includes ambient light objects that move slowly within the scene. These ambient light objects can be rendered using one of the environment layer methods disclosed herein, for example, by not assigning spatial priority to any user's viewpoint. In some examples, haptic content can be spatially focused in a lower temporal frequency spectrum and can be rendered solely by vibrating haptic motors in the mat.

[0096] Based on this example, a fireworks event during a music stream corresponds to multi-sensory content including the following: • The light object that spatially corresponds to the location of the fireworks in the event; and • A tactile object that enhances the dynamic feel of fireworks through shockwave effects.

[0097] In this example, MS Renderer 001 renders both light and tactile objects spatially. For instance, light objects can be rendered in a car so that if the fireworks content is on the left side of the scene, everyone in the car will perceive the light object as coming from the left. In this example, only the lights on the left side of the car are actuated. Tactile objects can be rendered on both seats and floor mats in a way that conveys directionality to each user separately.

[0098] At the end of the concert, fireworks will appear in the audio content, and fireworks and confetti will also appear in the video content. In addition to rendering the light and tactile objects corresponding to the fireworks as described above, airflow modalities can be used to render the effect of confetti spray. For example, individually controllable airflow vents in an HVAC system can be pulsed.

[0099] Living room example In this embodiment, in addition to an audio / visual (AV) system including multiple loudspeakers and a television, the following actuators and related controls are available in the living room: • A haptic vest worn by the user (also known as the player); • A haptic vibrator installed on the seat where the player is sitting; • (Haptic) controllable smartwatch; • Smart lights distributed throughout the room; • Wireless controller; and • Addressable airflow bar (AFB), which includes an array of individually controllable fans directed at the user (similar to HVAC vents in a car's dashboard).

[0100] In this example, the user is playing a first-person shooter game that includes a scene where a destructive hurricane moves through the level. As the hurricane moves, in-game objects are thrown around, and some objects hit the player. Haptic objects rendered by MS Renderer 001 enable the delivery of shockwave effects across all haptic devices the user can perceive. The actuator control signals sent to each device can be optimized based on the impact intensity of the in-game objects, the direction(s) of the impact, and the capabilities and positioning (as previously described) of each actuator.

[0101] Before the user is hit by an in-game object, the multisensory content includes tactile objects corresponding to non-spatial rumble, one or more airflow objects corresponding to directional airflow, and one or more light objects corresponding to lightning. The MS renderer 001 renders the non-spatial rumble to the haptic devices. The actuator control signals sent to each haptic device can be rendered such that the set of actuator control signals across the entire haptic array is consistent in the timing, intensity, and frequency of the perceived rumble. In some examples, the frequency content of the actuator control signals sent to the smartwatch can be low-pass filtered to match the limited frequency capabilities of the vest near the watch. The MS renderer 001 can render one or more airflow objects as actuator control signals for AFB, such that the airflow in the room is consistent with the player's position and line of sight in the game, as well as the direction of the hurricane itself. Lightning can be rendered in all modalities as (1) a white flash produced by a light source located in a suitable position (e.g., in or on the ceiling); and (2) a pulsed rumble in the user's wearable haptic and seat vibrator.

[0102] When a user is hit by an in-game object, a directional shockwave can be rendered to the haptic device. In some examples, a corresponding airflow pulse can be rendered. According to some examples, a damage absorption effect can be rendered by a light, indicating the amount of damage the player takes from being hit by an in-game object.

[0103] In some such examples, the signal can be spatially rendered to the haptic devices, causing the perceived shockwave to move across the player's body and within the room. The MS Renderer 001 can provide this effect based on actuator positioning information that indicates the haptic devices' positioning relative to each other. In addition to actuator capability information, the MS Renderer 001 can also provide the shockwave vector and position based on the actuator positioning information. According to some examples, non-directional airflow pulses can be rendered; for example, all AFB vents can be briefly enlarged to enhance the haptic modality. In some examples, a red vignette can be rendered simultaneously onto the light strip around the TV to indicate to the player that they have taken damage in the game.

[0104] Figure 5 This shows example components of another system used for creating and playing MS experiences. (Compared to other systems provided in this article...) Figure 1 Sample, Figure 5 The types and quantities of elements shown are provided by way of example only. Other implementations may include more, fewer, and / or different types and quantities of elements. According to some examples, system 500 may be or may include one or more devices configured to perform at least some of the methods disclosed herein. In some examples, system 500 may include devices configured to perform at least some of the methods disclosed herein. Figure 1 One or more instances of the control system 110.

[0105] Based on this example, Figure 5 The system shown is Figure 4 An example of the system shown. In this example, Figure 5 The system shown is an example of "Light and Shadow," in which visual (video), audio, and lighting effects are combined to create an MS experience.

[0106] In this example, system 500 includes a scene creation tool 100, which is a reference... Figure 4 An example of the described content creation tool 000. The lighting creation tool 100 is configured to design and output object-based lighting data 505', which, depending on a specific implementation, is output individually or in combination with corresponding audio data 111' and / or video data 112'. The object-based lighting data 505' may include timestamp information and information indicating the characteristics of the lighting object, etc. In some cases, timestamp information may be used to synchronize effects associated with the object-based lighting data 505' with the audio data 111' and / or video data 112', where timestamp information may also be included.

[0107] In this example, object-based light data 505' includes light objects and corresponding light metadata. For example, object-based light data may include light object location metadata, light object color metadata, light object size metadata, light object intensity metadata, light object shape metadata, light object diffusion metadata, light object gradation metadata, light object priority metadata, light object layer metadata, or a combination thereof. Although in this example, content creation tool 100 is shown as providing a stream of object-based light data 505' to experience player 102, in alternative examples, content creation tool 100 may generate object-based light data 505' that is stored for later use. An example of a graphical user interface for a light object-based content creation tool is described below.

[0108] In this example, system 500 includes experience player 102, which is configured to receive object-based light data 505', audio data 111', and video data 112', provide the object-based light data 505 to a lighting renderer 501, provide the audio data 111 to an audio renderer 106, and provide the video data 112 to a video renderer 107. According to some examples, experience player 102 may be a media player, game engine, or a component of a personal computer or mobile device, or integrated into a television, DVD player, soundbar, set-top box, or service provider media device (such as Chromecast, Apple TV, or Amazon Fire TV). In some examples, experience player 102 may be configured to receive encoded object-based light data 505', as well as encoded audio data 111' and / or encoded video data 112'. In some such examples, the encoded object-based light data 505' may be received as part of the same bitstream as the encoded audio data 111' and / or encoded video data 112'. Some examples are described in more detail below. According to some examples, experience player 102 can be configured to extract object-based light data 505 from the content bitstream and provide the decoded object-based light data 505 to a light scene renderer 501, provide decoded audio data 111 to an audio renderer 106, and provide decoded video data 112 to a video renderer 107. In some examples, experience player 102 can be configured to allow control over configurable parameters in the light scene renderer 501, such as immersion intensity. Some examples are described below.

[0109] In some examples, the room descriptor of the environment and lighting data 104 can describe the size and orientation of the playback environment itself to establish a relative or absolute coordinate system to which all objects are positioned. For example, in a living room, the display screen can be considered as the front, or in some cases as the center of the front, and the floor and ceiling can be considered as vertical boundaries. In some such examples, the room descriptor can also indicate the boundaries corresponding to the left, right, front, and back walls relative to the front position. According to some examples, the room descriptor can also be provided in the form of a matrix (such as a 3 × 3 matrix). This room descriptor information is used to describe the physical dimensions of the playback environment (e.g., expressed in physical distance units such as meters). In some such examples, sensory object positioning, sensory object size, and sensory object orientation can be described in units relative to the room size, such as in the range of -1 to 1. In some cases, the room descriptor can also describe the preferred viewing position based on the matrix.

[0110] According to this example, system 500 includes a lighting renderer 501 configured to render object-based light data 505 into luminaire control signals 515, at least in part based on environment and actuator data 104. In this example, lighting renderer 501 is configured to output the luminaire control signals 515 to lamp controllers 103 configured to control luminaires 108. Luminaires 108 may include a single controllable light source, a group of controllable light sources (such as controllable light strips), or a combination thereof. In some examples, lighting renderer 501 may be configured to manage various types of light object metadata layers, examples of which are provided herein. According to some examples, lighting renderer 501 may be configured to render the luminaire actuator signals at least in part based on the viewer's perspective. If the viewer is in a living room that includes a television (TV) screen, in some examples, lighting renderer 501 may be configured to render the actuator signals relative to the TV screen. However, in virtual reality (VR) use cases, the lighting renderer 501 can be configured to render actuator signals relative to the position and orientation of the user's head. In some examples, the lighting renderer 501 can receive input from the playback environment (such as light sensor data corresponding to ambient light, camera data corresponding to the person's position or orientation, etc.) to enhance the rendering effect.

[0111] In some examples, the lighting renderer 501 is configured to receive object-based lighting data 505, which includes light objects and object-based lighting metadata indicating the expected lighting environment, as well as environment and lighting data 104 corresponding to other features of the luminaires 108 and the local playback environment. These features may include, but are not limited to, reflective surfaces, windows, uncontrollable light sources, shading features, etc. In this example, the local playback environment includes one or more speakers 109 and one or more display devices 510.

[0112] According to some examples, the lighting renderer 501 is configured to calculate, at least in part, how to activate various controllable luminaires 108 based on object-based lighting data 505 and environment and luminaire data 104. For example, environment and luminaire data 104 may indicate the geometric location of the luminaires 108 in the environment, luminaire type information, etc. In some examples, the lighting renderer 501 may be configured to determine, at least in part, which luminaires will be activated based on position metadata and size metadata associated with each light object, for example, by determining which luminaires are located within a volume of the playback environment, corresponding to the position and size of the light object at a specific time indicated by light object timestamp information. In this example, the lighting renderer 501 is configured to send luminaire control signals 515 to the luminaire controller 103 based on environment and luminaire data 104 and object-based lighting data 505. The luminaire control signals 515 may be sent via one or more transport mechanisms, application programming interfaces (APIs), and protocols. For example, these protocols may include Hue API, LIFX API, DMX, Wi-Fi, Zigbee, Matter, Thread, Bluetooth Mesh, or other protocols.

[0113] In some examples, the lighting renderer 501 can be configured to determine the drive level of the lighting environment expected by the authors of the object-based light data 505 for each of one or more controllable light sources. According to some examples, the lighting renderer 501 can be configured to output a drive level to at least one of the controllable light sources.

[0114] According to some examples, the lighting renderer 501 can be configured to collapse one or more portions of a lighting map based on content metadata, user input (selection mode), lighting fixture limitations and / or configurations, other factors, or combinations thereof. For example, the lighting renderer 501 can be configured to render the same control signals to two or more different lights in the playback environment. In some such examples, the two or more lights can be positioned close to each other. For example, the two or more lights can be different lights with the same actuator, such as different bulbs in the same light bulb. Instead of calculating slightly different control signals for each bulb, the lighting renderer 501 can be configured to reduce computational overhead, improve rendering speed, etc., by rendering the same control signals to two or more different but closely spaced lights.

[0115] In some examples, the lighting renderer 501 can be configured to spatially upmix the object-based lighting data 505. For example, if the object-based lighting data 505 is generated for a single plane (such as a horizontal plane), in some cases, the lighting renderer 501 can be configured to project the light objects of the object-based lighting data 505 onto an upper hemisphere surface (e.g., above the user's actual or intended head position) to enhance the experience.

[0116] According to some examples, the lighting renderer 501 can be configured to apply one or more thresholds, such as one or more spatial thresholds, one or more brightness thresholds, etc., when rendering actuator control signals to the light actuators of the playback environment. In some cases, such thresholds may prevent some light objects from activating some lights.

[0117] Light objects can be used for a variety of purposes, such as creating room atmosphere, providing spatial information about people or objects, enhancing special effects, creating greater interactivity and immersion, diverting the viewer's attention, emphasizing content, and so on. Content creators may use object metadata that indicates the location and size of a sensory object (which is generally applicable to various types of sensory objects) to express some of these purposes (at least in part).

[0118] For example, the priority of sensor objects (including but not limited to light objects) can be indicated via sensor object priority metadata. In some such examples, sensor object priority metadata is considered when multiple sensor objects are simultaneously mapped to the same light fixture in the playback environment. This priority can be indicated via light priority metadata. In some examples, priority may not need to be indicated via metadata. For example, MS Renderer 001 might prioritize moving sensor objects (including but not limited to light objects) over stationary sensor objects.

[0119] A light object may trigger the activation of multiple lights, depending on its location and size, as well as the position of the light fixture within the playback environment. In some examples, when the size of a light object contains multiple lights, the renderer may apply one or more thresholds (such as one or more spatial thresholds or one or more brightness thresholds) to prevent the object from activating some of the contained lights.

[0120] Example of using lighting mapping In some implementations, a lighting map (an instance of an actuator map (AM) that includes a description of lighting in the playback environment) may be provided to the scene renderer 501. In some such examples, Figure 5The environmental and lighting data shown may include a lighting map. According to some examples, the lighting map may be allocentric, for example, indicating light attenuation based on absolute spatial coordinates; while in other examples, the lighting map may be egocentric, for example, projecting light onto a sphere representing the intended viewing position and orientation. In the case of a sphere, in some examples, the lighting map may be projected onto a two-dimensional (2D) surface, for example, to utilize a 2D image texture in the processing. In any case, the lighting map should indicate the capabilities and lighting settings of the playback environment (such as a room). In some embodiments, the lighting map may not be directly related to physical room characteristics, for example, if certain adjustments based on user preferences have been made.

[0121] In some examples, each luminaire or light in the playback environment may have a lighting map. According to some examples, the intensity of the light indicated by the lighting map may be inversely correlated with the distance to the center of the light, or may be approximately inversely correlated with the distance to the center of the light (e.g., within ±5%, ±10%, ±15%, ±20%, etc.). The intensity value of the lighting map can indicate the intensity or effect of a light object on a luminaire. For example, when a light object approaches a light bulb, the lighting renderer 501 can be configured to determine that the intensity of the light bulb will increase as the distance between the light object and the light bulb decreases. The lighting renderer 501 can be configured to determine the rate of this transition based at least in part on the light intensity indicated by the lighting map.

[0122] In this general rendering space, in some examples, the lightscape renderer 501 can be configured to calculate the light activation index for each light using a dot product multiplication between the light object and the light map, for example, as shown below:

[0123] In the aforementioned equation, Y represents the light activation index, LM represents the illumination map, and Obj represents the mapping of the light object. The light activation index indicates the relative light intensity of the actuator control signal output by the lighting renderer 501 based on the overlap between the light object and the light diffusion from the luminaire. In some examples, the lighting renderer 501 may use the maximum or nearest distance from the light object to the luminaire, or other geometric measures, as part of determining the light intensity. In some implementations, the lighting renderer 501 does not calculate the light activation index but may instead determine it by referring to a lookup table.

[0124] The lighting renderer 501 can repeat one of the above processes to determine the light activation metrics for all light objects and all controllable lights in the playback environment. Thresholding light objects that have a minimal impact on the lights can help reduce complexity. For example, if the effect of a light object would cause the light fixture activation to fall below a certain threshold percentage (such as below 10%, below 5%, etc.), the lighting renderer 501 might ignore the effect of that light object.

[0125] Then, the lighting renderer 501 can use the generated light activation matrix Y, along with various other properties such as the selected panning rules (indicated by the light object metadata or renderer configuration) or the priority of the light objects, to determine which objects are rendered by which lights and how. Rendering light objects into lighting control signals can involve: • Adjust the brightness of the light source based on the distance between the light source and the light fixture; • Mix the colors of multiple light objects rendered simultaneously (multiplexed) by a single light fixture; or • Change any of the above options based on the light object priority.

[0126] Rendering parameters In addition to the information carried by the light object's metadata, the rendering of a light object can be a function of the settings or parameters of the light renderer 501 itself. These settings or parameters may include: • Speed ​​Priority - When this parameter is set, moving light objects have higher priority than stationary objects. Having a speed priority parameter set enhances the dynamism of the rendered scene; • Color Priority - Light objects with higher saturation values ​​will be given priority; • Activation threshold - The minimum light activation Y that must be reached to activate the luminaire; • Accessibility - Certain colors can be selected instead of others to best represent the experience for colorblind users. For light-sensitive users, certain flash rates can be avoided.

[0127] Rendering configuration (mode) In some implementations, in addition to the information carried by the light object metadata, the lighting renderer 501 can be configured according to different modes. As used herein, the term "mode" differs from "parameter" because a mode can, for example, involve entirely different signal paths, while a parameter can simply parameterize those signal paths. For example, one mode might involve casting all light objects onto a lighting map before determining how / what to render to the lights, while another mode might align only the highest priority lights to the nearest lights. Modes may include: • Supports modes with low light counts. In these modes, rendering parameters and light object metadata are used to determine which subset of light objects to render and how to render them. Here, "how" refers to the trade-off between spatial, color, and temporal fidelity of the most prominent light objects in the scene; • Supports modes for different content types (such as music and games); • Multiple light objects can be rendered using a pattern that employs color mixing by a single light fixture (or a single light); • A single light fixture (or a single lamp) can only render a single light object; • A mode that changes the brightness of a light object based on the geometric or other distance between the light object and the luminaire.

[0128] As described above, according to some embodiments, system 500 may include methods configured to perform at least some of the methods disclosed herein. Figure 1 One or more instances of the control system 110. In some such examples, one instance of the control system 110 may implement the scene creation tool 100, and another instance of the control system 110 may implement the experience player 002. In some examples, one instance of the control system 110 may implement an audio renderer 006, a video renderer 007, a scene renderer 501, or a combination thereof. According to some examples, an instance of the control system 110 configured to implement the experience player 002 may also be configured to implement an audio renderer 006, a video renderer 007, a scene renderer 501, or a combination thereof.

[0129] Figure 6 It shows that it can be made by Figure 5 This is an example of a graphical user interface (GUI) displayed on a display device for a scene creation tool. (Compared to other tools provided in this article...) Figure 1 Sample, Figure 6 The types and quantities of components shown are provided as examples only. Other GUIs rendered by the scene creation tool may include more, fewer, and / or different types and quantities of components. Based on some examples, GUI 600 can be derived from... Figure 1 Commands for an instance of the control system 110 are presented on a display device; this control system is configured to implement... Figure 5 Scene creation tool 100.

[0130] In this example, the user can interact with GUI 600 to create light objects and assign them properties, which can be associated with the light objects as metadata. According to this example, the user is selecting properties for light object 630. In this example, GUI 600 displays light object 630 in three-dimensional space 631, which represents the playback environment. Element 634 illustrates the coordinate system of three-dimensional space 631. Therefore, in this example, light object 630 and three-dimensional space 631 are viewed from the top left corner.

[0131] Users can interact with the GUI 600 to select the position and size of the light object 630. In some examples, users can select the position of the light object 630 by dragging it to the desired location within the three-dimensional space 631 (e.g., by touching a touchscreen, using a cursor, etc.). According to some examples, users can select the size of the light object 630 by selecting the size of a circle (or other shape) displayed on the GUI 600 to indicate the outline of the light object 630. In some such examples, users can decrease the size of the light object 630 by pinching its outline with two fingers, increase its size by spreading their fingers, and so on.

[0132] Specifying the position and size of a light object within an abstract three-dimensional space (such as 3D space 631 of GUI 600) allows content creators to generalize the location and extent of corresponding light effects without prior knowledge of the specific playback environment that will provide the light effect. This is an advantage of the object-oriented approach in various disclosed implementations. For example, GUI 600 allows content creators to specify the position and size of a light object 630 within 3D space 631, thereby allowing content creators to generalize the location and extent of corresponding light effects without prior knowledge of the specific size of any particular playback environment that will provide the light effect, the number, type, and location of lights within the playback environment that will provide the light effect, etc. Lights that will potentially be actuated at a specific time in response to the presence of light object 630 will be lights within the volume of the playback environment corresponding to the position and size / extension of light object 630.

[0133] According to this example, a user can interact with the color wheel 635 of the GUI 600 to select the hue and color saturation of the current light object, and can interact with the slider 636 to select the brightness of the current light object. These and other selectable properties of the light object 630 are displayed in area 632 of the GUI 600. According to this example, the properties of the light object 630 that can be selected via the GUI 600 also include intensity, diffusivity, "feathering", whether the light object is hidden or not, saturation, priority, and layer. Light object layers and priorities are described in more detail below. In general, light object layers can be used to group light objects into categories such as "ambience" and "dynamics". Light object priorities can be assigned by the content creator and used by the renderer to determine, for example, which light object(s) will be rendered when two or more light objects are active simultaneously and simultaneously surround an area including the same light fixture.

[0134] Area 640 of GUI 600 indicates the corresponding time information for each of the multiple light objects created via the lightscape creation tool. In this example, the light objects are listed along the vertical axis on the left side of area 640, and the time is shown along the horizontal axis. In this example, the four-second time interval is depicted by a vertical line. Here, the time information for each light object is shown as isolated or connected diamond symbols or lines along a series of horizontal rows, each of which corresponds to one of the light objects indicated on the left side of area 640. For example, line 633 indicates that light object 3 will begin to appear between 39 and 40 seconds and will appear continuously until almost 1 minute and 6 seconds. The diamond symbol to the right of line 633 indicates that light object 3 will appear discontinuously over the next few seconds.

[0135] Figure 7A It shows that it can be made by Figure 5 Another example of a graphical user interface (GUI) presented on a display device for a scene creation tool. (This is in contrast to other tools provided in this article.) Figure 1 Sample, Figure 7A The types and quantities of components shown are provided as examples only. Other GUIs rendered by the scene creation tool may include more, fewer, and / or different types and quantities of components. Based on some examples, GUI 700 can be customized according to... Figure 1 Commands for an instance of the control system 110 are presented on a display device; this control system is configured to implement... Figure 5 Scene creation tool 100.

[0136] In this example, GUI 700 indicates that it has been passed Figure 5The light creation tool 100 creates light objects to control the timing of lighting fixtures in the actual playback environment. An image of the playback environment is shown in area 705 of the GUI 700. Various lighting fixtures 708 and a television 715 are shown in the playback environment of area 705. Specific moments are indicated by a vertical line 742 in area 740. At this time, the vertical line 742 intersects with horizontal lines 744a, 744b, 744c, and 744d, indicating that the light provided in the corresponding light objects 1, 4, 5, and 7 is being played. Area 732 indicates the characteristics of the light objects.

[0137] It can be observed that, Figure 7A At the moment depicted, the left side of the playback environment shown in area 705 is illuminated by blue light. This corresponds at least in part to the effect of the light object 730 shown within three-dimensional space 731.

[0138] According to this example, video and audio data are also played in an audio environment, and the playback of the rendered light object is synchronized with the playback of the video and audio data. In this example, an image of the playing video is shown in area 710 of the GUI 700. The video can be played, for example, by a television 715.

[0139] In some examples, users may be able to interact with the GUI 700 to adjust light object properties, add or delete light objects, etc. For example, a user can pause playback to adjust light object properties. In some alternative examples, users may need to return to the GUI (such as...). Figure 6 The GUI 600 allows users to adjust light object properties, add or delete light objects, etc.

[0140] Reference Figure 7A In the described example, although GUI 700 is presented as... Figure 5 The scene creation tool 100 corresponds to a display device, but renders light objects, audio, and video in a real-world environment. Therefore, in some implementations, reference is made to... Figure 7A The examples described may also involve those that can be derived from... Figure 5 The other boxes provide at least some of the "downstream" rendering and playback capabilities, including but not limited to a lightscape renderer 501, a light controller API 103 (in some cases, these light controller APIs may be implemented by the same device that implements the lightscape renderer 501), a light fixture 108, an audio renderer 106, a loudspeaker 109, a video renderer 107, and (multiple) display devices 510. In some such examples, refer to Figure 7A The description process can also involve Figure 5 Experience the features of Player 102.

[0141] In some alternative implementations, reference Figure 7AThe examples described may also involve those that can be derived from... Figure 4 The other boxes provide at least some of the "downstream" rendering and playback functions, including but not limited to MS renderer 001, MS controller API 003 (in some cases, these rendering and playback functions may be implemented by the same device that implements MS renderer 001), lighting fixtures 008, audio renderer 006, loudspeakers 009, video renderer 007, and (multiple) display devices 010. In some such examples, refer to Figure 7A The description process can also involve Figure 4 Experience the features of Player 002.

[0142] Figure 7B This is a flowchart outlining an example of a method that can be performed by a device or system such as the apparatus or system disclosed herein. As with other methods described herein, the blocks of method 750 need not be performed in the indicated order. In some embodiments, one or more blocks of method 750 may be performed simultaneously. Furthermore, some embodiments of method 750 may include more or fewer blocks than those shown and / or described. The blocks of method 750 may be performed by one or more devices, which may be (or may include) a control system (such as those described above). Figure 1 One or more instances of the control system 110 shown and described. For example, at least some aspects of method 750 can be implemented by a system configured to perform... Figure 4 The experience player 002 is executed by the control system 110 instance. Some other aspects of method 750 can be configured to implement Figure 4 The MS renderer 003's control system 110 instance is executed.

[0143] In this example, box 755 relates to the control system receiving a content bitstream comprising encoded object-based sensory data. In this case, the object-based sensory data includes sensory objects and corresponding sensory metadata, and corresponds to sensory effects to be provided by multiple sensory actuators in the environment. In some examples, the environment can be an actual real-world environment, such as a room environment or a car environment. Sensory effects can include sensory effects to be provided by multiple sensory actuators in the environment, such as lighting, tactile sensation, airflow, one or more position actuators, or combinations thereof. According to some examples, the environment can be or can include a virtual environment. In some such examples, method 750 can relate to providing a virtual environment (such as a game environment) while providing corresponding sensory effects in a real-world environment.

[0144] In some examples, the object-based sensory metadata may include sensory spatial metadata, which at least indicates the spatial location within the environment used to render the object-based sensory data, the region within the environment used to render the object-based sensory data, or both. According to some examples, the object-based sensory data does not correspond to a specific sensory actuator in the environment. For example, as referenced... Figure 6 and Figure 7A As described, a sensory object based on object-based sensory data can correspond to a portion of a three-dimensional region representing a playback environment. When creating a sensory object, the actual playback environment in which the sensory object will be rendered does not need to be known, and typically is not. Therefore, object-based sensory data includes abstract sensory reproduction information (in this example, the sensory object and its corresponding sensory metadata) that allows the sensory renderer to reproduce the created sensory effects via various sensory actuator types, via various numbers of sensory actuators, and from various sensory actuator locations in the environment.

[0145] According to this example, box 760 relates to extracting object-based sensory data from a content bitstream by a control system. In some such examples, the content bitstream may also include encoded audio objects synchronized with the encoded object-based sensory data. According to some such examples, the audio objects may include audio signals and corresponding audio object metadata. In some such examples, method 750 may also involve extracting audio objects from the content bitstream by a control system. According to some such examples, method 750 may also involve providing these audio objects to an audio renderer by the control system. In some examples, the audio object metadata may at least include audio object spatial metadata indicating the spatial location of the audio object used to render the audio signal within the environment.

[0146] In this example, box 765 relates to providing object-based sensory data from a control system to a sensory renderer. According to some such examples, method 750 may also involve the sensory renderer receiving object-based sensory data and receiving environmental descriptor data corresponding to the positioning of sensory actuators in the environment. In some such examples, method 750 may also involve the sensory renderer receiving actuator descriptor data corresponding to the characteristics of sensory actuators in the environment. In some examples, the environmental descriptor data and actuator descriptor data may be or may be included in a reference... Figure 4The described environment and actuator data 004. According to some examples, method 750 may also involve providing actuator control signals by the sensory renderer for controlling these sensory actuators in the environment to produce sensory effects indicated by the object-based sensory data. In some such examples, the MS renderer 001 may provide actuator control signals 310 to the MS controller API 003, and the MS controller API 003 may provide actuator-specific control signals to the actuator 008. In some alternative examples, Figure 4 The MS controller API 003 shown can be implemented via MS renderer 001, and actuator-specific signals can be provided to actuator 008 by MS renderer 001. In some examples, method 750 may also involve providing sensory effects by sensory actuators in the environment.

[0147] In some examples, method 750 may also involve receiving an audio object by an audio renderer and receiving loudspeaker data corresponding to a loudspeaker in the environment by the audio renderer. According to some such examples, method 750 may also involve providing loudspeaker control signals by the audio renderer for controlling a loudspeaker in the environment to play audio corresponding to the audio object and synchronized with a sensory effect. Synchronization may be based, for example, on time information included in or with the sensory object and audio object, such as a timestamp. In some examples, method 750 may also involve playing audio corresponding to the audio object by a loudspeaker in the environment.

[0148] According to some examples, the content bitstream includes encoded video data synchronized with the encoded audio object and the encoded object-based sensory data. In some such examples, method 750 may also involve the control system extracting video data from the content bitstream and providing the video data to a video renderer. According to some such examples, method 750 may also involve the video renderer receiving the video data and providing video control signals to control one or more display devices in the environment to render an image corresponding to the video control signals and synchronized with the audio object and sensory effects. In some examples, method 750 may also involve rendering by one or more display devices in the environment. The image may correspond to the video control signals.

[0149] Expected lighting environment metadata Some publicly disclosed examples involve adding lighting metadata to video and / or audio tracks, or using it as a standalone lighting-based sensory experience. This lighting metadata describes the expected lighting environment to be reproduced during playback. There are many ways to represent lighting metadata, several of which are described in detail in this disclosure.

[0150] In some examples, the lighting environment is expected to be transmitted as one or more image-based lighting (IBL) objects. IBL is a technique previously used to capture environment and lighting. IBL can be described as the process of illuminating a scene and (real or synthetic) objects with images from real-world lighting. IBL originates from previously disclosed reflection mapping techniques, used as texture mapping on computer graphics models in panoramic images to display shiny objects reflecting real or synthetic environments. Some aspects of IBL are similar to image-based modeling, where the geometry of a 3D scene can be derived from an image. Other aspects of IBL are similar to image-based rendering, where the rendered appearance of a scene can be generated based on how the scene appears in an image.

[0151] Previously, IBL objects were used in rendering computer graphics to create realistic reflections and lighting effects, such as rendering objects that appear to reflect parts of the real-world environment. In some previously published examples of virtual worlds / computer-generated graphics, IBL involves the following process: • Capture real-world illumination as an omnidirectional image; • Map illumination onto a representation of the environment; • Place computer-generated 3D objects within an environment; and • Simulate ambient light illuminating computer graphics objects.

[0152] There are various methods for capturing omnidirectional images. One method is to photograph a reflective sphere placed in the environment using a camera. Another method for obtaining an omnidirectional image is to obtain a mosaic based on multiple camera images obtained from different directions / viewpoints and to combine these images using an image stitching procedure. In some such examples, a fisheye lens can be used to obtain the image, which requires only two images to cover the entire field of view. Another method for obtaining an omnidirectional image is to use a scanning panoramic camera, such as a “rotating line camera,” which is configured to assemble digital images as the camera rotates to scan a 360-degree field of view. Further details of previously disclosed IBL methods are disclosed in “Image-Based Lighting” by Paul Debevec (IEEE, March / April 2002, pp. 26–34), which is incorporated herein by reference.

[0153] Some embodiments of this disclosure build upon previously disclosed methods involving IBL objects to achieve a new objective: reproducing a desired environment using dynamically controllable ambient lighting. The desired lighting environment and the endpoint environment may change over time. For example, the walls of the endpoint environment may be painted, actuators may be moved, furniture may be moved or replaced, and new furniture, shelves, and / or cabinets may be added.

[0154] Several types of mapping or projection can be used to map ambient lighting on a sphere around the intended viewing position onto a two-dimensional (2D) plane. Projecting onto a 2D plane simplifies the process of compressing the lit object using 2D image or video codecs and reduces the computational overhead required for more efficient transmission.

[0155] Figure 8A , Figure 8B and Figure 8C Three examples of projecting the lighting of a viewing environment onto a two-dimensional (2D) plane are shown. In these examples, spherical projection is used. Other representations are also possible. In these examples, the lighting of the viewing environment is shown from the intended viewing position, but the viewing angle exceeds 360 degrees. In these examples, the x-axis represents the horizontal angle (from left to right) from the viewer's perspective, and the y-axis represents the vertical angle (from top to bottom) from the viewer's perspective. In these examples, the centers of projections 805, 810, and 815 correspond to the direction in front of the viewer. The leftmost and rightmost sides of projections 805, 810, and 815 correspond to each other because the projections “wrap” around the viewing position. For example, projections 805, 810, and 815 could correspond to three different contents, or to three different times of the same content. Projection 805 represents studio-style lighting in an environment dominated by darkness and monochrome. Projection 810 includes a dark blue floor 812 and a main blue area 814 directly in front of the viewer. Projection 815 includes a bright red light 819 located slightly to the left of the viewer.

[0156] As described above, in some disclosed implementations, lighting metadata provided along with a light object can be used to generate video showing the environment from a intended viewing location. In some examples, the lighting metadata may include multiple metadata units, each of which may include a header describing how the metadata is represented for playback purposes. In some such examples, the lighting metadata may include information that can be provided using different metadata units: • Environment metadata version: This information enables the decoder to correctly interpret multiple representations of lighting and select the most suitable method for the lighting capabilities of a particular playback environment; • Environment metadata mapping type: This information enables the decoder to correctly convert metadata into real-world coordinates; • Environmental metadata timecode: This information enables the decoder to correctly synchronize the metadata with the audio or video track; • Environmental metadata location code: This information enables the decoder to correctly align the viewing position with a reference viewing position. In some examples, there may be multiple sets of metadata corresponding to different viewing positions, allowing viewers to experience content from multiple viewing positions, and adjusting the lighting accordingly by selecting the nearest suitable position or interpolating between nearby positions. For example, Locations X1, Y1, Z1 have environmental metadata payload EMP1, while locations X2, Y2, Z2 have EMP2; o Determine the user's location X', Y', Z'; The geometric distance is calculated using the following formula from the distances between X1, Y1, Z1 and X', Y', Z': D1 = sqrt((X1 - X') 2 + (Y1-Y') 2 + (Z1-Z') 2 ), and D2 = sqrt((X2-X') 2 + (Y2-Y') 2 + (Z2-Z') 2 ); Calculate the distance ratio to determine the "Alpha" value: Alpha = D1 / (D1 + D2); o Use alpha to interpolate the environmental metadata payload applicable to the user location EMP' between two reference viewing locations, for example, via EMP' = EMP1 (1-Alpha) + EMP2 (alpha); This simple example illustrates one method of linear interpolation between two points. If the viewer's position extends beyond either reference point, additional clamping may be expected. In some cases, triangulation may be more suitable. • Environmental metadata compression methods; • Size of the environmental metadata payload; and • Environment metadata payload.

[0157] In some examples, IBL indicates that depth information can be used to augment the relative distance between an ambient light source and a reference viewing position. This depth information can be used to adjust the position of the light source as the viewer moves around in the environment. For example, depth information can be obtained directly from an RGB image using a trained neural network, for instance, through monocular depth estimation. Alternatively or additionally, many consumer devices, such as the iPhone, are capable of directly measuring depth using infrared imaging techniques such as LiDAR and structured light.

[0158] Depending on the use case requirements, the spatial resolution of IBL technology may vary, depending on the available bit rate and the required compression quality. IBL can be compressed using known image-based compression methods such as Joint Picture Experts Group (JPG), JPG2000, Portable Network Graphics (PNG), etc., or using known video-based compression methods such as Advanced Video Coding (AVC), also known as H.264, H.265, Universal Video Coding (VVC), AOMedia Video 1 (AV1), etc.

[0159] However, the methods disclosed herein are not limited to the IBL-based examples. Other examples representing the intended lighting environment treat each light source as a unique light source object with a defined location and size. In some such methods, additional information can be used to define or describe each light source, including but not limited to the directionality of the light emitted by each light source object. Some methods may also include information about the reflectivity of one or more surfaces, one or more intended room dimensions, etc. This information can be used to support implementations that allow viewers to move freely to new locations. According to some examples, light source objects can be defined directly using computer programs specifically written for this task, or inferred by analyzing video images of light sources in the environment. Potential advantages of light source object-based methods include a smaller metadata payload size for relatively simple lighting scenarios.

[0160] Play lighting environment rendering Figure 9A Example elements of a lighting renderer are shown. (Compared to other elements provided in this article...) Figure 1 Sample, Figure 9A The types and quantities of components shown are provided as examples only. Other implementations may include more, fewer, and / or different types and quantities of components. According to this example, the scene renderer 501 is... Figure 5 An instance of a lighting renderer 501 is shown. In this example, the lighting renderer 501 is composed of... Figure 1 An example of a control system 110 is used for implementation. According to this example, the scene renderer 501 includes a scaling factor calculation box 925 and a digital drive value calculation box 930.

[0161] According to this example, during playback, the lighting renderer 501 receives object-based lighting data 505 (including expected ambient lighting metadata) as input. For example, the object-based lighting data 505 may have been generated by... Figure 5The scene creation tool 100 generates the lighting. In this example, the rendering engine also receives environment and lighting data 104 and calculates appropriate lighting control signals 515 for controllable lighting fixtures 108 (not shown) in the playback environment. In this example, the environment and lighting data 104 is shown as including environment and lighting data 104a (which includes information about the lighting fixtures in the playback environment and their capabilities) and environment and lighting data 104b (which includes information about the ambient light and / or uncontrollable lighting fixtures in the playback environment). In some examples, the scene renderer 501 may be configured to output the lighting control signals 515 to lighting controllers 103, which are configured to control the lighting fixtures 108. In this document, the lighting fixtures 108 in the playback environment may also be referred to as “endpoint dynamic lighting elements”.

[0162] In some examples, the environmental and luminaire data 104a may include information about a set of N dynamic lighting elements, where N represents the total number of controllable dynamic lighting elements in the playback environment. According to this example, for each dynamic lighting element... n The environmental and luminaire data 104a includes: (1) an IBL mapping 905 of the ambient lighting produced by the dynamic (controllable) luminaire at maximum intensity; and (2) a mapping 910 of the light intensity produced by the controllable luminaire to the digital drive signal provided to the controllable luminaire. In this example, the environmental and luminaire data 104b includes ambient light information 915 and ambient light level information 920. In some examples, the ambient light information 915 may include a basic IBL mapping of the basic ambient lighting that is not controllable by the system. According to some examples, the ambient light information 915 may include information about other light sources, such as windows in the playing environment, the direction the windows face, the amount of outdoor light received through the windows at different times of day in the environment, information about controllable curtains (if any), etc. In some examples, the ambient light level information 920 may include, for example, scale values ​​of the basic ambient lighting obtained by an optical light sensor.

[0163] In some implementations, the lighting renderer (such as...) Figure 5 The scene renderer (501) can perform the following operations: 1. Receive the IBL (Information Board) regarding the intended ambient lighting. ref The information is used as input, here as part of object-based optical data 505; 2. For example, the base lighting IBL can be calculated based on a constant estimate, or by scaling the base IBL mapping with the estimated ambient light. base ; 3. Calculate the IBL for each dynamic lighting element. n Linear scaling n This makes each scaled IBL nThe sum of these, plus the base lighting, most closely matches the expected ambient lighting, for example as follows: Minimize abs(IBL) ref - (IBL base + sum(Scale n IBL n ))) Each linear value can be encoded using a nonlinear function, and then the difference can be calculated. The nonlinear function can correspond to the human visual sensitivity to color and light intensity. 4. Calculate the scaling value (Scale). n The digital drive signal for each dynamic lighting element in the Drive n (exist Figure 9A The element is represented as 515 because they are... Figure 5 (Example of the lighting control signal 515); and 5. Transmit digital drive values ​​to dynamic lighting elements, or to... Figure 5 The lighting controller API 103.

[0164] In some examples, calculating the scaling factor may involve adjusting the IBL (Intensity Limiting Light) of the desired illumination. ref Subtract ambient light IBL base Then use deconvolution based on dynamic lighting IBL n Calculate the scaling factor. n Traditional deconvolution methods can be used, with regularization included as needed to improve robustness.

[0165] In some alternative examples, the process of calculating the scaling factor can be performed iteratively; for example, the scaling factor is initialized to an initial value, and then the scaling value is adjusted sequentially based on a reference value (expected illumination IBL). ref Evaluate the output. Based on some examples, traditional gradient descent and function minimization methods can be used during the calculation of the scaling factor.

[0166] In some examples of calculating numeric driven values ​​based on scaling factors, lookup tables (LUTs) can be used (such as...). Figure 9A The Dynamic Light LUT 910 (LUT) is used to determine the correct digital drive value required to obtain the desired light intensity from a specific light source. Alternatively, the functional scaling factor can be obtained from the measurement data.

[0167] Calibration and system configuration An important aspect of this application is generating dynamic lighting IBL data for a given playback environment. In one embodiment, the process is as follows: 1) Establish a connection between the host and the dynamic lighting fixtures (also referred to as controllable lighting fixtures in this article); 2) Install the IBL measurement equipment at the preferred viewing location. Commonly used techniques include: cameras that image the shiny sphere, cameras with fisheye lenses, cameras that pan around the scene, or fixed devices with multiple cameras capturing the scene from various directions (360-degree cameras); 3) Capture basic ambient light (IBL) base and light sensor readings; 4) For each dynamic luminaire, perform the following operations: a. Set the drive value to the maximum level; b. Capture IBL n ; c. Repeat multiple levels (driving values); d. Construct representative IBLs suitable for multiple drive levels n ;as well as e. Establishing the relationship between linear light and drive values. "Linear light" refers to a space where linear changes in values ​​are perceived as linear changes by humans. Most devices do not have a linear response in this respect. In other words, if the actuator's codeword (drive value) is doubled, the perceived change in brightness will not double. Working in linear light space is convenient and ensures, where applicable, a finite resolution is uniformly distributed within the range of human perception. After performing operations in linear light space and obtaining the corresponding values, these values ​​can be converted into drive values / codewords for controlling physical devices / light.

[0168] The above process applies to situations where the base ambient light is relatively fixed and only increases or decreases overall. For example, a corner of the playback environment might have a window, causing the ambient light to increase or decrease overall depending on the weather and time of day. In many viewing environments, uncontrollable lighting may be more dynamic, such as manually controllable light fixtures (not part of a dynamic setup) or automotive environments. In some examples, multiple base lighting scenes can be captured, and during playback, the captured base lighting scene that best matches the actual current base lighting conditions can be estimated using measurements of ambient light in the environment. The captured IBL can be... n IBL base The relationship between the driving values ​​is stored in a configuration file accessible during rendering.

[0169] Content creation / Master production The output of the content creation / mastering process is the expected lighting environment map or IBL (Intended Background Map). refIn some examples, the expected lighting environment map can be measured directly using the measurement methods described in the calibration section above. In other examples, the expected lighting environment map can be rendered using computer graphics software. Some publicly available examples involve dynamically changing the reference lighting to create a video sequence of the expected reference lighting environment.

[0170] Figure 9B This is a flowchart outlining an example of a method that can be performed by a device or system such as the apparatus or system disclosed herein. As with other methods described herein, the blocks of method 950 need not be performed in the indicated order. In some embodiments, one or more blocks of method 950 may be performed simultaneously. Furthermore, some embodiments of method 950 may include more or fewer blocks than those shown and / or described. The blocks of method 950 may be performed by one or more devices, which may be (or may include) a control system (such as those described above). Figure 1 One or more instances of the control system 110 shown and described. For example, at least some aspects of method 950 can be implemented by means of a system configured to at least implement Figure 5 Scene Renderer 501 or Figure 9A The rendering engine 501 is executed by an instance of the control system 110. In some examples, method 950 can be executed by one or more instances of the control system 110, which are configured to implement... Figure 5 The scene renderer 501 and the light controller API.

[0171] In this example, box 955 relates to receiving object-based light data indicating a desired lighting environment by a control system configured to implement a lighting environment renderer. In this case, the object-based light data includes light objects and lighting metadata. In some examples, the environment can be an actual real-world environment, such as a room environment or a car environment. According to some examples, the environment can be or may include a virtual environment. In some such examples, method 950 may relate to providing a virtual environment (such as a game environment) while simultaneously providing corresponding lighting effects in the real-world environment.

[0172] According to this example, box 960 relates to receiving lighting information about the local lighting environment by a control system. In this example, the lighting information includes one or more characteristics of one or more controllable light sources in the local lighting environment. For example, box 960 may relate to receiving information referenced herein. Figure 5 or Figure 9A The described environment and lighting data is 104.

[0173] In this example, box 965 relates to the control system determining a drive level for each of one or more controllable light sources that approximates a desired lighting environment. Here, box 970 relates to the control system outputting a drive level to at least one of the controllable light sources.

[0174] In some examples, object-based lighting metadata includes time information. In some such examples, box 965 may involve determining one or more drive levels for one or more time intervals corresponding to the time information.

[0175] According to some examples, method 950 may involve receiving viewing location information. In some such examples, box 965 may involve determining one or more drive levels corresponding to the viewing location information.

[0176] In some examples, lighting information may include one or more characteristics of one or more base light sources in the local lighting environment that are not controllable by the control system. In some such examples, box 965 may involve determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.

[0177] Based on some examples, object-based lighting metadata may include at least lighting object location information and lighting object color information. In some examples, the lighting environment is expected to be transmitted as one or more image-based lighting (IBL) objects. In some such examples, the determination process of box 965 may be based at least in part on the local lighting environment. n One controllable light source (IBL) n IBL mapping for each ambient light produced at maximum intensity in ).

[0178] In some examples, method 950 may involve receiving basic ambient lighting (IBL) that cannot be controlled by a control system. base The underlying IBL mapping. In some such examples, the determination process of box 965 can be based at least in part on the underlying IBL mapping.

[0179] In some examples, the determination process for box 965 can be based at least in part on each dynamic lighting element IBL. n Linear scaling (Scale) n ), so that each scaled IBL n The sum of light plus basic lighting IBL base IBL (Identical Lighting Environment) ref In some such examples, the determination process for box 965 can be based at least in part on minimizing IBL. ref With each scaled IBL n The sum of light plus basic lighting IBLbase The differences between them.

[0180] Examples of MS object properties The following is a non-exhaustive list of possible properties of an MS object: • Priority; •layer; • Hybrid mode; •Persistence; • Effects; and • Spatial sound and image laws.

[0181] Effect As used herein, the term "effect" for an MS object is a synonym for the MS object type. An "effect" is, or indicates, the sensory effect provided by the MS object. If the MS object is a light object, its effect will involve providing direct or indirect light. If the MS object is a tactile object, its effect will involve providing some type of tactile feedback. If the MS object is an airflow object, its effect will involve providing some type of airflow. Some examples involve other "effect" categories, as described in more detail below.

[0182] Persistence Some MS objects may contain persistent characteristics in their metadata. For example, when a movable MS object moves around the scene, it can persist for a period of time at the locations it has passed through. This period of time can be indicated by persistent metadata. In some implementations, the MS renderer is responsible for building and maintaining persistent state.

[0183] layer Based on some examples, individual MS objects can be assigned to "layers" where MS objects are grouped together based on one or more shared characteristics. For example, layers can be grouped together based on the expected effects or types of MS objects, which may include, but are not limited to, the following: - Atmosphere / Environment -Informative - Emphasis / Attention Alternatively or additionally, in some examples, layers can be used to group MS objects together based on shared characteristics, which may include, but are not limited to, the following: -color; -strength; -size; -shape; - Location; and - A zone in space.

[0184] Priority In some examples, MS objects can have a priority attribute, which allows the renderer to determine which(s) should have priority in an environment where MS objects compete for limited actuators. For example, if multiple light objects overlap with a single light fixture when all light objects are scheduled to be rendered, the renderer can refer to the priority of each light object to determine which(s) to render. In some examples, priority can be defined between or within layers. According to some examples, priority can be associated with specific characteristics such as intensity. In some examples, priority can be defined by time: for example, the most recently rendered MS object may take precedence over previously rendered MS objects. According to some examples, priority can be used to specify MS objects or layers that should be rendered regardless of the limitations of a particular actuator system in the playback environment.

[0185] Spatial sound image law The spatial acoustic-image law can define how MS objects move in space and how MS objects affect actuators when they move between actuators.

[0186] Hybrid mode Blend patterns specify how multiple objects can be reused on a single actuator. In some examples, blend patterns may include one or more of the following: -Maximum Mode: Selects the MS object that activates the actuator the most times; - Blend mode: Blend some or all objects according to a set of rules, such as by summing activation levels, taking the average of activation levels, or blending colors according to activation level or priority level; -MaxNmix: Mix the first N MS objects according to the rule set (by activation level).

[0187] Based on some examples, instead of (or in addition to) per-object metadata, more general metadata can be defined for the entire multi-sensory content file. For example, an MS content file may include metadata such as the trimming process or the mastering environment.

[0188] Repair and control In the context of Dolby Vision™, a feature called "Trim Controls" can serve as guidance on how to adjust the default rendering algorithm for specific environments or conditions at endpoints. Trim Controls can specify ranges and / or default values ​​for various attributes, including saturation, tonal detail, gamma, etc. For example, automotive trim controls can exist that provide specific default values ​​and / or sets of rules for rendering in automotive environments, such as guidance on including only objects of a specific priority or layer. Other examples include trim controls for environments with limited, complex, or sparse multi-sensory actuators.

[0189] Mastering environment A single multi-sensory content item can include metadata about the characteristics of the mastering environment, such as room size, reflectivity, and ambient bias lighting levels. Specific characteristics can vary depending on the desired endpoint actuator. Mastering environment information can help provide reference points for rendering in the playback environment.

[0190] Scene metadata layer When creating lighting for media content, it can be useful to identify at least two different methods (or layers) to be created. These layers can be used during the process of rendering the created MS object based on the available and controllable lighting fixtures in the playback environment. These layers can help capture artistic intent and allow flexibility in the constraints of the playback environment (such as due to the number of lighting fixtures or light occlusions), so that the main intent of (multiple) authors can still be rendered, but can be scaled or otherwise modified.

[0191] In some examples, direct lighting and indirect lighting can be assigned to different lighting metadata layers.

[0192] direct light object Light objects in the Direct Light Object layer (also referred to as "direct light objects" throughout this document) are light objects that represent light directly visible to the content creator or end user. Examples of direct light objects can include lamps in a scene, the sun, the moon, headlights from an approaching car, lightning during a storm, traffic lights, etc. Direct light objects can also be used to represent light sources that are part of a scene but are typically or temporarily invisible in the associated video content, for example, because they are outside a video frame or because they have moved outside a video frame. In some examples, direct light objects can be used to amplify or enhance auditory events, such as explosions, or to visually guide the trajectory of moving objects outside a video frame. The use of direct light objects is often dynamic. For example, associated metadata such as intensity, color, saturation, and position will typically change over time within the scene of the media content.

[0193] Indirect light object Light objects in the Indirect Light Object layer (also referred to as “indirect light objects” throughout this document) are light objects that represent the effects of indirect lighting. For example, an indirect light object can be used to represent the effect of light radiated by a luminaire as observed when light is reflected by one or more surfaces. Some examples of using indirect light objects include changing the observed color of the walls, ceiling, or floor of an environment to a color that matches the content, such as green for a forest scene or blue for the sky or water. Indirect light objects can also be used to set the atmosphere of a scene and environment in a similar way to how color-coded video content is achieved, but in a more immersive way. For example, science fiction films often use very specific (blue or green) video color grading palettes to enhance the feeling of being in outer space. Flashback scenes often use reduced saturation, muted color, or sepia overlays in the video content to enhance the effect of timeline changes. All these effects can be replicated or approximated outside of video frames by adjusting the light control signals accordingly. Compared to the lighting effects corresponding to indirect light objects, the lighting effects corresponding to indirect light objects are generally more static within a scene and are typically less localized and less dynamic.

[0194] Layer abstraction Some examples involve further abstracting the direct light object layer and the indirect light object layer into layers that include aspects of both. In some such examples, these layers may include one or more environment layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or combinations thereof. In some examples, these layers may be used for, or correspond to, linear or event-based triggering within content.

[0195] (Multiple) environmental layers Similar to indirect lighting, an environment layer can be used to set the atmosphere and tone of a space by playing thin layers of color on surfaces within the environment. An environment layer can serve as a base layer upon which to build a lighting scene. In some examples, an environment layer can be represented by a light object covering a relatively large area. In other examples, an environment layer can be represented by a light object covering a relatively small area (e.g., using one or more images). Depending on the example, an environment layer can be divided into zones. In some such examples, a particular lighting effect always occupies a specific spatial area. For example, the walls, ceiling, and floor in the creation or playback environment can each be considered a separate environment layer zone.

[0196] (Multiple) dynamic layers In some implementations, a dynamic layer can be used to represent spatial and temporal changes in MS objects, such as light objects. Within the dynamic layer, individual MS objects can also have priorities, such that, for example, one light object may take precedence over another when rendered by a light fixture. Within the dynamic layer, in some examples, individual MS objects can be linked to other objects, such as linking to audio objects or 3D world MS objects (from spatial audio).

[0197] (Multiple) Customized Layers In some examples, custom layers can be used to design light sequences that can be freely assigned to lights for functional purposes. These sequences may not be spatial in nature, but rather provide further information to the user. For example, in a game, light strips could be assigned to display the player's remaining health.

[0198] (Multiple) overlays Based on some examples, overlays can be used to render persistent light with sequential priority. Overlays can also be used, for example, to create "watermarks" on all other elements in a lighting scene.

[0199] Creation and distribution of light layer data Create direct and indirect light objects In some examples, direct light objects can be created by determining or setting the light source position, intensity, hue, saturation, and spatial extent of one or more light objects as a function of time. In some such examples, the creation process can create corresponding metadata, which can be distributed along with the direct light object and the audio and / or video content used for content rendering. Ideally, the direct light object is rendered as a direct light source.

[0200] In some examples, indirect lighting effects can be created as a dedicated group or category within the scene's metadata content, allowing for a greater focus on overall color and environment rather than dynamic effects. Indirect lighting effects can also be defined by intensity, hue, saturation, or a combination thereof as a function of time, but are typically associated with a large area of ​​the scene's rendering environment. Indirect lighting effects are ideally (but not necessarily) rendered to indirect light sources (if available).

[0201] Figure 10 This shows another example of a GUI that can be rendered by a display device used in scene creation tools. (Compared to other tools provided in this article...) Figure 1 Sample, Figure 10 The types and quantities of components shown are provided as examples only. Other GUIs rendered by the scene creation tool may include more, fewer, and / or different types and quantities of components. Based on some examples, GUI 1000 can be customized based on... Figure 1Commands for an instance of the control system 110 are presented on a display device; this control system is configured to implement... Figure 5 Scene creation tool 100.

[0202] In this example, a user can interact with GUI 1000 to create light objects and assign light object properties, which can be associated with the light objects as metadata. According to this example, GUI 1000 includes a direct light object metadata editor section 1005, which the user can interact with to define the metadata properties of direct light objects; and an indirect light object metadata editor section 1010, which the user can interact with to define the metadata properties of indirect light objects.

[0203] Users can interact with the direct light object metadata editor section 1005 to select the position, size, and other properties of direct light objects. In this example, the direct light object metadata editor section 1005 includes a hue-saturation-luminance (HSL) color wheel 1035a, which users can interact with to select the HSL attributes of the selected direct light object. According to this example, the direct light object metadata editor section 1005 represents direct light objects A, B, and C in a three-dimensional space 1031, which represents the playback environment. In this example, direct light objects A, B, and C are viewed along the z-axis from the top of the three-dimensional space 1031.

[0204] According to this example, the user has selected a direct light object A and is currently selecting its properties. Since the user has selected direct light object A, the coordinates (X, Y, Z), object extent (E), and corresponding time-automatic channel showing the changes of the HSL value of direct light object A over time are displayed in area 1025 and can be edited. In some examples, with... Figure 10 The time intervals corresponding to the time automation channels shown can be approximately 1 second or longer, for example, 1 second, 2 seconds, 3 seconds, 4 seconds, 5 seconds, etc. In this example, only the x and y dimensions of the three-dimensional space 1031 are shown, but area 1025 of the Direct Light Object Metadata Editor section 1005 still allows the user to indicate the x, y, and z coordinates.

[0205] In this example, the indirect light object metadata editor section 1010 includes an HSL color wheel 1035b, which the user can interact with to select the HSL attribute of the selected indirect light object. According to this example, the indirect light object metadata editor section 1010 also includes an intensity control 1030, which the user can interact with to select the intensity of the selected indirect light object.

[0206] Based on some examples, lighting creation tools or MS content creation tools can allow content creators to set indirect lighting effects associated with video scene boundaries, thereby setting indirect lighting effects for a specific scene. Alternatively or additionally, content creators can choose to modify intensity, hue, saturation, etc., over time. In some examples, lighting creation tools or MS content creation tools can allow content creators to use video color overlay information used during video content creation to determine indirect lighting effect metadata. Based on some examples, the indirect lighting settings in lighting creation tools or MS content creation tools can be used as a color / hue / saturation / intensity overlay on the direct lighting object metadata, causing the direct lighting objects to more closely conform to the indirect lighting characteristics.

[0207] Layered abstract creation In some examples, layers can be created in the scene creation tool or in an MS content creation tool configured to create linear-based content, where individual "objects" can be assigned to layers with characteristics such as color, intensity, shape, and position. In some such examples, positions can be specified only for dynamic objects, while sub-regions can be used for environment layer objects.

[0208] Based on some examples, MS content such as light-based content can be created for 3D worlds. Some such examples allow for event-based triggering, such as linking events to existing light metadata and triggering new scene creation.

[0209] Based on some examples, scene creation tools or MS content creation tools can allow blending between layers or objects. For example, it might be desirable to overlay and blend layers / objects with the same priority, while objects with higher priority should occlude all other objects. Scene creation tools or MS content creation tools can allow content creators to define blending rules that correspond to their intentions.

[0210] Rendering of light layer data Layer attributes and their metadata can be obtained using a scene renderer (such as...) Figure 5 A lighting renderer 501 is used for rendering. In some examples, the lighting renderer can be configured to send lighting control signals 515 to the lighting fixtures in the playback environment. In some examples, the lighting renderer 501 can be configured to output the lighting control signals 515 to lighting controllers 103, which are configured to control lighting fixtures 108. According to some examples, the lighting renderer uses environment and lighting fixture data 104 to determine the capabilities and spatial location of each lighting fixture. In some examples, layer priority can be a determining factor in the final content rendered to the lighting fixtures.

[0211] Rendering direct and indirect light objects In some implementations, the environmental and lighting data 104 received by the lighting renderer includes data about whether the lighting fixture is a direct or indirect light source from the viewing position. In some examples, if no indirect lighting fixture is available, the indirect light data may instead be sent to the direct lighting fixture, potentially with reduced brightness.

[0212] Direct light objects are preferably rendered to visible light fixtures, such as ceiling downlights, wall-mounted lights, table lamps, etc. Indirect light metadata is ideally targeted at light fixtures that are not directly visible, such as LED light strips illuminating walls, ceilings, shelves, and furniture, and spotlights illuminating walls or ceilings. If no such indirect light is available, indirect light metadata can be used instead to control direct light. In some such examples, the rendering engine can overlay direct and indirect light object metadata when rendering to light fixtures that function as both indirect and direct light sources.

[0213] Figure 11 This is a flowchart outlining an example of a method that can be performed by a device or system such as the apparatus or system disclosed herein. As with other methods described herein, the blocks of method 1100 need not be performed in the indicated order. In some embodiments, one or more blocks of method 1100 may be performed simultaneously. Furthermore, some embodiments of method 1100 may include more or fewer blocks than those shown and / or described. The blocks of method 1100 may be performed by one or more devices, which may be (or may include) a control system (such as those described above). Figure 1 One or more instances of the control system 1100 shown and described. For example, at least some aspects of method 1100 can be implemented by a system configured to perform... Figure 4 MS Renderer 001 Figure 5 Scene renderer 501 and / or Figure 9A The instance execution of the control system 110 of the scene renderer 501.

[0214] In this example, box 1105 relates to a control system configured to implement a sensory renderer receiving one or more sensory objects and corresponding sensory object metadata, the one or more sensory objects and corresponding sensory object metadata indicating the expected sensory effects to be provided in a sensory actuator playback environment. In some examples, the environment can be an actual real-world environment, such as a room environment or a car environment. According to some examples, the environment can be or may include a virtual environment. In some such examples, method 1100 may relate to providing a virtual environment (such as a game environment) while simultaneously providing corresponding lighting effects in a real-world environment.

[0215] According to this example, box 1110 relates to receiving playback environment information by a control system. In this example, the playback environment information includes sensory actuator positioning information and sensory actuator characteristic information, which relate to one or more controllable sensory actuators in the sensory actuator playback environment. For example, box 1110 may relate to receiving information referenced herein. Figure 4 The described environment and actuator data 004, or receive the references in this document. Figure 5 or Figure 9A The described environment and lighting data is 104.

[0216] In this example, box 1115 relates to the control system determining a sensory actuator control command or sensory actuator control signal based on playback environment information, sensory objects, and sensory object metadata, for controlling one or more controllable sensory actuators in the sensory actuator playback environment. Here, box 1120 relates to the control system outputting a sensory actuator control command or sensory actuator control signal for at least one of the controllable sensory actuators in the sensory actuator playback environment.

[0217] In some examples, one or more sensory objects may include one or more light objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or combinations thereof. Based on some examples of sensory object metadata including light object and lighting metadata, lighting metadata may include direct light object metadata, indirect light metadata, or combinations thereof. Alternatively or additionally, lighting metadata may be organized into one or more layers, which may include one or more environment layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or combinations thereof.

[0218] According to some examples, sensory object metadata includes time information. In some such examples, box 1115 may involve determining one or more drive levels for one or more time intervals corresponding to the time information.

[0219] In some examples, method 1100 may involve receiving viewing location information. In some such examples, box 1115 may involve determining one or more drive levels corresponding to the viewing location information.

[0220] In some examples, lighting information may include one or more characteristics of one or more base light sources in the local lighting environment that are not controllable by the control system. In some such examples, box 1115 may involve determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.

[0221] According to some examples, sensory object metadata may include sensory object size information. In some examples, each of the sensory objects may have a sensory object effect characteristic, which indicates the type of effect the sensory object is providing. According to some examples, one or more of these sensory objects may have a persistence characteristic, which indicates that the sensory object will persist for a period of time in the sensory actuator playback environment.

[0222] In some examples, one or more of these sensory objects may be assigned to one or more layers, where the sensory objects are grouped based on shared sensory object characteristics. In some such examples, one or more layers may group sensory objects based on atmosphere, environment, information, attention, color, intensity, size, shape, location, spatial zone, or a combination thereof.

[0223] According to some examples, at least some of these sensory objects may have a priority characteristic, which indicates the relative importance of each sensory object. In some examples, one or more of these sensory objects may have a spatial imaging characteristic, which indicates how the sensory objects can move within the sensory actuator playback environment, how the sensory objects will affect the controllable sensory actuators or combinations thereof within the sensory actuator playback environment. According to some examples, at least some of these sensory objects may have a hybrid mode characteristic, which indicates how multiple sensory objects can be reproduced by a single controllable sensory actuator.

[0224] Some examples of method 1100 may involve providing and / or processing more general metadata for the entire multi-sensory content file, rather than (or in addition to) metadata for each object. This more general metadata may be referred to as “overall sensory object metadata.” Some examples of method 1100 may involve receiving overall sensory object metadata, including trim control information, mastering environment information, or a combination thereof, by a control system. In some such examples, box 1115 may involve determining sensory actuator control commands or sensory actuator control signals based at least in part on trim control information, mastering environment information, or a combination thereof.

[0225] Bitstream including multi-sensory objects This section discloses various types of encoded bitstreams for carrying object-based multisensory data for rendering on any number of actuators. Throughout this document, such bitstreams may be referred to as containing encoded object-based sensory data, or as containing encoded object-based sensory data streams. Some encoded object-based sensory data streams may be delivered along with other media content, or as part of other media content. In some such examples, object-based sensory data streams may be interleaved or multiplexed with audio and / or video bitstreams. According to some examples, object-based sensory data streams may be arranged according to the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) to provide encoded object-based sensory data along with corresponding audio data, video media, or both, as encoded ISOBMFF bitstreams. Therefore, in some examples, the encoded bitstream may include encoded object-based sensory data streams, encoded audio data streams, and / or encoded video data streams. In some examples, the encoded object-based sensory data stream and (multiple) other associated data streams include associated synchronization data (such as timestamps) to allow synchronization between different types of content. For example, if the encoded bitstream includes encoded audio data, encoded video data, and encoded object-based sensory data, then the encoded audio data, encoded video data, and encoded object-based sensory data may each include associated synchronization data.

[0226] Just as there are channel-based surround sound formats (such as Dolby Digital (AC3) and Dolby Digital Plus) and object-based surround sound formats (such as Dolby Atmos (DD+ AJOC and AC4-JOC)), an object-based sensory data format is introduced here. These concepts are summarized in the table below:

[0227] Here are some examples of artistic intentions that can be conveyed through encoded, object-based sensory data streams: 1. Now turn all the lights in the room to bright white.

[0228] 2. At the demonstration timestamp of 1:15, turn off all the lights in the room.

[0229] 3. Over the next 10 seconds, the light at the back of the room (behind the audience seats) slowly turned red.

[0230] 4. Simulate a helicopter with searchlights flying over the audience, and in the next 15 seconds, turn on and off any available overhead lights in the room in sequence from front to back.

[0231] 5. Now, an orange light is emanating from the front left corner of the room.

[0232] 6. An airflow of 5 sections is generated from the front right corner of the room.

[0233] Readers will notice that these examples do not require knowledge of a specific set of actuators present in the playback environment or the location of these actuators during content creation. In contrast, MS renderers (such as...) Figure 4 The MS renderer 001 will be configured to control a specific set of actuators in a particular playback environment based on: (a) general instructions present in the object's sensory data 005; and (b) information in the specific environment and actuator data. In some examples (such as...) Figure 4 As shown), object-based sensory data 005 can be provided by experience player 002 to MS renderer 001, which may include bitstream decoder configured to extract object-based sensory data 005 from bitstream that also includes encoded audio data and / or video data.

[0234] Linear audio / video media content is typically packaged in a container format in the form of frames, where each frame contains information needed to render the media during a specific duration of the content (e.g., a 60 ms time period from 10 minutes 3.2 seconds to 10 minutes 3.8 seconds relative to the start of the content).

[0235] Some format streams can include separate streams for each modality. For example, one element stream may contain video information encoded using High Efficiency Video Coding (HEVC) (also known as H.265), another element stream may contain audio information encoded using Advanced Audio Coding (AAC) or Dolby AC4, and a third element stream may contain closed caption (subtitle) information. In some cases, multiple element streams can be selected or combined during rendering, such as audio tracks in different languages, director's commentary (which can be selectively mixed with one or more other audio tracks during playback), closed captions in multiple languages, etc.

[0236] This disclosure extends and summarizes prior bitstream encoding and decoding methods to include multiple element streams that convey multi-sensory information that is not channel-based (such as object-based or spherical harmonic function-based), suitable for demultiplexing, frame reassembly, and rendering using multiple actuators, and in some cases, for synchronous rendering with the audio and / or video modalities of the media stream.

[0237] Figure 12 Example components of a system for creating and playing multisensory (MS) experiences are shown. (This is in contrast to other systems provided in this article.) Figure 1 Sample, Figure 12The types and quantities of elements shown are provided by way of example only. Other implementations may include more, fewer, and / or different types and quantities of elements. According to some examples, system 1200 may be or may include one or more devices configured to perform at least some of the methods disclosed herein. In some examples, system 1200 may include devices configured to perform at least some of the methods disclosed herein. Figure 1 One or more instances of the control system 110. In this example, system 1200 includes a reference... Figure 4 Instances of some elements described.

[0238] In this example, system 1200 includes the following elements: 1200: A system configured to receive and process encoded bitstreams comprising multiple data frames, including encoded audio data, encoded video data, and encoded object-based sensory data; 1201: Encoded bitstream. In some examples, the data frames of the encoded bitstream may be arranged according to the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) "container" and / or encoded according to the Moving Picture Experts Group (MPEG) standard; 1201A-C: Multiplexed packet / part / frame sequence of encoded bitstream 1201; 1202A-B: Elements of the audio data stream. In some examples, the audio data stream may be encoded according to Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4, or Dolby Atmos codecs; 1203A-B: Elements of a video data stream. In some examples, the video data stream may be encoded according to Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Universal Video Coding (VVC), or AOMedia Video 1 (AV1) codec; 1204A-B: Elements of an encoded sensory data stream, in this example, an object-based sensory data stream. In this example, the encoded object-based sensory data includes sensory objects and corresponding sensory object metadata, which indicate the expected sensory effect to be provided via sensory actuators in a playback environment. In some implementations, each modality may have an object-based sensory data stream (also referred to as an "element stream") (e.g., an object-based lighting data stream, an object-based temperature data stream, an object-based airflow data stream, etc.). In other implementations, these modalities may be combined into a single encoded sensory data stream. In this example, the encoded audio data stream, the encoded video data stream, and the encoded object-based sensory data stream all include associated synchronization data, such as timestamps; 002: Figure 4 Examples of the 002 player experience include: 1206: Bitstream demultiplexer; 1207: Audio stream decoder; 1208: Video stream decoder; and 1209: Multi-sensory data stream decoder; 001: Multi-sensory renderer—uses environmental and actuator data 004 and a decoded multi-sensory stream to generate control signals to drive multiple actuators 008. In this example, the MS renderer 001 includes a reference... Figure 4 The description describes the functionality of the MS controller API 003. According to this example, the MS renderer 001 can determine how to synchronize the playback of audio, video, and sensory data based on the synchronization data in the encoded bitstream 1201; 004: Environmental and Actuator Data – Includes information about the physical location of actuators in the environment, visibility / area of ​​influence information, and descriptions of how viewers perceive each actuator (e.g., whether and how each luminaire is visible to (multiple) viewers). 1213A, 1213B, 1213C...: Multiple independent control flows for each actuator 008; 008: Multiple actuators under the control of the multi-sensory renderer 001; 008A and 008B: Smart lights; 008C: Intelligent RGB Light Emitting Diode (LED) Light Strip; and 008D: Other actuators in the playback environment.

[0239] Example 1: One element flow per mode In some examples, the multisensory data stream can be transmitted as multiple element streams (e.g., one element stream for lighting information, one element stream for airflow information, one element stream for tactile information, one element stream for temperature information, etc.). According to some implementations, multiple versions of one or more multisensory data streams may exist within the encoded bitstream 1201 to allow selection based on user preferences or user requirements (e.g., providing a default or standard lighting data stream for a typical viewer, and a separate lighting data stream containing more subtle lighting information, designed to be safe for photosensitive viewers).

[0240] Example 2: Combining Multi-Sensory Element Flow In some alternative examples, all multisensory modalities (e.g., lighting, airflow, touch, temperature) can be combined into a unified multisensory data stream.

[0241] Example 3: Combine multi-sensory information with existing element flows, or arrange it in an existing container format.

[0242] In another alternative embodiment, the multisensory data stream can be embedded within one of existing element streams. For example, some audio streaming formats may include the ability to encapsulate streaming synchronization metadata. In some examples, the multisensory data stream can be embedded within existing audio metadata data transmission mechanisms, such as within fields or sequences of fields reserved for audio metadata. In some alternative examples, existing audio, video, or container formats can be modified or adapted to allow the inclusion of synchronized multisensory information.

[0243] As described above, in some examples, the data frames of the encoded bitstream can be arranged according to an International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF) file or "container". According to some implementations, encoded object-based sensory data can reside in a timestamped metadata track, as defined in, for example, section 12.9 of the ISO / IEC 14496-12:2022 standard, which is incorporated herein by reference. According to some such examples, encoded audio data can reside in the audio track of an ISOBMFF file, and / or encoded video data can reside in the video track. Such examples have several potential advantages, including but not limited to the following: • The same timestamped metadata track can be associated with more than one track. In other words, a timestamped metadata track corresponding to encoded object-based sensory data may be unrelated to the content of the associated audio / video track; • It makes it easier to append timestamped metadata tracks to files; and • The duration of the timestamped metadata sample does not need to match the duration of the associated audio and / or video data.

[0244] Figure 13 This is a flowchart outlining an example of a method that can be performed by a device or system such as the apparatus or system disclosed herein. As with other methods described herein, the blocks of method 1300 need not be performed in the indicated order. In some embodiments, one or more blocks of method 1300 may be performed simultaneously. Furthermore, some embodiments of method 1300 may include more or fewer blocks than those shown and / or described. The blocks of method 1300 may be performed by one or more devices, which may be (or may include) a control system (such as those described above). Figure 1 One or more instances of the control system 110 shown and described. For example, at least some aspects of method 1300 can be implemented by a system configured to perform... Figure 4 or Figure 12 The experience player 002's control system 110 instance execution.

[0245] In this example, box 1305 relates to a control system configured to implement a demultiplexing module receiving an encoded bitstream comprising multiple data frames. For example, the demultiplexing module could be... Figure 12 An example of demultiplexer 1206 is provided. In this example, the data frame includes encoded audio data, encoded video data, and encoded object-based sensory data. According to this example, the encoded object-based sensory data includes sensory objects and corresponding sensory object metadata, which indicate the expected sensory effects to be provided in the sensory actuator playback environment. In this example, the encoded audio data stream, the encoded video data stream, and the encoded object-based sensory data stream all include associated synchronization data.

[0246] According to this example, box 1310 relates to the extraction of encoded audio data streams, encoded video data streams, and encoded object-based sensory data streams from an encoded bitstream by a control system. For example, box 1310 may relate to parsing and / or demultiplexing references. Figure 12 The described encoded bitstream is 1201.

[0247] In this example, box 1315 relates to the control system providing an encoded audio data stream to the audio decoder. For example, box 1315 could involve... Figure 12 The demultiplexer 1206 provides an encoded audio data stream to the audio decoder 1207.

[0248] Here, block 1320 relates to the control system providing an encoded video data stream to a video decoder. For example, block 1320 could relate to a demultiplexer 1206 providing an encoded video data stream to a video decoder 1208.

[0249] According to this example, box 1325 relates to providing an encoded object-based sensory data stream from a control system to a sensory data decoder. For example, box 1325 could relate to a demultiplexer 1206 providing an object-based sensory data stream to a multisensory data stream decoder 1209.

[0250] In some examples, sensory object metadata may include sensory object location information, sensory object size information, or both. According to some examples, each of the sensory objects may have a sensory object effect characteristic indicating the type of effect the sensory object is providing. In some examples, one or more of these sensory objects may have a persistence characteristic indicating that the sensory object will persist for a period of time in the sensory actuator playback environment. According to some examples, one or more of these sensory objects may be assigned to one or more layers, in which sensory objects are grouped based on common sensory object characteristics. In some examples, one or more layers may group sensory objects based on one or more of atmosphere, environment, information, attention, color, intensity, size, shape, location, spatial zone, or combinations thereof. According to some examples, at least some of these sensory objects may have a priority characteristic indicating the relative importance of each sensory object.

[0251] Based on some examples, object-based sensory data streams may include light objects, tactile objects, airflow objects, position actuator objects, or combinations thereof. Based on some examples of sensory object metadata including light objects and lighting metadata, lighting metadata may include direct light object metadata, indirect light metadata, or combinations thereof. Alternatively or additionally, lighting metadata may also be organized into one or more layers, which may include one or more environment layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or combinations thereof.

[0252] In some examples, method 1300 may involve decoding an encoded object-based sensory data stream and providing the decoded object-based sensory data stream to a sensory data renderer, the decoded object-based sensory data stream including associated sensory synchronization data. In some such examples, method 1300 may involve receiving the decoded object-based sensory data stream by the sensory data renderer and receiving playback environment information by the sensory data renderer. The playback environment information may be an instance of the environment and actuator data 004 described herein. Thus, the playback environment information may include sensory actuator positioning information and sensory actuator characteristic information relating to one or more controllable sensory actuators in the sensory actuator playback environment.

[0253] According to some such examples, method 1300 may involve a sensory data renderer determining sensory object metadata and associated sensory synchronization data from a decoded object-based sensory data stream, at least in part based on (a) the sensory object; and determining a sensory actuator control command or sensory actuator control signal, at least in part based on (b) playback environment information, for controlling one or more controllable sensory actuators in the sensory actuator playback environment. According to some such examples, method 1300 may involve a sensory actuator control command or sensory actuator control signal output by the sensory data renderer for at least one of the controllable sensory actuators in the sensory actuator playback environment. For example, it may be directed to a reference... Figure 4 The described MS controller API 003 provides sensory actuator control commands. For example, sensory actuator control signals can be provided to actuator 008 of the playback environment.

[0254] In some examples, each of the multiple data frames may include encoded audio data subframes, encoded video data subframes, and encoded object-based sensory data subframes. According to some examples, the bitstream data frames may be encoded according to the Moving Picture Experts Group (MPEG) standard. In some examples, the bitstream data frames may be arranged according to the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF). According to some examples, the encoded object-based sensory data may reside in a timestamped metadata track.

[0255] Based on some examples, audio data streams can be encoded using Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4, or Dolby Atmos codecs. In some examples, video data streams can be encoded using Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Universal Video Coding (VVC), or AOMedia Video 1 (AV1) codecs.

[0256] The various features and aspects will be understood from the following enumerated example embodiments (EEE): EEE 1. A method for rendering a desired lighting environment, comprising: The control system, configured to implement a lighting environment renderer, receives object-based light data indicating the expected lighting environment, the object-based light data including light objects and lighting metadata; The control system receives lighting information about the local lighting environment, wherein the lighting information includes one or more characteristics of one or more controllable light sources in the local lighting environment; The control system determines the drive level of each of the one or more controllable light sources to approximate the desired lighting environment; and The control system outputs the drive level to at least one of the controllable light sources.

[0257] EEE 2. As described in EEE 1, wherein: The object-based lighting metadata includes time information; and The determination involves determining one or more drive levels for one or more time intervals corresponding to the time information.

[0258] EEE 3. The method as described in EEE 1 or EEE 2, further comprising receiving viewing location information, wherein the determination involves determining one or more drive levels corresponding to the viewing location information.

[0259] EEE 4. The method as described in any one of EEE 1 to 3, wherein: The lighting information includes one or more characteristics of one or more basic light sources in the local lighting environment, which are not controlled by the control system; and The determination involves determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.

[0260] EEE 5. The method of any one of EEE 1 to 4, wherein the object-based lighting metadata indicates at least lighting object location information and lighting object color information.

[0261] EEE 6. The method of any one of EEE 1 to 5, wherein the intended lighting environment is transmitted as one or more image-based lighting (IBL) objects.

[0262] EEE 7. The method as described in EEE 6, wherein the determination is based at least in part on the local lighting environment. n One controllable light source (IBL) nIBL mapping for each ambient light produced at maximum intensity in ).

[0263] EEE 8. The method as described in EEE 7, wherein the determination is based at least in part on basic ambient lighting (IBL) that is not controllable by the control system. base The basic IBL mapping.

[0264] EEE 9. The method as described in EEE 8, wherein the determination is based at least in part on each dynamic lighting element IBL. n Linear scaling (Scale) n ), so that each scaled IBL n The sum of light plus the basic illumination IBL base The closest match to the expected lighting environment (IBL) ref .

[0265] EEE 10. The method as described in EEE 9, wherein the determination is based at least in part on minimizing IBL. ref With each scaled IBL n The sum of light plus basic lighting IBL base The differences between them.

[0266] EEE 11. An apparatus configured to perform an audio processing method as described in any one of EEE 1 to 10.

[0267] EEE 12. A system configured to perform an audio processing method as described in any one of EEE 1 to 10.

[0268] EEE 13. One or more non-transitory computer-readable media having instructions stored thereon for controlling one or more devices to perform a method as described in any one of EEE 1 to 10.

[0269] EEE 14. A method for providing an intended sensory experience, the method comprising: A control system configured to implement a sensor renderer receives one or more sensory objects and corresponding sensory object metadata, the one or more sensory objects and corresponding sensory object metadata indicating the expected sensory effects to be provided in the sensory actuator playback environment; The control system receives playback environment information, wherein the playback environment information includes sensory actuator positioning information and sensory actuator feature information, and the sensory actuator positioning information and sensory actuator feature information are related to one or more controllable sensory actuators in the sensory actuator playback environment; The control system determines sensory actuator control commands or sensory actuator control signals based on the playback environment information, the sensory object, and the sensory object metadata, for controlling one or more controllable sensory actuators in the playback environment; and The control system outputs a sensory actuator control command or sensory actuator control signal for at least one of the controllable sensory actuators in the sensory actuator playback environment.

[0270] EEE 15. The method as described in EEE 14, wherein the one or more sensory objects include one or more light objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or combinations thereof.

[0271] EEE 16. The method as described in EEE 14 or EEE 15, wherein the sensory object metadata includes light object and lighting metadata, and wherein the lighting metadata includes direct light object metadata, indirect light metadata, or a combination thereof.

[0272] EEE 17. The method as described in EEE 16, wherein the lighting metadata is organized into one or more layers, the layers including one or more environment layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.

[0273] EEE 18. The method of any one of EEE 14 to 17, wherein the sensory object metadata includes sensory object location information.

[0274] EEE 19. The method of any one of EEE 14 to 18, wherein the sensory object metadata includes sensory object size information.

[0275] EEE 20. The method of any one of EEE 14 to 19, wherein each of the sensory objects has a sensory object effect characteristic indicating the type of effect being provided by the sensory object.

[0276] EEE 21. The method of any one of EEE 14 to 20, wherein one or more of the sensory objects have a persistence characteristic indicating that the sensory object will persist for a period of time in the sensory actuator playback environment.

[0277] EEE 22. The method of any one of EEE 14 to 21, wherein one or more of the sensory objects are assigned to one or more layers, in which the sensory objects are grouped according to shared sensory object characteristics.

[0278] EEE 23. The method as described in EEE 22, wherein the one or more layers group sensory objects based on atmosphere, environment, information, attention, color, intensity, size, shape, location, spatial zone, or a combination thereof.

[0279] EEE 24. The method of any one of EEE 14 to 23, wherein at least some of the sensory objects have a priority characteristic indicating the relative importance of each sensory object.

[0280] EEE 25. The method of any one of EEE 14 to 24, wherein one or more of the sensory objects have spatial acoustic characteristics indicating how the sensory objects can move in the sensory actuator playback environment and how the sensory objects will affect the controllable sensory actuator or combination thereof in the sensory actuator playback environment.

[0281] EEE 26. The method of any one of EEE 14 to 25, wherein at least some of the sensory objects have a mixed-mode characteristic indicating how multiple sensory objects can be reproduced by a single controllable sensory actuator.

[0282] EEE 27. The method of any one of EEE 14 to 26 further includes receiving, by the control system, overall sensory object metadata including trimming control information, mastering environment information, or a combination thereof.

[0283] EEE 28. An apparatus configured to perform an audio processing method as described in any one of EEE 14 to 27.

[0284] EEE 29. A system configured to perform an audio processing method as described in any one of EEE 14 to 27.

[0285] EEE 30. One or more non-transitory computer-readable media having instructions stored thereon for controlling one or more devices to perform a method as described in any one of EEE 14 to 27.

[0286] EEE 31. A method for decoding a bitstream, the method comprising: A control system configured to implement a demultiplexing module receives an encoded bitstream comprising multiple data frames, the data frames including encoded audio data, encoded video data, and encoded object-based sensory data, the encoded object-based sensory data including sensory objects and corresponding sensory object metadata, the sensory objects and corresponding sensory object metadata indicating the expected sensory effect to be provided in the sensory actuator playback environment, and the encoded audio data stream, the encoded video data stream, and the encoded object-based sensory data stream all including associated synchronization data; The control system extracts encoded audio data streams, encoded video data streams, and encoded object-based sensory data streams from the encoded bitstream; The encoded audio data stream is provided to the audio decoder by the control system. The encoded video data stream is provided by the control system to the video decoder; and The encoded object-based sensory data stream is provided by the control system to the sensory data decoder.

[0287] EEE 32. The method of EEE 31 further includes decoding the encoded object-based sensory data stream and providing the decoded object-based sensory data stream to a sensory data renderer, the decoded object-based sensory data stream including associated sensory synchronization data.

[0288] EEE 33. The method as described in EEE 32, further comprising: The decoded object-based sensory data stream is received by the sensory data renderer; The sensory data renderer receives playback environment information, wherein the playback environment information includes sensory actuator positioning information and sensory actuator feature information, and the sensory actuator positioning information and sensory actuator feature information are related to one or more controllable sensory actuators in the sensory actuator playback environment; The sensory data renderer determines, at least in part, the sensory object metadata and the associated sensory synchronization data from the decoded object-based sensory data stream based on (a) the sensory object; and at least in part, based on (b) the playback environment information, a sensory actuator control command or sensory actuator control signal for controlling one or more controllable sensory actuators in the sensory actuator playback environment; and The sensory data renderer outputs sensory actuator control commands or sensory actuator control signals for at least one of the controllable sensory actuators in the sensory actuator playback environment.

[0289] EEE 34. The method of any one of EEE 31 to 33, wherein each of the plurality of data frames includes an encoded audio data subframe, an encoded video data subframe, and an encoded object-based sensory data subframe.

[0290] EEE 35. The method of any one of EEE 31 to 34, wherein the data frames of the bitstream are encoded according to the Moving Picture Experts Group (MPEG) standard.

[0291] EEE 36. The method of any one of EEE 31 to 35, wherein the data frames of the bitstream are arranged in accordance with the International Organization for Standardization (ISO) Basic Media File Format (ISOBMFF).

[0292] EEE 37. The method as described in EEE 35 or EEE 36, wherein the encoded object-based sensory data resides in a timestamped metadata track.

[0293] EEE 38. The method of any one of EEE 31 to 37, wherein the sensory object comprises one or more illumination objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or combinations thereof.

[0294] EEE 39. The method of any one of EEE 31 to 38, wherein the sensory object metadata includes lighting metadata, and wherein the lighting metadata includes direct light object metadata, indirect light metadata, or a combination thereof.

[0295] EEE 40. The method as described in EEE 39, wherein the lighting metadata is organized into one or more layers, the layers including one or more environment layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.

[0296] EEE 41. The method of any one of EEE 31 to 40, wherein the sensory object metadata includes sensory object location information, sensory object size information, or both.

[0297] EEE 42. The method of any one of EEE 31 to 41, wherein each of the sensory objects has a sensory object effect characteristic indicating the type of effect being provided by the sensory object.

[0298] EEE 43. The method of any one of EEE 31 to 42, wherein one or more of the sensory objects have a persistence characteristic indicating that the sensory object will persist for a period of time in the sensory actuator playback environment.

[0299] EEE 44. The method of any one of EEE 31 to 43, wherein one or more of the sensory objects are assigned to one or more layers, in which the sensory objects are grouped according to common sensory object characteristics.

[0300] EEE 45. The method as described in EEE 44, wherein the one or more layers group sensory objects based on one or more of atmosphere, environment, information, attention, color, intensity, size, shape, location, spatial zone, or combinations thereof.

[0301] EEE 46. The method of any one of EEE 31 to 45, wherein at least some of the sensory objects have a priority characteristic indicating the relative importance of each sensory object.

[0302] EEE 47. The method of any one of EEE 31 to 46, wherein the audio data stream is encoded according to an Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4, or Dolby Atmos codec.

[0303] EEE 48. The method of any one of EEE 31 to 46, wherein the video data stream is encoded according to an Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Universal Video Coding (VVC), or AOMedia Video 1 (AV1) codec.

[0304] EEE 49. An apparatus configured to perform an audio processing method as described in any one of EEE 31 to 48.

[0305] EEE 50. A system configured to perform an audio processing method as described in any one of EEE 31 to 48.

[0306] EEE 51. One or more non-transitory computer-readable media having instructions stored thereon for controlling one or more devices to perform a method as described in any one of EEE 31 to 48.

[0307] The foregoing description illustrates various embodiments of this disclosure and examples of how aspects of this disclosure may be implemented. The foregoing examples and embodiments should not be considered as limited embodiments, but are presented to illustrate the flexibility and advantages of this disclosure as defined by the appended claims. Other arrangements, embodiments, implementations, and equivalents will be apparent to those skilled in the art based on the foregoing disclosure and the appended claims, and may be employed without departing from the spirit and scope of this disclosure as defined by the claims.

Claims

1. A method comprising: The control system receives a content bitstream including encoded object-based sensory data, the encoded object-based sensory data including one or more sensory objects and corresponding sensory metadata, the encoded object-based sensory data corresponding to sensory effects, including lighting, touch, airflow, one or more position actuators or combinations thereof, the sensory effects being provided by one or more sensory actuators in the environment; The control system extracts the object-based sensory data from the content bitstream; as well as The control system provides the object-based sensory data to the sensory renderer.

2. The method as described in claim 1, wherein, The object-based sensory metadata includes sensory spatial metadata, which at least indicates the spatial location for rendering the object-based sensory data within the environment, the region for rendering the object-based sensory data within the environment, or a combination thereof.

3. The method as claimed in claim 1 or claim 2, wherein, The object-based sensory data does not correspond to a specific sensory actuator in the environment.

4. The method according to any one of claims 1 to 3, wherein, The object-based sensory data includes abstract sensory reproduction information that allows the sensory renderer to reproduce one or more created sensory effects via one or more sensory actuator types, via a variety of sensory actuators, and from various sensory actuator locations in the environment.

5. The method according to any one of claims 1 to 4, further comprising: The object-based sensory data is received by the sensory renderer; The sensor renderer receives environmental descriptor data corresponding to the location of one or more sensor actuators in the environment; The sensor renderer receives actuator descriptor data corresponding to the characteristics of the sensor actuators in the environment; as well as The sensory renderer provides one or more actuator control signals to control the sensory actuators in the environment to produce one or more sensory effects indicated by the object-based sensory data.

6. The method of claim 5, further comprising providing the one or more sensory effects by the one or more sensory actuators in the environment.

7. The method according to any one of claims 1 to 6, wherein, The content bitstream also includes one or more encoded audio objects synchronized with the encoded object-based sensory data, each audio object comprising one or more audio signals and corresponding audio object metadata, and the method further includes: The control system extracts audio objects from the content bitstream; and The control system provides the audio object to the audio renderer.

8. The method of claim 7, wherein, The audio object metadata includes at least audio object space metadata, which indicates the audio object space location used to render the one or more audio signals within the environment.

9. The method of claim 7 or claim 8, further comprising: The audio renderer receives the one or more audio objects; The audio renderer receives loudspeaker data corresponding to one or more loudspeakers in the environment; as well as The audio renderer provides one or more amplifier control signals to control the one or more amplifiers in the environment to play audio corresponding to the one or more audio objects and synchronized with the one or more sensory effects.

10. The method of claim 9, further comprising playing the audio corresponding to the one or more audio objects by the one or more loudspeakers in the environment.

11. The method of claim 9 or claim 10, wherein, The content bitstream includes encoded video data synchronized with the encoded audio object and the encoded object-based sensory data, and wherein the method further includes: The control system extracts video data from the content bitstream; The control system provides the video data to the video renderer; The video data is received by the video renderer; and The video renderer provides one or more video control signals to control one or more display devices in the environment to present one or more images corresponding to the one or more video control signals and synchronized with the one or more audio objects and the one or more sensory effects.

12. The method of claim 11, further comprising presenting the one or more images corresponding to the one or more video control signals by the one or more display devices in the environment.

13. The method according to any one of claims 1 to 12, wherein, The environment described is a virtual environment.

14. The method according to any one of claims 1 to 11, wherein, The environment described is the physical reality environment.

15. The method of claim 14, wherein, The environment referred to is either a room environment or a vehicle environment.

16. An apparatus configured to perform the method as claimed in any one of claims 1 to 15.

17. A system configured to perform the method as claimed in any one of claims 1 to 15.

18. One or more non-transitory computer-readable media having instructions stored thereon for controlling one or more devices to perform the method as described in any one of claims 1 to 15.