Providing object-based, multi-sensory experiences.
Patent Information
- Application Number
- JP2026502209
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-10
- Filing Date
- 2024-07-15
- Publication Date
- 2026-09-01
Smart Images

Figure 2026529511000001_ABST
Abstract
Description
[Technical Field]
[0001] [Cross-reference of related applications] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 669,232 filed on 10 July 2024, U.S. Provisional Patent Application No. 63 / 514,107 filed on 17 July 2023, U.S. Provisional Patent Application No. 63 / 514,096 filed on 17 July 2023, and U.S. Provisional Patent Application No. 63 / 514,094 filed on 17 July 2023, each of which is incorporated herein by reference in its entirety.
[0002] [Technical field] This disclosure relates to providing multisensory experiences, some of which include lighting (light)-based experiences. [Background technology]
[0003] Unless otherwise indicated herein, the methods described in this section are not prior art to the claims of this application and will not be recognized as prior art by their inclusion in this section.
[0004] Media content distribution generally focuses on audio and video experiences. The distribution of multisensory content is limited due to the bespoke nature of its operation. For example, lighting fixtures are widely used as artistic and functional expressions for concerts. However, each installation is specifically designed for a particular set of lighting fixtures. It is generally not feasible for a system to provide lighting designs beyond the set of installations for which it was designed. Other systems attempting to communicate lighting experiences more broadly do so easily by algorithmically extending the visual aspects of a screen, but these are not specifically authored. Haptic content is designed for specific haptic devices. If a different device is used, such as a game controller, mobile phone, or a haptic device of a different brand, there is no way to translate the creative intent of the content to a different actuator. [Overview of the Initiative]
[0005] At least some aspects of this disclosure may be implemented by methods such as audio processing methods. In some examples, such methods may be implemented at least in part by a control system such as those disclosed herein. Some methods may include the step of the control system receiving a content bitstream containing encoded object-based sensory data. The encoded object-based sensory data may include one or more sensory objects and corresponding sensory metadata. The encoded object-based sensory data may correspond to sensory effects provided by one or more sensory actuators in the environment. The sensory effects may include lighting, touch, airflow, one or more position actuators, or a combination thereof. Some methods may include the step of the control system extracting object-based sensory data from the content bitstream, and the step of the control system providing the object-based sensory data to a sensory renderer.
[0006] In some examples, object-based sensory metadata may include sensory spatial metadata indicating at least the spatial location for rendering object-based sensory data in the environment, the region for rendering object-based sensory data in the environment, or a combination thereof. According to some examples, object-based sensory data does not have to correspond to a specific sensory actuator in the environment. In some examples, object-based sensory data may include abstracted sensory reproduction information that enables a sensory renderer to reproduce one or more authored sensory effects from various sensory actuator locations in the environment via one or more sensory actuator types and via a variety of sensory actuators.
[0007] Some methods may include the step of having a sensory renderer receive object-based sensory data and environment descriptor data corresponding to one or more locations of sensory actuators in the environment. Some methods may include the step of having a sensory renderer receive actuator descriptor data corresponding to properties of sensory actuators in the environment. Some methods may include the step of having a sensory renderer provide one or more actuator control signals to control sensory actuators in the environment to produce one or more sensory effects indicated by object-based sensory data. Some methods may include the step of having one or more sensory actuators in the environment provide one or more sensory effects.
[0008] In some examples, the content bitstream may also include one or more encoded audio objects synchronized with encoded object-based sensory data. The audio objects may include one or more audio signals and corresponding audio object metadata. Some such methods may include the steps of: extracting audio objects from the content bitstream by a control system; and providing the audio objects to an audio renderer by the control system. In some examples, the audio object metadata may include at least audio object spatial metadata indicating the audio object spatial location for rendering one or more audio signals in the environment. Some methods may include the steps of: receiving one or more audio objects by an audio renderer; receiving loudspeaker data corresponding to one or more loudspeakers in the environment by an audio renderer; and providing one or more loudspeaker control signals by an audio renderer to control one or more loudspeakers in the environment to play audio corresponding to one or more audio objects and synchronized with one or more sensory effects. Some methods may include the step of playing audio corresponding to one or more audio objects by one or more loudspeakers in the environment.
[0009] According to some examples, the content bitstream may include encoded video data synchronized with encoded audio objects and encoded object-based sensory data. Some methods may include the step of extracting video data from the content bitstream by a control system, and the step of providing the video data to a video renderer by the control system. Some methods may include the step of receiving the video data by the video renderer, and the step of providing, by the video renderer, one or more video control signals for controlling one or more display devices in an environment to present one or more images that correspond to the one or more video control signals and are synchronized with one or more audio objects and one or more sensory effects. Some methods may include the step of presenting, by one or more display devices in the environment, one or more images corresponding to the one or more video control signals.
[0010] In some examples, the environment may be a virtual environment. According to some examples, the environment may be a physical real-world environment. For example, the environment may be a room environment or a vehicle environment.
[0011] Some or all of the operations, functions and / or methods described herein may be performed by one or more devices in accordance with instructions (e.g., software) stored in one or more computer-readable non-transitory media. Such non-transitory media may include, but are not limited to, one or more random access memory (RAM) devices, read-only memory (ROM) devices, and may include one or more memory devices such as those described herein. Accordingly, some innovative aspects of the subject matter described in the present disclosure can be implemented in one or more computer-readable non-transitory media having software stored thereon.
[0012] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of at least partially performing the methods disclosed herein. In some implementations, the apparatus may include an interface system and a control system. The control system may include one or more general-purpose single or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or combinations thereof. The control system may be configured to perform part or all of the disclosed methods.
[0013] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Note that relative dimensions of the following drawings may not be drawn to scale. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The disclosed embodiments are described herein, by way of example only, with reference to the accompanying drawings. [Figure 1] FIG. 1 is a block diagram illustrating an example of components of an apparatus capable of implementing various aspects of the present disclosure. [Figure 2] FIG. 2 shows exemplary elements of an endpoint. [Figure 3] FIG. 3 shows an example of actuator elements. [Figure 4] FIG. 4 shows exemplary elements of a system for creation and reproduction of multi-sensory (MS) experiences. [Figure 5]This section illustrates exemplary elements of another system for creating and recreating MS experiences. [Figure 6] Figure 5 shows an example of a graphical user interface (GUI) that can be presented by the display device of the lightscape creation tool. [Figure 7A] Figure 5 shows another example of a graphical user interface (GUI) that can be presented by the display device of the lightscape creation tool. [Figure 7B] This flowchart outlines an example of a method that may be performed by an apparatus or system such as those disclosed herein. [Figure 8A] This example shows how to project the lighting of the observation environment onto a two-dimensional (2D) plane. [Figure 8B] This example shows how to project the lighting of the observation environment onto a two-dimensional (2D) plane. [Figure 8C] This example shows how to project the lighting of the observation environment onto a two-dimensional (2D) plane. [Figure 9A] Exemplary elements of a lightscape renderer are shown. [Figure 9B] This flowchart outlines an example of a method that may be performed by an apparatus or system such as those disclosed herein. [Figure 10] Here is another example of a GUI that may be presented by the display device of a lightscape creation tool. [Figure 11] This flowchart outlines an example of a method that may be performed by an apparatus or system such as those disclosed herein. [Figure 12] This document illustrates exemplary elements of a system for creating and recreating multi-sensory (MS) experiences. [Figure 13] This flowchart outlines an example of a method that may be performed by an apparatus or system such as those disclosed herein. [Modes for carrying out the invention]
[0015] The technologies described herein relate to providing multisensory media content. For illustrative purposes, numerous examples and specific details are provided in the following description to provide a complete understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure as defined by the claims may include some or all of the features in these examples, either individually or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0016] In the following description, various methods, processes, and procedures are detailed. Certain steps may be described in a specific order, but such order is primarily for convenience and clarity. Certain steps may be repeated more than once, may be performed before or after other steps (even if those steps are described in a different order), or may be performed concurrently with other steps. A second step is required to follow a first step only if the first step must be completed before the second step begins. Such situations are specifically noted when they are not evident from the context.
[0017] In this specification, the terms “and,” “or,” and “and / or” are used. Such terms should be read as having an inclusive meaning. For example, “A and B” may mean at least “both A and B,” or “at least both A and B.” As another example, “A or B” may mean at least “at least A,” “at least B,” “both A and B,” or “at least both A and B.” As yet another example, “A and / or B” may mean at least “A and B,” or “A or B.” When an exclusive OR is intended, such an OR is specified (e.g., “either A or B,” or “at least one of A and B”).
[0018] This specification describes various processing functions associated with structures such as blocks, elements, components, circuits, etc. Generally, these structures may be implemented by one or more processors controlled by one or more computer programs.
[0019] As mentioned above, media content delivery generally focuses on audio and video experiences. Due to the customized nature of its operation, the delivery of multi-sensory (MS) content is limited.
[0020] This application describes methods for extending the creative palette for content creators and enabling the creation and distribution of spatial MS experiences at scale. Some such methods involve introducing a new layer of abstraction to enable authored MS experiences to be delivered to different endpoints using different types of equipment or actuators. As used herein, the term “endpoint” is synonymous with “playback environment” or simply “environment” and means an environment that includes one or more actuators that may be used to deliver an MS experience. Such endpoints may include rooms such as a living room in a house, a car, a cinema, a nightclub, or other venue. Some disclosed methods involve creating, delivering, and / or rendering object-based sensory data, which may include sensory objects and corresponding sensory metadata. This abstraction enables the implementation of creative intent in an object-based format that does not require prior knowledge of specific controller operations, thereby enabling greater flexibility and scalability of equipment and actuators between endpoints. An MS experience delivered via object-based sensory data may be referred to herein as a “flexibly scaled MS experience.”
[0021] [Acronym] MS - Multisensory MSIE - MS Immersive Experience AR - Augmented Reality VR - Virtual Reality PC - Personal computer
[0022] Figure 1 is a block diagram showing examples of components of a device capable of implementing various embodiments of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 1 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, device 101 may be, or include, a device configured to perform at least some of the methods disclosed herein, such as a smart audio device, a laptop computer, a cellular phone, a tablet device, a smart home hub, etc. In some such implementations, device 101 may be, or include, a server configured to perform at least some of the methods disclosed herein.
[0023] In this example, the device 101 includes at least an interface system 105 and a control system 110. In some implementations, the control system 110 may be configured to perform at least partially the methods disclosed herein. In some implementations, the control system 110 may be configured to receive a content bitstream containing encoded object-based sensory metadata via the interface system 105. The encoded object-based sensory metadata may correspond to sensory effects such as lighting, touch, airflow, one or more position actuators, or a combination thereof, provided by multiple sensory actuators in the environment. In some implementations, the control system 110 may be configured to extract the object-based sensory metadata from the content bitstream and provide the object-based sensory metadata to a sensory renderer.
[0024] In some examples, object-based sensory metadata may include sensory spatial metadata indicating at least the spatial location for rendering object-based sensory metadata in the environment, the region for rendering object-based sensory metadata in the environment, or a combination thereof. In some implementations, object-based sensory metadata does not correspond to any particular sensory actuator in the environment. In some examples, object-based sensory metadata may include abstracted sensory reproduction information that enables a sensory renderer to reproduce authored sensory effects, which may be referred to herein as intended sensory effects, from various sensory actuator locations in the environment, via various sensory actuator types and via various numbers of sensory actuators.
[0025] In some examples, the content bitstream may also include encoded audio objects synchronized with encoded object-based sensory metadata. The audio objects may include an audio signal and corresponding audio object metadata. In some such implementations, the control system 110 may be configured to extract audio objects from the content bitstream and provide the audio objects to the audio renderer. According to some examples, the audio objects may include an audio signal and corresponding audio object metadata. The audio object metadata may include at least audio object spatial metadata indicating the audio object spatial position for rendering the audio signal in the environment.
[0026] The interface system 105 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 105 may include one or more wireless interfaces. The interface system 105 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 105 may include one or more interfaces between the control system 110 and a memory system, such as the optional memory system 115 shown in Figure 1. However, the control system 110 may include a memory system in some examples.
[0027] The control system 110 may include, for example, a general-purpose single or multi-chip processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gates or transistor logic, and / or discrete hardware components.
[0028] In some implementations, the control system 110 may reside on more than one device. For example, part of the control system 110 may reside on a device within the environment (such as a laptop computer, tablet computer, or smart audio device), while another part of the control system 110 may reside on a device outside the environment, such as a server. In other examples, part of the control system 110 may reside on a device within the environment, while another part of the control system 110 may reside on one or more other devices in the environment.
[0029] Some or all of the methods described herein may be executed by one or more devices in accordance with instructions (e.g., software) stored on one or more non-temporary media. Such non-temporary media may include, but are not limited to, memory devices such as random access memory (RAM) devices and read-only memory (ROM) devices, as described herein. One or more non-temporary media may reside, for example, in an optional memory system 115 and / or control system 110 shown in Figure 1. Thus, various innovative embodiments of the subject matter described herein can be implemented on one or more non-temporary media on which software is stored. The software may include, for example, instructions for controlling at least one device to process audio data. The software may also be executable by one or more components of a control system, for example, the control system 110 in Figure 1.
[0030] In some examples, the device 101 may include an optional microphone system 120 as shown in Figure 1. The optional microphone system 120 may include one or more microphones. In some implementations, one or more of the microphones may be part of or associated with another device, such as a speaker in a speaker system, a smart audio device, etc.
[0031] In some implementations, the apparatus 101 may include an optional actuator system 125 as shown in Figure 1. The optional actuator system 125 may include one or more loudspeakers, one or more haptic devices, one or more lighting fixtures, also referred to herein as illuminators, one or more fans or other air-moving devices, one or more display devices including but not limited to one or more televisions, one or more position actuators, one or more other types of devices for providing an MS experience, or a combination thereof. As used herein, the term “lighting fixtures” generally refers to various types of light sources, including individual light sources, groups of light sources, light strips, etc. “Lighting fixtures” may be movable, and therefore, the term “fixtures” in this context does not necessarily mean that lighting fixtures are in a fixed position in space. As used herein, the term “position actuators” generally refers to devices configured to change the position or orientation of a person or object, such as a motion simulator seat. Loudspeakers may, in some cases, be referred to herein as “speakers.” In some implementations, the optional actuator system 125 may include a display system comprising one or more displays, such as one or more light-emitting diode (LED) displays, one or more organic light-emitting diode (OLED) displays, etc. In some examples where the apparatus 101 includes a display system, the optional sensor system 130 may include a touch sensor system and / or a gesture sensor system adjacent to one or more displays of the display system. According to some such implementations, the control system 110 may be configured to control the display system to present a graphical user interface (GUI), such as a GUI, relating to implementing one of the methods disclosed herein.
[0032] In some implementations, the device 101 may include an optional sensor system 130 as shown in Figure 1. The optional sensor system 130 may include a touch sensor system, a gesture sensor system, one or more cameras, etc.
[0033] This application describes a method for creating a flexibly scaled multi-sensory (MS) immersive experience (MSIE) and delivering it to different playback environments, which may also be referred to herein as endpoints. Such endpoints may include rooms such as living rooms in homes, cars, cinemas, nightclubs or other venues, AR / VR headsets, PCs, mobile devices, and the like.
[0034] Figure 2 shows exemplary elements of an endpoint. In this example, the endpoint is a living room 1001 containing multiple actuators 008, several pieces of furniture 1010, and a person 1000 (also referred to herein as a user) consuming a flexibly scaled MS experience. The actuators 008 are devices capable of changing the environment 1001 in which the user 1000 is located. The actuators 008 may include one or more televisions or other display devices, one or more lighting fixtures (also referred to herein as lighting equipment), one or more loudspeakers, etc.
[0035] The number of actuators 008, their arrangement, and the capacity of the actuators 008 within space 1001 may vary considerably between different endpoint types. For example, the number, arrangement, and capacity of actuators 008 in a car are generally different from those in a living room, nightclub, etc. In many implementations, the number, arrangement, and / or capacity of actuators 008 may vary considerably between different instances of the same type, for example, between a small living room with two actuators 008 and a large living room with sixteen actuators 008. This disclosure describes various methods for creating a flexibly scaled MSIE and delivering it to these non-homogeneous endpoints.
[0036] Figure 3 shows an example of actuator elements. In this example, the actuator is a luminaire 1100, which includes a network module 1101, a control module 1102, and a light emitter 1103. According to this example, the light emitter 1103 includes one or more light-emitting devices, such as light-emitting diodes, configured to emit light into the environment in which the luminaire 1100 is located. In this example, the network module 1101 is configured to provide network connectivity to one or more other devices in space, such as a device that transmits commands to control the emission of light by the luminaire 1100. According to this example, the control module 1102 is configured to receive signals via the network module 1101 and control the light emitter 1103 accordingly.
[0037] Other examples of actuators may include a network module 1101 and a control module 1102, but may also include other types of actuation elements. Some such actuators may include one or more loudspeakers, one or more tactile devices, one or more fans or other air-moving devices, one or more position actuators, one or more display devices, and so on.
[0038] Figure 4 shows exemplary elements of a system for creating and reproducing multi-sensory (MS) experiences. As with other figures provided herein, the types and number of elements shown in Figure 4 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In some examples, system 300 may be one or more devices configured to perform at least some of the methods disclosed herein, or may include such devices. In some examples, system 300 may include one or more instances of the control system 110 of Figure 1 configured to perform at least some of the methods disclosed herein.
[0039] According to examples of this disclosure, creating and providing an object-based MS Immersive Experience (MSIE) method involves applying a set of techniques for creating, delivering, and rendering object-based sensory data, which may include sensory objects and corresponding sensory metadata, to actuator 008. Several examples are described in the following paragraphs.
[0040] Object-Based Representation: In various disclosed implementations, multi-sensory (MS) effects are represented using what may be referred to herein as “sensory objects.” According to some such implementations, properties such as layer type and priority may be assigned to, associated with, and attached to each sensory object, allowing the content creator’s intent to be expressed in the rendered experience. Detailed examples of sensory object properties are given below.
[0041] In this example, system 300 includes a content creation tool 000 configured to design multi-sensory (MS) immersive content, depending on a specific implementation, and to output object-based sensory data 005 separately from or together with the corresponding audio data 011 and / or video data 012. The object-based sensory data 005 may include timestamp information and information indicating the type of sensory object, sensory object properties, etc. In this example, the object-based sensory data 005 is not "channel-based" data corresponding to one or more specific sensory actuators in the playback environment, but rather generalized to a wide range of playback environments having a wide range of actuator types, a wide number of actuators, etc. In some examples, the object-based sensory data 005 may include object-based lighting data, object-based tactile data, object-based airflow data or object-based position actuator data, object-based olfactory data, object-based smoke data, object-based data for one or more other types of sensory effects, or a combination thereof. According to some examples, the object-based sensory data 005 may include sensory objects and corresponding sensory metadata. For example, if object-based sensory data 005 includes object-based lighting data, the object-based lighting data may include lighting object position metadata, lighting object color metadata, lighting object size metadata, lighting object intensity metadata, lighting object shape metadata, lighting object diffusion metadata, lighting object gradient metadata, lighting object priority metadata, lighting object layer metadata, or a combination thereof. In this example, the content creation tool 000 is shown to provide a stream of object-based sensory data 005 to the experience player 002, but in an alternative example, the content creation tool 000 may generate object-based sensory data 005 that is stored for later use.An example of a graphical user interface for a lighting object-based content creation tool is provided below.
[0042] MS Object Renderer: Various disclosed implementations provide a renderer configured to render MS effects to actuators within a replay environment. In this example, system 300 includes an MS renderer 001 configured to render object-based sensory data 005 to actuator control signals 310, at least in part based on environment and actuator data 004. In this example, the MS renderer 001 is configured to output the actuator control signals 310 to an MS controller 003, which is configured to control actuators 008. In some examples, the MS renderer 001 may be configured to receive lighting objects and object-based lighting metadata indicating an intended lighting environment, and lighting information about a local lighting environment. The lighting information may be one general type of environment and actuator data 004, which may include one or more characteristics of one or more controllable light sources in the local lighting environment. In some examples, the MS renderer 001 may be configured to determine the respective drive levels of one or more controllable light sources that approximate an intended lighting environment. Some alternatives may include separate renderers for each type of actuator 008, such as one renderer for lighting fixtures, another for tactile devices, another for airflow devices, and so on. According to some examples, the MS renderer 001 (or one of the MS controllers 003) may be configured to output a drive level to at least one of the controllable light sources. In some implementations, the MS renderer 001 may be configured to adapt to changing conditions. Some examples of implementations of the MS renderer 001 are described in more detail below.
[0043] The environment and actuator data 004 may include what is referred to herein as a “room descriptor” that describes the actuator position (for example, according to an x,y,z coordinate system or a spherical coordinate system). In some examples, the environment and actuator data 004 may indicate the orientation and / or positioning properties of the actuator (e.g., direction and north-facing, omnidirectional, occluding information, etc.). According to some examples, the environment and actuator data 004 may indicate the orientation and / or positioning properties of the actuator according to a 3x3 matrix, where three elements (e.g., elements in the first row) represent the spatial position (x,y,z), three other elements (e.g., elements in the second row) represent the orientation (roll, pitch, yaw), and three other elements (e.g., elements in the third row) represent the scale or size (sx,sy,sz). In some examples, the environment and actuator data 004 may include device descriptors that describe actuator properties related to the MS renderer 001, such as the intensity range and color gamut of lighting equipment, and the airflow velocity range and direction of air transport devices.
[0044] In this example, system 300 includes an experience player 002 configured to receive object-based sensory data 005', audio data 011', and video data 012', and to provide the object-based sensory data 005 to the MS renderer 001, the audio data 011 to the audio renderer 006, and the video data 012 to the video renderer 007. In this example, the reference numbers of the object-based sensory data 005', audio data 011', and video data 012' received by the experience player 002 include a prime code (') to suggest that the data may be encoded in some examples. Similarly, the object-based sensory data 005, audio data 011, and video data 012 output by the experience player 002 do not include a prime code to suggest that the data may be decoded by the experience player 002 in some examples. In some examples, the experience player 002 may be a media player, a game engine, a personal computer or mobile device, or an element built into a television, DVD player, soundbar, set-top box, or a service provider media device such as Chromecast, Apple TV device, or Amazon Fire TV. In some examples, the experience player 002 may be configured to receive encoded object-based sensory data 005 together with encoded audio data 011 and / or encoded video data 012. In some such examples, the encoded object-based sensory data 005' may be received as part of the same bitstream as the encoded audio data 011' and / or encoded video data 012'. Several examples are described in more detail below. In some examples, the experience player 002 may be configured to extract object-based sensory data 005' from the content bitstream, provide the decoded object-based sensory data 005 to the MS renderer 001, provide the decoded audio data 011 to the audio renderer 006, and provide the decoded video data 012 to the video renderer 007.In some examples, timestamp information within object-based sensory data 005' may be used, for example, by the experience player 102, MS renderer 001, audio renderer 106, video renderer 107, or all of them, to synchronize the effects associated with object-based sensory data 005' with audio data 111' and / or video data 112', which may also include timestamp information.
[0045] In this example, system 300 includes an MS controller 003 configured to communicate with various actuator types using an application program interface (API) or one or more similar interfaces. Generally speaking, each actuator requires a specific type of control signal to produce a desired output from a renderer. In this example, the MS controller 003 is configured to map the output from the MS renderer 001 to the control signals of each actuator. For example, a Philips Hue® light bulb receives control information in a specific format for turning on the light, having a digital representation of a specific saturation, brightness and hue, as well as a desired drive level.
[0046] In some examples, the room descriptor may also describe the size and orientation of the reproduction environment itself in order to establish a relative or absolute coordinate system in which all objects are positioned. For example, in a living room, the display screen may be considered front, and possibly front and center, and the floor and ceiling may be considered vertical boundaries. In some such examples, the room descriptor may also indicate boundaries corresponding to the left, right, front, and rear walls relative to the front position. According to some examples, the room descriptor may also be provided in terms of a matrix, such as a 3x3 matrix. This room descriptor information is useful for describing the physical dimensions of the reproduction environment in units of physical distance, such as meters. In some such examples, the position, size, and orientation of sensory objects may be described in units relative to the size of the room, for example, in the range of -1 to 1. In some examples, the room descriptor may also describe preferred observation positions according to a matrix.
[0047] The type, number, and arrangement of actuators 008 generally vary according to the specific implementation. In some examples, actuators 008 may include lighting and / or light strips (also referred to herein as “light fixtures”), vibration motors, airflow generators, position actuators, or combinations thereof.
[0048] Similarly, the type, number, and arrangement of the loudspeakers 009 and display devices 010 generally vary according to the specific implementation. In the example shown in Figure 4, audio data 011 and video data 012 are rendered to the loudspeakers 009 and display devices 010, respectively, by the audio renderer 006 and video renderer 007.
[0049] As described above, according to some implementations, system 300 may include one or more instances of control system 110 of Figure 1, configured to perform at least some of the methods disclosed herein. In some such examples, one instance of control system 110 may implement content creation tool 000, and another instance of control system 110 may implement experience player 002. In some examples, one instance of control system 110 may implement audio renderer 006, video renderer 007, multisensory renderer 001, or a combination thereof. According to some examples, an instance of control system 110 configured to implement experience player 002 may also be configured to implement audio renderer 006, video renderer 007, multisensory renderer 001, or a combination thereof.
[0050] [Multi-sensory rendering synchronization] Object-based MS rendering includes different modalities that render flexibly to endpoints / playback environments. Endpoints have different capabilities according to various factors, including but not limited to the following: • Number of actuators, • Modalities of those actuators (e.g., lighting equipment, airflow control devices, tactile devices), • The type of actuators (e.g., white smart lights, RGB smart lights, or tactile vests, tactile seat cushions), and • The position / layout of those actuators.
[0051] To render object-based sensory content to any endpoint, some processing of the object signal, such as intensity, color, or pattern, is generally required. The processing of the signal paths for each modality should not alter the relative phase of any particular feature within the object signal. For example, suppose lightning is presented in both the haptic and litscape modalities. The signal processing chain for the corresponding actuator control signals should not introduce a time delay in either type of sensory object signal (haptic or illuminating) sufficient to alter the perceived synchronization of the two modalities. The required level of synchronization may depend on various factors, such as whether the experience is interactive and which other modalities are involved in the experience. The maximum time difference may range, for example, from approximately 10ms to 100ms, depending on the specific context.
[0052] [Touch] Rendering of object-based haptic content Object-based haptic content conveys the sensory aspects of a scene not only through channel-based methods but also through abstract sensory representations. For example, instead of defining haptic content solely as a single-channel time-dependent amplitude signal played from a specific haptic actuator, such as a vibrating haptic motor in a vest worn by the user, object-based haptic content may be defined by the sensation it is intended to convey. More specifically, one example might have a haptic object representing a collision haptic sensation effect. This object might have the following associated with it: • Spatial position of tactile objects, • Spatial direction / vector of tactile effect, • Intensity of tactile effect, • Tactile spatial and temporal frequency data, and ·Time-dependent amplitude signal.
[0053] In some examples, this type of haptic object may be automatically created in interactive experiences such as video games, for example, when another car crashes into the player's car from behind in a car racing game. In this example, the MS renderer decides how to render the spatial modality of this effect for a set of haptic actuators in the endpoint. In some examples, the renderer does this according to the following information: • Types of haptic devices available, e.g., haptic vests, haptic gloves, haptic seat cushions, haptic controllers, • Location of each haptic device relative to the user (some haptic devices do not need to be attached to the user, for example, shakers mounted on the floor or seat) • The type of action each tactile device provides, e.g., kinesthetic, vibratory, and tactile sensations. • Onset (start) and offset delay of each haptic device (in other words, how quickly each haptic device can be turned on and off), • The dynamic response of each haptic device (how much the amplitude can change), • The time-frequency response of each haptic device (what time frequencies the haptic device can provide), • The spatial distribution of addressable actuators within each haptic device (for example, a haptic vest may have dozens of addressable haptic actuators distributed across the user's torso), and • The time response of any tactile sensor (e.g., an active force feedback kinesthetic tactile device) used to render a closed-loop tactile effect.
[0054] These attributes of the endpoint's haptic modality inform the renderer how best to render a particular haptic effect. Consider the car collision effect example again. In this example, the player is wearing a haptic vest, haptic armband, and haptic gloves. In this example, the haptic shock wave effect is spatially positioned where the car collides with the player. The shock wave vector is determined by the relative velocity between the player's car and the car that hit the player. The spatial and temporal frequency spectrum of the shock wave effect is authored according to the type of material from which the virtual car is intended to be made, among other virtual world properties. The renderer then renders this shock wave through a set of haptic devices in the endpoint, according to the shock wave vector and the physical position of the haptic devices relative to the user.
[0055] The signals transmitted to each specific actuator are preferably provided so that the sensory effect matches across all (potentially heterogeneous) actuators for which it is available. For example, the renderer does not need to render a very high frequency to only one of the tactile actuators (e.g., a tactile armband) because other actuators are not capable of doing so. Otherwise, as a shock wave travels through the player's body, there will be a degradation of the tactile effect perceived by the user, because the tactile vest and gloves worn by the user are not capable of rendering such high frequencies, and the wave will pass through the vest, enter the armband, and finally enter the gloves.
[0056] Some types of abstract tactile effects include: • The shock wave effect described above, • Barrier effects, such as haptic effects used in video games to represent spatial limitations in a virtual world. If the input device has active or resistive kinesthetic actuators (e.g., force feedback on a steering wheel or joystick), such effects can be rendered through resistive forces applied to user input. If such actuators are not available at the endpoint, in some examples, vibratory haptic feedback corresponding to a collision between an in-game avatar and a barrier may be rendered. • Indicates the presence of a large object approaching the scene, such as a train. This type of tactile effect may be rendered using low-time-frequency vibrations of actuators of some tactile devices. This type of tactile effect may also be rendered through contact space feedback applied as pressure from an air cuff. • User interface feedback, such as clicks from virtual buttons. For example, this type of haptic effect may be rendered to the nearest actuator on the user's body that performed the click, such as a haptic glove worn by the user. Alternatively, or even further, this type of haptic effect may be rendered to a shaker coupled to the chair the user is sitting in. This type of haptic effect may be defined, for example, using a time-dependent amplitude signal. However, such a signal may be modified (modulated, frequency shifted, etc.) to best suit the haptic device providing the haptic effect. • Movement. These haptic effects are designed to make the user perceive some form of movement. These haptic effects may be rendered by actuators that actually move the user, e.g., a moving platform / seat. In some examples, the actuators may provide secondary modalities (e.g., via video) to enhance the rendered movement, and • Triggered sequence. These tactile effects are primarily characterized by their time-dependent amplitude signals. Such signals may be rendered to multiple actuators and may be augmented when doing so. Such augmentation may involve splitting the signal across multiple actuators in either time or frequency. Some examples may involve augmenting the signal itself so that the sum of the tactile actuator outputs does not match the original signal.
[0057] [Spatial and non-spatial effects] Spatial effects are configured to convey some spatial information of the rendered multisensory scene. For example, if the playback environment is a room, shock waves traveling through the room will be rendered differently for each haptic device, based on the location and size of one or more haptic objects being rendered at a given time, assuming their location within the room.
[0058] Non-spatial effects, in some examples, may target a specific location on the user, regardless of the user's position or orientation. One example is a haptic device that provides increasing vibrations on the user's back to indicate immediate danger. Another example is a haptic device that provides sharp vibrations to indicate injury to a specific area of the body.
[0059] Some effects may be non-digestive effects. Such effects are typically associated with user interface feedback, such as haptic feedback indicating that a user has completed a level or clicked a button on a menu item. Non-digestive effects may be either spatial or non-spatial.
[0060] [Haptic device type] Receiving information about different types of haptic devices available at the endpoint allows the renderer to determine what types of perceptual effects and rendering strategies are available to it. For example, local haptic device data indicating that the user is wearing both haptic gloves and a vibrating haptic vest (or at least that both haptic gloves and a vibrating haptic vest are present in the playback environment) allows the renderer to render a combined recoil effect across both devices when the user fires a gun in the virtual world. The actual actuator control signals sent to the haptic devices may differ from those in situations where only a single device is available. For example, if the user is wearing only the vest, the actuator control signals used to activate the vest may differ in terms of the timing of the start of the actuator control signal, the maximum amplitude, the frequency and decay time, or a combination thereof.
[0061] [Device Location] Knowledge of the location of haptic devices across endpoints allows the renderer to render spatial effects in a consistent manner. For example, knowledge of the location of shaker motors in a lounge allows the renderer to generate actuator control signals to each of the shaker motors in the lounge so as to convey spatial effects, such as shock waves propagating through the room. Furthermore, knowledge of the location of wearable haptic devices may be used by the renderer to convey spatial effects in addition to non-spatial effects, although implicitly due to their type, for example, a glove being on the user's hand.
[0062] [Types of actions provided by haptic devices] Haptic devices can provide a variety of different actions, and therefore, perceived sensations. These typically fall into two basic categories. 1. Vibratory tactile sensations, e.g., vibration, or 2. Kinesthetic sense, such as resistant or active force feedback.
[0063] The operation in any category may be static or dynamic, and the dynamic effect is modified in real time according to several sensor inputs. Examples include a touchscreen that renders textures using a vibratory tactile actuator, and a position sensor that measures the position of the user's finger.
[0064] Furthermore, the physical structure of such actuators can vary considerably and affect many other attributes of the device. An example of this is the onset delay or time-frequency response, which varies considerably across the following haptic device types. ·Eccentric rotating mass, • Linear resonant actuator, • Piezoelectric actuator, and Linear magnetic ram.
[0065] When rendering signals activated by haptic devices within an endpoint, the renderer should be configured to account for the onset delay of a particular haptic device type.
[0066] [Onset and offset delay of haptic devices] The onset delay of a haptic device represents the delay between the time the actuator control signal is sent to the device and the device's physical response. The offset delay represents the delay between the time the actuator control signal is sent to zero out the device's output and the time the device stops operating.
[0067] [Time-frequency response] The time-frequency response indicates the frequency range of the signal amplitude as a function of time over which the haptic device can operate in a steady state.
[0068] [Spatial frequency response] The spatial frequency response describes the frequency range of the signal amplitude as a function of the spacing between the actuators of a tactile device. Devices with actuators positioned closer together have a higher spatial frequency response.
[0069] [Dynamic Range] Dynamic range represents the difference between the minimum and maximum amplitudes of physical operation.
[0070] [Sensor characteristics in closed-loop tactile devices] Some dynamic effects involve using sensors to update the operating signal as a function of several observed states. Temporal and spatial sampling frequencies, along with noise characteristics, limit the control loop's ability to update the actuator, thus providing the dynamic effects.
[0071] [air current] Another modality that some multi-sensory immersive experiences (MSIEs) may utilize is airflow. Airflow may be rendered in conjunction with one or more other modalities, such as audio, video, lighting effects, and / or tactile sensations. Some airflow effects may be provided at other endpoints that typically involve airflow, such as a car or living room, rather than being limited to dedicated (e.g., channel-based) settings for 4D experiences in cinemas that may include "wind effects." Rather than channel-based systems, airflow sensory effects may be represented as airflow objects that may include properties such as: ·Spatial location, • The intended direction of the airflow effect, • Intensity / airflow velocity, and / or ·temperature.
[0072] Some examples of airflow objects may be used to represent the movement of passing birds. To render to the airflow actuator at the endpoint, MS Renderer 001 may provide information regarding the following: • Types of airflow devices, e.g., fans, air conditioners, heaters, • The position of each airflow device relative to the user's location or the user's expectations. • Capabilities of airflow devices, e.g., the ability of airflow devices to control direction, airflow, and temperature. • Control levels for each actuator, e.g., airflow velocity, temperature range, and • The response time of each actuator, for example, how long it takes to reach the selected speed.
[0073] [Several examples of airflow usage at different endpoints] In vehicles like cars, object-based metadata can be used to create experiences such as the following: • Using the airflow under the chair to mimic the "spine-chilling" feeling during horror movie or game content. • Simulating the movement of passing birds, and / or To create a gentle breeze in a seascape.
[0074] In the small, enclosed space of a typical vehicle, temperature changes can be achieved over relatively short periods compared to temperature changes in a larger environment such as a living room. For example, MS Renderer 001 may raise the temperature when the player enters a "lava level" or other hot area during gameplay. Some examples may include other elements, such as confetti in a vent, to celebrate events like a goal scored by the user's favorite football team.
[0075] In a living room space or another room, in one example, the airflow may be synchronized with the breathing rhythm of guided meditation. In another example, the airflow may be synchronized with the intensity of training, such that the airflow increases or the temperature decreases as the intensity increases. In some examples, there may be relatively little control over the spatial aspects during rendering. For example, many existing airflow actuators are optimized for heating and / or air conditioning rather than to provide spatially diverse sensory operation.
[0076] [Combination of lighting, airflow, and tactile sensations] Car example The following examples are written with reference to cars, but are applicable to other vehicles such as trucks and vans. In some examples, the user interface may be located on the steering wheel, near the dashboard, or on a touchscreen within the dashboard. According to some examples, the following actuators may be located inside the vehicle. 1. Individually addressable lighting spatially distributed around the car, as shown below. ○On the dashboard, ○Below the space under your feet, ○ Above the door, and ○Inside the center console. 2. Individually controllable air conditioning / heating vents distributed around the vehicle as shown below. ○Inside the front dashboard, ○Below the space under your feet, ○Inside the center console facing the rear seats, ○On the side pillar, ○Inside the seats, ○ Pointed towards the windshield (for de-fogging) 3. Individually controllable seats with vibratory tactile feedback, and 4. Individually controllable floor mats with vibratory tactile feedback.
[0077] In this example, the modalities supported by these actuators include: • In addition to individually addressable LED lighting throughout the vehicle, indicator lights on the dashboard and steering wheel, • Airflow through controllable air vents, • The following are tactile sensations ○ Steering wheel: Haptic vibration feedback, ○ Dash Touchscreen: Haptic vibration feedback and texture rendering, ○Seats: Tactile vibration and movement.
[0078] In one example, a live music stream is being rendered for four users seated in the front row. In this example, MS Renderer 001 attempts to optimize the experience for multiple viewing positions. As the artist takes the stage and prepares before the previous act ends, the content includes: • Interlude music, • Low-intensity lighting, and • Haptic content representing a crowd dancing.
[0079] In addition to rendered audio and video streams, lighting content includes ambient light objects that move slowly around the scene. These may be rendered using one of the ambient layering methods disclosed herein, for example, so that no spatial priority is given to any user's viewpoint. In some examples, haptic content may be spatially concentrated in a lower time-frequency spectrum and may be rendered only by vibrating haptic motors in a floor mat.
[0080] According to this example, a fireworks event during a music stream corresponds to multi-sensory content including the following: • Lighting objects that spatially correspond to the position of fireworks in an event, and • Tactile objects to enhance the dynamism of fireworks through a shockwave effect.
[0081] In this example, MS Renderer 001 spatially renders both lighting and haptic objects. For example, if the fireworks content is located on the left side of the scene, the lighting objects may be rendered inside the car so that each person inside the car perceives the lighting objects as coming from the left. In this example, only the left-side lighting of the car is activated. Haptics may be rendered across both the seats and floor mats to provide individual directional information to each user.
[0082] At the end of the concert, fireworks are present in the audio content, and both fireworks and confetti are present in the video content. In addition to rendering lighting and tactile objects corresponding to the fireworks as described above, the effect of confetti firing may be rendered using airflow modalities. For example, individually controllable vents in an HVAC system may be pulse-controlled.
[0083] Living room example In this implementation, in addition to an audio / visual (AV) system including multiple loudspeakers and a television, the following actuators and associated controls are available in the living room. • The tactile vest worn by the user (also called the player) • A tactile shaker attached to the seat where the player is sitting. • A smartwatch with (tactile) control capabilities. • Smart lights spatially distributed around the room, • Wireless controller, and Addressable airflow bars (AFBs) (similar to HVAC vents in a car's front dashboard) including an array of individually controllable fans directed towards the user.
[0084] In this example, the user is playing a first-person shooter game, which includes a scene where a destructive hurricane moves through the level. During this scene, in-game objects are thrown around, some hitting the player. The tactile objects, rendered by MS Renderer 001, provide a shockwave effect through all tactile devices the user can perceive. The actuator control signals sent to each device may be optimized according to the intensity and direction of the impact of the in-game objects, as well as the capabilities and position of each actuator (as described above).
[0085] Prior to the user being struck by an in-game object, the multisensory content includes a tactile object corresponding to a non-spatial rumble (vibrating sound), one or more airflow objects corresponding to directional airflow, and one or more lighting objects corresponding to electric lights. MS Renderer 001 renders the non-spatial rumble to the haptic device. Actuator control signals sent to each haptic device may be rendered such that the set of actuator control signals across the haptic array matches in perceived onset time, intensity, and frequency. In some examples, the frequency content of actuator control signals sent to a smartwatch may be low-pass filtered so that they match the best frequency limiting capability in proximity to the watch. MS Renderer 001 may render one or more airflow objects into the actuator control signals for AFB such that the airflow in the room matches the player's position and line of sight in the game, as well as the hurricane direction itself. Electric lights may be rendered across all modalities as (1) a white flash across lighting located in appropriate places in or above the ceiling, and (2) an impactful rumble in the user's wearable haptics and seat shaker.
[0086] When a user is struck by an in-game object, a directional shockwave may be rendered to the haptic device. In some examples, a corresponding airflow shock may be rendered. In some examples, a damage gain effect indicating the amount of damage inflicted on the player by being struck by an in-game object may be rendered by lighting.
[0087] In some such examples, the signal may be spatially rendered to the haptic device so that the perceived shock wave travels across the player's body and the room. The MS renderer 001 may provide such an effect according to actuator position information indicating the positions of the haptic devices relative to each other. In addition to actuator capability information, the MS renderer 001 may provide shock wave vectors and positions according to actuator position information. According to some examples, the shock of a non-directional airflow may be rendered, for example, all vents of the AFB may be temporarily strengthened to enhance the haptic modality. In some examples, a red vignette may be rendered simultaneously on a light strip surrounding the TV to indicate to the player that the player has taken damage in the game.
[0088] Figure 5 shows exemplary elements of another system for creating and reproducing MS experiences. As with other figures provided herein, the types and number of elements shown in Figure 5 are provided for illustrative purposes only. Other implementations may include more, fewer, and / or different types and numbers of elements. In some examples, system 500 may be one or more devices configured to perform at least some of the methods disclosed herein, or may include such devices. In some examples, system 500 may include one or more instances of control system 110 of Figure 1 configured to perform at least some of the methods disclosed herein.
[0089] In this example, the system shown in Figure 5 is an instance of the system shown in Figure 4. In this example, the system shown in Figure 5 is an embodiment of a "lightscape" in which visual (video), audio, and lighting effects are combined to create an MS experience.
[0090] In this example, system 500 includes a lightscape creation tool 100, which is an instance of the content creation tool 000 described with reference to Figure 4. Depending on the specific implementation, the lightscape creation tool 100 is configured to design and output object-based lighting data 505' separately from or together with the corresponding audio data 111' and / or video data 112'. The object-based lighting data 505' may include timestamp information and information indicating lighting object properties, etc. In some examples, the timestamp information may be used to synchronize the effects associated with the object-based lighting data 505' with the audio data 111' and / or video data 112', which may also include timestamp information.
[0091] In this example, the object-based lighting data 505' includes lighting objects and their corresponding lighting metadata. For example, the object-based lighting data may include lighting object position metadata, lighting object color metadata, lighting object size metadata, lighting object intensity metadata, lighting object shape metadata, lighting object diffuse metadata, lighting object gradient metadata, lighting object priority metadata, lighting object layer metadata, or a combination thereof. In this example, the content creation tool 100 is shown to provide a stream of object-based lighting data 505' to the experience player 102, but in an alternative example, the content creation tool 100 may generate object-based lighting data 505' that is stored for later use. An example of a graphical user interface for a lighting object-based content creation tool is shown below.
[0092] In this example, system 500 includes an experience player 102 configured to receive object-based lighting data 505', audio data 111', and video data 112', and to provide the object-based lighting data 505 to the lightscape renderer 501, the audio data 111 to the audio renderer 106, and the video data 112 to the video renderer 107. According to some examples, the experience player 102 may be a media player, a game engine, or a personal computer or mobile device, or an element built into a television, DVD player, soundbar, set-top box, or a service provider media device such as Chromecast, Apple TV device, or Amazon Fire TV. In some examples, the experience player 002 may be configured to receive the encoded object-based lighting data 505' together with the encoded audio data 111' and / or the encoded video data 112'. In some such examples, the encoded object-based lighting data 505' may be received as part of the same bitstream as the encoded audio data 111' and / or the encoded video data 112'. Several examples are described in more detail below. In some examples, the experience player 102 may be configured to extract object-based lighting data 505 from the content bitstream, provide the decoded object-based lighting data 505 to the lightscape renderer 501, provide the decoded audio data 111 to the audio renderer 106, and provide the decoded video data 112 to the video renderer 107. In some examples, the experience player 002 may be configured to allow control of configurable parameters within the lightscape renderer 501, such as immersion intensity. Several examples are described below.
[0093] In some examples, the room descriptor of the environment and lighting equipment data 104 may describe the size and orientation of the reproduction environment itself in order to establish a relative or absolute coordinate system in which all objects are placed. For example, in a living room, the display screen may be considered front, and possibly front and center, and the floor and ceiling may be considered vertical boundaries. In some such examples, the room descriptor may also indicate boundaries corresponding to the left, right, front and rear walls relative to the front position. According to some examples, the room descriptor may also be provided in terms of a matrix, such as a 3x3 matrix. This room descriptor information is useful for describing the physical dimensions of the reproduction environment in units of physical distance, such as meters. In some such examples, the position, size, and orientation of sensory objects may be described in units relative to the size of the room, for example, in the range of -1 to 1. In some examples, the room descriptor may also describe preferred observation positions according to a matrix.
[0094] In this example, system 500 includes a lightscape renderer 501 configured to render object-based lighting data 505 to lighting equipment control signals 515, at least partially based on environment and actuator data 104. In this example, the lightscape renderer 501 is configured to output the lighting equipment control signals 515 to a lighting controller 103 configured to control a lighting equipment 108. The lighting equipment 108 may include individual controllable light sources, groups of controllable light sources (such as controllable light strips), or combinations thereof. In some examples, the lightscape renderer 501 may be configured to manage various types of lighting object metadata layers, examples of which are provided herein. According to some examples, the lightscape renderer 501 may be configured to render actuator signals for the lighting equipment, at least partially based on the observer's viewpoint. If the observer is in a living room including a television (TV) screen, the lightscape renderer 501 may, in some examples, be configured to render actuator signals to the TV screen. However, in virtual reality (VR) use cases, the lightscape renderer 501 may be configured to render actuator signals for the position and orientation of the user's head. In some examples, the lightscape renderer 501 may receive inputs from the playback environment, such as light sensor data corresponding to ambient light, or camera data corresponding to the position or orientation of a person, in order to enhance the rendering.
[0095] In some examples, the lightscape renderer 501 is configured to receive object-based lighting data 505, which includes lighting objects and object-based lighting metadata that indicate the intended lighting environment, and environment and lighting equipment data 104, which corresponds to lighting equipment 108 and other features of the local playback environment, which may include, but are not limited to, reflective surfaces, windows, uncontrollable light sources, light occlusion features, etc. In this example, the local playback environment includes one or more loudspeakers 109 and one or more display devices 510.
[0096] In some examples, the lightscape renderer 501 is configured to calculate how to excite various controllable lighting fixtures 108 based at least partially on object-based lighting data 505 and environment and lighting fixture data 104. The environment and lighting fixture data 104 may indicate, for example, the geometric location of the lighting fixtures 108 in the environment, lighting fixture type information, etc. In some examples, the lightscape renderer 501 may be configured to determine which lighting fixtures are activated based at least partially on location metadata and size metadata associated with each lighting object, for example, by determining which lighting fixtures are in the volume of the regenerated environment corresponding to the location and size of the lighting object at a particular time indicated by lighting object timestamp information. In this example, the lightscape renderer 501 is configured to send a lighting fixture control signal 515 to the lighting controller 103 based on the environment and lighting fixture data 104 and object-based lighting data 505. The lighting fixture control signal 515 may be sent via one or more of various transmission mechanisms, application program interfaces (APIs), and protocols. The protocol may include, for example, Hue API, LIFX API, DMX, Wi-Fi, Zigbee, Matter, Thread, Bluetooth Mesh, or other protocols.
[0097] In some examples, the lightscape renderer 501 may be configured to determine the drive level of each of one or more controllable light sources that approximate the lighting environment intended by the creator of the object-based lighting data 505. According to some examples, the lightscape renderer 501 may be configured to output a drive level to at least one of the controllable light sources.
[0098] In some examples, the lightscape renderer 501 may be configured to scale down one or more parts of the lighting map according to content metadata, user input (mode selection), lighting equipment limitations and / or configuration, other factors, or a combination thereof. For example, the lightscape renderer 501 may be configured to render the same control signal to two or more different lights in the playback environment. In some such examples, the two or more lights may be placed in close proximity to each other. For example, the two or more lights may be different lights of the same actuator, or different bulbs in the same lamp, for example. The lightscape renderer 501 may be configured to reduce computational overhead and increase rendering speed, etc., by rendering the same control signal to two or more different but close-proximity lights, rather than calculating very slightly different control signals for each bulb.
[0099] In some examples, the lightscape renderer 501 may be configured to spatially upmix object-based lighting data 505. For example, if object-based lighting data 505 is generated for a single plane, such as a horizontal plane, in some examples, the lightscape renderer 501 may be configured to project the lighting objects of the object-based lighting data 505 onto an upper hemisphere (e.g., above the actual or expected position of the user's head) to enhance the experience.
[0100] In some examples, the lightscape renderer 501 may be configured to apply one or more thresholds, such as one or more spatial thresholds, one or more luminous intensity thresholds, etc., when rendering actuator control signals to lighting actuators in the replay environment. In some examples, such thresholds may prevent some lighting objects from triggering the operation of some lighting fixtures.
[0101] Lighting objects may be used for a variety of purposes, such as setting the atmosphere of a room, providing spatial information about a character or object, enhancing special effects, creating greater interaction and immersion, shifting the observer's attention, or puncturing content. Some such purposes may be expressed, at least in part, by the content creator in accordance with sensory object metadata types and / or properties that are generally applicable to various types of sensory objects, such as object metadata indicating the position and size of the sensory object.
[0102] For example, the priority of sensory objects, including but not limited to lighting objects, may be indicated by sensory object priority metadata. In some such examples, sensory object priority metadata is considered when multiple sensor objects are simultaneously mapped to the same equipment in the replay environment. Such priority may be indicated by lighting priority metadata. In some examples, priority may not need to be indicated via metadata. For example, MS renderer 001 may give priority to moving sensory objects (including but not limited to lighting objects) over stationary sensory objects.
[0103] A lighting object may potentially trigger the excitation of multiple lights, depending on its position and size, as well as the location of lighting equipment within the regeneration environment. In some examples, when the size of a lighting object encompasses multiple lights, the renderer may apply one or more thresholds (such as one or more spatial thresholds or one or more luminous intensity thresholds) to gate the object from activating some of the encompassed lights.
[0104] [Example using a lighting map] In some implementations, a lighting map, which is an instance of an actuator map (AM) containing a description of the lighting in the replay environment, may be provided to the lightscape renderer 501. In some such examples, the environment and lighting equipment data shown in Figure 5 may include a lighting map. In some examples, the lighting map may be other-centered, for example, showing the attenuation of light based on absolute space coordinates, while in other examples, the lighting map may be egocentered, for example, a lighting projection mapped onto a sphere at the intended observation position and orientation. In the case of a sphere, the lighting map may, in some examples, be projected onto a 2D surface, for example, to utilize a two-dimensional (2D) image texture in processing. In any case, the lighting map should indicate the capabilities and lighting settings of the replay environment, such as a room. In some embodiments, the lighting map may not be directly related to the physical characteristics of the room, for example, if adjustments have been made based on the preferences of a particular user.
[0105] In some examples, a lighting map may exist for each lighting fixture or for each light fixture in the playback environment. In some examples, the light intensity shown by the lighting map may be inversely correlated with the distance to the center of the light fixture, or approximately inversely correlated with the distance to the center of the light fixture (e.g., within plus or minus 5%, within plus or minus 10%, within plus or minus 15%, within plus or minus 20%, etc.). The intensity values of the lighting map may indicate the intensity or influence of the lighting object on the lighting fixture. For example, when a lighting object approaches a light bulb, the lightscape renderer 501 may be configured to determine that the bulb intensity increases as the distance between the lighting object and the bulb decreases. The lightscape renderer 501 may be configured to determine the rate of this transition based at least in part on the intensity of the lighting shown by the lighting map.
[0106] Within this common rendering space, in some examples, the lightscape renderer 501 may be configured to calculate the illumination activation metric using a dot product multiplication between the lighting object and the illumination map for each light, for example, as follows:
number
[0107] In the above formula, Y represents the illumination activation metric, LM represents the illumination map, and Obj represents the map of the illumination object. The illumination activation metric indicates the relative light intensity of the actuator control signal output by the lightscape renderer 501 based on the overlap between the light distribution from the illumination object and the lighting fixture. In some examples, the lightscape renderer 501 may use the maximum or minimum distance from the illumination object to the lighting fixture or other geometric metrics as part of determining the light intensity. In some implementations, the lightscape renderer 501 may determine the illumination activation metric by referring to a lookup table instead of calculating it.
[0108] The lightscape renderer 501 may repeat one of the above steps to determine the lighting activation metric for all lighting objects and all controllable lighting in the replay environment. Setting thresholds for lighting objects that have a very low impact on the lighting system can help reduce complexity. For example, if the effect of a lighting object causes activation below a threshold percentage for the lighting system activation, such as less than 10% or less than 5%, the lightscape renderer 501 may ignore the effect of that lighting object.
[0109] The lightscape renderer 501 may then use the resulting illumination activation matrix Y, along with various other properties such as the selected panning law (indicated by either the illumination object metadata or the renderer configuration) or the priority of the illumination objects, to determine which objects are rendered by which illumination and how. Rendering illumination objects to illumination control signals may include the following: • Changing the brightness of a lighting object as a function of the distance from the lighting equipment. - Mixing the colors of multiple lighting objects rendered simultaneously (multiplexed) by a single lighting fixture, or • Modify one of the above based on the priority of the lighting object.
[0110] [Rendering parameters] In addition to the information carried by the lighting object metadata, the rendering of lighting objects can be a function of the settings or parameters of the lightscape renderer 501 itself. These may include the following: • Speed Priority - When this parameter is set, moving lighting objects are given higher priority than stationary lighting objects. The speed priority parameter set improves the dynamism of the rendered scene. • Color priority - Lighting objects with higher saturation values are given priority. • Activation threshold - the minimum illumination activation value Y that must be achieved to activate the lighting equipment. • Accessibility - Certain colors may be selected over others to best represent the experience of colorblind users. Certain flash rates may be avoided for people with photosensitivity.
[0111] [Rendering Configuration (Mode)] In addition to the information carried by the lighting object metadata, the lightscape renderer 501 may be configured according to different modes in some implementations. The term “mode” as used herein is different from “parameter,” in that a mode may involve, for example, entirely different signal paths, while a parameter may simply parameterize these signal paths. For example, one mode may include the projection of all lighting objects onto the lighting map before determining how / what to render to the lighting fixtures, while another mode may simply snap the highest priority lighting to the nearest lighting fixture. Modes may include: • Modes that support low-lighting equipment counts. In these modes, rendering parameters and lighting object metadata are used to determine which subset of lighting objects should be rendered and how they should be rendered. Here, "how" refers to the trade-offs between spatial, chromatic, and temporal fidelity of the most prominent lighting objects in the scene. • Modes that support different content types such as music and games. • A mode in which multiple lighting objects can be rendered by a single lighting fixture (or single light) using color mixing. • A mode in which only a single lighting object can be rendered by a single lighting fixture (or single light). A mode in which the brightness of a lighting object is changed as a function of the geometric distance or other distance between the lighting object and the lighting fixture.
[0112] As described above, according to some implementations, system 500 may include one or more instances of control system 110 of Figure 1, configured to perform at least some of the methods disclosed herein. In some such examples, one instance of control system 110 may implement a lightscape creation tool 100, and another instance of control system 110 may implement an experience player 002. In some examples, one instance of control system 110 may implement an audio renderer 006, a video renderer 007, a lightscape renderer 501, or a combination thereof. According to some examples, an instance of control system 110 configured to implement an experience player 002 may also be configured to implement an audio renderer 006, a video renderer 007, a lightscape renderer 501, or a combination thereof.
[0113] Figure 6 shows an example of a graphical user interface (GUI) that may be presented by the display device of the lightscape creation tool of Figure 5. As with other figures provided herein, the types and number of elements shown in Figure 6 are provided for illustrative purposes only. Other GUIs presented by the lightscape creation tool may include more, fewer, and / or different types and numbers of elements. According to some examples, GUI 600 may be presented to the display device in accordance with commands from an instance of the control system 110 of Figure 1 configured to implement the lightscape creation tool 100 of Figure 5.
[0114] In this example, the user may interact with GUI600 to create a lighting object and assign lighting object properties that can be associated with the lighting object as metadata. In this example, the user is selecting properties for lighting object 630. In this example, GUI600 displays lighting object 630 in three-dimensional space 631, the latter representing the playback environment. Element 634 indicates the coordinate system of three-dimensional space 631. Thus, in this example, lighting object 630 and three-dimensional space 631 are observed from the top left.
[0115] The user may interact with the GUI 600 to select the position and size of the lighting object 630. In some examples, the user may select the position of the lighting object 630 by dragging it to a desired position in three-dimensional space 631, for example by touching a touchscreen or using a cursor. In some examples, the user may select the size of the lighting object 630 by selecting the size of a circle (or other shape) shown in the GUI 600 to indicate the outline of the lighting object 630. In some such examples, the user may reduce the size of the lighting object 630 by pinching its outline with two fingers, or increase its size by spreading it with two fingers, and so on.
[0116] Specifying the location and size of MS objects in an abstracted 3D space, such as the 3D space 631 of GUI600, allows content creators to generalize the location and range of corresponding MS effects without prior knowledge of the specific playback environment in which the MS effect is provided. This is an advantage of MS object-oriented approaches in various disclosed implementations. For example, GUI600 allows content creators to specify the location and size of a lighting object 630 in the 3D space 631, thereby allowing content creators to generalize the location and range of corresponding lighting effects without prior knowledge of the specific size of any particular playback environment in which the lighting effect is provided, or the number, type, and location of lighting fixtures within the playback environment in which the lighting effect is provided. Lighting fixtures that are potentially activated in response to the presence of the lighting object 630 at a particular time are lighting fixtures within the volume of the playback environment corresponding to the location and size / range of the lighting object 630.
[0117] In this example, the user may interact with the color circle 635 of GUI600 to select the hue and saturation of the current lighting object, and may interact with the slider 636 to select the brightness of the current lighting object. These and other selectable properties of the lighting object 630 are displayed in area 632 of GUI600. In this example, the properties of the lighting object 630 that can be selected via GUI600 also include intensity, diffuseness, "feathering", whether the lighting object is hidden or not, saturation, priority, and layer. Lighting object layers and priority are described in more detail below. Generally speaking, lighting object layers may be used to group lighting objects into categories such as "surrounding" and "dynamic". Lighting object priority is assigned by the content creator and may be used by the renderer to determine, for example, which lighting object is presented when two or more lighting objects are active simultaneously and simultaneously contain areas with the same lighting equipment.
[0118] Area 640 of GUI600 displays the time information corresponding to each of the multiple lighting objects created via the Lightscape Creation Tool. In this example, the lighting objects are listed on the left side of Area 640 along the vertical axis, and the time is shown along the horizontal axis. In this example, a 4-second time interval is drawn by a vertical line. Here, the time information for each lighting object is shown as an isolated or connected diamond symbol or line along a series of horizontal rows, and each diamond symbol or line corresponds to one of the lighting objects shown on the left side of Area 640. For example, line 633 indicates that lighting object 3 begins to appear between 39 and 40 seconds and is displayed continuously until approximately 1 minute and 6 seconds. The diamond symbol to the right of line 633 indicates that lighting object 3 is displayed discontinuously for the next few seconds.
[0119] Figure 7A shows another example of a graphical user interface (GUI) that may be presented by the display device of the lightscape creation tool of Figure 5. As with other figures provided herein, the types and number of elements shown in Figure 7A are provided merely as examples. Other GUIs presented by the lightscape creation tool may include more, fewer, and / or different types and numbers of elements. According to some examples, GUI 700 may be presented to the display device in accordance with commands from an instance of the control system 110 of Figure 1 configured to implement the lightscape creation tool 100 of Figure 5.
[0120] In this example, GUI700 represents a point in time when the lighting fixtures in the actual playback environment are controlled according to the lighting objects created by the implementation of the lightscape creation tool 100 in Figure 5. An image of the playback environment is shown in area 705 of GUI700. Various lighting fixtures 708 and televisions 715 are shown in the playback environment in area 705. A specific point in time is indicated by a vertical line 742 in area 740. At this point, the vertical line 742 intersects with the horizontal lines 744a, 744b, 744c, and 744d, indicating that the lighting provided to the corresponding lighting objects 1, 4, 5, and 7 is being reproduced. Area 732 shows the lighting object properties.
[0121] At the point shown in Figure 7A, it can be observed that the left side of the regeneration environment shown in region 705 is illuminated by blue light. This corresponds at least partially to the effect of the illumination object 730 shown in three-dimensional space 731.
[0122] In this example, video and audio data are also played within the audio environment, and the playback of the rendered lighting objects is synchronized with the playback of the video and audio data. In this example, the image of the played video is shown in area 710 of GUI 700. The video may be played by, for example, a television 715.
[0123] In some examples, the user may be able to interact with GUI700 to adjust lighting object properties, add or remove lighting objects, etc. For example, the user may pause playback to adjust lighting object properties. In some alternative examples, the user may need to return to a GUI like GUI600 in Figure 6 to adjust lighting object properties, add or remove lighting objects, etc.
[0124] In the example described with reference to Figure 7A, the GUI 700 is presented on a display device corresponding to the lightscape creation tool 100 in Figure 5, but the lighting objects, audio, and video are rendered within a real-world environment. Thus, in some implementations, the example described with reference to Figure 7A may also include at least some of the “downstream” rendering and playback functions that may be provided by other blocks in Figure 5, including but not limited to the functions of the lightscape renderer 501, the lighting controller API 103 (which in some examples may be implemented by the same device that implements the lightscape renderer 501), the lighting equipment 108, the audio renderer 106, the loudspeaker 109, the video renderer 107, and the display device 510. In some such examples, the process described with reference to Figure 7A may also include the functions of the experience player 102 in Figure 5.
[0125] In some alternative implementations, the example described with reference to Figure 7A may also include at least some of the “downstream” rendering and playback functions that may be provided by other blocks in Figure 4, including but not limited to the functions of the MS renderer 001, MS controller API 003 (which in some examples may be implemented by the same device that implements the MS renderer 001), lighting equipment 008, audio renderer 006, loudspeaker 009, video renderer 007, and display device 010. In some such examples, the process described with reference to Figure 7A may also include the functions of the experience player 002 in Figure 4.
[0126] Figure 7B is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of Method 750, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 750 may be performed simultaneously. Furthermore, some implementations of Method 750 may include more or fewer blocks than those shown and / or described. The blocks of Method 750 may be performed by one or more devices, which may be (or include) one or more instances of a control system, such as the control system 110 shown in Figure 1 and described above. For example, at least some aspects of Method 750 may be performed by an instance of the control system 110 configured to implement the experience player 002 of Figure 4. Some other aspects of Method 750 may be performed by an instance of the control system 110 configured to implement the MS renderer 003 of Figure 4.
[0127] In this example, block 755 includes receiving a content bitstream containing encoded object-based sensory data from a control system. In this case, the object-based sensory data includes sensory objects and corresponding sensory metadata, corresponding to sensory effects provided by multiple sensory actuators in the environment. In some examples, the environment may be a real-world environment such as a room environment or a car environment. The sensory effects may include lighting, touch, airflow, one or more position actuators, or a combination thereof, provided by multiple sensory actuators in the environment. According to some examples, the environment may be a virtual environment, or may include a virtual environment. In some such examples, method 750 may also include providing a virtual environment such as a game environment while providing corresponding sensory effects in a real-world environment.
[0128] In some examples, object-based sensory metadata may include sensory spatial metadata that at least indicates a spatial location for rendering object-based sensory data in an environment, or indicates a region for rendering object-based sensory data in an environment, or both. According to some examples, object-based sensory data does not correspond to a specific sensory actuator in an environment. For example, as illustrated with reference to Figures 6 and 7A, a sensory object in object-based sensory data may correspond to a portion of a three-dimensional region representing the reproduction environment. The actual reproduction environment in which the sensory object is rendered does not need to be known at the time the sensory object is authored, and is generally unknown. Therefore, object-based sensory data includes abstracted sensory reproduction information (in this example, the sensory object and the corresponding sensory metadata), enabling the sensory renderer to reproduce the authored sensory effects from various sensory actuator locations in the environment via various sensory actuator types and via various numbers of sensory actuators.
[0129] According to this example, block 760 includes the control system extracting object-based sensory data from the content bitstream. In some such examples, the content bitstream may also include encoded audio objects synchronized with the encoded object-based sensory data. According to some such examples, the audio objects may include an audio signal and corresponding audio object metadata. In some such examples, method 750 may also include the control system extracting audio objects from the content bitstream. According to some such examples, method 750 may also include the control system providing the audio objects to an audio renderer. In some examples, the audio object metadata may include at least audio object spatial metadata indicating the audio object spatial location for rendering the audio signal in the environment.
[0130] In this example, block 765 includes the control system providing object-based sensory data to the sensory renderer. According to some such examples, method 750 may also include the sensory renderer receiving object-based sensory data and the sensory renderer receiving environment descriptor data corresponding to the location of a sensory actuator in the environment. In some such examples, method 750 may also include the sensory renderer receiving actuator descriptor data corresponding to the properties of a sensory actuator in the environment. In some examples, the environment descriptor data and actuator descriptor data may be or may be included in the environment and actuator data 004 described with reference to Figure 4. According to some examples, method 750 may also include the sensory renderer providing actuator control signals for controlling the sensory actuator in the environment to produce sensory effects indicated by the object-based sensory data. In some such examples, the MS renderer 001 may provide actuator control signals 310 to the MS controller API 003, and the MS controller API 003 may provide actuator-specific control signals to actuator 008. In some alternative examples, the MS controller API 003 shown in Figure 4 may be implemented via the MS renderer 001, and actuator-specific signals may be provided to the actuator 008 by the MS renderer 001. In some examples, method 750 may also include providing sensory effects by sensory actuators in the environment.
[0131] In some examples, Method 750 may also include, by an audio renderer, receiving audio objects and by an audio renderer, receiving loudspeaker data corresponding to loudspeakers in the environment. According to some such examples, Method 750 may also include, by an audio renderer, providing loudspeaker control signals to control loudspeakers in the environment to play audio corresponding to audio objects and synchronized with sensory effects. Synchronization may be based on time information contained in or included with sensory objects and audio objects, such as timestamps. In some examples, Method 750 may also include, by a loudspeaker in the environment, playing audio corresponding to audio objects.
[0132] According to some examples, the content bitstream includes encoded video data synchronized with encoded audio objects and encoded object-based sensory data. In some such examples, Method 750 may also include, by a control system, extracting video data from the content bitstream and by the control system providing the video data to a video renderer. According to some such examples, Method 750 may also include, by a video renderer, receiving the video data and by the video renderer providing video control signals to control one or more display devices in the environment to present images corresponding to video control signals and synchronized with audio objects and sensory effects. In some examples, Method 750 may also include presentation by one or more display devices in the environment. The images may correspond to video control signals.
[0133] [Intended lighting environment metadata] Some disclosed examples include including lighting metadata for use with video and / or audio tracks, or as a standalone lighting-based sensory experience. This lighting metadata describes the intended lighting environment to be reproduced during playback. Various methods exist for representing lighting metadata, some of which are described in detail in this disclosure.
[0134] In some examples, the intended lighting environment may be transmitted as one or more Image-Based Lighting (IBL) objects. IBL is a technique previously used to capture environments and lighting. IBL may be described as a process of illuminating scenes and objects (real or composite) with images of light from the real world. Evolving from reflection mapping techniques previously disclosed in panoramic images, IBL is used as a texture map on a computer graphics model to represent glossy objects that reflect a real or composite environment. Some aspects of IBL are analogous to image-based modeling, where the geometric structure of a three-dimensional scene can be derived from an image. Other aspects of IBL are analogous to image-based rendering, where the rendered appearance of a scene can be generated from the appearance of the scene in an image.
[0135] Previously, IBL objects were used to render computer graphics to create realistic reflection and lighting effects, such as the effect where the rendered object itself appears to reflect a part of the real-world environment. In some previously disclosed virtual world / computer-generated examples, IBL involves the following process: • Capturing real-world lighting as an omnidirectional image. • Mapping lighting to represent the environment, • Placing computer-generated 3D objects within the environment, • Simulating light from the environment that illuminates computer graphics objects.
[0136] Various methods exist for capturing omnidirectional images. One method is to use a camera to photograph a reflective sphere placed in the environment. Another method for obtaining omnidirectional images is to obtain a mosaic based on many camera images taken from different directions / viewpoints and combine the images using an image stitching program. In some such examples, the image may be acquired using a fisheye lens that can cover the entire field of view with just two images. Another method for obtaining omnidirectional images is to scan across a 360° field of view using a scanning panoramic camera, such as a “rotating line camera,” configured to assemble a digital image as the camera rotates. Further details of previously disclosed IBL methods are disclosed in Debevec, Paul, “Image-Based Lighting” (IEEE, March / April 2002, pp.26–34), which is incorporated herein by reference.
[0137] Some implementations of this disclosure are based on previously disclosed methods, including IBL objects, for the novel purpose of recreating an intended environment using dynamically controllable surround lighting. Both the intended lighting environment and the endpoint environment may change over time. For example, walls in the endpoint environment may be painted, actuators may be moved, furniture may be moved or replaced, and new furniture, shelves, and / or cabinets may be added.
[0138] Several types of mapping or projection may be used to map ambient illumination from a sphere surrounding the intended observation position onto a two-dimensional (2D) plane. Projection onto a 2D plane facilitates the compression of the illumination object using a 2D image or video codec for more efficient transmission (reducing the computational overhead required).
[0139] Figures 8A, 8B, and 8C show three examples of projecting the illumination of an observation environment onto a two-dimensional (2D) plane. These examples use a spherical projection; other representations are possible. In these examples, the illumination of the observation environment is shown across a 360-degree observation angle from the intended observation position. In these examples, the x-axis represents the horizontal angle (left to right) from the observer's viewpoint, and the vertical axis represents the vertical angle (up and down) from the observer's viewpoint. In these examples, the center of projections 805, 810, and 815 corresponds to the direction in front of the observer. The leftmost point of projections 805, 810, and 815 corresponds to the rightmost point of projections 805, 810, and 815, as the projection "wraps" around the observation position. Projections 805, 810, and 815 may, for example, correspond to three different contents, or to three different times of the same content. Projection 805 represents studio-type lighting in a mostly dark, monochromatic environment. Projection 810 includes a dark blue floor 812 and a dominant blue region 814 directly in front of the observer. Projection 815 includes a bright red light 819 in front of and slightly to the left of the observer.
[0140] As described above, in some disclosed implementations, lighting metadata provided with a lighting object may be used to generate a video showing the environment from an intended observation position. In some examples, the lighting metadata may include multiple metadata units, each of which may include a header describing how the metadata is presented to facilitate playback. In some such examples, the lighting metadata may include the following information, which may be provided in different metadata units: • Environment metadata version: This information allows the decoder to correctly interpret multiple ways of representing lighting and select the one most appropriate for the capabilities of the lighting equipment in a particular reproduction environment. • Environment metadata mapping type: This information allows the decoder to correctly translate metadata into real-world coordinates. • Environmental metadata timecode: This information allows the decoder to accurately synchronize the metadata to the audio or video track. • Environmental metadata location code: This information allows the decoder to correctly align the observation location with a reference observation location. In some cases, multiple sets of metadata exist corresponding to different observation locations, allowing the observer to experience content from multiple observation locations, and the lighting adapts accordingly by selecting the nearest appropriate location or by interpolating between nearby locations. For example, Locations X1, Y1, and Z1 have EnvironmentMetadataPayload EMP1, and locations X2, Y2, and Z2 have EMP2. ○The user's position X', Y', Z' is determined. ○The geometric distance is D1=sqrt((X1-X') 2 +(Y1-Y') 2 +(Z1-Z') 2 ) and D2=sqrt((X2-X') 2 +(Y2-Y') 2 +(Z2-Z') 2 It is calculated by the distance between X1,Y1,Z1 and X',Y',Z'. ○The ratio of distances is used to determine the value of "alpha": alpha = D1 / (D1+D2). The EnvironmentMetadataPayload suitable for the user position EMP' is interpolated between two reference observation positions using Alpha, for example, by EMP' = EMP1*(1-Alpha) + EMP2*(Alpha). This simple example demonstrates one method of linear interpolation between two points. Further clamping may be desirable if the observer's position exceeds either of the reference points. Triangular interpolation may be preferable in some cases. • Environmental metadata compression method, • Environmental metadata payload size, and • Environmental metadata payload.
[0141] In some examples, IBL representations may be augmented with depth information indicating the relative distance of ambient light sources from a reference observation position. This depth information can be used to adapt the position of light sources as the observer moves around the environment. Depth information may be obtained directly from RGB images, for example, through monocular depth estimation using a trained neural network. Alternatively, or even further, many consumer devices such as the iPhone® can directly measure depth using infrared imaging techniques such as LiDAR and structured illumination.
[0142] The spatial resolution of IBL technology may be variable depending on the requirements of the use case, depending on the available bitrate and the required compression quality. IBL may be compressed using known image-based compression methods such as Joint Photographic Experts Group (JPG), JPG2000, Portable Network Graphic (PNG), etc., or known video-based compression methods such as Advanced Video Coding (AVC), Versatile Video Encoding (VVC), AOMedia Video 1 (AV1), also known as H.264, H.265.
[0143] However, the methods disclosed herein are not limited to IBL-based examples. Several other examples representing the intended lighting environment treat each light source as a unique light source object with a defined position and size. In some such methods, other information, including but not limited to the direction of light emitted from each light source object, may be used to define or describe each light source. Some methods may also include information such as the reflectivity of one or more surfaces, one or more intended room dimensions, etc. Such information may be used to support implementations in which an observer can freely move to a new position. According to some examples, light source objects may be defined directly using a computer program written specifically for this task, or they may be inferred by analyzing video images of light sources in the environment. Potential advantages of light source object-based methods include a smaller metadata payload size for relatively simple lighting scenarios.
[0144] [Replay Lighting Environment Rendering] Figure 9A shows exemplary elements of a litescape renderer. As with other figures provided herein, the types and number of elements shown in Figure 9A are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. In this example, the litescape renderer 501 is an instance of the litescape renderer 501 shown in Figure 5. In this example, the litescape renderer 501 is implemented by an instance of the control system 110 in Figure 1. In this example, the litescape renderer 501 includes a scale factor calculation block 925 and a digital drive value calculation block 930.
[0145] In this example, during replay, the lightscape renderer 501 receives object-based lighting data 505 as input, which includes the intended environmental lighting metadata. The object-based lighting data 505 may be generated, for example, by the lightscape creation tool 100 in Figure 5. In this example, the rendering engine also receives environment and lighting equipment data 104 and calculates appropriate lighting equipment control signals 515 for controllable lighting equipment 108 (not shown) in the replay environment. In this example, the environment and lighting equipment data 104 is shown to include environment and lighting equipment data 104a, which includes information about the lighting equipment in the replay environment and their capabilities, and environment and lighting equipment data 104b, which includes information about the ambient light and / or uncontrollable lighting equipment in the replay environment. In some examples, the lightscape renderer 501 may be configured to output lighting equipment control signals 515 to a lighting controller 103 configured to control the lighting equipment 108. The lighting equipment 108 in the replay environment may also be referred to herein as “endpoint dynamic lighting elements”.
[0146] In some examples, environment and lighting equipment data 104a may include information about a set of N dynamic lighting elements, where N represents the total number of controllable dynamic lighting elements in a reproduction environment. According to this example, for each dynamic lighting element n, the environment and lighting equipment data 104a includes (1) an IBL map 905 of ambient lighting generated by dynamic (controllable) lighting equipment at maximum intensity, and (2) a mapping 910 that maps light intensity generated by the controllable lighting equipment to a digital drive signal provided to the controllable lighting equipment. In this example, environment and lighting equipment data 104b includes ambient light information 915 and ambient light level information 920. In some examples, the ambient light information 915 may include a base IBL map of base environment lighting that is not controllable by the system. According to some examples, the ambient light information 915 may include information about other light sources, such as windows in the reproduction environment, the direction the windows face, the amount of outdoor light received in the environment through the windows at various dates and times, and information about controllable window shades if such shades are present. In some examples, the ambient light level information 920 may include, for example, a scale value of the base environment lighting acquired by an optical sensor.
[0147] In some implementations, a lightscape renderer, such as the lightscape renderer 501 in FIG. 5, may perform the following operations. 1. Receiving, as input (herein as part of object-based lighting data 505), information related to an intended environment lighting IBL ref . 2. Calculating a base lighting IBL base , for example, according to a constant estimation value, or by scaling the base IBL map with estimated ambient light. 3. For each dynamic lighting element IBL n , calculating a linear scaling Scale n such that a sum of each corresponding scaled IBL n and the base lighting most closely matches the intended environment lighting, for example, as follows. abs(IBL ref -(IBL base +sum(Scalen *IBL n Minimize ))) A nonlinear function may be used to encode each of the linear values before calculating the difference. The nonlinear function may correspond to the sensitivity of human vision to color and light intensity. 4. For each dynamic lighting element, the scaled value Scale n From digital drive signal n (This is an instance of the lighting equipment control signal 515 in Figure 5, and is therefore represented as element 515 in Figure 9A) is calculated. 5. The digital drive value is transmitted to the dynamic lighting element or the lighting controller API 103 shown in Figure 5.
[0148] In some cases, calculating the scale factor is the intended illumination IBL ref From ambient light IBL base Subtract, then, dynamic lighting IBL n Using deconvolution (reverse convolution) from the scale factor Scale n This may include calculating [the specified value]. Conventional deconvolution methods, including regularization as needed, can be used to improve robustness.
[0149] In some alternative examples, the process of calculating the scale factor involves, for example, initializing the scale factor to an initial value, and then using the reference (intended illumination IBL). ref This may be done iteratively by sequentially adjusting the scale value while evaluating the output against ). According to some such examples, conventional methods of gradient descent and function minimization can be used during the process of calculating the scale factor.
[0150] In some examples of calculating digital drive values from a scale factor, a look-up table (LUT), such as the dynamic illumination LUT910 in Figure 9A, may be used to determine the correct digital drive values required to achieve the desired light intensity from a particular light source. Alternatively, the functional scale factor may be derived from measured data.
[0151] [Calibration and System Configuration] One important aspect of this application is the generation of dynamic illumination IBL data for a given regeneration environment. In one embodiment, the process is as follows: 1) Establish a connection between the host and the dynamic lighting equipment (also referred to herein as controllable lighting equipment). 2) Position the IBL measurement device in a preferred observation location. Commonly used techniques include a camera that images a glossy sphere, a camera with a fisheye lens, a camera that pans around the scene, or a multi-camera setup (360° camera) that captures the scene from all directions. 3) Base ambient lighting (IBL) base ) and capture the readings from the light sensor. 4) For each dynamic lighting system, the following shall be performed: a. Set the drive value to the maximum level. b. IBL n Capture it. c. Repeat for several levels (drive values). d. Representative IBLs suitable for multiple drive levels n To construct. e. Establish a relationship between linear illumination and drive values. "Linear illumination" refers to a space where a linear change in value is perceived as a linear change by humans. Most devices do not have a linear response in this respect. In other words, if the codeword (drive value) to the actuator doubles, the perceived change in brightness does not double. Working in a linear illumination space is convenient and, when applicable, ensures that a finite amount of resolution is equally diffused across human perceptual responses. After the operation is performed in a linear illumination space and the corresponding values are derived, these values may be converted into drive values / codewords for controlling the physical device / illumination.
[0152] The above process is applicable when the base ambient lighting is relatively constant, except for overall increases and decreases. For example, there may be a window in one corner of the playback environment that causes an overall increase or decrease in ambient light, depending on the weather and time of day. In many observation environments, uncontrollable lighting can be more dynamic, such as manually controllable (not part of a dynamic setup) lighting fixtures, or automotive environments. In some examples, multiple base lighting scenarios may be captured, and during playback, ambient light measurements in the environment may be used to estimate the captured base lighting scenario that best matches the actual current base lighting conditions. Captured IBL n IBL base The relationship between the drive values may be stored in a configuration file accessible during rendering.
[0153] [Content Authoring / Mastering] The output of the content authoring / mastering step is the intended lighting environment map or IBL. ref In some examples, the intended lighting environment map may be measured directly using the measurement techniques described in the calibration section above. In other examples, the intended lighting environment map may be rendered from computer graphics software. Some disclosed examples involve dynamically changing the reference lighting to create a video sequence of the intended reference lighting environment.
[0154] Figure 9B is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of Method 950, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 950 may be performed simultaneously. Furthermore, some implementations of Method 950 may include more or fewer blocks than those illustrated and / or described. The blocks of Method 950 may be performed by one or more devices, which may include (or include) one or more instances of a control system, such as the control system 110 shown in Figure 1 and described above. For example, at least some aspects of Method 950 may be performed by an instance of the control system 110 configured to implement the lightscape renderer 501 of Figure 5 or the lightscape renderer 501 of Figure 9A. In some examples, Method 950 may be performed by one or more instances of the control system 110 configured to implement the lightscape renderer 501 of Figure 5 and the lighting controller API.
[0155] In this example, block 955 includes receiving object-based lighting data representing an intended lighting environment from a control system configured to implement a lighting environment renderer. In this case, the object-based lighting data includes lighting objects and lighting metadata. In some examples, the environment may be a real-world environment such as a room environment or a car environment. According to some examples, the environment may be a virtual environment or may include a virtual environment. In some such examples, method 950 may also include providing a virtual environment such as a game environment while providing corresponding lighting effects within a real-world environment.
[0156] In this example, block 960 includes receiving lighting information relating to the local lighting environment by the control system. In this example, the lighting information includes one or more characteristics of one or more controllable light sources in the local lighting environment. Block 960 may also include receiving environment and lighting equipment data 104 as described herein, for example with reference to Figure 5 or Figure 9A.
[0157] In this example, block 965 includes the control system determining the drive level for each of one or more controllable light sources that approximate the intended lighting environment. Here, block 970 includes the control system outputting a drive level to at least one of the controllable light sources.
[0158] In some examples, object-based lighting metadata includes time information. In some such examples, block 965 may include determining one or more drive levels for one or more time intervals corresponding to the time information.
[0159] According to some examples, method 950 may include receiving observation position information. In some such examples, block 965 may include determining one or more drive levels corresponding to the observation position information.
[0160] In some examples, the lighting information may include one or more characteristics of one or more base light sources in a local lighting environment, where the base light sources are not controllable by the control system. In some such examples, block 965 may include determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.
[0161] In some examples, object-based lighting metadata may include at least lighting object location information and lighting object color information. In some examples, the intended lighting environment may be transmitted as one or more image-based lighting (IBL) objects. In some such examples, the decision process in block 965 is to determine n controllable light sources (IBL) within the local lighting environment at maximum intensity. n This may be at least partially based on the IBL maps of ambient lighting generated by each of the following:
[0162] In some cases, Method 950 is used for base ambient lighting (IBL) that is not controllable by the control system. base This may include receiving a base IBL map of ). In some such examples, the decision process for block 965 may be based at least partially on the base IBL map.
[0163] In some examples, the decision process for block 965 is scaled IBL. n Lighting from and base lighting IBL base The sum of these is the intended lighting environment IBL ref Each dynamic lighting element IBL is designed to closely match the IBL. n Linear scaling of n ) may be based at least partially on the IBL. In some such examples, the decision process for block 965 is based on the IBL. ref And each scaled IBL n Lighting from and base lighting IBL base This may be based at least partially on minimizing the difference between the sum and the given value.
[0164] [Example of MS object properties] The following is a non-exhaustive list of possible properties of an MS object. ·Priority, ·layer, • Mixing mode, • Durability, • Effects, and • Spatial panning method.
[0165] effect As used herein, “effect” of an MS object is a synonym for the type of MS object. “Effect” is a sensory effect provided by or indicating such an MS object. If an MS object is a lighting object, its effect includes providing direct or indirect light. If an MS object is a tactile object, its effect includes providing some type of tactile feedback. If an MS object is an airflow object, its effect includes providing some type of airflow. Some examples include other “effect” categories, as will be described in more detail below.
[0166] Durability Some MS objects may include persistence properties within their metadata. For example, if a movable MS object moves around the scene, the movable MS object may persist for a certain period of time at the locations it passes through. This period may be indicated by persistence metadata. In some implementations, the MS renderer is responsible for constructing and maintaining the persistence state.
[0167] layer According to some examples, individual MS objects may be assigned to “layers” in which MS objects are grouped together according to one or more shared characteristics. For example, a layer may group MS objects together according to their intended effect or type, which may include, but is not limited to, the following: Mood / Atmosphere ·information • Rapid changes / caution Alternatively, or even further, in some examples, layers may be used to group MS objects together according to shared properties, which may include, but are not limited to, the following: ·color, ·Strength, ·size, ·shape, • Location, and • A region within space.
[0168] priority In some examples, MS objects may have a priority property that allows the renderer to determine which object should take precedence in an environment where MS objects are competing for a limited number of actuators. For example, if multiple lighting objects overlap with a single lighting fixture at a time when all lighting objects are scheduled to be rendered, the renderer may refer to the priority of each lighting object to determine which lighting object should be rendered. In some examples, priority may be defined between or within layers. According to some examples, priority may be associated with a specific property, such as intensity. In some examples, priority may be defined temporally, for example, the most recent MS object to be rendered may take precedence over previously rendered MS objects. According to some examples, priority may be used to specify which MS objects or layers should be rendered regardless of the limitations of a particular actuator system in the playback environment.
[0169] Spatial panning method The spatial panning method may define the movement of MS objects across space, how MS objects affect actuators as they move between them, and so on.
[0170] Mixed mode The mixed mode may specify how multiple objects are multiplexed onto a single actuator. In some examples, the mixed mode may include one or more of the following: • Maximum Mode: Selects the MS object that will activate the actuator the most. • Mixing Mode: Mixes some or all of the objects according to a rule set, for example, by summing the activation levels, taking the average of the activation levels, or mixing colors according to the activation level or priority level. • MaxNmix: Mixes the top N MS objects (by activation level) according to the rule set.
[0171] According to some examples, instead of (or in addition to) object-specific metadata, more general metadata may be defined for the entire multisensory content file. For example, an MS content file may include metadata such as a trim path or mastering environment.
[0172] Trim control In the context of Dolby Vision®, what is referred to as “trim control” may function as guidance on how to adjust the default rendering algorithm for a particular environment or condition at an endpoint. Trim control may specify ranges and / or default values for various properties, including saturation, tone detail, gamma, etc. For example, there may be automotive trim control that provides guidance for rendering in an automotive environment, such as including only objects of a specific priority or layer. Another example is providing trim control for environments with limited, complex, or sparse multisensory actuators.
[0173] Mastering environment A single multisensory content may include metadata about mastering environment properties such as room size, reflectivity, and ambient bias lighting levels. Specific properties may vary depending on the desired endpoint actuator. Mastering environment information can help provide a reference point for rendering within the playback environment.
[0174] [Lightscape Metadata Layer] When authoring the lightscape for media content, it can be useful to identify at least two different methods (or layers) of authoring. These layers may be used during the process of rendering MS objects authored according to lighting equipment available and controllable within the playback environment. These layers can help capture artistic intent and allow for flexibility in limitations within the playback environment, such as the number of lighting fixtures or lighting occlusions, so that the author's main intent can still be rendered, but can be scaled or otherwise modified.
[0175] In some examples, direct and indirect lighting may be assigned to different lighting metadata layers. Direct lighting object Lighting objects within the Direct Lighting Objects layer, also referred to herein as “Direct Lighting Objects,” are lighting objects that represent lighting directly visible to the content creator or end user. Examples of Direct Lighting Objects may include lighting fixtures in a scene, the sun, the moon, headlights from an approaching car, lightning during a storm, traffic lights, etc. Direct Lighting Objects may also be used to represent light sources that are part of the scene but are typically or temporarily invisible within the associated video content, for example, because they are outside the video frame or moving outside the video frame. In some examples, Direct Lighting Objects may be used to enhance or augment auditory events such as explosions, or to visually guide the trajectory of moving objects outside the video frame, etc. Relevant metadata, such as intensity, color, saturation, and position, often changes as a function of time within the media content scene.
[0176] Indirect lighting object In this specification, lighting objects in the indirect lighting object layer, also referred to as “indirect lighting objects,” are lighting objects that represent the effect of indirect lighting. For example, indirect lighting objects may be used to represent the effect of lighting emitted by a fixture, which is observed when light is reflected by one or more surfaces. Some examples of using indirect lighting objects include changing the observed color of the walls, ceiling, or floor of an environment to a color that matches the content, such as green for a forest scene, or blue for the sky or water. Indirect lighting objects may also be used to set the atmosphere of a scene and environment in a more immersive way, similar to how it is achieved by color grading video content. For example, science fiction films often use very specific (blue or greenish) video color grading palettes to enhance the sense of being in outer space. Flashback scenes often enhance the effect of timeline changes by desaturating, decolorizing, or overlaying sepia tones on the video content. All of these effects can be reproduced or approximated outside the video frame by adjusting the lighting control signals accordingly. Lighting effects corresponding to indirect lighting objects are often more static within the scene, typically less localized, and less dynamic than lighting effects corresponding to indirect lighting objects.
[0177] [Layer abstraction] Some examples involve further abstracting direct lighting object layers and indirect lighting object layers into layers that include both aspects. In some such examples, these layers may include one or more ambient layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof. In some examples, these layers may be used for, or correspond to, linear or event-based triggers within the content.
[0178] Surrounding layer Similar to indirect lighting, ambient layers can be used to set the atmosphere and tone in a space through a wash of surface colors within the replay environment. Ambient layers may also be used as a base layer for constructing a lighting scene. In some examples, ambient layers may be represented through lighting objects that cover a relatively large area. In other examples, ambient layers may be represented through lighting objects that cover a relatively small area, for example, using one or more images. According to some examples, ambient layers may be divided into zones. In some such examples, a particular lighting effect always occupies a specific area of the space. For example, the walls, ceiling, and floor within the authoring or replay environment may each be considered separate ambient layer zones.
[0179] Dynamic Layer In some implementations, dynamic layers may be used to represent the spatial and temporal variations of MS objects, such as lighting objects. Within a dynamic layer, individual MS objects may also have priorities, for example, so that one lighting object may take precedence over another when presented through a lighting fixture. Within a dynamic layer, individual MS objects may be linked to other objects, in some examples, such as audio objects (from spatial audio) or MS objects in the 3D world.
[0180] Custom Layer In some examples, custom layers can be used to design lighting sequences that can be freely assigned to lighting fixtures for functional purposes. These sequences do not necessarily have to be spatial in nature, but instead may provide additional information to the user. For example, in a game, a light strip may be assigned to indicate the player's remaining life.
[0181] Overlay Layer As some examples show, overlay layers can be used to present persistent lighting with a continuous priority. Overlay layers may also be used, for example, to create a "watermark" on top of all other elements in a lighting scene.
[0182] [Authoring and distribution of lightscape layer data] Authoring of direct lighting objects and indirect lighting objects In some examples, a direct lighting object may be authored by determining or setting the light source position, intensity, hue, saturation, and spatial range as a function of time for one or more lighting objects. In some such examples, this authoring process may create corresponding metadata that can be delivered with the direct lighting object along with the audio and / or video content of the content presentation. Ideally, the direct lighting object is rendered to a direct light source.
[0183] Indirect lighting effects may, in some cases, be authored as a dedicated group or class within the lightscape metadata content, focusing on overall color and atmosphere rather than dynamic effects. Indirect lighting effects may also be defined by intensity, hue, saturation, or a combination thereof as a function of time, but are typically associated with a significant portion of the lightscape rendering environment. Indirect lighting effects are rendered with indirect light sources, ideally (though not always), when available.
[0184] Figure 10 shows another example of a GUI that may be presented by a display device of a lightscape creation tool. As with other figures provided herein, the types and number of elements shown in Figure 10 are provided merely as examples. Other GUIs presented by a lightscape creation tool may include more, fewer, and / or different types and numbers of elements. According to some examples, GUI 1000 may be presented to a display device in accordance with commands from an instance of control system 110 of Figure 1 configured to implement the lightscape creation tool 100 of Figure 5.
[0185] In this example, the user may interact with GUI1000 to create lighting objects and assign lighting object properties that can be associated with the lighting objects as metadata. According to this example, GUI1000 includes a direct lighting object metadata editor section 1005 in which the user can interact to define metadata properties of direct lighting objects, and an indirect lighting object metadata editor section 1010 in which the user can interact to define metadata properties of indirect lighting objects.
[0186] The user may interact with the direct lighting object metadata editor unit 1005 to select the position, size, and other properties of the direct lighting object. In this example, the direct lighting object metadata editor unit 1005 includes a hue-saturation-lightness (HSL) color wheel 1035a, and the user may interact with this color wheel to select the HSL attributes of the selected direct lighting object. In this example, the direct lighting object metadata editor unit 1005 represents direct lighting objects A, B, and C in a three-dimensional space 1031 representing the playback environment. In this example, direct lighting objects A, B, and C are observed from above in the three-dimensional space 1031 along the z-axis.
[0187] In this example, the user has selected direct lighting object A and is currently selecting its properties. Because the user has selected direct lighting object A, the corresponding time automation lanes for the coordinates (X, Y, Z), object range (E), and HSL values over time for direct lighting object A are visible and editable in area 1025. In some examples, the time intervals corresponding to the time automation lanes shown in Figure 10 may be on the order of one second or more, for example, 1 second, 2 seconds, 3 seconds, 4 seconds, 5 seconds, etc. In this example, only the x and y dimensions of the three-dimensional space 1031 are shown, but nevertheless, area 1025 of the direct lighting object metadata editor section 1005 allows the user to specify the x, y, and z coordinates.
[0188] In this example, the indirect lighting object metadata editor unit 1010 includes an HSL color wheel 1035b, and the user may interact with the HSL color wheel 1035b to select the HSL attributes of the selected indirect lighting object. According to this example, the indirect lighting object metadata editor unit 1010 also includes an intensity control 1030, and the user may interact with the intensity control 1030 to select the intensity of the selected indirect lighting object.
[0189] In some examples, the Lightscape Creator or MS Content Creator tool may allow content creators to set up lightscape indirect lighting effects for specific scenes by setting up indirect lighting effects linked to video scene boundaries. Alternatively, or further, content creators may choose to modify intensity, hue, saturation, etc., as a function of time. In some examples, the Lightscape Creator or MS Content Creator tool may allow content creators to determine indirect lighting effect metadata using video color overlay information used during video content creation. In some examples, the indirect lighting settings of the Lightscape Creator or MS Content Creator tool may be used as a color / hue / saturation / intensity overlay on direct lighting object metadata so that direct lighting objects adhere more closely to the indirect lighting properties.
[0190] Layer abstraction authoring In some examples, layers may be authored in a lightscape creation tool or MS content creation tool configured for linear-based content creation, and individual “objects” may be assigned to layers having properties such as color, intensity, shape, and position. In some such examples, position may be specified only for dynamic objects, while subzones may be used for surrounding layer objects.
[0191] According to some examples, MS content, such as lighting-based content, may be created for a 3D world. Some such examples allow for event-based triggers, for example, triggers that link events to existing lighting metadata and the creation of new scenes.
[0192] In some examples, lightscape creation tools or MS content creation tools may allow blending between layers or individual objects. For example, it may be desirable for layers / objects with the same priority to blend additively, with higher priority objects obscuring all other objects. Lightscape creation tools or MS content creation tools may also allow content creators to define blending rules that correspond to their intentions.
[0193] [Rendering lightscape layer data] Layer attributes and their metadata may be rendered by a lightscape renderer, such as the lightscape renderer 501 in Figure 5. In some examples, the lightscape renderer may be configured to transmit lighting equipment control signals 515 to the lighting equipment of the reconstructed environment. In some examples, the lightscape renderer 501 may be configured to output lighting equipment control signals 515 to a lighting controller 103 configured to control the lighting equipment 108. According to some examples, the lightscape renderer uses environment and lighting equipment data 104 to determine the capacity and spatial location of each lighting equipment. In some examples, layer priority may also be a determinant in what is ultimately rendered to the lighting equipment.
[0194] Rendering of direct lighting objects and indirect lighting objects In some implementations, the environmental and lighting equipment data 104 received by the lightscape renderer includes data on whether the equipment is directly visible from the observation position or an indirect light source. In some examples, if indirect lighting equipment is not available, the indirect lighting data may be sent to the direct lighting equipment instead, with potentially reduced brightness.
[0195] Direct lighting objects are rendered for visible lighting fixtures such as ceiling downlights, wall-mounted lights, and table lamps. Indirect lighting metadata ideally targets lighting fixtures that are not directly visible, such as LED strips illuminating walls, ceilings, shelves, and furniture, and spotlights illuminating walls or ceilings. If such indirect lighting is not available, indirect lighting metadata may be used to control direct lighting instead. In some such examples, the lightscape renderer may overlay direct lighting object metadata and indirect lighting object metadata when rendering for lighting fixtures that function as both indirect and direct light sources.
[0196] Figure 11 is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of Method 1100, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 1100 may be performed simultaneously. Furthermore, some implementations of Method 1100 may include more or fewer blocks than those shown and / or described. The blocks of Method 1100 may be performed by one or more devices, which may be (or include) one or more instances of a control system, such as the control system 110 shown in Figure 1 and described above. For example, at least some aspects of Method 1100 may be performed by an instance of the control system 110 configured to implement the MS renderer 001 in Figure 4, the lightscape renderer 501 in Figure 5, and / or the lightscape renderer 501 in Figure 9A.
[0197] In this example, block 1105 comprises receiving, by a control system configured to implement a sensory renderer, one or more sensory objects indicating an intended sensory effect to be provided within a sensory actuator reproduction environment, and corresponding sensory object metadata. In some examples, the environment may be an actual real-world environment such as a room environment or an automobile environment. According to some examples, the environment may be a virtual environment, or may comprise a virtual environment. In some such examples, while method 1100 provides a virtual environment such as a game environment, it may also comprise providing a corresponding lighting effect within a real-world environment.
[0198] According to this example, block 1110 comprises receiving reproduction environment information by the control system. In this example, the reproduction environment information comprises sensory actuator position information and sensory actuator characteristic information regarding one or more controllable sensory actuators within the sensory actuator reproduction environment. Block 1110 may comprise, for example, receiving environment and actuator data 004 described herein with reference to FIG. 4, or receiving environment and lighting fixture data 104 described herein with reference to FIG. 5 or FIG. 9A.
[0199] In this example, block 1115 comprises determining, by the control system, based on the reproduction environment information, the sensory objects, and the sensory object metadata, a sensory actuator control command or a sensory actuator control signal for controlling one or more controllable sensory actuators within the sensory actuator reproduction environment. Here, block 1120 comprises outputting, by the control system, the sensory actuator control command or the sensory actuator control signal for at least one of the controllable sensory actuators within the sensory actuator reproduction environment.
[0200] In some examples, one or more sensory objects may include one or more lighting objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or a combination thereof. According to some examples, sensory object metadata includes lighting objects and lighting metadata, and lighting metadata may include direct lighting object metadata, indirect lighting metadata, or a combination thereof. Alternatively or further, lighting metadata may be organized into one or more layers, which may include one or more ambient layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.
[0201] According to some examples, sensory object metadata includes temporal information. In some such examples, block 1115 may include determining one or more drive levels for one or more temporal intervals corresponding to the temporal information.
[0202] In some examples, method 1100 may include receiving observation position information. In some such examples, block 1115 may include determining one or more drive levels corresponding to the observation position information.
[0203] In some examples, the lighting information may include one or more characteristics of one or more base light sources in a local lighting environment, where the base light sources are not controllable by the control system. In some such examples, block 1115 may include determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.
[0204] In some examples, sensory object metadata may include sensory object size information. In some examples, each sensory object may have a sensory object effect property that indicates the type of effect the sensory object provides. In some examples, one or more sensory objects may have a persistence property that indicates the duration for which the sensory object persists within the sensory actuator regeneration environment.
[0205] In some examples, one or more sensory objects may be assigned to one or more layers in which sensory objects are grouped according to shared sensory object characteristics. In some such examples, one or more layers may group sensory objects according to mood, atmosphere, information, attention, color, intensity, size, shape, location, area in space, or a combination thereof.
[0206] In some examples, at least some of the sensory objects may have a priority property indicating the relative importance of each sensory object. In some examples, one or more of the sensory objects may have a spatial panning property indicating how a sensory object can move within a sensory actuator regeneration environment, how a sensor object affects a controllable sensory actuator within a sensory actuator regeneration environment, or a combination thereof. In some examples, at least some of the sensory objects may have a mixed-mode property indicating how multiple sensory objects can be reproduced by a single controllable sensory actuator.
[0207] Some examples of Method 1100 may include providing and / or processing more general metadata for the entire multisensory content file instead of (or in addition to) per-object metadata. This more general metadata may be called “comprehensive sensory object metadata”. Some examples of Method 1100 may include receiving comprehensive sensory object metadata, including trim control information, mastering environment information, or a combination thereof, by a control system. In some such examples, block 1115 may include determining a sensory actuator control command or sensory actuator control signal based at least in part on the trim control information, mastering environment information, or a combination thereof.
[0208] [Bitstream containing multisensory objects] This section discloses various types of encoded bitstreams for carrying object-based multisensory data for rendering with any multiple actuators. Such bitstreams may be referred to herein as containing encoded object-based sensory data or containing encoded object-based sensory data streams. Some encoded object-based sensory data streams may be delivered together with and / or as part of other media content. In some such examples, the object-based sensory data stream may be interleaved or multiplexed with audio and / or video bitstreams. According to some examples, the object-based sensory data stream may be composed of an International Standards Organization (ISO) based media file format (ISOBMFF), and as a result, the encoded object-based sensory data can be provided within the encoded ISOBMFF bitstream together with corresponding audio data, video media, or both. Thus, in some examples, the encoded bitstream may include an encoded object-based sensory data stream and an encoded audio data stream and / or an encoded video data stream. In some examples, encoded object-based sensory data streams and other related data streams include relevant synchronization data, such as timestamps, to enable synchronization between different types of content. For example, if an encoded bitstream includes encoded audio data, encoded video data, and encoded object-based sensory data, then the encoded audio data, encoded video data, and encoded object-based sensory data may all include relevant synchronization data.
[0209] Just as there are channel-based surround sound formats such as Dolby Digital (AC3) and Dolby Digital Plus, and object-based surround sound formats such as Dolby Atmos (DD+AJOC and AC4-JOC), we introduce an object-based sensory data format here. These concepts are summarized in the table below. [Table 1]
[0210] The following are some examples of artistic intentions that can be communicated using encoded object-based sensory data streams. 1. Turn all the lights in the room to bright white. 2. Turn off all lights in the room at the timestamp 1:15. 3. Over the next 10 seconds, slowly transition the lighting at the back of the room (behind the audience) to red. 4. Over the next 15 seconds, simulate a helicopter flying over the audience with a searchlight by sequentially turning on any available overhead lights in the room from front to back, and then turning them off. 5. Generate an orange light in the front left corner of the room. 6. Create an airflow of 5 knots from the front right corner of the room.
[0211] Readers should note that these examples do not require any content authoring knowledge of a specific set of actuators present in the playback environment or the locations of those actuators. Instead, an MS renderer, such as MS renderer 001 in Figure 4, is configured to control a specific set of actuators in a particular playback environment, based in particular on (a) general instructions found in object-based sensory data 005 and (b) information in the environment and actuator data. In some examples, as shown in Figure 4, the object-based sensory data 005 may be provided to the MS renderer 001 by an experience player 002, which may include a bitstream decoder configured to extract the object-based sensory data 005 from a bitstream that also contains encoded audio and / or video data.
[0212] Linear audio / video media content is, as before, packed into a container format within a frame, with each frame containing the information necessary to render the media during a specific duration of the content (for example, a 60ms period from 10 minutes 3.2 seconds to 10 minutes 3.8 seconds from the start of the content).
[0213] Some format streams may include separate streams for each modality, for example, one base stream containing video information encoded using High Efficiency Video Coding (HEVC), also known as H.265; one base stream containing audio information encoded using Advanced Audio Coding (AAC) or Dolby AC4; and a third base stream containing closed caption (subtitle) information. In some examples, there may be multiple base streams that can be selected or combined at rendering time, such as audio tracks featuring different languages, director's commentary that can be optionally mixed with one or more other audio tracks during playback, and closed captions in multiple languages.
[0214] This disclosure extends and generalizes existing bitstream encoding and decoding methods to include, in some examples, multiple base streams that transmit non-channel-based (object-based or spherical harmonic-based, etc.) multisensory information suitable for demultiplexing, frame reassembly, and presentation using multiple actuators, synchronized with the audio and / or video modality of a media stream.
[0215] Figure 12 shows exemplary elements of a system for creating and reproducing multi-sensory (MS) experiences. As with other figures provided herein, the types and number of elements shown in Figure 12 are provided merely as examples. Other implementations may include more, fewer, and / or different types and numbers of elements. According to some examples, system 1200 may be one or more devices configured to perform at least some of the methods disclosed herein, or may include such devices. In some examples, system 1200 may include one or more instances of control system 110 of Figure 1 configured to perform at least some of the methods disclosed herein. In this example, system 1200 includes instances of some elements described with reference to Figure 4.
[0216] In this example, system 1200 includes the following elements: 1200: A system configured to receive and process an encoded bitstream containing multiple dataframes, wherein the dataframes contain encoded audio data, encoded video data, and encoded object-based sensory data. 1201: Symbolized bit stream. In some examples, data frames of the encoded bit stream may be placed in an International Standards Organization (ISO) base media file format (ISOBMFF) "container" and / or encoded in accordance with Moving Picture Experts Group (MPEG) standards. 1201A~C: Multiplexed sequence of packets / parts / frames of the encoded bit stream 1201 1202A~B: Elements of an audio data stream. In some examples, the audio data stream may be encoded in accordance with Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4 or Dolby Atmos codecs. 1203A~B: Elements of a video data stream. In some examples, the video data stream may be encoded in accordance with Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC) or AOMedia Video 1 (AV1) codecs. 1204A~B: Elements of the encoded sensory data stream, which in this example is an object-based sensory data stream. In this example, the encoded object-based sensory data includes sensory objects and corresponding sensory object metadata that represent the intended sensory effects to be provided via sensory actuators within the playback environment. In some implementations, there may be one object-based sensory data stream (which may also be called the "basic stream") for each modality (e.g., an object-based lighting data stream, an object-based temperature data stream, an object-based airflow data stream, etc.). In other implementations, these modalities may be combined into a single encoded sensory data stream. In this example, the encoded audio data stream, encoded video data stream, and encoded object-based sensory data stream all include relevant synchronization data, such as timestamps. 002: An instance of the experience player 002 in Figure 4, including the following. 1206: Bitstream demultiplexer, 1207: Audio stream decoder, 1208: Video stream decoder, and 1209:A Multisensory Data Stream Decoder 001: Multisensory renderer - Uses the environment and actuator data 004 and the decoded multisensory stream to generate control signals for driving multiple actuators 008. In this example, the MS renderer 001 includes the functionality of the MS controller API 003, which is described with reference to Figure 4. According to this example, the MS renderer 001 can determine how to synchronize the playback of audio, video, and sensory data according to synchronization data in the encoded bitstream 1201. 004: Environment and actuator data - including information about the physical location of the actuators within the environment, and visibility / impact zone information describing how each actuator is perceived by the audience (e.g., whether and how each lighting fixture is visible to the observer). 1213A, 1213B, 1213C...: Multiple individual control streams for each actuator 008 008: Multiple actuators under the control of the multi-sensory renderer 001 008A and 008B: Smart Lamp 008C: Smart RGB Light-Emitting Diode (LED) Strip 008D: Other actuators in the regeneration environment
[0217] [Embodiment 1: One basic stream for each modality] In some examples, the multisensory data stream may be transmitted in multiple base streams (e.g., one base stream for illumination information, one base stream for airflow information, one base stream for tactile information, one base stream for temperature information, etc.). According to some implementations, multiple versions of one or more multisensory data streams may reside in the encoded bitstream 1201 to allow selection based on user preference or user requirements (e.g., a default or standard illumination data stream for a typical observer, and a separate illumination data stream containing more subtle illumination information intended to be safe for light-sensitive observers).
[0218] [Embodiment 2: Combined multisensory basic streams] In some alternatives, all multisensory modalities (e.g., lighting, airflow, touch, temperature) may be combined into an integrated multisensory data stream.
[0219] [Embodiment 3: Multisensory information combined with an existing basic stream or placed in an existing container format] In another alternative embodiment, the multisensory data stream may be embedded within one of the existing base streams. For example, some audio stream formats may include functionality for encapsulating stream synchronization metadata within them. In some examples, the multisensory data stream may be embedded within an existing audio metadata transport mechanism, for example, within a field or sequence of fields reserved for audio metadata. In some alternative examples, existing audio, video, or container formats may be modified or adapted to allow the inclusion of synchronized multisensory information.
[0220] As described above, in some examples, the encoded bitstream data frame may be placed in an ISO-based media file format (ISOBMFF) file or "container". According to some implementations, the encoded object-based sensory data may reside in a timed metadata track, as defined, for example, in section 12.9 of the ISO / IEC 14496-12:2022 standard, which is incorporated herein by reference. According to some such examples, the encoded audio data may reside in an audio track, and / or the encoded video data may reside in a video track of the ISOBMFF file. Such examples have a variety of potential advantages, including, but not limited to, the following: • The same time-stamped metadata track may be associated with more than one track. In other words, a time-stamped metadata track corresponding to encoded object-based sensory data may be independent of the content of the associated audio / video track. • It may become easier to add time-based metadata tracks to files. • The duration of the time-specified metadata sample does not need to match the duration of the associated audio and / or video data.
[0221] Figure 13 is a flowchart outlining an example of a method that may be performed by an apparatus or system such as those disclosed herein. The blocks of Method 1300, as with other methods described herein, are not necessarily performed in the order shown. In some implementations, one or more blocks of Method 1300 may be performed simultaneously. Furthermore, some implementations of Method 1300 may include more or fewer blocks than those illustrated and / or described. The blocks of Method 1300 may be performed by one or more devices, which may be (or include) one or more instances of a control system, such as the control system 110 shown in Figure 1 and described above. For example, at least some aspects of Method 1300 may be performed by an instance of the control system 110 configured to implement the experience player 002 in Figure 4 or Figure 12.
[0222] In this example, block 1305 includes receiving an encoded bitstream containing multiple data frames by a control system configured to implement a demultiplexing module. The demultiplexing module may be, for example, an instance of the demultiplexer 1206 in Figure 12. In this example, the data frames include encoded audio data, encoded video data, and encoded object-based sensory data. According to this example, the encoded object-based sensory data includes sensory objects and corresponding sensory object metadata that represent the intended sensory effects to be provided within the sensory actuator playback environment. In this example, the encoded audio data stream, encoded video data stream, and encoded object-based sensory data stream all include associated synchronization data.
[0223] In this example, block 1310 includes the control system extracting an encoded audio data stream, an encoded video data stream, and an encoded object-based sensory data stream from the encoded bitstream. Block 1310 may also include parsing and / or demultiplexing the encoded bitstream 1201, as described with reference to Figure 12, for example.
[0224] In this example, block 1315 includes the control system providing the encoded audio data stream to the audio decoder. Block 1315 may also include, for example, the demultiplexer 1206 in Figure 12 providing the encoded audio data stream to the audio decoder 1207.
[0225] Here, block 1320 includes the control system providing the encoded video data stream to the video decoder. Block 1320 may also include, for example, the demultiplexer 1206 providing the encoded video data stream to the video decoder 1208.
[0226] In this example, block 1325 includes the control system providing an encoded object-based sensory data stream to a sensory data decoder. Block 1325 may also include, for example, the demultiplexer 1206 providing an object-based sensory data stream to a multisensory data stream decoder 1209.
[0227] In some examples, sensory object metadata may include sensory object location information, sensory object size information, or both. In some examples, each sensory object may have a sensory object effect property indicating the type of effect the sensory object provides. In some examples, one or more sensory objects may have a persistence property indicating the duration for which the sensory object persists within the sensory actuator playback environment. In some examples, one or more sensory objects may be assigned to one or more layers in which sensory objects are grouped according to common sensory object characteristics. In some examples, one or more layers may group sensory objects according to one or more of the following: mood, atmosphere, information, attention, color, intensity, size, shape, location, area in space, or a combination thereof. In some examples, at least some of the sensory objects may have a priority property indicating the relative importance of each sensory object.
[0228] According to some examples, an object-based sensory data stream may include lighting objects, tactile objects, airflow objects, position actuator objects, or a combination thereof. According to some examples, where the sensory object metadata includes lighting objects and lighting metadata, the lighting metadata may include direct lighting object metadata, indirect lighting metadata, or a combination thereof. Alternatively, or further, the lighting metadata may be organized into one or more layers, which may include one or more ambient layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.
[0229] In some examples, Method 1300 may include decoding an encoded object-based sensory data stream and providing the decoded object-based sensory data stream, including associated sensory synchronization data, to a sensory data renderer. In some such examples, Method 1300 may also include receiving the decoded object-based sensory data stream by the sensory data renderer and receiving playback environment information by the sensory data renderer. The playback environment information may be instances of the environment and actuator data 004 described herein. Thus, the playback environment information may include sensory actuator position information and sensory actuator characteristic information relating to one or more controllable sensory actuators in a sensory actuator playback environment.
[0230] According to some such examples, Method 1300 may include a sensory data renderer determining a sensory actuator control command or sensory actuator control signal for controlling one or more controllable sensory actuators in a sensory actuator playback environment, at least partially based on (a) sensory objects, sensory object metadata and associated sensory synchronization data from a decoded object-based sensory data stream, and (b) playback environment information. According to some such examples, Method 1300 may also include the sensory data renderer outputting a sensory actuator control command or sensory actuator control signal for at least one of the controllable sensory actuators in the sensory actuator playback environment. The sensory actuator control command may be provided to, for example, an MS controller API 003 as described with reference to Figure 4. The sensory actuator control signal may be provided to, for example, an actuator 008 in the playback environment.
[0231] In some examples, each data frame of a multi-dataframe may include an encoded audio data subframe, an encoded video data subframe, and an encoded object-based sensory data subframe. In some examples, the data frames of a bitstream may be encoded according to the Moving Picture Experts Group (MPEG) standard. In some examples, the data frames of a bitstream may consist of an ISO-based media file format (ISOBMFF). In some examples, the encoded object-based sensory data may reside in a time-specified metadata track.
[0232] In some examples, the audio data stream may be encoded according to the Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4, or Dolby Atmos codec. In some examples, the video data stream may be encoded according to the Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or AOMedia Video 1 (AV1) codec.
[0233] Various features and embodiments can be understood from the following listed exemplary embodiments ("EEE").
[0234] EEE1. A method for rendering an intended lighting environment. The steps include: receiving object-based lighting data representing the intended lighting environment by a control system configured to implement a lighting environment renderer, wherein the object-based lighting data includes lighting objects and lighting metadata; The steps include: receiving lighting information about a local lighting environment by a control system, wherein the lighting information includes one or more characteristics of one or more controllable light sources within the local lighting environment; The control system determines the drive level of one or more controllable light sources that approximate the intended lighting environment, The control system outputs a drive level to at least one of the controllable light sources. A method that includes this.
[0235] EEE2. Object-based lighting metadata includes time information. The method according to EEE1, wherein the determination step includes determining one or more drive levels for one or more time intervals corresponding to time information.
[0236] EEE3. Further includes the step of receiving observation location information, The method according to EEE1 or EEE2, wherein the determination step includes determining one or more drive levels corresponding to observation position information.
[0237] EEE4. Lighting information includes one or more characteristics of one or more base light sources in the local lighting environment, and the base light sources are not controllable by the control system. The method according to any one of EEE1 to EEE3, wherein the determination step includes determining one or more drive levels based at least in part on one or more characteristics of one or more uncontrollable light sources.
[0238] EEE5. Object-based lighting metadata is the method described in any one of EEE1 to EEE4, which includes at least lighting object location information and lighting object color information.
[0239] EEE6. The intended lighting environment is transmitted as one or more image-based lighting (IBL) objects, as described in any one of EEE1 to EEE5.
[0240] EEE7. The step to determine is to select n controllable light sources (IBL) within the local lighting environment at maximum intensity. n The method described in EEE6, which is at least partially based on IBL maps of ambient lighting generated by each of the following.
[0241] EEE8. The step to determine is that base ambient lighting (IBL) is not controllable by the control system. base The method described in EEE7, which is at least partially based on the base IBL map of ).
[0242] EEE9. The steps to be determined are each scaled IBL n Lighting from and base lighting IBL base The sum of these is the intended lighting environment IBL ref Each dynamic lighting element IBL is designed to closely match the IBL. n Linear scaling of n A method of EEE8, at least in part, based on ).
[0243] EEE10. The step to decide is IBL ref And each scaled IBL n Lighting from and base lighting IBL base A method according to EEE9, which is at least partially based on minimizing the difference between the sum and .
[0244] An apparatus configured to perform the method described in any one of the items EEE11.EEE1 to EEE10.
[0245] A system configured to perform any one of the following: EEE12.EEE1 to EEE10.
[0246] EEE13. One or more non-temporary computer-readable media storing instructions for controlling one or more devices to perform the actions described in any one of the items EEE1 through EEE10.
[0247] EEE14. A method for providing an intended sensory experience, A control system configured to implement a sensory renderer receives one or more sensory objects and corresponding sensory object metadata that represent the intended sensory effects to be provided within a sensory actuator playback environment. The control system receives playback environment information, the playback environment information includes sensory actuator position information and sensory actuator characteristic information relating to one or more controllable sensory actuators in the sensory actuator playback environment. The control system determines a sensory actuator control command or sensory actuator control signal for controlling one or more controllable sensory actuators in the sensory actuator playback environment, based on playback environment information, sensory objects, and sensory object metadata. The control system outputs a sensory actuator control command or sensory actuator control signal for at least one of the controllable sensory actuators in the sensory actuator regeneration environment. A method that includes this.
[0248] EEE15. The method according to EEE14, wherein one or more sensory objects include one or more lighting objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or a combination thereof.
[0249] EEE16. The method according to EEE14 or EEE15, wherein the sensory object metadata includes lighting objects and lighting metadata, and the lighting metadata includes direct lighting object metadata, indirect lighting metadata, or a combination thereof.
[0250] EEE17. The method described in EEE16, wherein the lighting metadata is organized into one or more layers, including one or more ambient layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.
[0251] EEE18. Sensory object metadata includes sensory object location information, as described in any one of EEE14 to EEE17.
[0252] EEE19. Sensory object metadata includes sensory object size information, as described in any one of EEE14 to EEE18.
[0253] EEE20. The method according to any one of EEE14 to EEE19, wherein each sensory object has a sensory object effect property that indicates the type of effect the sensory object provides.
[0254] EEE21. The method according to any one of EEE14 to EEE20, wherein one or more of the sensory objects have a persistence property that indicates the duration for which the sensory object persists within the sensory actuator playback environment.
[0255] EEE22. The method according to any one of EEE14 to EEE21, wherein one or more sensory objects are assigned to one or more layers in which sensory objects are grouped according to shared sensory object properties.
[0256] EEE23. The method according to EEE22, which groups sensory objects according to mood, atmosphere, information, attention, color, intensity, size, shape, location, area in space, or a combination thereof.
[0257] EEE24. The method described in any one of EEE14 to EEE23, wherein at least some of the sensory objects have a priority property indicating the relative importance of each sensory object.
[0258] EEE25. The method according to any one of EEE14 to EEE24, wherein one or more sensory objects have spatial panning properties that indicate how a sensory object can move within a sensory actuator regeneration environment, how a sensor object affects a controllable sensory actuator within a sensory actuator regeneration environment, or a combination thereof.
[0259] EEE26. The method according to any one of EEE14 to EEE25, wherein at least some of the sensory objects have a mixed-mode property indicating how multiple sensory objects can be reproduced by a single controllable sensory actuator.
[0260] EEE27. The method according to any one of EEE14 to EEE26, further comprising the step of receiving comprehensive sensory object metadata, including trim control information, mastering environment information, or a combination thereof, by a control system.
[0261] An apparatus configured to perform the method described in any one of the items EEE28.EEE14 to EEE27.
[0262] A system configured to perform any one of the following actions: EEE29.EEE14 to EEE27.
[0263] One or more non-temporary computer-readable media storing instructions for controlling one or more devices to perform the actions described in any one of the items EEE30, EEE14 through EEE27.
[0264] EEE31. A method for decoding a bitstream, The steps include: receiving an encoded bitstream containing multiple data frames by a control system configured to implement a demultiplexing module, wherein the data frames contain encoded audio data, encoded video data, and encoded object-based sensory data, the encoded object-based sensory data containing sensory objects and corresponding sensory object metadata that represent the intended sensory effects to be provided within the sensory actuator playback environment, and all encoded audio data streams, encoded video data streams, and encoded object-based sensory data streams contain associated synchronization data; and The control system extracts an encoded audio data stream, an encoded video data stream, and an encoded object-based sensory data stream from an encoded bitstream. The control system provides the encoded audio data stream to the audio decoder, The control system provides the encoded video data stream to the video decoder, The control system provides an encoded object-based sensory data stream to a sensory data decoder. A method that includes this.
[0265] EEE32. The method according to EEE31, further comprising the step of decoding an encoded object-based sensory data stream and providing the decoded object-based sensory data stream, including associated sensory synchronization data, to a sensory data renderer.
[0266] EEE33. The step of receiving the decoded object-based sensory data stream by the sensory data renderer, The steps include: receiving playback environment information by a sensory data renderer, wherein the playback environment information includes sensory actuator position information and sensory actuator characteristic information relating to one or more controllable sensory actuators within the sensory actuator playback environment; The steps include: determining a sensory actuator control command or sensory actuator control signal for controlling one or more controllable sensory actuators in a sensory actuator playback environment, based at least partially on (a) sensory objects, sensory object metadata and associated sensory synchronization data from a decoded object-based sensory data stream, and (b) playback environment information, using a sensory data renderer; The steps include: outputting a sensory actuator control command or sensory actuator control signal for at least one of the controllable sensory actuators in the sensory actuator playback environment using a sensory data renderer; The method described in EEE32, further including the above.
[0267] EEE34. The method according to any one of EEE31 to EEE33, wherein each data frame of a plurality of data frames includes an encoded audio data subframe, an encoded video data subframe, and an encoded object-based sensory data subframe.
[0268] EEE35. The bitstream data frame is encoded according to the Moving Picture Expert Group (MPEG) standard, as described in any one of EEE31 to EEE34.
[0269] EEE36. A bitstream data frame is constructed in the International Organization for Standardization (ISO)-based media file format (ISOBMFF) as described in any one of EEE31 through EEE35.
[0270] EEE37. Encoded object-based sensory data is present in a time-specified metadata track as described in any one of EEE35 to EEE36.
[0271] EEE38. A sensory object is a method according to any one of EEE31 to EEE37, including one or more lighting objects, one or more tactile objects, one or more airflow objects, one or more position actuator objects, or a combination thereof.
[0272] EEE39. The method described in any one of EEE31 to EEE38, wherein the sensory object metadata includes lighting metadata, and the lighting metadata includes direct lighting object metadata, indirect lighting metadata, or a combination thereof.
[0273] EEE40. The method described in EEE39, wherein the lighting metadata is organized into one or more layers, including one or more ambient layers, one or more dynamic layers, one or more custom layers, one or more overlay layers, or a combination thereof.
[0274] EEE41. Sensory object metadata includes sensory object location information, sensory object size information, or both, as described in any one of EEE31 to EEE40.
[0275] EEE42. The method according to any one of EEE31 to EEE41, wherein each sensory object has a sensory object effect property that indicates the type of effect the sensory object provides.
[0276] EEE43. The method according to any one of EEE31 to EEE42, wherein one or more of the sensory objects have a persistence property that indicates the period of time that the sensory object persists within the sensory actuator playback environment.
[0277] EEE44. The method according to any one of EEE31 to EEE43, wherein one or more sensory objects are assigned to one or more layers in which sensory objects are grouped according to common sensory object characteristics.
[0278] EEE45. The method of EEE44, in which one or more layers group sensory objects according to one or more of the following: mood, atmosphere, information, attention, color, intensity, size, shape, location, area in space, or a combination thereof.
[0279] EEE46. The method according to any one of EEE31 to EEE45, wherein at least some of the sensory objects have a priority property indicating the relative importance of each sensory object.
[0280] EEE47. The audio data stream is encoded according to any one of the methods described in EEE31 to EEE46, using Advanced Audio Coding (AAC), Dolby AC3, Dolby EC3, Dolby AC4, or Dolby Atmos codecs.
[0281] EEE48. The video data stream is encoded according to any one of the following methods from EEE31 to EEE46: Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), or AOMedia Video 1 (AV1) codec.
[0282] An apparatus configured to perform the method described in any one of the items EEE49, EEE31, or EEE48.
[0283] A system configured to perform the actions described in any one of the items EEE50, EEE31, or EEE48.
[0284] One or more non-temporary computer-readable media storing instructions for controlling one or more devices to perform the actions described in any one of the items EEE51.EEE31 through EEE48.
[0285] The above description illustrates various embodiments of the Disclosure, along with examples of how aspects of the Disclosure may be implemented. The above examples and embodiments should not be considered sole embodiments, but are presented to illustrate the flexibility and advantages of the Disclosure as defined by the following claims. Based on the above disclosure and the following claims, other configurations, embodiments, implementations and equivalents will become apparent to those skilled in the art and may be adopted without departing from the spirit and scope of the Disclosure as defined by the claims.
Claims
1. The steps include: receiving a content bitstream containing encoded object-based sensory data by a control system, wherein the encoded object-based sensory data includes one or more sensory objects and corresponding sensory metadata, and the encoded object-based sensory data corresponds to sensory effects including lighting, touch, airflow, one or more position actuators, or a combination thereof, and the sensory effects are provided by one or more sensory actuators in the environment; The control system includes the steps of extracting the object-based sensory data from the content bitstream, The control system provides the object-based sensory data to the sensory renderer. A method that includes this.
2. The method according to claim 1, wherein the object-based sensory metadata includes sensory spatial metadata indicating at least a spatial location for rendering the object-based sensory data within the environment, a region for rendering the object-based sensory data within the environment, or a combination thereof.
3. The method according to claim 1, wherein the object-based sensory data does not correspond to a specific sensory actuator in the environment.
4. The method according to claim 1, wherein the object-based sensory data includes abstracted sensory reproduction information that enables the sensory renderer to reproduce one or more authored sensory effects from various sensory actuator locations in the environment via one or more sensory actuator types and via a variety of sensory actuators.
5. The steps include receiving the object-based sensory data using the sensory renderer, The steps include: receiving environment descriptor data corresponding to one or more locations of sensory actuators in the environment using the sensory renderer; The steps include: receiving actuator descriptor data corresponding to the properties of the sensory actuator in the environment using the sensory renderer; The steps include: providing the sensory renderer with one or more actuator control signals for controlling the sensory actuators in the environment to generate one or more sensory effects represented by the object-based sensory data; The method according to claim 1, further comprising:
6. The method according to claim 5, further comprising the step of providing the one or more sensory effects by the one or more sensory actuators in the environment.
7. The content bitstream also includes one or more encoded audio objects synchronized with the encoded object-based sensory data, the audio object including one or more audio signals and corresponding audio object metadata. This method is The control system includes the steps of extracting audio objects from the content bitstream, The control system provides the audio object to the audio renderer. The method according to claim 1, further comprising:
8. The method according to claim 7, wherein the audio object metadata includes at least audio object spatial metadata indicating an audio object spatial location for rendering one or more audio signals in the environment.
9. The steps include receiving one or more audio objects using the audio renderer, The steps include: receiving loudspeaker data corresponding to one or more loudspeakers in the environment using the audio renderer; The steps include: providing the audio renderer with one or more loudspeaker control signals to control one or more loudspeakers in the environment to play audio corresponding to one or more audio objects and synchronized with one or more sensory effects; The method according to claim 7, further comprising:
10. The method according to claim 9, further comprising the step of playing the audio corresponding to the one or more audio objects using one or more loudspeakers in the environment.
11. The content bitstream includes encoded video data synchronized with the encoded audio objects and the encoded object-based sensory data. This method is The control system performs the steps of: extracting video data from the content bitstream; The control system provides the video data to the video renderer, The video renderer receives the video data, The video renderer provides one or more video control signals for controlling one or more display devices in the environment to present one or more images that correspond to one or more video control signals and are synchronized with one or more audio objects and one or more sensory effects. The method according to claim 9, further comprising:
12. The method according to claim 11, further comprising the step of presenting the one or more images corresponding to the one or more video control signals by one or more display devices in the environment.
13. The method according to claim 1, wherein the environment is a virtual environment.
14. The method according to claim 1, wherein the environment is a physical, real-world environment.
15. The method according to claim 14, wherein the environment is a room environment or a vehicle environment.
16. An apparatus configured to perform the method described in any one of claims 1 to 15.
17. A system configured to perform the method described in any one of claims 1 to 15.
18. One or more non-temporary computer-readable media storing instructions for controlling one or more devices to perform the method according to any one of claims 1 to 15.