System and method for providing interactivity with a light source in a scene description
The integration of scene description data with trigger and action information in MPEG-I frameworks allows for user-specific interactions with virtual objects and light sources, addressing the lack of runtime interaction support in existing frameworks and enhancing immersive experiences.
Patent Information
- Application Number
- JP2024572709
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-16
- Filing Date
- 2023-06-07
- Publication Date
- 2025-07-30
AI Technical Summary
Existing MPEG-I scene description frameworks for immersive XR experiences lack support for user interaction with scene objects during runtime, limiting the ability to create user-specific experiences.
Introduce scene description data that includes scene element information, trigger conditions, action information, and behavior information to enable interactive actions on virtual objects and light sources, such as changing intensity, color, or position based on user interactions or environmental conditions.
Enables user-specific, interactive XR experiences by allowing real-time manipulation of virtual objects and light sources, enhancing the immersive experience.
Smart Images

Figure 2025524392000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application claims the benefit of European Patent Application No. 22305880.1, filed on June 16, 2022, which is incorporated herein by reference in its entirety.
Background Art
[0002] To generate, process, and render virtual three - dimensional (3D) scenes, various techniques are available. The information characterizing the 3D scene, called scene description, can be time - dependent, enabling the 3D scene to change in a way similar to the playback of a video. This kind of behavior can be achieved by relying on the framework defined in the ISO / IEC DIS 23090 - 14:2021(E) standard "Scene Description for MPEG media document, Information technology - Coded representation of immersive media - Part 14: Scene Description for MPEG media". A scene update mechanism based on the JSON patch protocol as defined in IETF RFC 6902 can be used to synchronize virtual content to an MPEG media stream.
[0003] The MPEG - I description (MPEG - I Scene Description) framework guarantees that time - domain media and the corresponding related virtual content are always available, but there is no description of how a user can interact with scene objects during the runtime of an immersive XR experience. Therefore, it does not provide support for user - specific XR experiences based on immersive media.
Summary of the Invention
[0004] The method according to some embodiments includes obtaining scene description data of a 3D scene. The scene description data includes scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, action information that describes at least one action to be performed on one or more scene elements associated with the action, and behavior information that associates at least one of the trigger conditions with at least one of the actions. In response to a determination that at least one of the trigger conditions is satisfied and at least a first action of the actions is associated with the trigger condition by the behavior information, the first action is performed on at least a first scene element associated with the first action. In some embodiments, at least one of the scene elements is a virtual light source, and at least one of the actions is an action performed on the virtual light source.
[0005] Actions that may be performed on a virtual light source according to some embodiments include multiplying a predetermined value by the intensity of the virtual light source, setting a value indicating whether a shadow is cast by the virtual light source, changing the color of the virtual light source, setting a distance cutoff for the virtual light source, defining an inner cone angle at which intensity attenuation begins for a virtual spotlight, defining an outer cone angle at which intensity attenuation ends for a virtual spotlight, changing the width of a virtual area light source, changing the height of a virtual area light source, applying a specified rotation to a virtual cube map light source, applying a multiplier to the irradiance factor of a virtual cube map light source, changing the resolution of an image of a virtual cube map light source, and defining an image used by a virtual cube map light source, including one or more of these.
[0006] In some embodiments, the trigger condition may be satisfied under one or more of a visibility condition that is satisfied when a specified node in the scene description is visible to a specified camera node, a proximity condition that is satisfied when the distance from a reference node (e.g., a user camera) to the specified node is within a specified range, a user input condition that is satisfied when a specific user interaction is detected, a time limit condition that is satisfied during a specified period, and a collider condition that is satisfied in response to the detection of a collision between specified nodes.
[0007] In some embodiments, the behavior information identifies at least one behavior, and the behavior information for each behavior identifies at least one of the triggers and at least one of the actions. In some embodiments, the behavior information includes, for at least one behavior, information indicating whether the identified action should be executed in response to at least one of the identified triggers, or whether the identified action should be executed only in response to all of the identified triggers. In some embodiments, the behavior information includes, for at least one behavior, information indicating the order in which the identified actions are to be executed. In some embodiments, the behavior information includes information indicating that the identified actions should be executed simultaneously.
[0008] Exemplary embodiments further include an apparatus comprising one or more processors configured to execute any of the methods described herein.
[0009] Exemplary embodiments further include a computer-readable medium including instructions for causing one or more processors to execute any of the methods described herein. The computer-readable medium may be a non-transitory storage medium.
[0010] Exemplary embodiments further include a computer program product including instructions that, when executed by one or more processors, may cause the one or more processors to perform any of the methods described herein.
[0011] Signals according to some embodiments include scene description data of a 3D scene. The scene description data includes scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, action information that describes at least one action to be performed on one or more scene elements associated with an action, and behavior information that associates at least one of the trigger conditions with at least one of the actions. In some embodiments, at least one of the scene elements is a virtual light source, and at least one of the actions is associated with the virtual light source.
[0012] A computer-readable medium according to some embodiments includes scene description data of a 3D scene. The scene description data includes scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, action information that describes at least one action to be performed on one or more scene elements associated with an action, and behavior information that associates at least one of the trigger conditions with at least one of the actions. In some embodiments, at least one of the scene elements is a virtual light source, and at least one of the actions is associated with the virtual light source.
Brief Description of the Drawings
[0013]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
[0014] An extended reality (XR) display device. An exemplary extended reality (XR) display device is shown in FIG. 1A. FIG. 1A is a schematic cross-sectional view of an operating waveguide display device. The image is projected by an image generator 102. The image generator 102 can use one or more of various techniques for projecting an image. For example, the image generator 102 can be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light emitting diode (LED) display (including an organic LED (OLED) or a micro LED (μLED) display), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other types of image generators or light engines.
[0015] The light representing the image 112 generated by the image generator 102 is coupled to the waveguide 104 by the diffraction coupler 106. The coupler 106 diffracts the light representing the image 112 into one or more diffraction orders. For example, a light ray 108, which is one of the light rays representing a part of the bottom of the image, is diffracted by the coupler 106, and one of the diffraction orders 110 (e.g., the second order) is at an angle such that it can propagate through the waveguide 104 by total internal reflection. The image generator 102 displays an image as instructed by a control module 124 that operates to render image data, video data, point cloud data, or other displayable data.
[0016] At least a portion of the light 110 coupled to the waveguide 104 by the diffraction coupler 106 is coupled out of the waveguide by the diffraction outcoupler 114. At least a portion of the light coupled out of the waveguide 104 replicates the angle of incidence of the light coupled to the waveguide. For example, in the figure, the external coupled light rays 116a, 116b, and 116c replicate the angle of the internal coupled light ray 108. Since the light exiting the outcoupler replicates the direction of the light entering the coupler, the waveguide substantially replicates the original image 112. The user's eye 118 can focus on the replicated image.
[0017] In the example of FIG. 1A, outcoupler 114 outcouples only a portion of the light with each reflection, enabling a single input beam (such as beam 108) to generate a plurality of parallel output beams (such as beams 116a, 116b, and 116c). In this way, at least a portion of the light generated from each part of the image is likely to reach the user's eye even if the eye is not perfectly aligned with the center of the outcoupler. For example, when the eye 118 moves downward, beam 116c can enter the eye even if beams 116a and 116b do not, and thus the user can still perceive the bottom of the image 112 despite the misalignment. Thus, outcoupler 114 operates partially as a vertical exit pupil expander. The waveguide may also include one or more additional exit pupil expanders (not shown in FIG. 1A) to expand the exit pupil horizontally.
[0018] In some embodiments, waveguide 104 is at least partially transparent to light originating outside the waveguide display. For example, at least a portion of the light 120 from a real-world object (such as object 122) traverses waveguide 104, enabling the user to view the real-world object while using the waveguide display. Since the light 120 from the real-world object also passes through diffraction grating 114, there are multiple diffraction orders and thus multiple images. To minimize the visibility of the multiple images, it is desirable for diffraction order 0 (no deviation by 114) to have a large diffraction efficiency for light 120 and order 0, while the higher diffraction orders have lower energy. Thus, in addition to magnifying and outcoupling the virtual image, outcoupler 114 is preferably configured to pass the zero-order real image. In such embodiments, the image displayed by the waveguide display may appear to be superimposed on the real world.
[0019] Figure 1B schematically illustrates another type of augmented reality head-mounted display that may be used in some embodiments. In XR head-mounted display device 1250, control module 1254 controls display 1256, which may be an LCD, to display an image. The head-mounted display includes a partially reflective surface 1258 that reflects the image displayed on the LCD (in some embodiments, performs both reflection and focusing) to make the image visible to the user. The partially reflective surface 1258 also allows at least a portion of the external light to pass through and enables the user to see the surroundings.
[0020] Figure 1C schematically illustrates another type of augmented reality head-mounted display that may be used in some embodiments. In XR head-mounted display device 1260, control module 1264 controls display 1266, which may be an LCD, to display an image. The image is focused by one or more lenses of display optics 1268 to make the image visible to the user. In the example of Figure 1C, the external light does not reach the user's eyes directly. However, in some such embodiments, an external camera 1270 may be used to capture an image of the external environment and display such an image on display 1266 together with any virtual content that may also be displayed.
[0021] The embodiments described herein are not limited to any particular type or structure of XR display device.
[0022] An augmented reality display device, together with its control electronics, may be implemented using a system such as the system of FIG. 1D. FIG. 1D is a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device that includes various components described below and is configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers, among other electronic devices. The elements of system 1000 can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices, for example, via a communication bus or through dedicated input ports and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.
[0023] System 1000 includes, for example, at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein. The processor 1010 can include an embedded memory, an input / output interface, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, which can include a non-volatile memory and / or a volatile memory, such as electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or optical disk drive, but is not limited thereto. The storage device 1040 can include, as non-limiting examples, an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network-accessible storage device.
[0024] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device for implementing an encoding function and / or a decoding function. As is known, a device can include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 1030 can be implemented as a separate element of the system 1000 or can be incorporated within the processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0025] The program code to be loaded into the processor 1010 or the encoder / decoder 1030 to execute the various aspects described herein can be stored in the storage device 1040 and subsequently loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of the various items during the execution of the processes described herein. Such stored items can include, but are not limited to, input video, decoded video, or a portion of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, expressions, operations, and operation logic.
[0026] In some embodiments, the memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, an external memory of the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding operations and decoding operations, such as MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding and is also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard under development by JVET).
[0027] Inputs to the elements of the system 1000 can be provided through various input devices as indicated by block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, an RF signal transmitted over the air by a broadcast station, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video, although not shown in FIG. 1C.
[0028] In various embodiments, the input device of block 1130 has respective input processing elements associated therewith known in the art. For example, the RF section can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a certain frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band that can be referred to as a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a stream of desired data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (such as an intermediate frequency or near baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. In various embodiments, the order of the above (and other) elements can be rearranged, some of these elements can be omitted, and / or other elements performing similar or different functions can be added. Adding elements can include, for example, inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0029] In addition, the USB terminal and / or the HDMI terminal can each include an interface processor for connecting the system 1000 to other electronic devices via a USB connection and / or an HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, can be implemented, for example, as needed, within a separate input processing IC or within the processor 1010. Similarly, aspects of USB or HDMI interface processing can be implemented, as needed, within a separate interface IC or within the processor 1010. Demodulation, error correction, and the demultiplexed stream are provided to various processing elements, such as the processor 1010 and an encoder / decoder 1030 that operates in combination with memory and storage elements, to process the data stream required to present to the output device.
[0030] The various elements of the system 1000 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected using an internal bus known in the art, including a suitable connection device 1140, such as an Inter-IC (I2C) bus, wiring, and a printed circuit board, and data can be transmitted between them.
[0031] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or a network card, and the communication channel 1060 can be implemented, for example, within a wired medium and / or a wireless medium.
[0032] In various embodiments, data is streamed to, or otherwise provided to, system 1000 using a Wi-Fi network, such as a wireless network like IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of such embodiments are received via communication channel 1060 and communication interface 1050 that are adapted for Wi-Fi communication. Typically, the communication channel 1060 of such embodiments is connected to an access point, or router, that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that distributes data via the HDMI connection of input block 1130 is used to provide the streamed data to system 1000. In yet other embodiments, the RF connection of input block 1130 is used to provide the streamed data to system 1000. As indicated above, various embodiments provide data in a non-streaming fashion. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0033] System 1000 can provide output signals to various output devices, including display 1100, speaker 1110, and other peripheral devices 1120. The display 1100 in various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a mobile phone, or other devices. Further, the display 1100 can be integrated with other components (e.g., like within a smartphone), or can be separate (e.g., an external monitor for a laptop). In various examples of embodiments, the other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (both terms are abbreviated as DVR), a disc player, a stereo system, and / or an illumination system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of the system 1000. For example, a disc player executes a function of playing back the output of the system 1000.
[0034] In various embodiments, the control signal is communicated between the system 1000 and the display 1100, the speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable control between devices, with or without user intervention. The output devices can be communicatively coupled to the system 1000 via dedicated connections through their respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to the system 1000 using the communication channel 1060 via the communication interface 1050. The display 1100 and the speaker 1110 may be integrated into a single unit with other components of the system 1000 in an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip, for example.
[0035] Alternatively, for example, if the RF portion of the input 1130 is part of a separate set-top box, the display 1100 and the speaker 1110 can be separated from one or more of the other components. In various embodiments where the display 1100 and the speaker 1110 are external components, an output signal can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0036] System 1000 may include one or more sensor devices 1095. Examples of sensor devices that may be used include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors may be used to determine information such as the user's position and orientation. When System 1000 is used as a control module (such as control modules 124, 1254, etc.) for an augmented reality display, the user's position and orientation may be used when determining how to render image data so that the user can perceive the correct part of the virtual object or virtual scene from the correct perspective. In the case of a head-mounted display device, the position and orientation of the device itself may be used to determine the user's position and orientation for the purpose of rendering virtual content. In the case of other display devices such as a phone, tablet, computer monitor, or television, other inputs may be used to determine the user's position and orientation for the purpose of rendering content. For example, the user may use a touch screen, keypad or keyboard, trackball, joystick, or other input to select and / or adjust the desired perspective and / or line of sight direction. If the display device has sensors such as an accelerometer and / or gyroscope, the perspective and orientation used for the purpose of rendering content may be selected and / or adjusted based on the movement of the display device.
[0037] The embodiments can be implemented by computer software implemented by the processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type suitable for the technical environment, and as a non-limiting example, can be implemented using any suitable data storage technology such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory devices. The processor 1010 can be of any type suitable for the technical environment, and as a non-limiting example, can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a multi-core architecture-based processor.
[0038] Scene description framework for XR. In an XR application, scene description is used to combine an explicit and easily analyzable description of the scene structure with some binary representations of media content.
[0039] In time-based media streaming, the scene description itself can be evolved over time to provide virtual content associated with each sequence of the media stream. For example, for advertising purposes, a virtual bottle can be displayed during a video sequence of a person drinking alcohol.
[0040] This kind of behavior can be achieved by relying on the framework defined in the ISO / IEC DIS 23090-14:2021(E) standard "Scene Description for MPEG media document, Information technology - Coded representation of immersive media - Part 14: Scene Description for MPEG media". A scene update mechanism based on the JSON Patch protocol as defined in IETF RFC 6902 can be used to synchronize virtual content to the MPEG media stream.
[0041] Figure 1E shows an example of the relationship between the scene description (stored as an item in glTF.json) within the ISOBMFF file, three video tracks, an audio track, and a JSON patch update track.
[0042] The MPEG-I description framework guarantees that the timed media and the corresponding associated virtual content are always available, but there is no description of how the user can interact with the scene objects during the runtime of the immersive XR experience. Therefore, there is no support for user-specific XR experiences for consuming immersive media.
[0043] The exemplary embodiments described herein can be used to provide a scene description that includes virtual objects or light sources but does not necessarily display or render the virtual objects or light sources even if they are available. In some embodiments, one or more of the following aspects may be considered when determining whether to display a virtual object or a light source.
[0044] When determining whether to display a virtual object or a light source, spatial aspects can be considered. For example, if the user environment is not suitable (e.g., the user is too far from the position of the rendered time-domain media), or the user is not looking in the correct direction, or the virtual object should be displayed in a user-specific area (e.g., on the user's left hand that has not yet been detected), the virtual object or the light source may not be displayed.
[0045] When determining whether to display a virtual object or a light source, temporal aspects can be considered. For example, if the user is not yet ready, or if the user wants to trigger the display of the object by themselves (e.g., using a specific gesture), the virtual object or the light source may not be displayed until an appropriate trigger is detected.
[0046] In some embodiments, in the scene description, it is specified which objects or light sources the user is permitted to operate or interact with through potential tactile feedback.
[0047] Runtime interactivity. In an exemplary embodiment, the time-evolving scene description is extended by adding information that identifies behaviors. These behaviors can be related to pre-defined virtual objects and enable runtime interactivity for the user-specific XR experience on those pre-defined virtual objects. An example of an MPEG-I node hierarchy that supports elements of scene interactivity is shown in Figure 2.
[0048] In some embodiments, these behaviors are time-evolving. In such embodiments, the behaviors can be updated through the existing scene description update mechanism.
[0049] In an exemplary embodiment, a behavior is characterized by one or more of the following properties. - One or more triggers that define the conditions to be met for its activation. - Trigger control parameters that define logical operations between defined triggers. - Actions performed in response to the activation of a trigger. - Action control parameters that define the execution order of defined actions. - Priority numbers that enable the selection of the highest-priority behavior when several behaviors occur simultaneously on the same virtual object. And - Optional interrupt actions that specify how to terminate this behavior when the behavior is no longer defined in the newly received scene update. For example, when the related object is removed, or when the behavior is no longer relevant to this current media (e.g., audio or video) sequence, the behavior is no longer defined.
[0050] By adding these behaviors, a way can be defined for the user to interact at runtime in immersive content for XR experiences.
[0051] Exemplary embodiments are described with reference to the scope of the MPEG-I scene description framework that uses the Khronos glTF extension mechanism to support additional scene description functions. However, the principles described herein are not limited to a particular scene description framework.
[0052] In an exemplary embodiment, the glTF scene description is extended to support interactivity. The interactivity extension is applied at the glTF scene level and is called MPEG_scene_interactivity (MPEG_scene_interactivity). The corresponding semantics are provided in Table 1.
[0053]
Table 1
[0054] In Table 1 and other semantic tables described in this specification, the "Usage" column indicates "M" for "Required" features and "O" for "Optional" features. However, it should be understood that such features can be "Required" or "Optional" only according to a specific proposed syntax. Features marked as "Required" are not necessarily features required to practice the present invention. For example, in some embodiments, features marked as "Required" are present to meet the expectations of certain types of parsing and rendering software. However, in other embodiments, without departing from the scope of the present disclosure, the feature may be optional or completely omitted, provided that the corresponding function is implemented using default values or not implemented at all.
[0055] Table 2 shows the parameters of an example of a trigger.
[0056]
Table 2-1
[0057]
Table 2-2
[0058] In some embodiments, instead of using string parameters related to the Khronos OpenXR interaction profile syntax to define the user body part and gesture for the USER_INPUT trigger, other syntax formats such as the syntax format used for the representation of haptic objects, where an array of vertices (geometric model) and a binary mask (body part mask) are used to specify where the haptic effect should be applied, may be used.
[0059] Exemplary parameters of actions are shown in Table 3.
[0060]
Table 3-1
[0061]
Table 3-2
[0062] Exemplary parameters of the behavior are shown in Table 4.
[0063]
Table 4
[0064] During runtime, the application can operate by iterating over each defined behavior and checking for the realization of related triggers according to the procedure shown in FIG. 3. FIG. 3 is a flowchart of a processing model for activating a trigger. At 302, a determination is made as to whether the trigger condition is satisfied for the related trigger. If the trigger condition is not satisfied, the result of that determination is stored, for example, at 304, the activation status variable is set to FALSE. If the trigger condition is satisfied, at 306, a determination is made as to whether the trigger has already been activated, for example, by checking whether the activation status is already set to TRUE. If not, at 308, the trigger activation status is set to TRUE, and at 312, the trigger is activated. If the determination at 306 indicates that the trigger has already been activated, at 310, a determination is made as to whether the related trigger is intended to be activated only once (e.g., as opposed to a trigger that is intended to be active as long as the trigger condition exists). If the trigger is intended to be activated only once and (as determined at 306) has already been activated, the trigger is not activated again. Otherwise, at 312, the trigger is activated again.
[0065] In response to the activation of a defined trigger of the behavior, the corresponding action is started. The behavior has a status in progress between the start and completion of its defined action.
[0066] In some embodiments, when several behaviors occur simultaneously and affect the same node (or other scene element) simultaneously, only the behavior with the highest priority is started.
[0067] In some embodiments, in response to receiving a new scene description, the application follows the procedure as shown in FIG. 4. FIG. 4 is a flowchart showing a processing model implemented in response to receiving a new scene description. After a new scene description is acquired (402), a determination is made as to whether the behavior has a status in progress (404). In response to the behavior not having a status in progress, the new scene description is applied (410). If the behavior has a status in progress, a determination is made as to whether the behavior is still defined (406). For example, if a related object is removed, or if the behavior is no longer relevant to this current media (e.g., audio or video) sequence, the behavior is no longer defined. If the behavior is no longer defined, an interrupt action is processed (412), the ongoing behavior is stopped (414), and then the new scene description is applied (410). In 406, if it is determined that the behavior is still defined, the ongoing behavior is continued (408) even when the new scene description is applied (410).
[0068] Examples of interactive objects. As an example of an interactive virtual object according to some embodiments, a virtual 3D advertisement object can be continuously displayed and transformed during a defined period (e.g., between 20 seconds and 40 seconds) within an intervening MPEG media sequence. In this example, when the user's left hand is detected, the virtual 3D object is placed on the user's left hand and continuously follows the user's left hand.
[0069] In this embodiment, two behaviors are defined to support this dialogue scenario. The first behavior with the following parameters: o A single trigger related to the time sequence of MPEG media between 20 seconds and 40 seconds with continuous startup ("activateOnce" = FALSE). And o Two consecutive actions (node 0) to enable and transform virtual 3D objects.
[0070] The second behavior with the following parameters: o A single trigger related to the detection of the user's left hand with continuous startup ("activateOnce" = FALSE). And o A single action to place a virtual object (node 0) on the user's left hand.
[0071] The two behaviors define the same interrupt action to disable the virtual object (node 0).
[0072] Since the two behaviors affect the same virtual object (node 0), a higher priority is set for the second behavior related to the user gesture (left hand pose) to execute the desired dialog scenario. In this embodiment, the desired behavior can be implemented using scene dialog information, and the scene dialog information can be given in the following JSON format or another format.
[0073]
Table 5-1
[0074]
Table 5-2
[0075] Runtime interactivity for the light source. In an XR application, a scene description is used to combine an explicit and easily analyzable description of the scene structure with several binary representations of media content. The above section describes the action mechanism for scene description. These behaviors are related to pre-defined virtual objects and enable runtime interactivity for the user-specific XR experience on those pre-defined virtual objects. Figure 5 shows the structure of an exemplary behavior mechanism.
[0076] Some embodiments extend the interactivity for virtual objects as described above to provide interactivity of lighting, such as the intensity of light that depends on the user's proximity (triggered, for example, by a proximity trigger), the color of light that depends on the time, and other types of interactivity. Exemplary embodiments enable special lighting interactions for any light in the scene description.
[0077] In some embodiments, the action semantics of the scene description include a SET_LIGHT field for dedicated light actions.
[0078] In Table 5 below, the term cube map refers to an array of image-based lights (also called IBL) described in the EXT_lights_image_based extension.
[0079] Exemplary embodiments include one or more of the following in the scene description extension.
[0080] ○ A parameter (Number) that multiplies the current intensity level of the light node. ○ A parameter (Boolean) that enables and disables casting shadows from the light. ○ A parameter (Array) that enables changing the color of the light. ○ A parameter (Number) that defines the distance at which the light intensity can be considered to have reached zero. ○ A parameter (Number) that changes the width of the area light. ○ Parameter (Number) for changing the height of the area light. ○ Parameter (Quaternion) that enables changing the rotation of the cube map. ○ Parameter (Array) for setting the irradiance coefficient. ○ Parameter (Number) for changing the resolution of the images that make up the cube map. ○ Parameter (Array) for replacing the images used for the cube map. And ○ Parameter (Array) indicating which light nodes are being referenced. In some embodiments, this parameter may be required.
[0081] In the example of Khronos GLTF, there are two relevant light extensions, namely, the KHR_lights_punctual extension that supports directional, point, and spot lights, and the EXT_lights_image_based extension that supports an array of image-based lights (in other words, these are cube maps that describe the spherical harmonic function coefficients of l = 2 for the specular radiance, diffuse irradiance, rotation, and intensity values of the scene).
[0082] The current implementation of KHR_lights_punctual does not yet support area lights, but this light type is efficient for integrating virtual objects in a real environment and describing real lights and exists in game engines such as Unity or Unreal Engine, so it is expected that the glTf format will add support for area lights.
[0083] This area light can be a square quad that has parameters for width and height, and whether it inherits the parent transformation like other lights (in addition to all existing light parameters).
[0084] The exemplary embodiments are described with reference to an MPEG-I scene description framework that uses the Khronos glTF extension mechanism to support additional scene description capabilities, however, the principles described herein are not limited to any particular scene description framework.
[0085] In an exemplary embodiment, a SET_LIGHT action is provided in addition to the actions already defined in Table 3. Example semantics for the SET_LIGHT action at the node level are provided in Table 5. An exemplary SET_LIGHT action indicates a change to at least one lighting characteristic of a virtual light source, such as the intensity, color, size, range, orientation, and / or shadowing characteristics of the virtual light source.
[0086] [Table 6]
[0087] As an example, information indicating runtime interactivity for a light source may be represented in a GLTF file using a JSON (JavaScript® Object Notation) format such as: However, the principles described herein are not limited to any particular syntax format.
[0088] [Table 7-1]
[0089] [Table 7-2]
[0090] [Table 7-3]
[0091] FIG. 6 is a flowchart showing a method executed according to some embodiments. Scene description data is obtained for a 3D scene. The scene description data may be in the GLTF format or in other formats. The scene description data can include scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, action information that describes at least one action to be executed on one or more scene elements associated with an action, and behavior information that associates at least one of the trigger conditions with at least one of the actions. In some embodiments, the action is executed on a scene element associated with a corresponding node in the hierarchical scene description graph. Alternatively, or in addition to this, one or more of the actions can be executed on a scene element not associated with a specific node, such as an animation or an MPEG_media element included in the scene description.
[0092] In the exemplary method shown in FIG. 6, scene description data of a 3D scene is acquired (602). The scene description data includes scene element information that describes each of a plurality of scene elements in the scene, and at least one of the scene elements is a virtual light source. The scene description data further includes trigger information that describes at least a first trigger condition. The scene description data further includes action information that describes at least a first action and associates one or more of the scene elements including the virtual light source with the at least first action. The scene description data further includes behavior information. The behavior information associates at least the first trigger condition with the at least first action. The behavior information can also identify additional behaviors that associate other trigger conditions with other actions. At 604, the scene is rendered based on the scene description data. At 606, which can be performed during the rendering of the scene, the scene (and / or the user's interaction with the scene) is monitored for trigger conditions, including monitoring for at least the first trigger condition. At 608, a determination is made as to whether the trigger condition is satisfied. If the trigger condition is satisfied, at 610, one or more actions associated with that trigger are identified. This identification can be made based on the behavior information that associates the trigger condition with the corresponding action. At 612, the identified action is executed on the node (or other scene element) identified in the action. For example, if the action is associated with a virtual light source, the action is executed on the virtual light source. The rendering of the scene 604 can continue, and the appearance or other characteristics of the scene are updated according to the one or more actions that have been executed (e.g., the characteristics of one or more virtual light sources are changed based on the action).
[0093] In some embodiments, user interactivity characteristics are monitored to detect whether trigger conditions are met. The characteristics to be monitored may include, for example, the position of the user camera or viewpoint (such as determined through user input and / or through sensors such as gyroscopes and accelerometers, among other possibilities), user gestures (such as detected through cameras, wrist-mounted or handheld accelerometers, or other sensors), or other user inputs.
[0094] In response to a determination that at least one of the trigger conditions is met and that at least a first action of the actions is associated with the trigger condition by behavior information, the first action is performed on at least a first scene element associated with the first action. In some embodiments, at least one of the scene elements is a virtual light source, and at least one of the actions is an action performed on the virtual light source.
[0095] In some embodiments, the action results in a modification to one or more scene elements within the scene description of the 3D scene. The method in some embodiments includes rendering the 3D scene according to the modified scene description. In other embodiments, the modified scene description is provided to a separate renderer to render the 3D scene according to conventional 3D rendering techniques. The 3D scene can be displayed on any display device, such as the display devices described herein. In some embodiments, the 3D scene can be displayed as an overlay with the real-world scene using an optical see-through or video see-through display. Note that the display of the 3D scene referred to herein includes displaying a 2D projection of the 3D scene using a 2D display device.
[0096] Further embodiments. This disclosure describes various aspects including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and often in ways that may seem limiting in order to show at least individual characteristics. However, this is for the purpose of clarifying the description and is not intended to limit the disclosure or scope of those aspects. In fact, all of the different aspects can be combined and replaced to provide further aspects. Additionally, these aspects can similarly be combined with and replace aspects described in previous applications.
[0097] The aspects described and contemplated in this disclosure can be implemented in many different forms. Although some embodiments are specifically shown, other embodiments are also contemplated, and the description of specific embodiments is not intended to limit the scope of implementation. At least one of the above aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These aspects, and other aspects, can be implemented as a computer-readable storage medium that internally stores instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium that stores within itself a bitstream generated according to any of the described methods.
[0098] Various methods are described herein, and each of the methods includes one or more steps or acts for achieving the described method. The order of the steps or acts, and / or the use of specific steps and / or acts, may be modified or combined, provided that a particular order of the steps or acts is not required for the proper operation of the method. Note that terms such as "first," "second," etc. may be used in various embodiments to modify elements, components, steps, acts, etc., such as, for example, "first decoding" and "second decoding." The use of such terms does not imply an ordering with respect to the modified acts, unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding and may occur, for example, before, during, or overlapping with the second decoding.
[0099] In the present disclosure, for example, various numerical values may be used. The specific values are for illustrative purposes only, and the described aspects are not limited to these specific values.
[0100] The embodiments described herein may be implemented by computer software implemented by a processor, or by other hardware, or by a combination of hardware and software. By way of non-limiting example, embodiments can be implemented by one or more integrated circuits. The processor can be of any type suitable for the technical environment and can include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a multi-core architecture-based processor.
[0101] When a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.
[0102] The implementations and aspects described herein can be implemented, for example, as a method or process, apparatus, software program, data stream, or signal. Even if considered only in the context of a single form of implementation (e.g., only considered as a method), the implementation of the considered features can also be implemented in other forms (e.g., an apparatus or program). For example, an apparatus can be implemented in appropriate hardware, software, and firmware. This method can be implemented by a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Further, the processor includes, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information among end users.
[0103] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation", and the appearance of any other variations, appear throughout this disclosure but do not necessarily all refer to the same embodiment.
[0104] Furthermore, this disclosure may refer to "determining" various pieces of information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0105] Furthermore, the present disclosure may refer to "accessing" various pieces of information. Accessing information can include, for example, one or more of receiving information, obtaining information (e.g., from memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0106] Furthermore, the present disclosure may refer to "receiving" various pieces of information. Receiving is intended to be a broad term, similar to "accessing". Receiving information can include, for example, one or more of accessing information or obtaining information (e.g., from memory). Further, "receiving" typically involves in some way during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, computing information, determining information, predicting information, or estimating information.
[0107] For example, in the case of "A / B", "A and / or B", and "at least one of A and B", it should be understood that any use of the following " / ", "and / or", and "at least one of" is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or selection of only the first and second listed options (A and B), or selection of only the first and third listed options (A and C), or selection of only the second and third listed options (B and C), or selection of all three options (A and B and C). This can be extended for as many items as are described.
[0108] Also, as used herein, the term "signaling" specifically means indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals a particular one of a plurality of parameters for selecting filter parameters based on regions for artifact removal filtering. Thus, in one embodiment, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send a particular parameter to the decoder (explicit signaling) so that the decoder can use the same particular parameter. In contrast, if the decoder already has other parameters along with that particular parameter, signaling (implicit signaling) can be used that does not perform the transmission but simply enables the decoder to know and select that particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It will be understood that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the term "signal", but the term "signal" may also be used as a noun herein.
[0109] Embodiments can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for implementing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiment. For example, such a signal can be formatted as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. As is known, signals can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium.
[0110] Some embodiments are described. The features of these embodiments can be provided singly or in any combination across various claim categories and types. Further, embodiments can include one or more of the following features, devices, or aspects, singly or in any combination, across various claim categories and types. ● The bitstream or signal includes one or more of the described syntactic elements or variants thereof. ● The bitstream or signal includes a syntax that carries information generated according to any of the described embodiments. ● Create and / or transmit and / or receive and / or decode a bitstream or signal that includes one or more of the described syntactic elements or variants thereof. ● Create and / or transmit, and / or receive and / or decode, a bitstream or signal according to any of the described embodiments. ● A method, process, apparatus, medium storing instructions, medium storing data or signals according to any of the described embodiments.
[0111] One or more of the various hardware elements of the described embodiments, in conjunction with their respective modules, are noted to be referred to as "modules" that perform (i.e., implement, execute, etc.) the various functions described herein. As used herein, a module can include hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) that is considered suitable for a given embodiment. Each of the described modules can also include executable instructions for performing one or more of the functions described as being performed by their respective modules, and those instructions can take the form of, or include, hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and can be stored on any suitable non-transitory computer-readable medium, or medium generally referred to as RAM, ROM, etc.
[0112] Features and elements are described above in specific combinations, but each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVD). A processor associated with software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method comprising: obtaining scene description data of a 3D scene, the scene description data including: scene element information describing each of a plurality of scene elements in the scene, where at least one of the scene elements is a virtual light source; trigger information describing at least a first trigger condition; action information describing at least a first action, the first action indicating a modification to at least one characteristic of the virtual light source; and behavior information associating at least the first trigger condition with at least the first action, and obtaining the scene description data of the 3D scene; and in response to a determination that at least the first trigger condition is satisfied, performing at least the first action on at least the virtual light source. A method comprising the above.
2. The method according to claim 1, wherein the first action includes multiplying a specified value by the intensity of the virtual light source.
3. The method according to claim 1 or 2, wherein the first action includes setting a value indicating whether a shadow is cast by the virtual light source.
4. The method according to any one of claims 1 to 3, wherein the first action includes changing the color of the virtual light source.
5. The method according to any one of claims 1 to 4, wherein the first action includes setting a distance cutoff for the virtual light source.
6. The method according to any one of claims 1 to 5, wherein the virtual light source is a virtual spot light source, and the first action includes defining an inner cone angle at which intensity attenuation begins for the virtual spot light source.
7. The method according to any one of claims 1 to 6, wherein the virtual light source is a virtual spot light source, and the first action includes defining an outer cone angle at which intensity attenuation ends for the virtual spot light source.
8. The method according to any one of claims 1 to 5, wherein the virtual light source is a virtual area light source, and the first action includes changing the width or height of the virtual area light source.
9. The method according to any one of claims 1 to 5, wherein the virtual light source is a virtual cube map light source, and the first action includes applying a specified rotation to the virtual cube map light source.
10. The virtual light source is a virtual cube map light source, and the first action includes applying a multiplier to the irradiance coefficient of the virtual cube map light source. The method according to any one of claims 1 to 5 or 9.
11. The virtual light source is a virtual cube map light source, and the first action includes changing the resolution of the image of the virtual cube map light source. The method according to any one of claims 1 to 5 or 9 to 10.
12. The virtual light source is a virtual cube map light source, and the first action includes defining the image used by the virtual cube map light source. The method according to any one of claims 1 to 5 or 9 to 11.
13. The first trigger condition is a visibility condition that is satisfied when a specified scene element is visible to a specified camera node. The method according to any one of claims 1 to 12.
14. The first trigger condition is a proximity condition that is satisfied when the distance from a reference node to a specified scene element is within a specified range. The method according to any one of claims 1 to 13.
15. The first trigger condition is a user input condition that is satisfied when a specified user interaction is detected. The method according to any one of claims 1 to 14.
16. The first trigger condition is a time limit condition that is satisfied during a specified period. The method according to any one of claims 1 to 15.
17. The first trigger condition is a collider condition that is satisfied in response to the detection of a collision between specified scene elements. The method according to any one of claims 1 to 16.
18. The behavior information associates the first action with a plurality of trigger conditions. The behavior information further includes information indicating whether the first associated action should be executed in response to (i) a determination that all of the plurality of trigger conditions are satisfied, or (ii) a determination that any one of the plurality of trigger conditions is satisfied. The method according to any one of claims 1 to 17.
19. The behavior information associates at least the first trigger condition with a plurality of actions. The behavior information further includes information indicating the order in which the plurality of actions are to be executed. The method according to any one of claims 1 to 18.
20. The method according to any one of claims 1 to 19, wherein the behavior information associates at least the first trigger condition with a plurality of actions, and the behavior information further includes information indicating that the plurality of actions should be executed simultaneously.
21. The method according to any one of claims 1 to 20, wherein at least one of the scene elements in the scene is a virtual object.
22. The method according to any one of claims 1 to 21, further including rendering the 3D scene according to the scene description data manipulated by the first action.
23. The method according to any one of claims 1 to 22, wherein the trigger information includes an array of all triggers in the 3D scene, and each trigger describes its respective trigger condition.
24. The method according to any one of claims 1 to 23, wherein the action information includes an array of all actions in the 3D scene, and the action information includes information associating each action with at least one respective scene element.
25. The method according to any one of claims 1 to 24, wherein the behavior information includes an array of all behaviors in the 3D scene.
26. Each action described in the action information includes action type information, and the action type information for at least the first action includes information indicating a dedicated optical action. The method according to any one of claims 1 to 25.
27. The method according to any one of claims 1 to 26, further including rendering the 3D scene, and the determination that the first trigger condition is satisfied is made during the rendering of the 3D scene.
28. An apparatus comprising one or more processors, the one or more processors being configured to execute the method according to any one of claims 1 to 27.
29. A computer-readable medium including instructions for causing one or more processors to execute the method according to any one of claims 1 to 27.
30. The computer-readable medium according to claim 29, wherein the computer-readable medium is a non-transitory storage medium.
31. A computer program product comprising instructions which, when the program is executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 29.
32. A signal comprising scene description data of a 3D scene, wherein the scene description data is scene element information describing each of a plurality of scene elements in the scene, wherein at least one of the scene elements is a virtual light source, the scene element information, trigger information describing at least a first trigger condition, and action information describing at least a first action, wherein the first action indicates a modification to at least one characteristic of the virtual light source, the action information, and behavior information associating at least the first trigger condition with at least the first action.
33. A computer-readable medium comprising scene description data of a 3D scene, wherein the scene description data is scene element information describing each of a plurality of scene elements in the scene, wherein at least one of the scene elements is a virtual light source, the scene element information, trigger information describing at least a first trigger condition, and action information describing at least a first action, wherein the first action indicates a modification to at least one characteristic of the virtual light source, the action information, and behavior information associating at least the first trigger condition with at least the first action.