Parameter selection for media playback
By adjusting the presentation rules and parameters of video content based on contextual factors within a 3D environment, the problem of not considering environmental and user factors in existing technologies is solved, achieving higher fidelity and user-friendly media content presentation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2023-09-21
- Publication Date
- 2026-04-14
AI Technical Summary
Existing electronic devices do not adequately consider the intended viewing environment, user, and other contextual factors when presenting media content, resulting in a poor user experience.
By selecting parameters based on contextual factors, the rendering rules and parameters of video content in a 3D environment are adjusted, including position, brightness, audio mode, etc., to match the specific conditions of the user and the environment.
It improves the fidelity of media content presentation and user experience in 3D environments, ensuring that the content aligns with the creator's intent and adapts to different viewing environments.
Smart Images

Figure CN117750136B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates in general to systems, methods, and apparatus for presenting media content (such as video content) in a three-dimensional (3D) environment. Background Technology
[0002] Electronic devices such as head-mounted displays (HMDs) include applications for watching movies and other media content. Such devices often display media content without adequately considering the intended viewing environment, the user, the context in which the content will be viewed, and other contextual factors. Summary of the Invention
[0003] The various specific embodiments disclosed herein include devices, systems, and methods for providing video content (e.g., television programs, recorded sports events, movies, 3D videos, etc.) within a 3D environment using parameters selected based on one or more contextual factors. Such contextual factors can be determined based on attributes of the content (e.g., its intended purpose or intended viewing environment), the user (e.g., the user's vision, interpupillary distance, viewing preferences, etc.), the 3D environment (e.g., current lighting conditions, spatial considerations, etc.), or other contextual attributes.
[0004] In some implementations, the device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method determines a context for rendering video content items within a view of a 3D environment, wherein the context is determined based on attributes of the video content items. The context may relate to the intended viewing environment or playback mode of the content (e.g., intended for use in a dark theater / cinema rather than a dimly lit room or a bright work environment, intended for HDR playback rather than SDR playback, etc.). The context may relate to attributes of the user and / or the 3D environment, such as the user's position within the physical environment upon which the 3D environment may be based. The method determines, based on the context, rendering rules for rendering the video content items within a view of the 3D environment, and determines, based on the context and rendering rules, parameters (e.g., 3D position, audio mode, brightness, etc.) for rendering the video content items within a view of the 3D environment. The method renders the video content items within a view of the 3D environment based on these parameters.
[0005] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein. Attached Figure Description
[0006] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.
[0007] Figure 1 These are exemplary physical environments in which the devices can provide a view, based on some specific implementations.
[0008] Figure 2 It is based on some specific implementation descriptions by Figure 1 The device provides a view of the physical environment.
[0009] Figure 3 It is a view that describes the physical environment and video content items using one or more context-based parameters based on some specific implementation.
[0010] Figure 4 It is a view that describes the physical environment and video content items using one or more context-based parameters based on some specific implementation.
[0011] Figure 5 It is a view that describes the physical environment and video content items using one or more context-based parameters based on some specific implementation.
[0012] Figures 6A to 6C This illustrates a view of the physical environment and video content items based on the expected viewing conditions of the video content items in some specific implementations.
[0013] Figure 7 This illustrates a view of the physical environment and video content items based on the expected viewing conditions of the video content items in some specific implementations.
[0014] Figure 8 This illustrates a method for providing views of the physical environment and video content items using context-based parameters, depending on some specific implementation.
[0015] Figure 9 Exemplary device configurations based on some specific implementations are shown.
[0016] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Detailed Implementation
[0017] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0018] Figure 1 This is a block diagram of an exemplary physical environment 100 that can be viewed according to some specific implementation, such as device 110. In this example, physical environment 100 includes walls (such as wall 120), doors 130, windows 140, plants 150, sofa 160, and table 170.
[0019] Electronic device 110 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about physical environment 100 and objects within it, as well as to capture information about user 102. Device 110 may use the information obtained from its sensors about its physical environment 100 or user 102 to provide visual and audio content. For example, such information may be used to determine one or more contextual attributes that can be used to configure parameters for video content and / or the depiction of other content within the 3D environment.
[0020] In some embodiments, device 110 is configured to present a generated view to user 102, including a view that may be based on physical environment 100 and one or more video content items. According to some embodiments, electronic device 110 generates and presents a view of an extended reality (XR) environment.
[0021] In some embodiments, device 110 is a handheld electronic device (e.g., a smartphone or tablet). In some embodiments, user 102 wears device 110 on their head. Thus, device 110 may include one or more displays provided for displaying content. For example, device 110 may surround the user 102's field of view.
[0022] In some implementations, the functionality of device 110 is provided by more than one device. In some implementations, device 110 communicates with a separate controller or server to manage and coordinate the user experience. Such a controller or server may be local or remote relative to the physical environment 100.
[0023] Device 110 obtains (e.g., by receiving information or making a determination) one or more contextual attributes it uses to provide a view to user 102. In some embodiments, video content may be displayed in different display modes (e.g., small screen mode, large screen mode, cinema / full screen mode, etc.), and the context includes the display mode already selected by the user. In some embodiments, the user selects the display mode by providing input to select the display mode. Parameters of the video content item are determined based on contextual attributes associated with the video content item, the user, or the 3D environment in which the video content item will be played. Unlike providing video content flatly on a device's screen (e.g., on a television, mobile device, or other conventional electronic device), video content provided within a view of a 3D environment or otherwise on a 3D-enabled device (such as an HMD) can be displayed in numerous different ways. Parameters controlling the display of such video content within a 3D environment on such devices can be selected desirably or optimally based on contextual factors. For example, one or more of the following factors can be determined based on contextual factors: the size of the video content item (e.g., its virtual screen), its position within the environment, its position relative to the viewer, its height, angle, brightness, color format / attribute, display frame rate, etc. In some implementations, the display of the surrounding 3D environment can optionally or additionally be adjusted based on context, such as changing brightness or color to better match or complement the video content.
[0024] Using context to determine the parameters that control how video content items or their surrounding 3D environment are presented in the view can provide a higher fidelity or otherwise more desirable user experience than simply using default or context-agnostic parameters. Furthermore, context can include information about the creator's intent regarding the video content item, such as how the creator wants the content to be experienced. Determining the parameters that control how video content items or their surrounding environment are presented in the view ensures that the video content item is experienced according to or otherwise based on that intent. For example, a video intended to be presented in a dark theater / cinema can be presented in an environment whose surroundings have been altered to appear dark. In another example, a video intended to be presented according to a specific color theme or color interpretation can be presented in an environment that matches or otherwise aligns with that color theme or color interpretation.
[0025] In some implementations, the context relates to the frame rate of the video content item. For example, a context attribute could be that the content item has a frame rate of 24 frames per second. This context can be used to adjust the view provided by the device, for example, by changing the display frame rate to match the display frame rate of the content or a multiple thereof (e.g., 24 frames per second, 48 frames per second, etc.).
[0026] In some implementations, the location of video content items (e.g., on a virtual screen) is determined based on context to optimize or otherwise provide the desired binocular viewing experience. For example, a 3D movie might be presented at a location relative to the user that is determined to provide a comfortable or desired experience. Such a location can depend on individual users, such as the user's visual quality, interpupillary distance (IPD), physiological attributes, or viewing preferences, which may individually or collectively provide contextual attributes for locating video content items relative to the viewer. The location of video content items may additionally or alternatively take into account the content's resolution or the display's resolution, for example, avoiding locations that would provide a pixelated appearance given the content's or display's resolution. Thus, for example, 4K content might be displayed larger than high-definition (HD) content. In some implementations, the content's location is chosen to occupy as large a portion of the displayed view as possible while meeting user comfort requirements, for example, not so close as to make it uncomfortable for the user. Subjective comfort can be estimated based on average / typical user physiology and preferences or based on user-specific physiology and preferences.
[0027] In some implementations, the brightness of video content items is determined based on context (e.g., on a virtual screen) to optimize or otherwise provide the desired viewing experience. Brightness parameters can be selected based on the type of content or the intended viewing environment. The content itself can be altered. For example, content intended for use in dark environments but viewed in a bright environment can be brightened. Video content items can be modified to achieve a desired effect on the user. For example, a video content item may have reduced brightness for a limited time period and then brighten for a specific scene to provide a beneficial user experience. Users may become accustomed to the lower brightness level and then experience the brightened scene as significantly brighter without actually brightening the scene beyond its original brightness level (which could be impossible given the brightness constraints of a given device).
[0028] In some implementations, video content items include metadata that provides contextual attributes, such as identifying frame rate, expected viewing environment, content type, color range / format, etc. In some implementations, video content items are examined to determine such contextual attributes. Such examinations can be performed locally (e.g., in real time on device 110 when it is ready to play video content) or remotely (e.g., on a server separate from device 110 during a previous examination process).
[0029] Figure 2 It is a description of Figure 1 The device 110 provides a view 200 of the physical environment 100. In this example, view 200 is a view depicting and enabling user interaction with real or virtual objects in an XR environment. Such a view may include optical perspective or passthrough video that provides a depiction of various parts of the physical environment 100. In one example, one or more outward-facing cameras on the device 110 capture images of the physical environment, which are passed through to provide at least some of the content depicted in view 200. In this example, view 200 includes depictions of walls (such as depiction 220 of wall 120), floors and ceilings, doors 130, windows 140, flowers 150, sofas 160, and tables 170.
[0030] Figure 3 This is a view 300 that describes the physical environment 100 and the video content item 380 using one or more context-based parameters according to some specific implementation. In this example, the context determines the parameters that define the position, orientation, and other attributes of the video content item 380 and other descriptions 220, 230, 250, 260, 270. The context may include a video content item with 4K resolution, a user who has selected a large-screen display mode, a user with an average / typical IPD, a video content item designed for a home viewing environment (not a theater / cinema), the physical environment on which the 3D environment is based, which has the spatial dimension that defines it as a large room, the location of potential viewing obstacles within the 3D environment, etc. In this example, based on such contextual attributes, device 110 determines to position the video content item 380 one foot in front of wall 220, two feet above the floor, and slightly downward at an angle (given the user's height and position). It also determines to give the video content item 380 a size that will provide a comfortable and non-pixelated view given the 4K content and the viewer's position. It also determines to maintain the brightness of the video content item 380 and the brightness of other depictions 220, 230, 250, 260, and 270, given the expected viewing environment that matches the actual 3D environment depicted in the view. Furthermore, device 110 determines to display a control panel 390 below the video content item 380.
[0031] Figure 4This is a view 400 that depicts the physical environment 100 and video content item 480 using one or more context-based parameters according to some specific implementation. In this example, the parameters that define the position, orientation, and other attributes of the video content item 480 and other depictions 220, 230, 250, 260, 270 are determined based on context. The context may include video content items with high-definition (less than 4K) resolution, users who have selected a standard screen display mode, users with average / typical IPD, video content items designed for home viewing environments (not theaters / cinemas), the physical environment on which the 3D environment is based, which has the spatial dimension that defines it as a large room, the location of potential viewing obstacles within the 3D environment, etc. In this example, based on such contextual attributes, device 110 determines to position the video content item 480 six inches in front of wall 220 and three feet above the floor (given the user's height and position). The position and size of the video content item 480 are set to block any depiction of window 140. Device 110 also determines to give the video content item 480 a size that will provide a comfortable and non-pixelated view, given the high-definition (less than 4K) resolution of the content and the viewer's position. It also determines to maintain the brightness of the video content item 480 and the brightness of other depictions 220, 230, 250, 260, and 270, given the expected viewing environment that matches the actual 3D environment depicted in the view. Additionally, device 110 determines to present a control panel 490 that overlaps with a portion of the video content item 480.
[0032] Figure 5This is a view 500 that describes the physical environment 100 and the video content item 580 using one or more context-based parameters according to some specific implementation. In this example, context-based parameters are used to determine the position, orientation, and other attributes that define the video content item 580 and other descriptions of 220, 260. Context may include a video content item with 4K resolution, a user who has selected a full-screen / cinema screen display mode, a user with average / typical IPD, a video content item designed for a home viewing environment (rather than a theater / cinema), the physical environment on which the 3D environment is based, which has the spatial dimension that defines it as a large room, the location of potential viewing obstacles within the 3D environment, etc. In this example, based on such contextual attributes, device 110 determines to position the video content item 580 directly in front of the user's view and occupy most (if not all) of the view 500. It also determines to give the video content item 580 a size that will provide an immersive view that makes full use of the device's display capabilities and is consistent with the user's attributes and preferences. It also determines to maintain the brightness of the video content item 580 and the brightness of other depictions 220 and 260 constant, given the expected viewing environment that matches the actual 3D environment depicted in the view. In various examples, the content may be intended for dim, dark, bright, or various types of viewing environments, and the brightness of the content or the surrounding environment may be adjusted accordingly. Furthermore, device 110 determines to present a control panel 590 below the video content item 580.
[0033] Figure 6A A view 630a is shown that provides a 3D environment 620a and a video content item 610a based on the expected viewing conditions of the video content item 610a. In this example, the video content item 610a is television content intended for use in dimly lit viewing environments (e.g., environments where television content is typically viewed, such as a family room with room lights providing dim lighting). However, the 3D environment 620a is based on the user's actual physical environment (which is dark). For example, physical environment 100 might allow the lights to be completely off. In some implementations, the brightness level of the 3D environment is determined based on sensors in the physical environment on which the 3D environment is based. For example, an ambient light sensor can be used to determine the brightness level of the environment. In another example, an image of the physical environment is evaluated (e.g., via an algorithm or machine learning model) to estimate the current brightness level. Parameters for providing view 630a are determined based on the fact that the 3D environment 620a is darker than the expected environment of the video content item 610a. In this example, the depiction of the physical environment in view 630a is made brighter to provide the expected dim environment. In other examples, only parts of the environment (e.g., those parts of view 630a that are close to video content item 610a or within a few feet of it) are made brighter.
[0034] Figure 6BA view 630b is shown that provides a 3D environment 620b and a video content item 610b based on the expected viewing conditions of the video content item 610b. In this example, the video content item 610b is television content intended for use in dimly lit viewing environments (e.g., environments where television content is typically viewed, such as a family room with room lights providing dim lighting). However, the 3D environment 620b is based on the user's actual physical environment (which is dark). For example, physical environment 100 might allow the lights to be completely off. Parameters for providing the view 630b are determined based on the fact that the 3D environment 620b is darker than the expected environment of the video content item 610b. In this example, the depiction of the television content in the view 630b is made darker to better match the actual dark viewing environment.
[0035] Figure 6C A view 630c is shown that provides a 3D environment 620c and a video content item 610c based on the expected viewing conditions of the video content item 610c. In this example, the video content item 610c is television content intended for use in dimly lit viewing environments (e.g., environments where television content is typically viewed, such as a family room with room lights providing dim lighting). However, the 3D environment 620c is based on the user's actual physical environment (which is dark). Parameters for providing the view 630c are determined based on the fact that the 3D environment 620c is darker than the expected environment of the video content item 610c. In this example, the depiction of the physical environment in the view 630c is made brighter, and the depiction of the television content in the view 630c is made darker to better match each other.
[0036] In some implementations, the intended viewing environment for the content is determined based on its type. For example, traditional / television content may be considered for use in dimly lit environments, cinematic content for use in dark environments, and computer / user interface content (e.g., XR content) for use in bright environments. The content can be presented within the view of a 3D environment by varying the brightness of the content, the brightness of the surrounding environment, or both, to provide an expected or otherwise consistent user experience, such as a user experience where the content is viewed in its intended viewing environment or matches the brightness of the viewing environment in which the content is experienced.
[0037] Figure 7A view 730 is shown that provides a 3D environment 720 and a video content item 710 based on the expected viewing conditions of the video content item 710 according to some specific implementations. In this example, the video content item 710 is cinema content designed for use in a dark viewing environment (e.g., where the cinema environment is illuminated by minimal lighting other than the cinema screen). However, the 3D environment 720 is based on the user's actual physical environment, which is illuminated by natural light from windows or room lights. For example, physical environment 100 may allow lights to be fully on and sunlight to illuminate the room through windows. In some specific implementations, the brightness level of the 3D environment is determined based on sensors in the physical environment on which the 3D environment is based. For example, an ambient light sensor may be used to determine the brightness level of the environment. In another example, an image of the physical environment is evaluated (e.g., via an algorithm or machine learning model) to estimate the current brightness level. Parameters for providing view 730 are determined based on the fact that the 3D environment 720 is brighter than the expected environment for the video content item 710. In this example, the depiction of the physical environment in view 730 is darkened and the video content item 710 is made slightly brighter. In other examples, only non-video content items are depicted darker (e.g., video content item 710 is not changed). In other examples, only portions of the depiction (e.g., those portions in view 730 that are close to video content item 710 / within a few feet of it) are darkened. In other examples, only video content item 710 is brightened, for example, without changing the depiction of the physical environment.
[0038] Figure 8 This is a flowchart representation of an exemplary method 800 for providing a view of the physical environment and video content items using context-based parameters. In some specific implementations, method 800 is provided by a device (e.g., Figure 1 The method 800 may be executed by a device 110, such as a mobile device, desktop computer, laptop computer, or server device. The method 800 may be executed on a device having a screen for displaying images and / or a screen for viewing stereoscopic images, such as a head-mounted display (HMD). In some embodiments, the method 800 is executed by processing logic components (including hardware, firmware, software, or a combination thereof). In some embodiments, the method 800 is executed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., memory).
[0039] At box 802, method 800 determines the context for rendering video content items within a view of the 3D environment, wherein the context is determined based on the attributes of the video content items. The context may be determined based on analyzing the video content items. In one example, the context is determined based on metadata associated with the video content items (e.g., stored with or otherwise associated with the video content items).
[0040] The context can be determined by processing video content items using an algorithm or machine learning model configured to classify video content items as having, for example, content type, brightness classification, expected viewing environment classification, color scheme / range classification, etc.
[0041] Context can be used to identify or is based on the content item type or the expected viewing environment.
[0042] Context can be based on the viewer and therefore can identify one or more viewer attributes. For example, context can identify the viewer's intent regarding the viewing of video content (e.g., which viewing mode the user has selected through input). As another example, context can identify that the viewer generally prefers to be at a specific distance (e.g., 7 to 10 feet) from the video content or a particular type of video content (e.g., sports, movies, concerts, theatrical films, etc.). As yet another example, context can identify that the viewer has a specific interpupillary distance (IPD) or other physiological characteristic or state. For example, context can identify the user's current activity, such as being engaged in other activities compared to focusing on watching video content.
[0043] The context can identify environmental attributes of the 3D environment, such as the type of objects in the surrounding environment, the size and shape of the room, the acoustic characteristics of the room, the identity and / or number of other people in the room, the furniture in the room where the expected viewer is sitting to watch video content items, the ambient lighting in the room, etc. The context can be based on object detection or scene understanding of the 3D environment or the physical environment on which it is based. For example, the object detection module can analyze RGB images from a light intensity camera and / or sparse depth maps from a depth camera (e.g., a time-of-flight sensor) and other sources of physical environment information (e.g., camera positioning information from a camera's SLAM system, VIO, etc., such as a position sensor) to identify objects (e.g., people, pets, furniture, walls, floors, etc.) in a sequence of light intensity images. In some implementations, the object detection module uses machine learning for object recognition. In some implementations, the machine learning model is a neural network (e.g., an artificial neural network), decision tree, support vector machine, Bayesian network, etc. For example, the object detection module can use an object detection neural network unit to identify objects and / or an object classification neural network to classify each type of object.
[0044] The context can recognize automatically or manually selected viewing modes, such as full-screen / movie, 2D, 3D, custom, etc. The context can also recognize user-specific changes to the viewing mode or the size of the video content. In some implementations, the user's selection of a viewing mode triggers the determination or re-determination of parameters used to display content in the newly selected viewing mode.
[0045] At box 804, method 800 determines rendering rules for presenting video content items within a view of a 3D environment based on context. Such rules specify one or more specific parameters that will be used to render the depiction of video content items or other objects in the 3D environment, given a particular context. In one example, the rendering rules position the virtual 2D screen for the video content to optimize viewing angle, maximize display screen occupancy (e.g., avoid wasting physical display pixels on non-video content), ensure user comfort, or avoid exceeding visible video display resolution limits. In one example, the rendering rules position the virtual 2D screen for the video content based on the viewer's (e.g., viewer IPD) physiological data. In one example, the rendering rules position the 3D video content for optimal binocular stereoscopic viewing.
[0046] Rendering rules can adjust brightness, or brightness can be used to adjust rendering rules. In one example, a rendering rule adjusts the brightness of video content by mapping / matching content brightness to ambient brightness. In another example, a rendering rule adjusts brightness by dimming or darkening the pass-through environment. In one example, a rendering rule adjusts the brightness of either the video content or the ambient content based on the type of video content (e.g., theater / cinema or television). In one example, a rendering rule adjusts the brightness of the video content over time to alter the viewer's light perception (e.g., reducing brightness for a limited time so that a normal image will appear brighter than it actually is at certain points during the playback of the video content to match the brightness requirements of the video content item).
[0047] Rendering rules can adjust the color gamut based on content type; for example, determining colors based on media and selecting the color gamut based on content or system details. Rendering rules can also adjust the color gamut based on the viewer; for example, using HDR if the user is staring / focused at the content, and using a non-HDR format otherwise.
[0048] Rendering rules can adjust the color space based on context. For example, the color space can be chosen based on the video content or the environment. In one example, the color space used to display the content is based on one or more colors in the environment, such as other user interface content (if any) displayed simultaneously with the video content or otherwise present in the visible portion of the environment. In another example, the color space is chosen to correspond to the content creator's intent for the video content, taking into account other colors in the viewing environment. In one example, if the user is only looking at the video content and there is no other user interface content, a content reference color is used to adjust the video content to match the creator's intent. On the other hand, if other user interface elements are displayed, a content reference color may not be used to allow the video to blend better with the overall environment.
[0049] Rendering rules can adjust the frame rate based on video content and viewing mode. For example, this could involve changing the device's display frame rate to match the content's frame rate when in cinema / full-screen mode.
[0050] Similarly, presentation rules can adjust the resolution of an electronic device's display based on standard viewer faculty and the viewer's 3D position.
[0051] Presentation rules can, for example, adjust audio spatiality based on context or viewing mode. In one example, the number and location of audio sources are determined based on context. Such audio parameters can be determined to match a selected screen size or viewing mode. For example, Figure 4 The relatively small screen size allows one or more audio sources to be positioned behind the screen, while Figure 3 The relatively large screen size can be used in conjunction with multi-channel audio sources positioned around the viewer. A full-screen / cinema viewing experience can provide maximum spatialization of sound within the device's capabilities and taking into account the 3D environment. In small rooms, the spatialization of audio can be limited to avoid providing an unrealistic experience, while ensuring at least a minimum quality audio is provided; for example, even if the user is in a small closet-sized room, the audio can have acoustic characteristics consistent with at least a medium-sized room to avoid providing an undesirable listening experience. In some specific implementations, the level of video immersion and the number and location of spatialization sources are determined based on context to provide a realistic but optimal user experience.
[0052] At box 806, method 800 determines parameters (e.g., 3D positioning, audio mode, brightness, color format, etc.) for rendering video content items within a view of the 3D environment based on context and rendering rules, and at box 808, method 800 renders the video content items within a view of the 3D environment based on the parameters.
[0053] Presenting a view can involve using content from sensors located in the physical environment (e.g., image sensors, depth sensors, etc.) to represent the physical environment. For example, an outward-facing camera (e.g., a light intensity camera) captures pass-through video of the physical environment. Therefore, if a user wearing an HMD is sitting in their living room, the representation could be a pass-through video of the living room displayed on the HMD screen.
[0054] In some implementations, method 800 determines the viewpoint for presenting video content within a view of a 3D environment based on context and presentation rules. This may involve intentionally altering the user's perceptual understanding to improve the user's perception of the content or otherwise enhance the user's experience.
[0055] In some implementations, a microphone (one of the I / O devices and sensors of device 110) can capture sound in the physical environment and can incorporate sound into the experience.
[0056] The view may include a virtual 2D screen or a virtual 3D content viewing area positioned within the surrounding 3D environment. Video content items may be presented as appearing on such a 2D screen or within such a 3D content viewing area. In one example, a user may be wearing an HMD and viewing a real-world physical environment (e.g., in a kitchen as a representation of the presented physical environment) via pass-through video (or optical pass-through video), and a virtual screen may be generated for the user to view video content items.
[0057] Figure 9 This is a block diagram of an example of a device 110 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 110 includes one or more processing units 902 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 906, one or more communication interfaces 908 (e.g., USB, Firewire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BlueTooth, ZigBee, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 910, one or more AR / VR displays 912, one or more internal and / or external image sensors 914, memory 920, and one or more communication buses 904 for interconnecting these components and various other components.
[0058] In some embodiments, the one or more communication buses 904 include circuitry for interconnecting system components and controlling communication between system components. In some embodiments, one or more I / O devices and sensors 906 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, an ambient light sensor (ALS), one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, or one or more depth sensors (e.g., structured light, time-of-flight, etc.).
[0059] In some embodiments, one or more displays 912 are configured to present an experience to a user. In some embodiments, one or more displays 912 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 912 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, device 110 includes a single display. As another example, device 110 includes displays for each of the user's eyes.
[0060] In some embodiments, one or more image sensor systems 914 are configured to acquire image data corresponding to at least a portion of the physical environment 105. For example, one or more image sensor systems 914 include one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), monochrome cameras, IR cameras, event-based cameras, etc. In various embodiments, one or more image sensor systems 914 also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 914 also include an on-camera image signal processor (ISP) configured to perform a plurality of processing operations on the image data, including at least a portion of the processes and techniques described herein.
[0061] Memory 920 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 920 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 920 optionally includes one or more storage devices remotely located to one or more processing units 902. Memory 920 includes a non-transitory computer-readable storage medium. In some embodiments, memory 920 or its non-transitory computer-readable storage medium stores programs, modules, and data structures, or subsets thereof, including an optional operating system 930 and one or more instruction sets 940.
[0062] Operating system 930 includes procedures for handling various basic system services and for performing hardware-related tasks. In some specific implementations, instruction set 940 is configured to manage and coordinate one or more experiences for one or more users (e.g., a single experience for one or more users, or multiple experiences for a corresponding group of one or more users).
[0063] Instruction set 940 includes a content rendering instruction set 942 configured with instructions executable by a processor to provide content on a display of an electronic device (e.g., device 110). For example, the content may include an XR environment comprising a depiction of a physical environment including real-world objects and virtual objects (e.g., a virtual screen overlaid on an image of a real-world physical environment). Content rendering instruction set 942 is further configured with instructions executable by a processor to obtain image data (e.g., light intensity data, depth data, etc.) using one or more techniques disclosed herein, generate virtual data (e.g., a virtual movie screen), and integrate (e.g., fuse) the image data and virtual data (e.g., mixed reality (MR)). Content rendering instruction set 942 may determine a context (e.g., one or more contextual factors / attributes) and apply one or more rendering rules based on the context to determine parameters for rendering video content items within a depiction of a 3D environment, as described herein.
[0064] Although these components are shown residing on a single device (e.g., device 110), it should be understood that in other specific implementations, any combination of components may reside in a separate computing device. Furthermore, Figure 9 This is used more as a functional description of various features present in a specific implementation, and differs from the structural diagrams of the specific implementations described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 9 Some functional modules shown individually (e.g., instruction set 940) may be implemented in a single module, and the various functions of a single functional block (e.g., in the instruction set) may be implemented in various specific implementations by one or more functional blocks. The actual number of modules and the division of specific functions and how features are allocated therein will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software and / or firmware selected for a particular implementation.
[0065] This document provides numerous specific details to enable those skilled in the art to thoroughly understand the claimed subject matter. However, the claimed subject matter can be practiced without these details. In other instances, methods, apparatus, or systems known to those of ordinary skill in the art have not been described in detail so as not to obscure the claimed subject matter.
[0066] Specific implementations of the methods disclosed herein can be performed in the operation of such a computing device. The order of the boxes presented in the above examples can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel.
[0067] The embodiments of the subject matter and operations described in this specification may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents) or in a combination thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. Alternatively or otherwise, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or may be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices), or included in one or more separate physical components or media.
[0068] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0069] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, systems-on-a-chip, or multiple or combinations thereof. The apparatus may include special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)). In addition to hardware, the apparatus may include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying" refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, that manipulate or convert data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0070] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific embodiments of the subject matter of this invention. The teachings contained herein can be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0071] Specific implementations of the methods disclosed herein can be performed in the operation of such a computing device. The order of the boxes presented in the above examples can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel.
[0072] The use of “applies to” or “configured to” in this document implies open and inclusive language, which does not exclude applicability to or configuration to devices performing additional tasks or steps. Similarly, the use of “based on” implies openness and inclusivity, as processes, steps, calculations, or other actions “based on” one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included in this document are for illustrative purposes only and are not intended to be restrictive.
[0073] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node can be called a second node, and similarly, a second node can be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0074] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and the appended claims, the singular forms “a” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising” as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0075] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.
[0076] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of different embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features of a claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.
[0077] Similarly, although operations are shown in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in a sequential order or the specific order shown, or requiring all shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the division of various system components in the above embodiments should not be construed as requiring such division in all embodiments, and it should be understood that the program components and the system may generally be integrated together in a single software product or packaged into multiple software products.
[0078] Therefore, specific embodiments of the subject matter have been described. Other embodiments are also within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes shown in the accompanying drawings do not necessarily require a specific order or sequence to achieve the desired result. In some specific embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method comprising: In electronic devices with processors: Determine a context for presenting a video content item within a view of a 3D environment, wherein the context is determined based on the attributes of the video content item; The rendering rules for presenting video content items within a view of the 3D environment are determined based on the context. The parameters for rendering the video content item within the view of the 3D environment are determined based on the context and the rendering rules. as well as A view of the 3D environment is presented, the view including the video content items presented based on the parameters.
2. The method of claim 1, wherein the context is determined based on analysis of the video content items.
3. The method of claim 1, wherein the context is determined based on metadata associated with the video content item.
4. The method of claim 1, wherein the context includes the content item type or the intended viewing environment of the content item.
5. The method of claim 1, wherein the context includes viewer attributes.
6. The method of claim 1, wherein the context includes environmental properties of the 3D environment.
7. The method of claim 1, wherein the context includes a viewing mode.
8. The method of claim 1, wherein the context includes user-specific changes to the viewing mode or the size of the video content.
9. The method of claim 1, wherein the presentation rules are positioned for a virtual 2D screen for video content to prioritize viewing angle, screen occupancy, or user comfort, or to avoid limitations on the visible video display resolution.
10. The method of claim 1, wherein the presentation rules are based on the viewer’s physiological data to locate a virtual 2D screen for the video content.
11. The method of claim 1, wherein the presentation rules are based on binocular stereoscopic viewing to locate 3D movie content.
12. The method of claim 1, wherein the presentation rule is: The brightness of the video content is adjusted by mapping the content brightness to the ambient brightness. Brightness can be adjusted by dimming or darkening the transparent environment; Adjust the brightness of the video content or ambient content based on the type of video content; or Adjusting the brightness of video content over time to alter the viewer's perception of light.
13. The method of claim 1, wherein the presentation rule is: Adjust the color range based on the content type; or Adjust the color gamut based on the viewer.
14. The method of claim 1, wherein the presentation rules adjust the frame rate based on video content and viewing mode.
15. The method of claim 1, wherein the presentation rules adjust the resolution of the display of the electronic device based on standard viewer capabilities and viewer 3D position.
16. The method of claim 1, wherein the presentation rules adjust the audio spatiality based on the viewing mode.
17. The method of claim 1, wherein the electronic device is a head-mounted device.
18. The method of claim 1, wherein the 3D environment is an extended reality (XR) environment.
19. The method of claim 1, wherein the 3D environment comprises a virtual screen positioned at a location depicted relative to a portion of the physical environment.
20. A system comprising: Non-transitory computer-readable storage medium; as well as One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the system to perform operations including: Determine a context for presenting a video content item within a view of a 3D environment, wherein the context is determined based on the attributes of the video content item; The rendering rules for presenting video content items within a view of the 3D environment are determined based on the context. The parameters for rendering the video content item within the view of the 3D environment are determined based on the context and the rendering rules. as well as A view of the 3D environment is presented, the view including the video content items presented based on the parameters.
21. The system of claim 20, wherein the context is determined based on analysis of the video content items.
22. The system of claim 20, wherein the context is determined based on metadata associated with the video content item.
23. The system of claim 20, wherein the context includes the content item type or the intended viewing environment of the content item.
24. A non-transitory computer-readable storage medium storing program instructions executable by one or more processors to perform operations, the operations including: Determine a context for presenting a video content item within a view of a 3D environment, wherein the context is determined based on the attributes of the video content item; The rendering rules for presenting video content items within a view of the 3D environment are determined based on the context. The parameters for rendering the video content item within the view of the 3D environment are determined based on the context and the rendering rules. as well as A view of the 3D environment is presented, the view including the video content items presented based on the parameters.
Citation Information
Patent Citations
Rendering virtual objects with realistic surface properties that match the environment
CN110866966A
Dynamic contextual media filter
CN113228693A