Content matching for a spatial 3D environment

The system addresses the limitations of 2D display in 3D environments by matching content elements to surrounding surfaces, ensuring seamless and adaptive content display in VR, AR, and MR systems.

JP7704801B2Active Publication Date: 2025-07-08MAGIC LEAP INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023076383
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-03-16
Filing Date
2023-05-04
Publication Date
2025-07-08
Estimated Expiration
2038-05-01

AI Technical Summary

Technical Problem

Conventional methods for displaying content in spatial 3D environments, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR), are suboptimal as they are limited to 2D screen displays and fail to effectively utilize the spatial dimensions of these environments.

Method used

A system and method for identifying content elements and matching them to surrounding surfaces within a spatially organized 3D environment, using sensors and computer vision to determine virtual surfaces and attributes, and displaying content on both physical and virtual surfaces based on matching scores and user preferences.

Benefits of technology

Enables seamless and dynamic display of content in 3D environments, adapting to user movements and environmental changes, while maintaining continuity and user interaction with content elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704801000006
    Figure 0007704801000006
  • Figure 0007704801000007
    Figure 0007704801000007
  • Figure 0007704801000008
    Figure 0007704801000008
Patent Text Reader

Abstract

To provide matching of content relative to spatial 3D environment.SOLUTION: The present disclosure relates to a system and a method for displaying content in spatial 3D environment. The system and the method are for matching a content element to a surface in spatially organized 3D environment. The method includes the steps of: receiving content; identifying one or more elements in the content; determining one or more surfaces; matching one or more elements to one or more surfaces; and displaying one or more elements on one or more surfaces as virtual content.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to systems and methods for displaying content within a spatial 3D environment.

Background Art

[0002] A typical way to view content is to open an application that will display the content on the display screen of a display device (e.g., a monitor of a computer, smartphone, tablet, etc.). The user will navigate the application and view the content. Usually, there is a fixed formatting regarding how the content is displayed within the application and on the display screen of the display device when the user is looking at the display screen of the display.

[0003] When using virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) systems (hereinafter collectively referred to as "mixed reality" systems), the application will display the content within a spatial three-dimensional (3D) environment. Conventional approaches for displaying content on a display screen do not work well when used in a spatial 3D environment. One reason is that when using conventional approaches, the display area of the display device is a 2D medium limited to the screen area of the display screen on which the content is displayed. As a result, conventional approaches are configured to only understand how to compose and display the content within the screen area of the display screen. In contrast, a spatial 3D environment is not strictly limited to the exact confines of the screen area of the display screen. Therefore, conventional approaches may perform sub-optimally when used in a spatial 3D environment because conventional approaches do not necessarily have the functionality or ability to utilize the spatial 3D environment for displaying content.

[0004] Therefore, there is a need for an improved approach for displaying content within a spatial 3D environment.

[0005] The subject matter discussed in the background section should not be assumed to be prior art merely as a result of its description in the background section. Similarly, the understanding of the problems and causes of the problems described in the background section, or associated with the subject matter of the background section, should not be assumed to have been previously recognized in the prior art. The subject matter in the background section may simply represent different approaches that may themselves also be disclosures.

Summary of the Invention

Means for Solving the Problems

[0006] Embodiments of the present disclosure provide an improved system and method for displaying information within a spatially organized 3D environment. The method includes receiving content, identifying elements within the content, determining surrounding surfaces, matching the identified elements to the surrounding surfaces, and displaying the elements as virtual content on the surrounding surfaces. Additional embodiments of the present disclosure provide an improved system and method for push delivering content to a user of a virtual reality or augmented reality system.

[0007] In one embodiment, the method includes receiving content. The method also includes identifying one or more elements within the content. The method further includes determining one or more surfaces. Additionally, the method includes matching the one or more elements to the one or more surfaces. In addition, the method includes displaying the one or more elements as virtual content on the one or more surfaces.

[0008] In one or more embodiments, the content includes at least one of pull-distributed content or push-distributed content. The step of identifying one or more elements may include, for each of the one or more elements, the step of determining one or more attributes. The one or more attributes include at least one of a priority attribute, an orientation attribute, an aspect ratio attribute, a dimension attribute, an area attribute, a relative visual recognition position attribute, a color attribute, a contrast attribute, a position type attribute, a margin attribute, a content type attribute, a focus attribute, a readability index attribute, or a surface type attribute for installation. The step of determining one or more attributes for each of the one or more elements is based on explicit indications within the content.

[0009] In one or more embodiments, for each of the one or more elements, the step of determining one or more attributes is based on the location of the one or more elements within the content. The method further includes the step of storing the one or more elements in one or more logical structures. The one or more logical structures include at least one of an ordered array, a hierarchical table, a tree structure, or a logical graph structure. The one or more surfaces include at least one of a physical surface or a virtual surface. The step of determining one or more surfaces includes the step of analyzing the environment and determining at least one of the one or more surfaces.

[0010] In one or more embodiments, the step of determining one or more surfaces includes receiving raw sensor data, simplifying the raw sensor data to produce simplified data, and creating one or more virtual surfaces based on the simplified data. The one or more surfaces include the one or more virtual surfaces. The step of simplifying the raw sensor data includes filtering the raw sensor data to produce filtered data, and grouping the filtered data into one or more groups per point cloud point. The simplified data includes the one or more groups. The step of creating one or more virtual surfaces includes processing each of the one or more groups in turn to determine one or more real-world surfaces, and creating one or more virtual surfaces based on the one or more real-world surfaces.

[0011] In one or more embodiments, the step of determining one or more surfaces includes, for each of the one or more surfaces, determining one or more attributes. The one or more attributes include at least one of a priority attribute, an orientation attribute, an aspect ratio attribute, a dimension attribute, an area attribute, a relative viewing position attribute, a color attribute, a contrast attribute, a position type attribute, a margin attribute, a content type attribute, a focus attribute, a readability index attribute, or a surface type attribute for installation. The method also includes storing the one or more surfaces in one or more logical structures. The step of matching one or more elements to one or more surfaces includes prioritizing the one or more elements, comparing, for each element of the one or more elements, one or more attributes of the element with one or more attributes of each of the one or more surfaces, calculating a matching score based on the one or more attributes of the element and the one or more attributes of each of the one or more surfaces, and identifying the best matching surface having the highest matching score. Additionally, for each of the one or more elements, it includes storing an association between the element and the best matching surface.

[0012] In one or more embodiments, one element is matched to one or more surfaces. Further, it includes the step of displaying each of the one or more surfaces to the user. Additionally, it includes the step of receiving a user selection indicating a winning surface from the one or more displayed surfaces. Further, it includes the step of storing surface attributes of the winning surface into a user preference data structure from the user selection. The content is streamed from a content provider. The one or more elements are displayed to the user through a mixed reality device.

[0013] In one or more embodiments, the method further includes the step of displaying one or more additional surface options for displaying one or more elements, at least in part, based on a changed view of the user. The display of the one or more additional surface options is based, at least in part, on a time threshold corresponding to the changed view. The display of the one or more additional surface options is based, at least in part, on a head pose change threshold.

[0014] In one or more embodiments, the method also includes the step of overriding the display of one or more elements on the one or more surfaces to which they were matched. The step of overriding the display of one or more elements on one or more surfaces is based, at least in part, on surfaces frequently used historically. The method further includes the step of moving one or more elements displayed on one or more surfaces to different surfaces, at least in part, based on selecting specific elements to be displayed on one or more surfaces to which the user is about to move to different surfaces. The specific elements to be moved to different surfaces are at least visible to the user.

[0015] In one or more embodiments, the method further includes, in response to a change in the user's view from a first view to a second view, slowly moving the display of one or more elements onto a new surface to follow the user's view change to the second view. The one or more elements may move to the direct front of the user's second view only in response to confirmation received from the user that the content has been moved to the direct front of the user's second view. The method includes, at least in part, based on the user moving from a first location to a second location, pausing the display of one or more elements on one or more surfaces at the first location and resuming the display of one or more elements on one or more other surfaces at the second location. The step of pausing the display of one or more elements is performed automatically, at least in part, based on the determination that the user is moving or has moved from the first location to the second location. The resumption of the display of one or more elements is performed automatically, at least in part, based on the identification of one or more other surfaces and the matching with one or more elements at the second location.

[0016] In one or more embodiments, the step of determining one or more surfaces includes identifying one or more virtual objects for displaying one or more elements. The step of identifying one or more virtual objects is based, at least in part, on data received from one or more sensors indicating the lack of a suitable surface. The elements of the one or more elements are TV channels. The user interacts with the elements of the one or more elements displayed by purchasing one or more items or services displayed to the user.

[0017] In one or more embodiments, the method also includes detecting a change in the environment from a first location to a second location; determining one or more additional surfaces at the second location; matching one or more elements currently displayed at the first location to the one or more additional surfaces; and displaying the one or more elements as virtual content on the one or more additional surfaces at the second location. The determination of the one or more additional surfaces is initiated after the change in the environment exceeds a time threshold. The user pauses the active content displayed at the first location, resumes the active content to be displayed at the second location, and the active content resumes at the same interaction point where the user paused the active content at the first location.

[0018] In one or more embodiments, the method also includes transitioning spatialized audio being delivered to the user from a location associated with the content displayed at the first location to an audio virtual speaker directed at the center of the user's head as the user leaves the first location; and transitioning from the audio virtual speaker directed at the center of the user's head to spatialized audio being delivered to the user from one or more additional surfaces at the second location where the one or more elements are displayed.

[0019] In another embodiment, a method for push delivering content to a user of a mixed reality system includes receiving one or more available surfaces from the user's environment. The method also includes identifying one or more content that match the dimensions of one available surface from the one or more available surfaces. The method further includes calculating a score based on comparing one or more constraints of the one or more content with one or more surface constraints of one available surface. Further, the method includes selecting the content from the one or more content having the highest score. Further, the method includes storing a one-to-one matching of the selected content and one available surface. Further, the method includes displaying the selected content to the user on the available surface.

[0020] In one or more embodiments, the user's environment is the user's personal home. One or more available surfaces from the user's environment are in the vicinity of the user's focused view area. One or more contents are advertisements. The advertisements are targeted at a specific group of users located in a specific environment. One or more contents are notifications from an application. The application is a social media application. One of the one or more constraints of the one or more contents is orientation. The selected content is 3D content.

[0021] In another embodiment, the augmented reality (AR) display system includes a head-mounted system including one or more sensors and one or more cameras with outward-facing cameras. The system also includes a processor for executing a set of program code instructions. Further, the system includes a memory for holding a set of program code instructions, the set of program code instructions including program code for performing the step of receiving content. The program code also performs the step of identifying one or more elements within the content. Further, the program code also performs the step of determining one or more surfaces. Additionally, the program code also performs the step of matching one or more elements to one or more surfaces. Further, the program code also performs the step of displaying one or more elements as virtual content on one or more surfaces.

[0022] In one or more embodiments, the content includes at least one of pull-distributed content or push-distributed content. The step of identifying one or more elements includes the step of analyzing the content and identifying one or more elements. The step of identifying one or more elements includes, for each of the one or more elements, the step of determining one or more attributes. Additionally, the program code also performs the step of storing one or more elements into one or more logical structures. The one or more surfaces include at least one of a physical surface or a virtual surface. The step of determining one or more surfaces includes the step of analyzing the environment and determining at least one of the one or more surfaces.

[0023] In one or more embodiments, the step of determining one or more surfaces includes the step of receiving raw sensor data, the step of simplifying the raw sensor data to produce simplified data, and the step of creating one or more virtual surfaces based on the simplified data, and the one or more surfaces include the one or more virtual surfaces. The step of determining one or more surfaces includes, for each of the one or more surfaces, the step of determining one or more attributes. Additionally, the program code also performs the step of storing one or more surfaces into one or more logical structures.

[0024] In one or more embodiments, the step of matching one or more elements to one or more surfaces includes the step of prioritizing the one or more elements, the step of comparing, for each element of the one or more elements, one or more attributes of the element with one or more attributes of each of the one or more surfaces, the step of calculating a matching score based on one or more attributes of the element and one or more attributes of each of the one or more surfaces, and the step of identifying the best matching surface having the highest matching score. One element is matched to one or more surfaces. The content is streamed from a content provider.

[0025] In one or more embodiments, the program code also performs, at least in part, the step of displaying one or more surface options for displaying one or more elements based on a user's changing view. The program code also performs the step of overriding the display of one or more elements on one or more surfaces on which the display had been matched. The program code also performs, at least in part, the step of moving one or more elements displayed on one or more surfaces to different surfaces based on selecting specific elements to be displayed on one or more surfaces that the user will move to different surfaces. The program code also performs the step of slowly moving the display of one or more elements onto a new surface and following the user's view change to the second view in response to the user's view change from the first view to the second view.

[0026] In one or more embodiments, the program code also performs, at least in part, the step of pausing the display of one or more elements on one or more surfaces at a first location and resuming the display of one or more elements on one or more other surfaces at a second location based on the user moving from the first location to the second location. The step of determining one or more surfaces includes the step of identifying one or more virtual objects for displaying one or more elements. The elements of the one or more elements are TV channels.

[0027] In one or more embodiments, the user interacts with the elements of one or more elements displayed by purchasing one or more items or services displayed to the user. The program code also performs the step of detecting a change in the environment from a first location to a second location, the step of determining one or more additional surfaces at the second location, the step of matching one or more elements currently displayed at the first location to the one or more additional surfaces, and the step of displaying the one or more elements as virtual content on the one or more additional surfaces at the second location.

[0028] In another embodiment, the augmented reality (AR) display system includes a head-mounted system that includes one or more sensors and one or more cameras with outward-facing cameras. The system also includes a processor for executing a set of program code instructions. The system further includes a memory for holding a set of program code instructions, the set of program code instructions comprising program code for performing the step of receiving one or more available surfaces from the user's environment. The program code also performs the step of identifying one or more contents that match the dimensions of one available surface from the one or more available surfaces. The program code further performs the step of calculating a score based on comparing one or more constraints of the one or more contents with one or more surface constraints of one available surface. The program code also performs the step of selecting the content from the one or more contents having the highest score. Additionally, the program code performs the step of storing a one-to-one mapping of the selected content and one available surface. The program code also performs the step of displaying the selected content to the user on the available surface.

[0029] In one or more embodiments, the user's environment is the user's personal home. One or more available surfaces from the user's environment are around the user's focus view area. One or more contents are advertisements. One or more contents are notifications from an application. One of the one or more constraints of the one or more contents is orientation. The selected content is 3D content.

[0030] In another embodiment, a computer-implemented method for decomposing 2D content includes the step of identifying one or more elements within the content. The method also includes the step of identifying one or more surrounding surfaces. The method further includes the step of mapping the one or more elements to the one or more surrounding surfaces. Additionally, the method includes the step of displaying the one or more elements as virtual content on the one or more surfaces.

[0031] In one or more embodiments, the content is a web page. The element of one or more elements is a video. One or more surrounding surfaces include physical surfaces within a physical environment or virtual objects that are not physically located within the physical environment. The virtual object is a multi-stack virtual object. The first set of results of the one or more identified elements and the second set of results of the one or more identified surrounding surfaces are stored in a database table within a storage device. The storage device is a local storage device. The database table stores the results of the one or more identified surrounding surfaces, including surface ID, width dimension, height dimension, orientation description, and position relative to a reference frame.

[0032] In one or more embodiments, the step of identifying one or more elements within the content includes identifying attributes from tags corresponding to the location of the elements, extracting hints from tags related to the one or more elements, and storing the one or more elements. The step of identifying one or more surrounding surfaces includes identifying the user's current surrounding surface, determining the user's posture, identifying the dimensions of the surrounding surface, and storing the one or more surrounding surfaces. The step of mapping one or more elements to one or more surrounding surfaces includes looking up pre-defined rules for identifying candidate surrounding surfaces for the mapping, and for each of the one or more elements, selecting the best-fit surface. The display of one or more elements on one or more surrounding surfaces is performed by an augmented reality device.

[0033] In another embodiment, a method of matching content elements of content to a spatial three-surrounding surface (3D) environment includes a content structuring process, an environment structuring process, and a synthesizing process.

[0034] In one or more embodiments, a content structuring process reads content and organizes and / or stores the content into a logical / hierarchical structure for accessibility. The content structuring process includes a parser for receiving content. The parser analyzes the received content and identifies content elements from the received content. The parser identifies / determines attributes and stores them for each content element into the logical / hierarchical structure.

[0035] In one or more embodiments, an environment structuring process analyzes environmental - related data and identifies surfaces. The environment structuring process includes one or more sensors, a computer vision processing unit (CVPU), a perception framework, and an environment analyzer. The one or more sensors provide raw data regarding real - world surfaces (e.g., a point cloud of objects and structures from the environment) to the CVPU. The CVPU simplifies and / or filters the raw data. The CVPU modifies the remaining data into grouped point cloud points by distance and planarity to extract / identify / determine surfaces by downstream processes. The perception framework receives the grouped point cloud points from the CVPU and prepares environmental data for the environment analyzer. The perception framework creates / determines structures / surfaces / planes and captures them into one or more data storage devices. The environment analyzer analyzes the environmental data from the perception framework and determines the surfaces within the environment. The environment analyzer uses object recognition to identify objects based on the environmental data received from the perception framework.

[0036] In one or more embodiments, a synthesis process matches content elements from a parser (e.g., a table of content elements stored within a logical structure) with surfaces from an environment analyzer (e.g., a table of surfaces stored within a logical structure) from the environment to determine content elements to be rendered / mapped / displayed on the surfaces of the environment. The synthesis process includes a matching module, a rendering module, a virtual object creation module, a display module, and a receiving module.

[0037] In one or more embodiments, the matching module pairs / matches content elements stored within a logical structure with a surface stored within the logical structure. The matching module compares the attributes of the content elements with the attributes of the surface. The matching module matches the content elements to the surface based on content elements and surfaces that share similar and / or opposing attributes. The matching module may access one or more preference data structures such as user preferences, system preferences, and / or passable preferences, and may use the one or more preference data structures in the matching process. The matching module matches one content element to one or more surfaces based at least in part on at least one of a content vector (e.g., an orientation attribute), a head pose vector (e.g., an attribute of a VR / AR device other than the surface), or a surface normal vector of one or more surfaces. The result may be stored in a cache memory or a persistent storage device for further processing. The result may be compiled and stored in a table to inventory the matching.

[0038] In one or more embodiments, an optional virtual object creation module creates a virtual object to display a content element based on a determination that it is an optimal option to create a virtual object to display the content element, and the virtual object is a virtual planar surface. The step of creating a virtual object to display a content element may be based on data received from a particular sensor or sensors or the lack of sensor input from a particular sensor or sensors. Data received from an environment-centered sensor of a plurality of sensors (such as a camera or depth sensor) indicates the lack of a suitable surface based on the user's current physical environment, or such sensors may not be able to determine the presence of a surface at all.

[0039] In one or more embodiments, a rendering module renders content elements onto individual matched surfaces, where the matched surfaces include real surfaces and / or virtual surfaces. The rendering module renders the content elements at an accurate scale and adapts them to the matched surfaces. Content elements matched to a surface (real and / or virtual) in a first room remain matched to the surface in the first room even when the user moves from the first room to a second room. Content elements matched to a surface in the first room are not mapped to a surface in the second room. As the user returns to the first room, the content elements rendered on the surface in the first room resume display, and other features such as audio playback and / or playback time will resume seamlessly as if the user had not exited the room.

[0040] In one or more embodiments, content elements matched to a surface in a first room are matched to a surface in a second room when the user exits the first room and enters the second room. A first set of content elements matched to a surface in the first room remains matched to the surface in the first room, while a second set of content elements matched to a surface in the first room may move with the device implementing the AR system to the second room. The second set of content elements moves with the device as the device progresses from the first room to the second room. The step of determining whether a content element is in the first set or the second set is based at least in part on at least one of the attributes of the content element, the attributes of one or more surfaces in the first room to which the content element is matched, user preferences, system preferences, and / or passable world preferences. A content element may be matched to a surface but not rendered on that surface when the user is not within the proximity of the surface or when the user is not within the field of view of the surface.

[0041] In some embodiments, the content element is displayed on all of the top three surfaces at once. The user may select a surface as the preferred surface from the top three surfaces. The content element is then displayed on only one of the top three surfaces at a time, along with an indication to the user that the content element can be displayed on two other surfaces. The user may then navigate through the other surface options, and as each surface option is activated by the user, the content element may be displayed on the activated surface. The user may then select a surface as the preferred surface from the surface options.

[0042] In some embodiments, the user extracts a channel for a television by using a totem to target a channel, pressing a trigger on the totem, selecting the channel, holding the trigger for a time period (e.g., about 1 second), moving the totem, identifying a desired location within an environment for displaying the extracted TV channel, pressing the trigger on the totem, and placing the extracted TV channel at the desired location within the environment. The desired location is a surface suitable for displaying the TV channel. A prism is created at the desired location, along with the selected channel content loaded and displayed within the prism. While moving the totem and identifying a desired location within the environment for displaying the extracted TV channel, a video is displayed to the user. The video may be at least one of a single image illustrating the channel, one or more images illustrating a preview of the channel, or a video stream illustrating the current content of the channel.

[0043] In some embodiments, the method includes identifying a first field of view of a user, generating one or more surface options for displaying content, changing from the user's first field of view to a second field of view, generating one or more additional surface options for displaying content corresponding to the second field of view, presenting one or more surface options corresponding to the first field of view and one or more additional surface options corresponding to the second field of view, receiving a selection from the user to display content on a surface corresponding to the first field of view while the user is looking within the second field of view, and displaying an indication in the direction of the first field of view and navigating the user in a direction indicated to return to the first field of view to view the surface option selected by the user.

[0044] In some embodiments, the method includes displaying content on a first surface within a first field of view of a user, the first field of view corresponding to a first head pose. The method also includes determining a change from the first field of view to a second field of view over a time period exceeding a time threshold, and displaying options for the user to change the display location of the content from the first surface within the first field of view to one or more surface options within the second field of view. The second field of view corresponds to a second head pose. In some embodiments, the system immediately displays options for the user to change the display location of the content once the user's field of view changes from the first field of view to the second field of view. The first head pose and the second head pose having a change in position exceed a head pose change threshold.

[0045] In some embodiments, the method includes rendering and displaying content on one or more first surfaces, and a user viewing the content has a first head pose. The method also includes rendering the content on one or more second surfaces in response to the user changing from the first head pose to a second head pose, and a user viewing the content has the second head pose. The method further includes providing the user with an option to change the display location of the content from the one or more first surfaces to the one or more second surfaces. When the head pose change exceeds a corresponding head pose change threshold, the user is provided with an option to change the display location. The head pose change threshold is greater than 90 degrees. The head pose change is maintained over a time period that exceeds the threshold. When the head pose change is less than the head pose change threshold, no option to change the display location of the content is provided.

[0046] In some embodiments, the method includes evaluating a list of surfaces viewable by a user as the user moves from a first location to a second location, the list of surfaces being suitable for displaying a certain type of content that can be pushed to the user's environment without the user having to search for or select the content. The method also includes determining user preference attributes indicating the times and locations at which a certain type of push - delivered content can be displayed. The method further includes displaying the push - delivered content on one or more surfaces based on the preference attributes.

[0047] In another embodiment, a method for push delivering content to a user of an augmented reality system includes determining one or more surfaces corresponding to attributes of the one or more surfaces. The method also includes receiving, at least in part, one or more content elements that match one or more surfaces based on one surface attribute. The method further includes calculating a matching score, at least in part, based on the extent to which the attributes of the content element match the attributes of the one or more surfaces, the matching score being based on a scale of 1 to 100, with a score of 100 being the highest score and a score of 1 being the lowest score. The method also includes selecting the content element from one or more content elements having the highest matching score. Additionally, the method includes storing the matching of the content element and the surface. In addition, the method includes rendering the content element on the matched surface.

[0048] In one or more embodiments, the content element includes a notification, and the matching score is calculated based on an attribute that may indicate the priority of the content element that needs to be notified as opposed to matching with a particular surface. The step of selecting the preferred content element is based on the extent to which the attributes of the content element match the attributes of the surface. The content element is selected based on the content type, and the content type is a notification from 3D content and / or social media contacts.

[0049] In some embodiments, a method for generating a 3D preview for a web link includes representing the 3D preview for the web link as a set of new HTML tags and properties associated with the web page. The method also includes defining the 3D model as an object and / or a surface and rendering the 3D preview. The method further includes generating the 3D preview and loading the 3D preview onto the 3D model. The 3D model is a 3D volume etched into the 2D web page.

[0050] Each of the individual embodiments described and illustrated in this specification has discrete components and features that can be easily separated from or combined with any of the components and features of some other embodiments. The present invention provides, for example, the following. (Item 1) A method comprising: Receiving content; Identifying one or more elements within the content; Determining one or more surfaces; Matching the one or more elements to the one or more surfaces; And displaying the one or more elements as virtual content on the one or more surfaces. A method. (Item 2) The method according to item 1, wherein the content includes at least one of pull-distributed content or push-distributed content. (Item 3) The method according to item 1, wherein identifying the one or more elements includes analyzing the content to identify the one or more elements. (Item 4) The method according to item 1, wherein identifying the one or more elements includes determining one or more attributes for each of the one or more elements. (Item 5) The method according to item 4, wherein the one or more attributes include at least one of a priority attribute, an orientation attribute, an aspect ratio attribute, a dimension attribute, an area attribute, a relative visual recognition position attribute, a color attribute, a contrast attribute, a position type attribute, a margin attribute, a content type attribute, a focus attribute, a readability index attribute, or a surface type attribute for installation. (Item 6) The method according to item 4, wherein determining the one or more attributes for each of the one or more elements is based on an explicit indication within the content. (Item 7) For each of the one or more elements, determining the one or more attributes is the method according to item 4 based on the location of the one or more elements in the content. (Item 8) The method according to item 1, further comprising storing the one or more elements in one or more logical structures. (Item 9) The method according to item 8, wherein the one or more logical structures include at least one of an ordered array, a hierarchical table, a tree structure, or a logical graph structure. (Item 10) The method according to item 1, wherein the one or more surfaces include at least one of a physical surface or a virtual surface. (Item 11) The method according to item 1, wherein determining the one or more surfaces includes analyzing the environment and determining at least one of the one or more surfaces. (Item 12) Determining the one or more surfaces includes receiving raw sensor data, simplifying the raw sensor data to produce simplified data, and creating one or more virtual surfaces based on the simplified data and the one or more surfaces include the one or more virtual surfaces, which is the method according to item 1. (Item 13) Simplifying the raw sensor data includes filtering the raw sensor data to produce filtered data, and grouping the filtered data into one or more groups by point cloud points and the simplified data includes the one or more groups, which is the method according to item 12. (Item 14) Creating the one or more virtual surfaces includes sequentially processing each of the one or more groups to determine one or more real-world surfaces, Creating the one or more virtual surfaces based on the one or more real-world surfaces The method according to item 13, comprising: (Item 15) Determining the one or more surfaces includes determining, for each of the one or more surfaces, one or more attributes, according to the method of item 1 (Item 16) The one or more attributes include at least one of a priority attribute, an orientation attribute, an aspect ratio attribute, a dimension attribute, an area attribute, a relative visibility position attribute, a color attribute, a contrast attribute, a position type attribute, a margin attribute, a content type attribute, a focus attribute, a readability index attribute, or a surface type attribute for installation, according to the method of item 15 (Item 17) The method according to item 1, further comprising storing the one or more surfaces in one or more logical structures (Item 18) Matching the one or more elements to the one or more surfaces Prioritizing the one or more elements For each of the one or more elements Comparing one or more attributes of the element with one or more attributes of each of the one or more surfaces Calculating a matching score based on one or more attributes of the element and one or more attributes of each of the one or more surfaces Identifying the best matching surface having the highest matching score The method according to item 1, comprising: (Item 19) For each of the one or more elements The method according to item 18, further comprising storing an association between the element and the best matching surface (Item 20) One element is matched to one or more surfaces, according to the method of item 1 (Item 21) Displaying each of the one or more surfaces to a user Receiving a user selection indicating a winning surface from the one or more surfaces being presented; Storing, within a user preference data structure, surface attributes of the winning surface from the user selection; The method of claim 20, further comprising: (Item 22) The method of claim 1, wherein the content is data streamed from a content provider. (Item 23) The method of claim 1, wherein the one or more elements are presented to the user through a mixed reality device. (Item 24) The method of claim 1, further comprising presenting one or more additional surface options for presenting the one or more elements, at least in part based on a changed field of view of the user. (Item 25) The method of claim 24, wherein the presentation of the one or more additional surface options is based at least in part on a time threshold corresponding to the changed field of view. (Item 26) The method of claim 24, wherein the presentation of the one or more additional surface options is based at least in part on a head pose change threshold. (Item 27) The method of claim 1, further comprising overriding the presentation of the one or more elements on the one or more surfaces that were being matched. (Item 28) The method of claim 27, wherein overriding the presentation of the one or more elements on the one or more surfaces is based at least in part on surfaces that have been frequently used historically. (Item 29) The method of claim 1, further comprising moving the one or more elements presented on the one or more surfaces to different surfaces, at least in part based on selecting specific elements presented on the one or more surfaces that the user is about to move to different surfaces. (Item 30) The method of claim 29, wherein the specific elements to be moved to the different surfaces are at least visible to the user. (Item 31) The method according to item 1, further comprising slowly moving the display of the one or more elements onto a new surface in response to a change in the user's field of view from a first field of view to a second field of view, and following the change in the user's field of view to the second field of view. (Item 32) The method according to item 31, wherein the one or more elements can move to the direct front of the user's second field of view only in response to confirmation received from the user to move the content to the direct front of the user's second field of view. (Item 33) The method according to item 1, further comprising at least partially suspending the display of the one or more elements on the one or more surfaces at a first location based on the user moving from the first location to a second location, and resuming the display of the one or more elements on one or more other surfaces at the second location. (Item 34) The method according to item 33, wherein suspending the display of the one or more elements is performed automatically at least partially based on a determination that the user is moving or has moved from the first location to the second location. (Item 35) The method according to item 33, wherein resuming the display of the one or more elements is performed automatically at least partially based on the identification of the one or more other surfaces and the matching of the one or more elements at the second location. (Item 36) The method according to item 1, wherein determining the one or more surfaces includes identifying one or more virtual objects for displaying the one or more elements. (Item 37) The method according to item 36, wherein identifying the one or more virtual objects is at least partially based on data received from one or more sensors indicating the lack of a suitable surface. (Item 38) The element of the one or more elements is a TV channel, the method according to item 1. (Item 39) The method according to item 1, wherein the user interacts with the elements of the one or more elements displayed by purchasing one or more items or services displayed to the user. (Item 40) Detecting a change in the environment from a first location to a second location; Determining one or more additional surfaces at the second location; Matching the one or more elements currently displayed at the first location to the one or more additional surfaces; Displaying the one or more elements as virtual content on the one or more additional surfaces at the second location; The method according to item 1, further comprising: (Item 41) The method according to item 40, wherein the determination of the one or more additional surfaces is initiated after the change in the environment exceeds a time threshold. (Item 42) The method according to item 40, wherein the user pauses the active content displayed at the first location and resumes the active content so as to be displayed at the second location, and the active content resumes at the same interaction point where the user paused the active content at the first location. (Item 43) As the user leaves the first location, transitioning the spatialized audio being delivered to the user from a location associated with the content displayed at the first location to an audio virtual speaker directed at the center of the user's head; Transitioning from the audio virtual speaker directed at the center of the user's head to the spatialized audio being delivered to the user from the one or more additional surfaces displaying the one or more elements at the second location; The method according to item 40, further comprising: (Item 44) A method for push delivering content to a user of a mixed reality system, the method comprising: Receiving one or more available surfaces from a user's environment; Identifying one or more contents that match the dimensions of one available surface from the one or more available surfaces; Calculating a score based on comparing one or more constraints of the one or more contents with one or more surface constraints of the one available surface; Selecting content from the one or more contents having the highest score; Storing a one-to-one matching of the selected content and the one available surface; Displaying the selected content to the user on the available surface A method comprising: (Item 45) The method according to item 44, wherein the user's environment is the user's private residence. (Item 46) The method according to item 44, wherein the one or more available surfaces from the user's environment are around the user's focused viewing area. (Item 47) The method according to item 44, wherein the one or more contents are advertisements. (Item 48) The method according to item 47, wherein the advertisement is targeted at a specific group of users located in a specific environment. (Item 49) The method according to item 44, wherein the one or more contents are notifications from an application. (Item 50) The method according to item 49, wherein the application is a social media application. (Item 51) The method according to item 44, wherein one of the one or more constraints of the one or more contents is orientation. (Item 52) The method according to item 44, wherein the selected content is 3D content. (Item 53) An augmented reality (AR) display system, A head-mounted system, One or more sensors, One or more cameras comprising outward-facing cameras A head-mounted system comprising A processor for executing a set of program code instructions, A memory for holding the set of program code instructions, the set of program code instructions comprising Receiving content, Identifying one or more elements within the content, Determining one or more surfaces, Matching the one or more elements to the one or more surfaces, Displaying the one or more elements as virtual content on the one or more surfaces, A memory comprising program code for performing A system comprising. (Item 54) The system according to item 53, wherein the content includes at least one of pull-distributed content or push-distributed content. (Item 55) The system according to item 53, wherein identifying the one or more elements includes analyzing the content and identifying the one or more elements. (Item 56) The system according to item 53, wherein identifying the one or more elements includes determining one or more attributes for each of the one or more elements. (Item 57) The system according to item 53, further comprising program code for storing the one or more elements in one or more logical structures. (Item 58) The system according to item 53, wherein the one or more surfaces include at least one of a physical surface or a virtual surface. (Item 59) The system according to item 53, wherein determining the one or more surfaces includes analyzing the environment and determining at least one of the one or more surfaces. (Item 60) Determining the one or more surfaces includes receiving raw sensor data, simplifying the raw sensor data to produce simplified data, creating one or more virtual surfaces based on the simplified data, and including the one or more surfaces include the one or more virtual surfaces, the system according to item 53. (Item 61) The system according to item 53, wherein determining the one or more surfaces includes determining one or more attributes for each of the one or more surfaces. (Item 62) The system according to item 53, further comprising program code for storing the one or more surfaces in one or more logical structures. (Item 63) Matching the one or more elements to the one or more surfaces includes prioritizing the one or more elements, for each element of the one or more elements, comparing one or more attributes of the element with one or more attributes of each of the one or more surfaces, calculating a matching score based on the one or more attributes of the element and the one or more attributes of each of the one or more surfaces, identifying the best matching surface having the highest matching score, and including (Item 64) The system according to item 53, wherein one element is matched to one or more surfaces. (Item 65) The system according to item 53, wherein the content is streamed from a content provider. (Item 66) The system according to item 53, further comprising program code for displaying one or more surface options for displaying the one or more elements, at least in part, based on the changed view of the user. (Item 67) The system according to item 53, further comprising program code for overriding the display of the one or more elements on the one or more surfaces on which the display was matched. (Item 68) The system according to item 53, further comprising program code for moving the one or more elements displayed on the one or more surfaces to different surfaces, at least in part, based on selecting specific elements to be displayed on the one or more surfaces on which the user is to be moved to different surfaces. (Item 69) The system according to item 53, further comprising program code for slowly moving the display of the one or more elements onto a new surface in response to a change in the user's view from a first view to a second view, and following the change in the user's view to the second view. (Item 70) The system according to item 53, further comprising program code for temporarily stopping the display of the one or more elements on the one or more surfaces at the first location and resuming the display of the one or more elements on one or more other surfaces at the second location, at least in part, based on the user moving from the first location to the second location. (Item 71) Determining the one or more surfaces includes identifying one or more virtual objects for displaying the one or more elements, in the system according to item 53. (Item 72) The elements of the one or more elements are TV channels, in the system according to item 53. (Item 73) The user interacts with the elements of the one or more elements displayed by purchasing one or more items or services displayed to the user, in the system according to item 53. (Item 74) detecting a change in the environment from a first location to a second location; determining one or more additional surfaces at the second location; matching the one or more elements currently displayed at the first location to the one or more additional surfaces; displaying the one or more elements as virtual content on the one or more additional surfaces at the second location The system according to item 53, further comprising program code for performing. (Item 75) An augmented reality (AR) display system, A head-mounted system, one or more sensors, one or more cameras comprising an outward-facing camera A head-mounted system comprising; a processor for executing a set of program code instructions; a memory for holding the set of program code instructions, the set of program code instructions comprising: receiving one or more available surfaces from a user's environment; identifying one or more contents that match the dimensions of one available surface from the one or more available surfaces; calculating a score based on comparing one or more constraints of the one or more contents with one or more surface constraints of the one available surface; selecting content from the one or more contents having the highest score; storing a one-to-one matching of the selected content and the one available surface; displaying the selected content to a user on the available surface A memory comprising program code for performing; A system comprising. (Item 76) The system according to item 75, wherein the user's environment is the user's private residence. (Item 77) The system according to item 75, wherein the one or more available surfaces from the user's environment are around the user's focus view area. (Item 78) The system according to item 75, wherein the one or more contents are advertisements. (Item 79) The system according to item 75, wherein the one or more contents are notifications from an application. (Item 80) The system according to item 75, wherein one of the one or more constraints of the one or more contents is orientation. (Item 81) The system according to item 75, wherein the selected content is 3D content.

[0051] Further details of the features, objectives, and advantages of the present disclosure will be described below in the mode for carrying out the invention, the drawings, and the claims. Both the foregoing general description and the following detailed description are exemplary and explanatory and are not intended as limitations regarding the scope of the present disclosure.

Brief Description of the Drawings

[0052] The drawings illustrate the design and usefulness of various embodiments of the present disclosure. It should be noted that the figures are not drawn to scale, and elements with similar structures or features are represented by the same reference numerals throughout the figures. To gain a deeper understanding of the foregoing and other advantages and objectives of various embodiments of the present disclosure and the method for achieving them, a more detailed description of the present disclosure briefly described above will be provided by referring to the specific embodiments illustrated in the accompanying drawings. These drawings depict only typical embodiments of the present disclosure and, therefore, are not to be considered as limiting its scope. Under the understanding that the present disclosure will be described and explained with additional specificity and detail through the use of the accompanying drawings.

[0053]

Figure 1A

Figure 1B

[0054]

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

[0055]

Figure 3A

Figure 3B

[0056]

Figure 4

[0057]

Figure 5

[0058]

Figure 6

[0059]

Figure 7A

Figure 7B

[0060]

Figure 7C

[0061]

Figure 8

[0062]

Figure 9

[0063]

Figure 10

[0064]

Figure 11

[0065]

Figure 12

[0066]

Figure 13

[0067]

Figure 14A

Figure 14B

[0068]

Figure 15

[0069]

Figure 16

[0070]

Figure 17

[0071]

Figure 18

[0072]

Figure 19

[0073]

Figure 20A

Figure 20B

Figure 20C

Figure 20D

Figure 20E

Figure 20F

Figure 20G

Figure 20H

Figure 20I

Figure 20J

Figure 20K

Figure 20L

Figure 20M

Figure 20N

Figure 20O

[0074]

Figure 21

[0075]

Figure 22

DETAILED DESCRIPTION OF THE INVENTION

[0076] Various embodiments will now be described in detail with reference to the drawings, which are provided as illustrative examples of the disclosure to enable those skilled in the art to practice the disclosure. It should be noted that the following figures and examples are not meant to limit the scope of the disclosure. Where an element of the disclosure may be implemented partially or fully using known components (or methods or processes), only that portion of such known components (or methods or processes) that is necessary for an understanding of the disclosure will be described, and detailed descriptions of the other portions of such known components (or methods or processes) will be omitted so as not to obscure the disclosure. Further, various embodiments include presently and future known equivalents of the components referred to herein by way of example.

[0077] Embodiments of the present disclosure display content or content elements within a spatially organized 3D environment. For example, the content or content elements may include push delivery content, pull delivery content, first-party content, and third-party content. Push delivery content is content that a server (e.g., a content designer) sends to a client (e.g., a user), and the initial request originates from the server. Examples of push delivery content may include (a) notifications from various applications such as stock price notifications, news feeds, etc., (b) prioritized content such as updates and notifications from, for example, social media applications, email updates, and the like, and / or (c) advertisements targeting broad target groups and / or specific target groups and the like. Pull delivery content is content that a client (e.g., a user) requests from a server (e.g., a content designer), and the initial request originates from the client. Examples of pull delivery content may include (a) web pages requested by a user, for example, by using a browser, (b) streaming data from a content provider requested by a user, for example, by using a data streaming application such as a video and / or audio streaming application, and / or (c) any digital formatting that a user may request / access / query. First-party content is content generated by a client (e.g., a user) on any device owned / used by the client (e.g., client devices such as mobile phones, tablets, cameras, head-mounted display devices, and the like). Examples of first-party content include photos, videos, and the like. Third-party content is content generated by a party other than the client (e.g., a television network, a movie streaming service provider, a web page created by someone other than the user, and / or any data not generated by the user).Examples of third-party content can include web pages generated by someone other than the user, data / audio / video streams received from one or more sources and associated content, any data generated by someone other than the user, and the like.

[0078] Content can originate from web pages and / or applications on a head-mounted system, mobile devices (e.g., cell phones), tablets, televisions, servers, and the like. In some embodiments, content may be received from a laptop computer, a desktop computer, an email application with a link to the content, an electronic message including a reference or link to the content, and other applications or devices such as the like. The following detailed description includes, as content, examples of web pages. However, the content can be any content, and the principles disclosed herein will apply.

[0079] Block Diagram FIG. 1A illustrates an exemplary system and computer-implemented method for matching content elements of content to a spatial three-dimensional (3D) environment, according to some embodiments. System 100 includes a content structuring process 120, an environment structuring process 160, and a synthesis process 140. System 100 or a portion thereof may be implemented on a device such as a head-mounted display device.

[0080] The content structuring process 120 is a process that reads the content 110, organizes / stores the content 110 into a logical structure, makes the content 110 accessible, and makes it easier to extract content elements from the content 110 programmatically. The content structuring process 120 includes a parser 115. The parser 115 receives the content 110. For example, the parser 115 receives the content 110 from an entity (e.g., a content designer). The entity may be, for example, an application. The entity may be external to the system 100. The content 110 may be, for example, push-distributed content, pull-distributed content, first-party content, and / or third-party content as described above. An external web server may provide the content 110 when the content 110 is requested. The parser 115 analyzes the content 110 and identifies the content elements of the content 110. The parser 115 may identify the content elements in order to inventory the content 110 and subsequently organize and store them within a logical structure such as a table of contents, which may be a tree structure such as a document tree or graph and / or a database table such as a relational database table.

[0081] The parser 115 may identify / determine and store the attributes for each content element. The respective attributes of the content elements may be explicitly indicated by the content designer of the content 110, or may be determined or inferred by the parser 115, for example, based on the location of the content elements within the content 110. For example, the respective attributes of the content elements may be determined or inferred by the parser 115 based on the location of the content elements within the content 110 relative to each other. The attributes of the content elements are further described in detail below. The parser 115 may generate the individual attributes parsed from the content 110, along with a list of all the content elements. After parsing and storing the content elements, the parser 115 may order the content elements based on the associated priorities (e.g., from highest to lowest).

[0082] Some of the advantages of organizing and storing content elements within a logical structure are that once the content elements are organized and stored within the logical structure, the system 100 can query and manipulate the content elements. For example, in a hierarchical / logical structure represented as a tree structure with nodes, if a node is deleted, everything under the deleted node can also be deleted. Similarly, if a node is moved, everything under the node can move with it.

[0083] The environmental structuring process 160 is a process that analyzes environmental - related data and identifies surfaces. The environmental structuring process 160 may include a sensor 162, a computer vision processing unit (CVPU) 164, a perception framework 166, and an environmental analyzer 168. The sensor 162 provides raw data regarding real - world surfaces (e.g., a point cloud of objects and structures from the environment) to the CVPU 164 for processing. Examples of the sensor 162 may include a global positioning system (GPS), wireless signal sensors (WiFi, Bluetooth®, etc.), cameras, depth sensors, an inertial measurement unit (IMU) including accelerometer three - axis symmetry and gyroscope three - axis symmetry, magnetometers, radars, barometers, altimeters, accelerometers, exposure meters, gyroscopes, and / or equivalents.

[0084] The CVPU 164 simplifies or filters the raw data. In some embodiments, the CVPU 164 may filter out noise from the raw data and produce simplified raw data. In some embodiments, the CVPU 164 may filter out data that cannot be used and / or is not relevant to the current environmental scanning task from the raw data and / or the simplified raw data, and produce filtered data. The CVPU 164 may modify the remaining data by distance and planarity among group point cloud points to facilitate downstream surface extraction / identification / determination. The CVPU 164 provides the processed environmental data to the perception framework 166 for further processing.

[0085] The perception framework 166 receives the grouped point cloud points from the CVPU 164 and prepares the environmental data for the environmental analyzer 168. The perception framework 166 creates / determines a structure / surface / plane (e.g., a list of surfaces) and incorporates it into one or more data storage devices such as, for example, an external database, a local database, a dedicated local storage device, local memory, and equivalents. For example, the perception framework 166 processes all the grouped point cloud points received from the CVPU 164 in sequence and creates / determines a virtual structure / surface / plane corresponding to the real-world surface. The virtual plane may be four vertices (picked up from the grouped point cloud points) that create a virtually constructed rectangle (e.g., divided into two triangles within the rendering pipeline). The structure / surface / plane created / determined by the perception framework 166 is referred to as environmental data. When rendered and overlaid across the real-world surface, the virtual surface is placed substantially across its corresponding one or more real-world surfaces. In some embodiments, the virtual surface is placed completely across its corresponding one or more real-world surfaces. The perception framework 286 may maintain a one-to-one or one-to-many matching / mapping of the virtual surface to the corresponding real-world surface. The one-to-one or one-to-many matching / mapping may be used for querying. The perception framework 286 may update the one-to-one or one-to-many matching / mapping when the environment changes.

[0086] The environment parser 168 analyzes the environmental data from the perception framework 166 and determines the surfaces within the environment. The environment parser 168 may use object recognition to identify objects based on the environmental data received from the perception framework 166. Further details regarding object recognition are described in U.S. Patent No. 9,671,566 entitled "PLANAR WAVEGUIDE APPARATUS WITH DIFFRACTION ELEMENT(S) AND SYSTEM EMPLOYING SAME" and U.S. Patent No. 9,761,055 entitled "USING OBJECT RECOGNIZERS IN AN AUGMENTED OR VIRUTAL REALITY SYSTEM" (incorporated by reference). The environment parser 168 may organize and store the surfaces within a logical structure such as a table of surfaces for inventorying the surfaces. The table of surfaces may be, for example, an ordered array, a hierarchical table, a tree structure, a logical graph structure, and / or the like. In one embodiment, the ordered array may be linearly iterated until a well-fitting surface is determined. In one embodiment, with respect to a tree structure ordered by specific parameters (e.g., maximum surface area), the best-fitting surface may be determined by continuously comparing whether each surface within the tree is smaller or larger than the required area. In one embodiment, in a logical graph data structure, the best-fitting surface may be searched based on relevant adjacency parameters (e.g., distance from the viewer), or have a table with a fast search for specific surface requirements.

[0087] The data structure described above can be a place where the environment analyzer 168 stores data corresponding to the surfaces determined at runtime (updating the data based on environmental changes if necessary), processes surface matching, and launches any other algorithms. In one embodiment, the data structure described above with respect to the environment analyzer 168 may not be a place where the data is stored more persistently. The data may be stored more persistently by the perception framework 166, which can be the runtime memory RAM, an external database, a local database, and the like when receiving and processing the data. Before processing the surface, the environment analyzer 168 may receive surface data from a persistent storage device, import them into a logical data structure, and then launch a matching algorithm on the logical data structure.

[0088] The environment analyzer 168 may determine and store an attribute for each surface. Each attribute of the surface may be meaningful with respect to the attributes of the content elements in the content table from the analyzer 115. The attributes of the surface are further described in detail below. The environment analyzer 168 may generate individual attributes analyzed from the environment along with a list of all surfaces. After analyzing and storing the surfaces, the environment analyzer 168 may order the surfaces based on the associated priorities (e.g., from highest to lowest). The associated priorities of the surfaces may be established when the environment analyzer 168 receives surface data from a persistent storage device and imports them into a logical data structure. For example, if the logical data structure includes a binary search tree, for each surface from the storage device (received in a normal enumerated list), the environment analyzer 168 may first calculate a priority (e.g., based on one or more attributes of the surface) and then insert the surface into its appropriate place in the logical data structure. The environment analyzer 168 may analyze through a point cloud and extract surfaces and / or planes based on the proximity of points / relationships within the space. For example, the environment analyzer 168 may extract horizontal and vertical planes and associate sizes with the planes.

[0089] The content structuring process 120 analyzes through the content 110 and arranges content elements into a logical structure. The environment structuring process 160 analyzes through the data from the sensor 162 and arranges surfaces from the environment into a logical structure. The logical structure including content elements and the logical structure including surfaces are used for matching and manipulation. The logical structure including content elements may be different (in type) from the logical structure including surfaces.

[0090] The synthesis process 140 is a process that matches content elements from an analyzer 115 (e.g., a table of content elements stored in a logical structure) with surfaces from the environment from an environment analyzer 168 (e.g., a table of surfaces stored in a logical structure) and determines which content elements should be rendered / mapped / displayed on which surfaces of the environment. In some embodiments, as illustrated in FIG. 1A, the synthesis process 140 may include a matching module 142, a rendering module 146, and an optional virtual object creation module 144. In some embodiments, as illustrated in FIG. 1B, the synthesis process 140 may further include a display module 148 and a receiving module 150.

[0091] The matching module 142 pairs / matches content elements stored within the logical structure to surfaces stored within the logical structure. The matching may be a one-to-one or one-to-many matching of a content element and a surface (e.g., one content element and one surface, one content element and two or more surfaces, two or more content elements and one surface, etc.). In some embodiments, the matching module 142 may pair / match a content element to a portion of a surface. In some embodiments, the matching module 142 may pair / match one or more content elements to one surface. The matching module 142 compares the attributes of the content elements and the attributes of the surfaces. The matching module 142 matches the content elements to the surfaces based on content elements and surfaces that share similar and / or opposing attributes. Having such an organized infrastructure of content elements stored within the logical structure and surfaces stored within the logical structure enables matching rules, policies, and constraints to be easily created, updated, and implemented, and supports and improves the matching process performed by the matching module 142.

[0092] The matching module 142 may access one or more preference data structures such as user preferences, system preferences, and / or passable preferences, and may use the one or more preference data structures in the matching process. User preferences may be, for example, a model based on aggregated preferences based on past actions, and may be specific to a particular content element type. System preferences may include the top two or more surfaces for one content element, and the user may have the ability to navigate through the two or more surfaces and select a preferred surface. The top two or more surfaces may be based on user preferences and / or passable preferences. Passable preferences may be read from a cloud database, and passable preferences may be, for example, a model based on a grouping of other users, similar users, all users, similar environments, content element types, and / or equivalents. The passable preference database may pre-import consumer data (e.g., aggregated consumer data, consumer test data, etc.) and provide reasonable matching even before a large dataset (e.g., a user dataset) is accumulated.

[0093] The matching module 142 matches one content element to one or more surfaces, at least in part, based on a content vector (e.g., an orientation attribute), a head pose vector (e.g., an attribute of a VR / AR device that is not a surface), and a surface normal vector of one or more surfaces. The content vector, head pose vector, and surface normal vector are described in detail below.

[0094] The matching module 142 generates a matching result having at least a one-to-one or one-to-many matching / mapping of a content element and a surface (e.g., one content element and one surface, one content element and two or more surfaces, two or more content elements and one surface, etc.). The result may be stored in a cache memory or a persistent storage device for further processing. The result may be compiled and stored in a table to inventory the matching.

[0095] In some embodiments, the matching module 142 may generate matching results, and one content element may be matched / mapped to multiple surfaces such that the content element can be rendered and displayed on any one of the multiple surfaces. For example, the content element may be matched / mapped to five surfaces. The user may then select, from the five surfaces, a preferred surface on which the content element should then be displayed. In some embodiments, the matching module 142 may generate matching results, and one content element may be matched / mapped to the top three of the multiple surfaces.

[0096] In some embodiments, when the user selects or chooses a preferred surface, the selection made by the user may update the user preference so that the system 100 can make more accurate and precise recommendations of content elements for the surface.

[0097] If the matching module 142 matches all content elements to at least one surface or discards the content elements (e.g., for mapping to other surfaces or if no suitable match is found), the composition process 140 may proceed to the rendering module 146. In some embodiments, for content elements without a matching surface, the matching module 142 may create a match / mapping for the content element and a virtual surface. In some embodiments, the matching module 142 may close content elements without a matching surface.

[0098] The random virtual object creation module 144 may create a virtual object for displaying content elements such as on the surface of a virtual plane. During the matching process of the matching module 142, it may be determined that the virtual surface can be an arbitrary surface for displaying a certain content element. This determination may be based on the texture attributes, occupancy attributes, and / or other surface attributes determined by the environment analyzer 168 and / or the attributes of the content element determined by the analyzer 115. The texture attributes and occupancy attributes of the surface are described in detail below. For example, the matching module 142 may determine that the texture attributes and / or occupancy attributes may be ineligible attributes for the potential surface. The matching module 142 may determine that the content element can alternatively be displayed on the virtual surface, at least based on the texture attributes and / or occupancy attributes. The position of the virtual surface may be relative to the position of one or more (real) surfaces. For example, the position of the virtual surface may be at a certain distance away from the position of one or more (real) surfaces. In some embodiments, the matching module 142 may determine that there is no suitable (real) surface, or the sensor 162 may not detect any surface at all. Therefore, the virtual object creation module 144 may create a virtual surface for displaying the content element.

[0099] In some embodiments, the step of creating a virtual object for displaying a content element may be based on data received from a specific sensor or a plurality of sensors of the sensor 162 or the lack of sensor input from a specific sensor or a plurality of sensors. Data received from the environmental center sensor of the sensor 162 (such as a camera or a depth sensor) may indicate the lack of a suitable surface based on the user's current physical environment, or such a sensor may not be able to distinguish the presence of a surface at all (for example, a highly absorbent surface may make surface identification difficult depending on the quality of the depth sensor, or the lack of connectivity may prevent access to a certain shareable map that can provide surface information).

[0100] In some embodiments, if the environment analyzer 168 does not receive data from the sensor 162 or the perception framework 166 within a certain time frame, the environment analyzer 168 may passively determine that there is no suitable surface. In some embodiments, the sensor 162 may actively confirm that the environment-centered sensor cannot determine the surface and may pass such a determination to the environment analyzer 168 or the rendering module 146. In some embodiments, if the environment structuring 160 does not have a surface for providing to the synthesis process 140 by either the passive determination by the environment analyzer 168 or the active confirmation by the sensor 162, the synthesis process 140 may create a virtual surface or access a stored or registered surface such as from the memory module 152. In some embodiments, the environment analyzer 168 may directly receive surface data from a hot spot or a third-party perception framework or memory module, etc., without the input from the device-specific sensor 162.

[0101] In some embodiments, a certain sensor such as a GPS can determine that the user is present in a location that does not have a suitable surface for displaying content elements such as a park or a beach in an open space, or the only sensor that provides data does not provide mapping information but instead provides orientation information (such as a magnetometer). In some embodiments, a certain type of display content element may require a type of display surface that may not be available or detectable within the user's physical environment. For example, the user may desire to view a map that displays the direction in which the user should walk from a certain location in the user's hotel room. As the user navigates to that location, in order for the user to maintain a view of the route map, the AR system, based on the data received (or not received) from sensor 162, will enable the user to continuously view the route map from the starting position of the user's room in the hotel to the destination location on the route map. Since there may not be a proper surface available or detectable by the environmental analyzer 168 to display the route map, it may be necessary to consider creating a virtual object such as a virtual surface or a screen to display the route map. For example, the user may need to walk in an open area such as a park where network connectivity may be limited or blocked, enter an elevator, leave the hotel, there may not be a suitable surface available for displaying content elements, or there may be too much noise for the sensor to accurately detect the desired surface. In this example, based on the content to be displayed and potential problems that may include lack of network connectivity or lack of a suitable display surface (e.g., based on the GPS data of the user's current location), the AR system may determine that it may be best to create a virtual object to display the content element, as opposed to relying on the environmental analyzer 168 to use the information received from sensor 162 to find a suitable display surface. In some embodiments, the virtual object created to display the content element may be a prism.Further details regarding the prism are described in the co-owned U.S. Provisional Patent Application No. 62 / 610,101, filed on December 22, 2017, entitled "METHODS AND SYSTEM FOR MANAGING AND DISPLAYING VIRTUAL CONTENT IN A MIXED REALITY SYSTEM", which is incorporated by reference in its entirety. Those skilled in the art can understand more embodiments when it may be more beneficial to create a virtual surface for displaying content elements as opposed to displaying the content elements on a (real) surface.

[0102] The rendering module 146 renders the content elements on their matched surfaces. The matched surfaces may include real surfaces and / or virtual surfaces. In some embodiments, the matching is performed between the content elements and the surfaces, but the matching may not be a perfect match. For example, a content element may require a 2D area of 1000×500. However, the best-matching surface may have dimensions of 900×450. In one embodiment, the rendering module 146 may render a 1000×500 content element and best-fit it to a 900×450 surface, which may include, for example, scaling the content element while maintaining a constant aspect ratio. In another embodiment, the rendering module 146 may crop a 1000×500 content element to fit within a 900×450 surface.

[0103] In some embodiments, the device implementing the system 100 may be movable. For example, the device implementing the system 100 may move from a first room to a second room.

[0104] In some embodiments, content elements that are matched to surfaces (physical and / or virtual) within the first room may remain matched to the surfaces within the first room. For example, the device implementing the system 100 may move from the first room to the second room, and the content elements that are matched to the surfaces within the first room will not be matched to the surfaces within the second room and thus will not be rendered thereon. If the device then moves from the second room back to the first room, the content elements that are matched to the surfaces within the first room will be rendered / displayed thereon on the corresponding surfaces within the first room. In some embodiments, the content will continue to be rendered within the first room but will not be displayed because it will be outside the device's field of view. However, certain features such as audio playback or playback time will continue to operate such that when the device returns to a state where the matched content is within the field of view, the rendering will resume seamlessly (similar to the effect when a user leaves a room while a movie is playing on a conventional TV),

[0105] In some embodiments, content elements that are matched to surfaces within the first room may be matched to surfaces within the second room. For example, the device implementing the system 100 may move from the first room to the second room, and after the device is present within the second room, the environment structuring process 160 and the composition process 140 may occur / start / run, and the content elements may be matched to surfaces (physical and / or virtual) within the second room.

[0106] In some embodiments, some content elements matched to the surfaces in the first room may remain in the first room, while other content elements matched to the surfaces in the first room may move to the second room. For example, a first set of content elements matched to the surfaces in the first room may remain matched to the surfaces in the first room, while a second set of content elements matched to the surfaces in the first room may move to the second room with the device implementing the system 100. The second set of content elements may move with the device as the device proceeds from the first room to the second room. Whether a content element is in the first set or the second set may be determined based on the attributes of the content element, the attributes of one or more surfaces in the first room to which the content element is matched, user preferences, system preferences, and / or passable world preferences. Underlying these various scenarios is the fact that matching and rendering can be exclusive. That is, content may be matched to a surface but not rendered. This can save computing cycles and power because the user device does not always need to keep the surface matched, and selective rendering can reduce the latency when resuming visible content on the matched surface.

[0107] FIG. 1B illustrates an exemplary system and computer-implemented method for matching content elements of content to a spatial 3D environment, according to some embodiments. The system 105, similar to FIG. 1A, includes a content structuring process 120, an environment structuring process 160, and a synthesis process 140. The synthesis process 140 of FIG. 1B includes additional modules including a display module 148 and a receiving module 150.

[0108] As described above, the matching module 142 may generate a matching result, and one content element may be matched / mapped to multiple surfaces such that the content element can be rendered and displayed on any one of the multiple surfaces. The display module 148 displays the content element, or the outline of the content element, or a reduced-resolution version of the content element (each referred to herein as a "candidate view") on multiple surfaces or within multiple portions of a single surface. In some embodiments, the multiple surface displays are continuous such that only a single candidate view is visible to the user at a time and can be cycled or scrolled through one by one via additional candidate view options. In some embodiments, all candidate views are displayed simultaneously and the user selects a single candidate view (e.g., via voice command, input to a hardware interface, eye tracking, etc.). The receiving module 150 receives from the user a selection of one candidate view on the surface of the multiple surfaces. The selected candidate view may be referred to as the preferred surface. The preferred surface may be stored in the memory module 152 as a user preference or a passable preference such that when matching the content element to the surface, future matching may benefit from such a preference as indicated by the information flow 156 from the receiving module 150 to the matching module 142 or the information flow 154 from the memory module 152 to the matching model 142. According to some embodiments, the information flow 156 may be an iterative process such that after several iterations, user preferences may begin to take precedence over system and / or passable preferences. By comparison, the information flow 154 may be a fixed output such that the matching priority will always be given to the user of the moment, or other users entering or desiring to display content therein, of the system 100 of FIG. 1A or 105 of FIG. 1B. System and / or passable preferences may take precedence over user preferences, but as more information flow 156 continues over the use of the user of the system 100 of FIG. 1A or 105 of FIG. 1B, user preferences may begin to be prioritized by the system through a natural learning process algorithm.Thus, in some embodiments, the content element will be rendered / displayed on a preferred surface regardless of the availability of other surfaces or environmental inputs that would otherwise cause the matching module 142 to place the content element elsewhere. Similarly, the information flow 154 may define a rendering / display matching with the preferred surface for a second user who has never been present within that environment and has not constructed a repetitive information flow 156 to the preferred surface that the first user has.

[0109] Advantageously, as sensor data and virtual models are generally stored in short-term computer memory, a persistent storage module that stores the preferred surface can cycle more quickly through the synthesis process 140 if a device shutdown occurs between content placement sessions. For example, if the sensor 162 collects depth information to match content in a first session, creates a virtual mesh reconstruction through the environmental structuring 160, and the system shutdown empties the random access memory that stores that environmental data, the system will need to repeat the environmental structuring pipeline upon restart for the next matching session. However, because the computing resources are saved by the storage module 152, the matching 142 is updated with the preferred surface information without a full iteration of the environmental structuring process 160.

[0110] In some embodiments, the content element may be displayed on all of the top three surfaces at once. The user may then select a surface as a preferred surface from the top three surfaces. In some embodiments, the content element may be displayed on only one of the top three surfaces at once, along with an indication to the user that the content element may also be displayed on two other surfaces. The user may then navigate through the other surface options, and as each surface option is activated by the user, the content element may be displayed on the activated surface. The user may then select a surface as a preferred surface from the surface options.

[0111] Figures 2A-2E depict content elements (e.g., monitors such as computers, smartphones, tablets, TVs, web browsers, screens, etc.) that are matched to three possible locations within the user's physical environment 1105. Figure 2A shows how the content elements are matched / mapped to the three possible locations as indicated by the viewing location proposal 214. The three white dots displayed on the upper left side within the viewing location proposal 214 indicate that three display locations may exist. The fourth white dot with an "x" may be a close button for closing the viewing location proposal 214 and indicating the selection of a preferred display location based on the selected / highlighted display location when the user selects the "x". The display location 212a is the first option for displaying the content element as indicated by the highlighted first white dot among the three white dots on the upper left side. Figure 2B shows the same user environment 1105 where the display location 212b is the second option for displaying the content element as indicated by the highlighted second white dot among the three white dots on the upper left side. Figure 2C shows the same user environment 1105 where the display location 212c is the third option for displaying the content element as indicated by the highlighted third white dot among the three white dots on the upper left side. Those skilled in the art can understand that there may be other approaches for indicating the display options for the user to select, and the examples illustrated in Figures 2A-2C are merely one example. For example, another approach may be to display all the display options at once and let the user select the preferred option using a VR / AR device (e.g., via a controller, line of sight, etc.).

[0112] It should be understood that the AR system has a certain field of view onto which virtual content can be projected, and such a field of view is typically less than the potential of the entire human field of view. A human can generally have a natural field of view of 110 - 120 degrees, and in some embodiments, the display field of view of the AR system 224 is less than this potential, as depicted in FIG. 2D, meaning that the surface candidate 212c is within the user's natural field of view but outside the device's field of view (e.g., the system can render content on its surface, but in reality, it will not display the content). In some embodiments, a field of view attribute (the attribute will be further described below) is assigned to the surface to indicate whether the surface can support the content to be displayed for the device's display field. In some embodiments, surfaces outside the display field of view of the display are not presented to the user for displaying options as described above.

[0113] In some embodiments, the default position 212 is in front of the user at a predefined distance (such as a certain focal distance specification of the device's display system) as a virtual surface, as depicted in FIG. 2E. The user can then adjust the default position 212 to a desired position within the environment (e.g., either a registered location from the storage device 285 or a matched surface from the synthesis process 140) by means of head pose or gesture or other input means measured by the sensor 162. In some embodiments, the default position may remain fixed for the user such that as the user moves through the environment, the default position 212 remains within substantially the same portion of the user's field of view (which is the same as the display field of view in this embodiment).

[0114] FIG. 2E also illustrates, as an example, a virtual television (TV) having three TV application previews (e.g., TV App1, TV App2, TV App3) associated with the virtual TV at a default position 212. The three TV applications may correspond to different TV channels or different TV applications corresponding to different TV channel / TV content providers. The user may extract a single channel for TV playback by selecting an individual TV application / channel shown below the virtual TV. The user may extract a channel for the TV by: (a) using the totem to target the channel, (b) pressing the trigger on the totem to select the channel and holding the trigger for a time period (e.g., about 1 second), (c) moving the totem to identify a desired location within the environment for displaying the extracted TV channel, and (d) pressing the trigger on the totem to place the extracted TV channel at the desired location within the environment. The step of selecting virtual content is further described in U.S. Patent Application No. 15 / 296,869, filed Oct. 20, 2015, and titled "SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE", the contents of which are incorporated herein by reference.

[0115] The desired location may be a surface suitable for displaying TV channels or other surfaces identified in accordance with the teachings of the present disclosure. In some embodiments, a new prism may be created at the desired location, along with selected channel content that is loaded and displayed within the new prism. Further details regarding the totem are described in U.S. Patent No. 9,671,566, entitled "PLANAR WAVEGUIDE APPARATUS WITH DIFFRACTION ELEMENT(S) AND SYSTEM EMPLOYING SAME" (incorporated by reference in its entirety). In some embodiments, the three TV applications may be "channel previews", which are small apertures for viewing the content being played on an individual channel by displaying a dynamic or static depiction of the channel content. In some embodiments, video may be presented to the user while moving the (c) totem and identifying the desired location within the environment for displaying the extracted TV channel. The video may be, for example, a single image illustrating the channel, one or more images illustrating a preview of the channel, a video stream illustrating the current content of the channel, and the like. The video stream may be, for example, low resolution or high resolution and may vary (such as resolution, frame rate, etc.) as a function of the available resources and / or bandwidth.

[0116] Figures 2A-2E illustrate different display options for presenting content (e.g., elements as virtual content) within the original field of view of a user and / or a device (e.g., based on a particular head pose). In some embodiments, the field of view of the user and / or device may change (e.g., the user moves their head from one field of view to another). As a result of the changed field of view, additional surface options for presenting the content may become available to the user, at least in part, based on the change in the user's field of view (e.g., a change in head pose). The additional surface options for presenting the content may also become available, at least in part, based on other surfaces that were not originally available within the original field of view of the user and / or device but are now visible to the user based on the change in the user's field of view. Thus, the viewing location option 214 of FIGS. 2A-2D may also depict additional options for presenting the content. For example, FIGS. 2A-2D depict three display options. As the user's field of view changes, more display options may become available, which may result in the viewing location option 214 displaying more dots to indicate the additional display options. Similarly, if the new field of view has fewer surface options, the viewing location option 214 may display less than three dots to indicate some of the display options available for presenting the content in the new field of view. Thus, based on the user's changed field of view, one or more additional surface options for presenting the content may be presented to the user for selection, and the changed field of view corresponds to a change in the user's head pose and / or the device.

[0117] In some embodiments, a user and / or a device may have a first view. The first view may be used to generate surface options for displaying content. For example, three surface options within the first view may be available for displaying content elements. The user may then change the view from the first view to a second view. The second view may then be used to generate additional surface options for displaying content. For example, two surface options within the second view may be available for displaying content elements. There may be a total of five surface options between the surfaces within the first view and the surfaces within the second view. The five surface options may be presented to the user as view location suggestions. If the user is looking within the second view and selects a view location suggestion within the first view, the user may receive an indication (e.g., an arrow, a glow, etc.) in the direction of the first view indicating that the user should navigate in the direction shown to return to the first view to view the surface option / view location selected by the user.

[0118] In some embodiments, a user may view content displayed on a first surface within a first field of view. The first field of view may have an associated first head pose. When the user changes their field of view from the first field of view to a second field of view, after a certain time period, the system may provide the user with an option to change the display location of the content from the first surface within the first field of view to one or more surface options within the second field of view. The second field of view may have an associated second head pose. In some embodiments, once the user's field of view has changed from the first field of view to the second field of view, and thus the first head pose of the user and / or device has changed to the second head pose of the user and / or device, and the first head pose and the second head pose have a change in position that exceeds a head pose change threshold, the system may immediately provide the user with an option to move the content. In some embodiments, a time threshold (e.g., 5 seconds) during which the user remains in the second field of view, and thus the state of the second head pose, may determine whether the system provides the user with an option to change the display location of the content. In some embodiments, the change in the field of view that triggers the system to provide an option to change the display location of the content may be a slight change, such as less than a corresponding head pose change threshold (e.g., less than 90 degrees in any direction relative to the direction of the first field of view, and thus the first head pose). In some embodiments, the change in the head pose may exceed the head pose change threshold (e.g., more than 90 degrees in any direction) before the system provides the user with an option to change the display location of the content. Thus, one or more additional surface options for displaying the content based on the changed field of view may be displayed at least in part based on a time threshold corresponding to the changed field of view. In some embodiments, one or more additional surface options for displaying the content based on the user's changed field of view may be displayed at least in part based on the head pose change threshold.

[0119] In some embodiments, the system may render / display the content on one or more first surfaces where the user viewing the content has a first head pose. The user viewing the content may change his and / or the device's head pose from the first head pose to a second head pose. In response to the change in head pose, the system may render / display the content on one or more second surfaces where the user viewing the content has a second head pose. In some embodiments, the system may provide the user with an option to change the rendering / display location of the content from one or more first surfaces to one or more second surfaces. In some embodiments, once the user's head pose changes from the first head pose to the second head pose, the system may immediately provide the user with an option to move the content. In some embodiments, the system may provide the user with an option to change the rendering / display location of the content if the head pose change exceeds a corresponding head pose change threshold (e.g., 90 degrees). In some embodiments, the system may provide the user with an option to change the rendering / display location of the content if the head pose change is maintained over a threshold time period (e.g., 5 seconds). In some embodiments, the change in head pose that triggers the system to provide an option to change the rendering / display location of the content may be a slight change, such as less than a corresponding head pose change threshold (e.g., less than 90 degrees).

[0120] Attribute General Attributes As described above, the parser 115 may identify / determine and store attributes for each content element, and the environment parser 168 may determine and store attributes for each surface. The attributes of the content elements may be explicitly indicated by the content designer of the content 110 or may be determined or otherwise inferred by the parser 115. The attributes of the surface may be determined by the environment parser 168.

[0121] Attributes that both the content element and the surface may have include, for example, orientation, aspect ratio, dimension, area (e.g., size), relative viewing position, color, contrast, readability index, and / or time. Further details regarding these attributes are provided below. One of ordinary skill in the art can understand that the content element and the surface may have additional attributes.

[0122] Regarding the content element and the surface, the orientation attribute indicates the orientation. The orientation value may include vertical, horizontal, and / or a specific angle (e.g., 0 degrees with respect to horizontal, 90 degrees with respect to vertical, or anywhere between 0 and 90 degrees with respect to an angled orientation). The specific angle orientation attribute may be defined / determined in degrees or radians, or may be defined / determined with respect to the x-axis or y-axis. In some embodiments, an inclined surface may be defined, for example, to display a water flow of content flowing at a certain inclination angle and to show different artistic works. In some embodiments, regarding the content element, the navigation bar of the application may be defined to be in a horizontal orientation but inclined at a specific angle.

[0123] Regarding the content element and the surface, the aspect ratio attribute indicates the aspect ratio. The aspect ratio attribute may be defined, for example, as a 4:3 or 16:9 ratio. The content element may be scaled based on the aspect ratio attribute of the content element and one or more corresponding surfaces. In some embodiments, the system may determine the aspect ratio of a content element (e.g., a video) based on the attributes of other content elements (e.g., dimension and / or area), and scale the content element based on the determined aspect ratio. In some embodiments, the system may determine the aspect ratio of a surface based on the attributes of other surfaces.

[0124] Within the aspect ratio attribute, there may be specific properties that can be used by the content designer of a content element to recommend a specific aspect ratio for which the content element is to be maintained or changed. In one example, when this specific property is set to "maintain" or a similar keyword or phrase, the aspect ratio of the content element will be maintained (i.e., not changed). In one example, when this specific attribute is set to, for example, "free" or a similar keyword or phrase, the aspect ratio of the content element may be changed (e.g., scaled or otherwise), for example, to match the aspect ratio of one or more surfaces to which the content element is being matched. The default value of the aspect ratio attribute may be for maintaining the original aspect ratio of the content element, and the default value of the aspect ratio attribute may be overridden when the content designer specifies some other value or keyword for the aspect ratio attribute of the content element and / or when the system determines that the aspect ratio attribute should be overridden so as to better match the content element to one or more surfaces.

[0125] Regarding the content element and the surface, the dimension attribute indicates the dimensions. The dimension attribute of the content element may indicate the dimensions of the content element as a function of pixels (e.g., 800 pixels × 600 pixels). The dimension attribute of the surface may indicate the dimensions of the surface as a function of meters (e.g., 0.8 meters × 0.6 meters) or any other unit of measurement. The dimension attribute of the surface may indicate the measurable range of the surface, and the measurable range may include length, scope, depth, and / or height. Regarding the content element, the dimension attribute may be defined by the content designer so as to propose a certain shape and external size of the surface for displaying the content element.

[0126] Regarding content elements and surfaces, the area attribute indicates area or size. The area attribute of a content element may indicate the area of the content element as a function of pixels (e.g., 480,000 square pixels). The attribute of an area surface may indicate the area of the surface as a function of meters (e.g., 48 square meters) or any other unit of measurement. Regarding a surface, the area may be a perceived area perceived by a user or an absolute area. The perceived area may be defined by an absolute area that increases with the angle and distance of the content element displayed from the user such that when the content element is displayed closer to the user, the content element is perceived to be of a smaller size, and when the content element is farther away from the user, the content element is still perceived by the user to be of the same specific size, and vice versa when the content element comes closer to the user. The absolute area may simply be defined, for example, in square meters regardless of the distance from the content element displayed in the environment.

[0127] Regarding content elements and surfaces, the relative viewing position attribute relates to the position relative to the user's head pose vector. The head pose vector may be a combination of the position and orientation of a head-mounted device worn by the user. The position may be the fixed point of a device worn on the user's head that is tracked within a real-world coordinate system using information received from the environment and / or a user sensing system. The orientation component of the user's head pose vector may be defined by the relationship between a 3D device coordinate system local to the head-mounted device and a 3D real-world coordinate system. The device coordinate system may be defined by three orthogonal directions, namely, a forward-facing viewing direction approximating the line of sight in front of the user through the device, the upright direction of the device, and the right direction of the device. Other reference directions may also be selected. Information obtained by sensors within the environment and / or the user sensing system may be used to determine the orientation of the local coordinate system with respect to the real-world coordinate system.

[0128] To further illustrate the device coordinate system, when a user is wearing the device and it is suspended upside down, the upright direction with respect to the user and the device is actually the direction facing the ground (e.g., the downward direction of gravity). However, from the user's perspective, the relative upright direction of the device still aligns with the user's upright direction. For example, if the user is reading a book in the typical up / down left / right manner while suspended upside down, the user would be seen by others in the real-world coordinate system, who are standing normally and not suspended upside down, as holding the book upside down. However, with respect to a local device coordinate system that approximates the user's perspective, the book is oriented upright.

[0129] Regarding content elements, the relative visual position attribute may indicate the position at which the content element should be displayed with respect to the head pose vector. Regarding a surface, the relative visual position attribute may indicate the position of the surface within the environment with respect to the user's head pose vector. It should be understood that component vectors of the head pose vector, such as a forward-facing visual direction vector, may also be used as a reference for determining the relative visual position attribute of a surface and / or for determining the relative visual position attribute regarding a content element. For example, a content designer may indicate that a content element such as a search bar should always be at most 30 degrees to the left or right of the user's head pose vector, and that when the user moves more than 30 degrees to the left or right, the search bar should be adjusted so that it remains within 30 degrees to the left or right of the user's head pose vector. In some embodiments, the content is adjusted instantaneously. In some embodiments, the content is adjusted once a time threshold is met. For example, when the user moves more than 30 degrees to the left or right, the search bar should be adjusted after a 5-second time threshold has elapsed.

[0130] The relative viewing angle may be defined to maintain a certain angle or range of angles with respect to the user's head pose vector. For example, content elements such as videos may have a relative viewing position attribute indicating that the video should be displayed on a surface that is substantially orthogonal to the viewing vector facing the user. When the user stands in front of a surface such as a wall and is looking straight ahead, the relative viewing position attribute of the wall with respect to the viewing vector facing the user may meet the relative viewing position attribute requirements of the content element. However, when the user is looking down at the floor, the relative viewing position attribute of the wall changes, and the relative viewing position attribute of the floor better meets the relative viewing position attribute requirements of the content element. In such a scenario, the content element may be moved to be projected onto the floor instead of the wall. In some embodiments, the relative viewing position attribute may be the depth or distance from the user. In some embodiments, the relative viewing position attribute may be the relative position with respect to the user's current viewing position.

[0131] Regarding the content element and the surface, the color attribute indicates color. Regarding the content element, the color attribute may indicate one or more colors, whether the color can be changed, opacity, and the like. Regarding the surface, the color attribute may indicate one or more colors, color gradients, and the like. The color attribute may be associated with readability and / or the perception of the content element / the way the content element will be perceived on the surface. In some embodiments, the content designer may define the color of the content element, for example, as white or a light color. In some embodiments, the content designer may not desire the system to change the color of the content element (e.g., a company logo). In these embodiments, the system may change the background of one or more surfaces on which the content element is displayed to create the contrast necessary for readability.

[0132] With respect to a content element and a surface, a contrast attribute indicates contrast. With respect to a content element, the contrast attribute may indicate the current contrast, whether the contrast can be changed, the direction regarding how the contrast can be changed, and the like. With respect to a surface, the contrast attribute may indicate the current contrast. A contrast preference attribute may be associated with readability and / or the perception of the content element / how the content element will be perceived on the surface. In some embodiments, a content designer may desire that the content element be displayed with high contrast against the background of the surface. For example, a version of a content element may be presented within a web page as white text on a black background on a monitor of a computer, smartphone, tablet, etc. A white wall may be matched to display a text content element that is also white. In some embodiments, the system may change the text content element to a darker color (e.g., black) to provide contrast and satisfy the contrast attribute.

[0133] In some embodiments, the system may change the background color of the surface to provide color and / or contrast and satisfy the color and / or contrast attribute without changing the content element. The system may change the background color of the (real) surface by creating a virtual surface in the place of the (real) surface, and the color of the virtual surface is the desired background color. For example, if the color of a logo should not be changed, the system may change the background color of the surface to provide color contrast and satisfy the color and / or contrast preference attribute while preserving the logo to provide appropriate contrast.

[0134] With respect to a content element and a surface, the readability index attribute may indicate a readability metric. With respect to a content element, the readability index attribute indicates a readability metric that should be maintained with respect to the content element. With respect to a content element, the system may use the readability index attribute to determine a priority order with respect to other attributes. For example, the system may set the priority order with respect to these attributes to "high" if the readability index is "high" with respect to the content element. In some embodiments, even if the content element is in focus and there is proper contrast, the system may scale the content element based on the readability index attribute to ensure that the readability metric is maintained. In some embodiments, a high readability index attribute value for a particular content element may preempt or take precedence over other explicit attributes for other content elements if the priority order for the particular content element is set to "high". With respect to a surface, the readability index attribute may indicate how a content element including text will be perceived by a user when displayed on the surface.

[0135] Text legibility is a difficult problem to solve for a pure VR environment. This problem becomes even more complex in an AR environment because real-world colors, brightness, lighting, reflections, and other capabilities directly affect the user's ability to read text rendered by an AR device. For example, web content rendered by a web browser can be mainly text-driven. As an example, a set of JavaScript® APIs (e.g., via new extensions to the current W3C Camera API) can provide content designers with a current world palette and a contrast alternative palette for fonts and background colors. The set of JavaScript® APIs can provide content designers with a unique ability to adjust web content color schemes according to actual word color schemes to improve content contrast and text legibility (e.g., readability). Content designers can use this information by setting font colors to provide better legibility for web content. These APIs may be used to track this information in real time, and thus, a web page may appropriately adjust its contrast and color scheme in response to changes in the environment's light. For example, FIG. 3A illustrates how web content 313 is adjusted for the light and color conditions of a dark real-world environment by at least adjusting the text of web content 313 to be displayed in a light color scheme, making it easier to read against the dark real-world environment. As illustrated in FIG. 3A, the text within web content 313 has a light color, and the background within web content 313 has a dark color. FIG. 3B illustrates how web content 315 is adjusted for the light and color conditions of a bright real-world environment by at least adjusting the text of web content 315 to be displayed in a dark color scheme, making it easier to read against the bright real-world environment. As illustrated in FIG. 3B, the text within web content 315 has a dark color, and the background within web content 313 has a light color.One skilled in the art can understand that other factors, such as the background color of web content 313 (e.g., a darker background and lighter text) or web content 315 (e.g., a lighter background and darker text), may also be adjusted to provide color contrast so that the text can be made more legible, at least in part, based on the light and color conditions of the real-world environment.

[0136] Regarding a content element, the time attribute indicates the time at which the content element should be displayed. The time attribute may be short (e.g., less than 5 seconds), medium (e.g., 5 seconds to 30 seconds), long (e.g., more than 30 seconds). In some embodiments, the time attribute may be infinite. When the time attribute is infinite, the content element may remain until it is closed and / or another content element is loaded. In some embodiments, the time attribute may be a function of an input. In one example, if the content element is an article, the time attribute may be a function of an input indicating that the user has reached the end of the article and may remain there for a threshold time period. In one example, if the content element is a video, the time attribute may be a function of an input indicating that the user has reached the end of the video.

[0137] With respect to a surface, the time attribute indicates the time when the surface will be available. The time attribute may be short (e.g., less than 5 seconds), medium (e.g., 5 seconds to 30 seconds), long (e.g., greater than 30 seconds). In some embodiments, the time attribute may be infinite. In some embodiments, the time attribute may be, for example, a function of the sensor input from sensor 162. Sensor input from sensor 162, e.g., an IMU, accelerometer, gyroscope, and the like, may be used to predict the availability of the surface with respect to the field of view of the device. In one example, when the user is walking, a surface near the user may have a short time attribute, a surface a short distance away from the user may have a medium time attribute, and a surface at a greater distance may have a long time attribute. In one example, when the user is sitting still on a bench, the wall in front of the user may have an infinite time attribute until a change in data above a threshold is received from sensor 162, and then the time attribute of the wall in front of the user may change from infinite to another value.

[0138] Content element attribute The content element may have attributes specific to the content element, such as, for example, priority, surface type, position type, margin, content type, and / or focus attribute. Further details regarding these attributes are provided below. One of ordinary skill in the art may understand that the content element may have additional attributes.

[0139] The priority attribute indicates a priority value for a content element (e.g., video, photo, or text). The priority value may include high, medium, or low priority, e.g., a numerical value ranging from 0 to 100, and / or a required or not required indicator. In some embodiments, the priority value may be defined with respect to the content element itself. In some embodiments, the priority value may be defined with respect to a specific attribute. For example, a readability index attribute for a content element may be set high, indicating that the content designer has placed an emphasis on the readability of the content element.

[0140] The surface type attribute or "surface type" attribute indicates the type of surface to which a content element should match. The surface type attribute may be based on semantics such as whether a content element should be placed within a location and / or on a particular surface. In some examples, a content designer may suggest that a particular content element not be displayed across a window or painting. In some examples, a content designer may always suggest that a particular content element be displayed on the largest vertical surface substantially in front of the user.

[0141] The position type attribute indicates the position of a content element. The position type attribute may be dynamic or fixed. A dynamic position type may, for example, assume that a content element is attached to the user's hand such that as the user's hand moves, the content element moves dynamically with the user's hand. A fixed position type may, for example, assume that a content element is fixed relative to a surface, a specific location within an environment, or a virtual world relative to the user's body or head / viewing position, examples of which are described in more detail below.

[0142] The term "fixation" can also have different levels such as (a) world fixation, (b) object / surface fixation, (c) body fixation, and (d) head fixation. Regarding (a) world fixation, the content element is fixed with respect to the world. For example, when the user moves around in the world, the content element does not move and remains fixed at its location with respect to the world. Regarding (b) object / surface fixation, the content element is fixed to an object or surface such that when the object or surface is moved, the content element moves with the object or surface. For example, the content element may be fixed to a memo pad held by the user. In this case, the content is an object fixed to the surface of the memo pad and moves with the memo pad as appropriate. Regarding (c) body fixation, the content element is fixed with respect to the user's body. When the user moves their body, the content element moves with the user and maintains a fixed position with respect to the user's body. Regarding (d) head fixation, the content element is fixed with respect to the user's head or posture. When the user rotates their head, the content element will move in response to the user's head movement. Also, when the user walks, the content element will also move with respect to the user's head.

[0143] The margin (or padding) attribute indicates the margin around a content element. The margin attribute is a layout attribute that describes the location of a content element relative to other content elements. For example, the margin attribute represents the distance from the content element boundary to the nearest allowable boundary of another content element. In some embodiments, the distance may be an x, y, z coordinate-based margin measured from the vertex of the content element boundary or other specified location. In some embodiments, the distance may be a polar coordinate-based margin measured from the center of the content element or other specified location such as the vertex of the content element. In some embodiments, the margin attribute defines the distance from the content element to the actual content inside the content element. In some embodiments, with respect to a content element to be decomposed, etc., the margin attribute represents the amount of margin that should be maintained with respect to the boundary of the surface to which the content element to be decomposed is matched so that the margin serves as an offset between the content element and the matching surface. In some embodiments, the margin attribute may be extracted from the content element itself.

[0144] The content type attribute or "content type" attribute indicates the type of the content element. The content type may include a reference to and / or link to the corresponding media. For example, the content type attribute may define the content element as an image, video, music file, text file, video image, 3D image, 3D model, container content (e.g., any content that can be wrapped within a container), advertisement, and / or a content designer-defined rendering canvas (e.g., a 2D canvas or a 3D canvas). The content designer-defined rendering canvas may include, for example, games, rendering, maps, data visualization, and the like. The advertisement content type may include attributes that define the sound or advertisement to be presented to the user when the user focuses on a particular content element or in its vicinity. The advertisement may be (a) auditory such as a ringing sound, (b) visual such as a video / image / text, and / or (c) tactile indicators such as vibrations in the user's controller or headset and the like.

[0145] The focus attribute indicates whether the content element should be in focus. In some embodiments, the focus may be a function of the distance of the user and the content element from the surface on which it is displayed. If the focus attribute for the content element is set to always be in focus, the system keeps the content element in focus regardless of how far the user is from the content element. If the focus attribute for the content element is not defined, the system may take the content out of focus when the user is at a certain distance from the content element. This may depend on other attributes of the content element such as, for example, the dimension attribute, area attribute, relative visual recognition position attribute, and the like.

[0146] Surface attribute The surface may have attributes specific to the surface such as, for example, surface contour, texture, and / or occupancy attribute. Further details regarding these attributes are provided below. One skilled in the art may understand that the surface may have additional attributes.

[0147] In some embodiments, the environmental analyzer 168 may determine surface profile attributes (and associated attributes) such as surface normal vectors, orientation vectors, and / or upright vectors for one and / or all surfaces. In the case of 3D, the surface normal to the surface at point P, or simply the normal, is a vector perpendicular to the tangent plane to the surface at point P. The term "normal" may also be used as an adjective, i.e., a line normal to a plane, the normal component of a force, a normal vector, and the like.

[0148] The surface normal vector of the environmental surface surrounding at least one component of the user and head pose vectors discussed above may be such that certain attributes of the surface (e.g., size, texture, aspect ratio, etc.) may be ideal for displaying certain content elements (e.g., video, 3D model, text, etc.), but as such a surface is approximated by at least one component vector of the user's head pose vector, it may have an improper positioning of the corresponding surface normal with respect to the user's line of sight, which may be important for the matching module 142. By comparing the surface normal vector with the user's head pose vector, a surface that might otherwise be suitable for the content to be displayed may be disqualified or filtered.

[0149] For example, the surface normal vector of the surface may be in substantially the same direction as the user's head pose vector. This means that the user and the surface are facing in the same direction rather than towards each other. For example, if the user's forward direction faces north, a surface with a normal vector facing north either faces the user's back or the user faces the back of the surface. If the user cannot see the surface because the surface is facing away from the user, that particular surface will not be an optimal surface for displaying content, despite having beneficial attribute values for that surface that might otherwise be presented.

[0150] A comparison between the device's forward-facing viewing vector, approximating the user's forward-facing viewing direction, and the surface normal vector may provide a numerical value. For example, a dot product function may be used to compare two vectors and determine a numerical relationship that describes the relative angle between the two vectors. Such a calculation can result in a number between 1 and -1, with more negative values corresponding to a more favorable relative angle for viewing, as the surface approaches being orthogonal to the user's forward-facing viewing direction so that the user would be able to comfortably view virtual content placed on the surface. Thus, based on the identified surface normal vector, the characteristics for good surface selection may be with respect to the user's head pose vector or its components such that the content should be displayed on a surface facing towards the user's forward-facing viewing vector. It should be understood that constraints may be imposed on an acceptable relationship between the head pose vector components and the surface normal components. For example, all surfaces that result in a negative dot product with the user's forward-facing viewing vector may be selected such that they can be considered for content display. Depending on the content, content providers or algorithms or user preferences that affect the acceptable range may be considered. In instances where video needs to be displayed substantially normal to the user's forward direction, a smaller range of dot product outputs may be enabled. Those skilled in the art will understand that many design options are possible depending on other surface attributes, user preferences, content attributes, and the like.

[0151] In some embodiments, the surface may be an excellent fit from the perspective of size, location, and head pose. However, the surface may not be a good option for selection because the surface may include attributes such as texture attributes and / or occupancy attributes. The texture attribute may include materials and / or designs that can change the simple appearance of a clean and clear surface for presentation into a cluttered surface that is not ideal for presentation. For example, a brick wall may have a large open area that is ideal for displaying content. However, due to the red stacked bricks within the brick wall, the system may consider the brick wall undesirable for directly displaying content thereon. This is because the texture of the surface has a non-neutral red color that can induce roughness variations between the bricks and mortar and stronger contrast complexity with the content. Another undesirable texture example may include a surface having a wallpaper design with imperfections such as blowouts or non-uniform applications that create surface roughness variations, as well as background design patterns and colors. In some embodiments, the wallpaper design may include a very large number of patterns and / or colors such that directly displaying content on the wallpaper may not result in the content being displayed in a preferred view. The occupancy attribute may indicate that the surface is currently occupied by another content such that displaying additional content on a particular surface having a value indicating that the surface is occupied may result in the new content not being displayed across the occupied content, or vice versa. In some embodiments, the occupancy attribute draws attention to the presence of small real-world defects or objects that occupy the surface. Such occupying real-world objects may include items of negligible surface area (such as cracks or nails) that are indistinguishable to the depth sensors within the sensor array 162 but may be prominent to the cameras within the sensors 162.Other occupied real-world objects may include a photograph or poster hanging from a wall that has low texture variation associated with the surface on which it is placed, and may not be distinguished as different from the surface by some sensors 162, but the camera of 162 can recognize, and the occupancy attribute appropriately updates the surface, excluding the system's determination that the surface is an "empty canvas".

[0152] In some embodiments, the content may be displayed on a virtual surface whose relative position is related to the (real) surface. For example, if the texture attribute indicating the surface is not simple / clean and / or the occupancy attribute indicates that the surface is occupied, the content may be displayed, for example, on the virtual surface in front of the (real) surface within the margin attribute tolerance. In some embodiments, the margin attribute for a content element is a function of the texture attribute and / or the occupancy attribute of the surface.

[0153] Flow Matching of Content Element and Surface (Overview) FIG. 4 is a flowchart illustrating a method for matching content elements to a surface according to some embodiments. The method includes, at 410, receiving content; at 420, identifying content elements within the content; at 430, determining a surface; at 440, matching the content elements within the surface; and at 450, rendering the content elements as virtual content on the matched surface. The parser 115 receives the content 110 at 410. The parser 115 identifies the content elements within the content 110 at 420. The parser 115 may identify / determine and store an attribute for each content element. The environment analyzer 168 determines the surface within the environment at 430. The environment analyzer 168 may determine and store an attribute for each surface. In some embodiments, the environment analyzer 168 continuously determines the surface within the environment at 430. In some embodiments, the environment analyzer 168 determines the surface within the environment at 430 as the parser 115 receives the content 110 at 410 and / or identifies the content elements within the content 110 at 420. The matching module 142 matches the content elements to the surface based on the attributes of the content elements and the attributes of the surface at 440. The rendering module 146 renders the content elements on their matched surface at 450. The storage module 152 registers the surface for future use, such as by user specification, for installing future content elements thereon. In some embodiments, the storage module 152 may be within the perception framework 166.

[0154] Identification of Content Elements within Content FIG. 5 is a flowchart illustrating a method for identifying content elements within content according to some embodiments. FIG. 5 is a detailed flow that discloses steps for identifying elements within the content at 420 of FIG. 4 according to some embodiments. The method includes, at 510, a step of identifying content elements within the content, similar to the step of identifying elements within the content at 420 of FIG. 4. The method proceeds to the next step 520 of identifying / determining attributes. For example, the attributes may be identified / determined from tags regarding the location of the content. For example, the content designer may define where and how to display content elements using the attributes (described above) while designing and configuring the content. The attributes may be related to the location of content elements at specific locations relative to each other. In some embodiments, the step 520 of identifying / determining attributes may include a step of inferring the attributes. For example, the respective attributes of content elements may be determined or inferred based on the location of content elements within the content relative to each other. A step of extracting hints / tags from each content element is performed at 530. The hint or tag may be a formatting hint or formatting tag provided by the content designer of the content. A step of looking up / searching for alternative display forms for the content elements is performed at 540. A certain formatting rule may be defined regarding content elements to be displayed on a specific visual device. For example, a certain formatting rule may be defined regarding an image on a web page. The system may access the alternative display forms. A step of storing the identified content elements is performed at 550. The method may store the identified elements in a non-transitory storage medium for use in the synthesis process 140 to match the content elements to the surface. In some embodiments, the content elements may be stored in a transitory storage medium.

[0155] Determination of the surface within the environment FIG. 6 is a flowchart illustrating a method for determining a surface from a user's environment, according to some embodiments. FIG. 6 is an exemplary detailed flow that discloses the step of determining the surface at 430 of FIG. 4. FIG. 6 begins at 610 with the step of determining a surface. The step of determining a surface at 610 may include the step of collecting depth information of the environment from the depth sensor of sensor 162 and the step of performing reconstruction and / or surface analysis. In some embodiments, sensor 162 provides a map of points, and system 100 reconstructs a series of connected vertices between the points to create a virtual mesh representing the environment. In some embodiments, plane extraction or analysis is performed to determine mesh properties that indicate an interpretation of a common surface or the content of the surface (e.g., wall, ceiling, etc.). The method proceeds to the next step of determining the user's pose at 620, which may include the step of determining a head pose vector from sensor 162. In some embodiments, sensor 162 collects inertial measurement unit (IMU) data to determine the rotation of the device on the user. In some embodiments, sensor 162 collects camera images to determine the position of the device on the user relative to the real world. In some embodiments, the head pose vector is derived from one or both of the IMU and camera image data. The step of determining the user's pose at 620 is an important step for identifying the surface because the user's pose will provide a viewpoint regarding the user relative to the surface. At 630, the method determines the attributes of the surface. Each surface is tagged and categorized with a corresponding attribute. This information will be used when matching content elements with the surface. In some embodiments, sensor 162 from FIG. 1 provides raw data to CVPU 164 for processing, and CVPU 164 provides the processed data to perception framework 166 to prepare data for environment analyzer 168. Environment analyzer 168 analyzes the environmental data from perception framework 166 to determine the surfaces and corresponding attributes within the environment.At 640, the method stores an inventory of surfaces in a non-transitory storage medium for use by a composition process / matching / mapping routine to match / map the extracted elements to a particular surface. The non-transitory storage medium may include a data storage device. The determined surface may be stored in a particular table, such as the table disclosed in FIG. 15 described below. In some embodiments, the identified surface may be stored in a transitory storage medium. In some embodiments, the storing step at 640 includes designating the surface as a preferred surface for future matching of content elements.

[0156] Content Element and Surface Matching (Details) FIGS. 7A-7B are flow diagrams illustrating various ways to match a content element to a surface.

[0157] FIG. 7A depicts a flow diagram illustrating a method for matching a content element to a surface, according to some embodiments. FIG. 7A is a detailed flow disclosing the step of matching the content element at 440 of FIG. 4 to a surface.

[0158] At 710, the method determines whether the identified content element contains a hint provided by a content designer. The content designer may provide a hint regarding the location for best displaying the content element.

[0159] In some embodiments, this may be accomplished by further defining how content elements may be displayed using existing tag elements (e.g., HTML tag elements) when a 3D environment is available. As another example, a content designer may provide a hint stating that a 3D image is available as a resource for a particular content element instead of a 2D image. For example, in the case of a 2D image, in addition to providing a basic tag to identify a resource for a content element, the content designer may provide other less frequently used tags to identify resources including a 3D image corresponding to the 2D image, and in addition, provide a hint for prominently displaying it in the front of the user's view when the 3D image is used. In some embodiments, if the display device rendering the content may not have 3D display functionality for leveraging 3D images, the content designer may simply provide this additional "hint" for resources for 2D images.

[0160] At 720, the method determines whether to use hints provided by the content designer or a set of predefined rules for matching / mapping content elements to a surface. In some embodiments, for a particular content element, if there are no hints provided by the content designer, the system and method may use a set of predefined rules to determine the best way to match / map the particular content element to the surface. In some embodiments, even when there may be hints regarding content elements provided by the content designer, the system and method may determine that it may be best to use a set of predefined rules for matching / mapping the content elements to the surface. For example, if the content provider provides a hint for displaying video content on a horizontal surface, but the system is configured with predefined rules for displaying video content on a vertical surface, the predefined rules may override the hint. In some embodiments, the system and method may determine that the hints provided by the content designer are sufficient and thus use the hints to determine to match / map the content elements to the surface. Ultimately, the determination of whether to use hints provided by the content designer or predefined rules for matching / mapping content elements to the surface is up to the system's final decision.

[0161] At 730, if the system utilizes hints provided by the content designer, the system and method analyze the hints and search for a logical structure that includes the identified surrounding surfaces that can be used, at least in part, based on the hints, to display the particular content element.

[0162] At 740, the system and method activate the best-fit algorithm and select the best-fit surface for a particular content element based on the provided hint. The best-fit algorithm may, for example, obtain a hint regarding a particular content element that proposes a direct view and attempt to identify the surface that is front and center with respect to the current field of view of the user and / or device.

[0163] At 750, the system and method store the matching result having the matching of the content element and the surface. A table may be stored in a non-transitory storage medium for use by a display algorithm to display the content element on its respective matched / mapped surface.

[0164] FIG. 7B depicts a flow diagram illustrating a method for matching / mapping elements from a content element to a surface according to some embodiments. FIG. 7B is a flow illustrating the matching / mapping of a content element stored in a logical structure and a surface stored in the logical structure as disclosed in step 440 of FIG. 4 with reference to the various elements of FIG. 1.

[0165] At 715, the content elements stored in the logical structure resulting from the content structuring process 120 from FIG. 1 are ordered based on the associated priority. In some embodiments, the content designer may define a priority attribute for each content element. It may be beneficial for the content designer to set a priority for each content element to ensure that a particular content element is prominently displayed within the environment. In some embodiments, the content structuring process 120 may determine the priority for a content element, for example, if the content designer has not defined a priority for the content element. In some embodiments, the system will make the dot product function of the surface orientation the default priority attribute if no content element has a developer-provided priority attribute.

[0166] At 725, the attributes of the content element are compared to the attributes of the surface to identify whether there is a surface that matches the content element and to determine the best-matching surface. For example, starting with the content element with the highest association priority (e.g., the "main" or parent element ID, as described in more detail below with respect to FIG. 14A), the system compares the attributes of the content element to the attributes of the surface, identifies the best-matching surface, and then proceeds to the content element with the second highest association priority, etc., and thus may sequentially traverse the logical structure that includes the content elements.

[0167] At 735, a matching score is calculated based on the degree to which the attributes of the content element match the attributes of the best-matching surface to which they correspond. One of ordinary skill in the art can understand that many different scoring algorithms and models may be used to calculate the matching score. For example, in some embodiments, the score is the simple sum of the attribute values of the content element and the surface. FIG. 8 illustrates various matching score methodologies.

[0168] FIG. 8 depicts three hypothetical content elements and three hypothetical surfaces with attributes that could be within the logical structure described in more detail in FIGS. 14A - 14B below. Element A may have a preference for dot product orientation surface relationships for surface selection that is heavier than texture or color, Element B may have a preference for a smooth texture but is multi-color content with few contrast constraints and no prioritization of color, and Element C may be a virtual painting and may have a preference for a higher color attribute than other attributes. One of ordinary skill in the art will understand that the values within the content element structure can reflect the content itself (e.g., Element C places a high weight on color) or can reflect the desired surface attributes (e.g., Element B prefers a smoother surface for rendering). Further, although depicted as numerical values, other attribute values such as explicit colors within the color gamut or precise sizes or positions for the room / user, etc., are of course also possible as considerations.

[0169] At 745, the surface having the highest matching score is identified. Returning to the summation example illustrated in FIG. 8, element A has the highest score with respect to surfaces A and C, element B has the highest score with respect to surface B, and element C has the highest score with respect to surface C. In such an illustrative embodiment, the system may render element A on surface A, element B on surface B, and element C on surface C. Although element A is scored equally with respect to surfaces A and C, the highest score of element C with respect to surface C prompts the assignment of element A to surface A. In other words, the system iterates the second summation of the matching scores and determines the combination of content elements and surfaces that produce the highest aggregate matching score. Note that the number of samples for the dot product in FIG. 8 is not an objective measurement but reflects the attribute values. For example, although a -1 dot product result is a preferred mathematical relationship, to avoid introducing negative numbers into the equation, surface attributes are scored as a positive 1 for a -1 dot product relationship with respect to the surface attribute.

[0170] In some embodiments, the identification of the highest score at 745 is done either by marking the surface having the highest matching score as the surface list is evaluated and not marking previously marked surfaces, or by tracking the highest matching score and the link to the surface having the highest matching score, or by tracking the highest matching score of all content elements that match the surface. In one embodiment, once a surface with a sufficient matching score is identified, it may be removed from the surface list and thus excluded from further processing. In one embodiment, once the surface having the highest matching score is identified, it may remain in the surface list along with an indication that it matches the content element. In this embodiment, some surfaces may match a single content element and each matching score may be stored.

[0171] In some embodiments, as each surface from the ambient surface list is evaluated, a matching score is calculated one by one, a match is determined (e.g., 80% or more of the listed attributes of the content element are supported by the surface that constitutes the match), and if applicable, the individual surface is marked as the best match and the next surface is continued. If the next surface is a better match, it is marked as the best match. Once all surfaces have been evaluated with respect to the content element, the surface remains marked as the best match that is the best match assuming that surface. In some embodiments, the highest matching score may need to exceed a predetermined threshold to be qualified as a good match. For example, if it is determined that the best matching surface is only 40% matching (either by the number of supported attributes or the percentage of the target matching score), and the threshold for being qualified as a good match is greater than 75%, it may be best to create a virtual object for displaying the content element as opposed to relying on the surface from the user's environment. This may be applicable especially when the user's environment is a beach, for example, a beach without indistinguishable surfaces other than the beach, ocean, and sky. One of ordinary skill in the art can understand that there are many different matching / mapping algorithms that can be defined for this process, and this is merely one example of many different types of algorithms.

[0172] At 750, the matching / mapping results are stored as disclosed above. In one embodiment, if a surface is removed from the surface list at 745, the matching stored may be considered final. In one embodiment, if a surface remains within the surface list at 745 and some surfaces match a single content element, an algorithm may be launched on the colliding content elements and surfaces to remove the collisions and have a one-to-one matching instead of a one-to-many or many-to-one matching. If a high-priority content element does not match a surface, the high-priority content element may be matched / mapped to a virtual surface. If a low-priority content element does not match a surface, the rendering module 146 may elect not to render the low-priority content element. The matching results may be stored in a specific table, such as the table disclosed in FIG. 18 described below.

[0173] Referring back to FIG. 7A, assuming that at 760 it is determined that a predetermined rule should be used in the way to proceed, the method queries a database containing the content element and surface matching rules and determines, for a particular content element, the types of surfaces to be considered for matching the content element. At 770, a set of predefined rules may launch an optimal matching algorithm to select one or more surfaces from the available candidate surfaces that are the best fit for the content element. At least in part based on the optimal matching algorithm, it is determined that a particular one of the candidate surfaces, whose attributes best match the attributes of the content element, is the surface to which the content element should be matched / mapped. Once the matching of the content element and surface is determined at 750, the method stores the matching results for the content element and surface in a table within a non-transitory storage medium as described above.

[0174] In some embodiments, the user may override the matched surface. For example, the user may select a location to override the surface for displaying content even when the surface is determined by the matching algorithm to be the optimal surface for the content. In some embodiments, the user may select a surface from one or more surface options provided by the system, and the one or more surface options may include surfaces that are sub-optimal surfaces. The system may present the user with one or more display surface options, and the display surface options may include physical surfaces within the user's physical environment, virtual surfaces for displaying content within the user's physical environment, and / or virtual screens. In some embodiments, a stored screen (e.g., a virtual screen) may be selected by the user for displaying content. For example, with respect to a particular physical environment where the user is currently located, the user may have a preference for displaying a certain type of content (e.g., a video) on a certain type of surface (e.g., a stored screen having a default screen size, a location from the user, etc.). The stored screen may be a surface that has been frequently used in the history, or the stored screen may be a stored screen identified within the user's profile or preferences for displaying a certain type of content. Thus, the step of overriding the display of one or more elements on one or more surfaces may be based, at least in part, on surfaces that have been frequently used in the history and / or stored screens.

[0175] FIG. 7C illustrates an example where a user may be able to move content 780 from a first surface to any surface available to the user. For example, the user may be able to move content 780 from the first surface to a second surface (i.e., vertical wall 795). The vertical wall 795 may have a work area 784. The work area 784 of the vertical wall 795 may be determined, for example, by the environmental analyzer 168. The work area 784 may have a display surface 782 where the content 780 can be displayed, for example, without being obstructed by other content / objects. The display surface 782 may be determined, for example, by the environmental analyzer 168. In the example illustrated in FIG. 7C, the work area 784 includes a photo frame and a lamp, which may make the display surface of the work area 784 smaller than the entire work area 784 as illustrated by the display surface 782. The movement of the content to the vertical wall 795 (e.g., movements 786a - 786c) may not need to be a perfect placement of the content 780 at the center of the display surface 782, within the work area 784, and / or onto the vertical wall 795. Instead, the content may be moved within at least a portion of the surrounding work space of the vertical wall 795 (e.g., work area 784 and / or display surface 782) based on the user's gesture to move the content to the vertical wall. As long as the content 780 enters the vertical wall 795, the work area 784, and / or the display surface 782, the system will display the content 780 within the display surface 782.

[0176] In some embodiments, the surrounding work space is an abstract boundary that encloses a target display surface (e.g., display surface 782). In some embodiments, the user's gesture may be a selection of the content 780 on the first surface by a totem / controller 790 to move the content 780 such that at least a portion is within the surrounding work space of the display surface 782. The content 780 may then be aligned with the contour and orientation of the display surface 782. The step of selecting virtual content further claims priority on October 20, 2015, to "SELECTING VIRTUAL OBJECTS" Described in U.S. Patent Application No. 15 / 296,869 entitled "IN A THREE-DIMENSIONAL SPACE", the step of aligning content to a selected surface further claims priority on August 11, 2016, and is described in U.S. Patent Application No. 15 / 673,135 entitled "AUTOMATIC PLACEMENT OF A VIRTUAL OBJECT IN A THREE-DIMENSIONAL SPACE" (the content of each is incorporated herein by reference).

[0177] In some embodiments, the user gesture may be a hand gesture that may include (a) selection of content from a first surface, (b) movement of the content from the first surface to a second surface, and (c) indication of placement of the content on the second surface. In some embodiments, the movement of the content from the first surface to the second surface is with respect to a specific portion of the second surface. In some embodiments, the content, when placed on the second surface, conforms to / fills the second surface (e.g., is scaled to conform to, fill, etc.). In some embodiments, the content placed on the second surface maintains the size it had when on the first surface. In these embodiments, the second surface may be larger than the first surface and / or the second surface may be larger than the size required to display the content. The AR system may display the content at or near a position on the second surface where the user indicated to display the content. In other words, the movement of the content from the first surface to the second surface may not require the system to perfectly place the content within the entire workable space of the second surface. The content may only need to end up in at least a first peripheral area of the second surface that is at least visible to the user.

[0178] Environment-Driven Content What has been disclosed so far has been content-driven for where content elements in the environment should be displayed. In other words, a user can select various contents (e.g., pull-distributed content from a web page) for display in the user's environment. However, in some embodiments, the environment can drive the content presented to the user, at least in part, based on the user's environment and / or surfaces within the environment. For example, a list of surfaces is constantly being evaluated by the environment analyzer 168 based on data from the sensor 162. Since the list of surfaces is constantly being evaluated by the environment analyzer 168 as the user moves from one environment to another or moves around within an environment, push distribution (e.g., push-distributed content) can occur within the user's environment without the user having to search for or select on the content, and new / additional surfaces that may be suitable for displaying certain types of content that cannot originate from a web page may become available. For example, certain types of push-distributed content may include (a) notifications from various applications such as stock price notifications, news feeds, etc., (b) prioritized content such as updates and notifications from, for example, social media applications, email updates, and the like, and / or (c) advertisements targeting broad target groups and / or specific target groups and the like. Each of these types of push-distributed content may have associated attributes such as size, dimension, orientation, and the like in order to display the advertisement in its most effective form. Depending on the environment, a certain surface may present an opportunity for these environment-driven contents (e.g., push-distributed content) to be displayed. In some embodiments, pull-distributed content may first be matched / mapped to a surface within the environment, and push-distributed content may be matched / mapped to any surface within the environment that does not have pull-distributed content matched / mapped thereto.

[0179] As an example of a push - delivered content type, consider an advertisement and a scenario where a user is present in an environment where there can be many surfaces having various dimensions and orientations. A particular advertisement may be best displayed on a surface having a certain dimension and orientation and at a particular location (e.g., a geographical location such as home, workplace, baseball field, grocery store, and the like, and an item location such as the front of a certain physical item / product within the environment). In these situations, the system may search through a database of push - delivered content and determine the push - delivered content that best matches the surfaces of the environment. If a match is found, the content may be displayed on the matching surface at the particular location. In some embodiments, the system provides a list of surfaces to an advertisement server, which uses built - in logic to determine the push - delivered content.

[0180] Unlike conventional online advertisements that rely on the layout of a web page that a user views to determine a portion on a web page window having available space for an online advertisement to be displayed, the present disclosure includes an environment analyzer 168 that identifies surfaces within an environment and determines candidate surfaces for certain push - delivered content, such as an advertisement.

[0181] In some embodiments, the user may define preference attributes regarding the time and location at which a certain type of push - delivered content may be displayed. For example, the user may indicate a preference attribute to prominently display high - priority content from certain people or groups on the front surface of the user, while displaying other types of push - delivered content, such as advertisements, on a smaller surface perimeter relative to the user's primary focus view area. The user's primary focus view area is an area of the view that is substantially forward in the direction the user is looking, as opposed to a peripheral view, which is to the side of the user's primary focus view area. In some embodiments, the high - priority content elements (e.g., pull - delivered content as opposed to push - delivered content) selected by the user are displayed on the most prominent surface within the user's environment (e.g., within the user's focused viewing area), while other non - matching / non - mapping surfaces around the user's focused viewing area may be available for push - delivered content.

[0182] In some embodiments, a world - location context API may be provided to content designers / web developers / advertisers to create location - aware content. The world - location context API may provide a set of capabilities to describe the local context specific to the particular location where the user currently exists. The world - location context API may include location context information that includes identification of specific types of rooms (e.g., living room, gym, office, kitchen), specific queries performed by the user from various locations (e.g., the user tends to search for movies from the living room, music from the gym, recipes from the kitchen, etc.), and specific services and applications used by the user at various locations (e.g., the email client is used from the office, Netflix is used from the living room). A content designer may associate an action with respect to the world - location context as an attribute of a content element.

[0183] The content provider may use the search history, object recognition device, and application data along with this information to provide location-specific content. For example, when a user launches a search in the kitchen, the ads and search results will likely be mainly food-related because the search engine will be aware that the user is launching the search from the user's kitchen. The information provided by the World Location Context API can provide accurate information for each room or location, which is more accurate than geographical location and more context-aware than geofencing. FIG. 9 is an example of a method in which the World Location Context API can be used to provide location-specific context. As an example, the user's kitchen 905 may include location-specific content 915a, 915b, and / or 915c that is displayed on a surface within the user's current physical location. Content 915a may be a recipe for a particular meal, content 915b may be an advertisement for a meal, and / or 915c may be a suggestion for a meal to prepare in the kitchen.

[0184] FIG. 10 is an example of a method 1000 for push-delivering content to a user of a VR / AR system. At 1010, one or more surfaces and their attributes are determined. The one or more surfaces can be determined from the environment structuring process 160, and the environment analyzer 168 analyzes the environmental data, determines the surfaces within the environment, and organizes and stores the surfaces within a logical structure. The user's environment can carry location attributes regarding surfaces such as the user's private residence, a specific room within a dwelling, the user's workplace, and the like. The one or more surfaces may be in the periphery of the user's focus view area. In some embodiments, the one or more surfaces may be within the user's focus view area in response to push-deliverable content (e.g., an emergency notification from an authority entity, a whitelisted application / notification, etc.) that the user may desire to be notified of. The dimensions of the surface may be 2D and / or 3D dimensions.

[0185] At 1020, one or more content elements that match one or more surfaces are received. In some embodiments, the step of receiving a content element or a single content element is based on at least one surface attribute. For example, the location attribute of "kitchen" may prompt for the push delivery of content elements corresponding to food items. In another example, a user may be viewing a first content element on a first surface, and that content element may have child content elements that will be displayed on a second surface only if a surface with a certain surface attribute is available.

[0186] At 1030, a matching score is calculated based on the degree to which the attributes of the content element match the attributes of the surface. In some embodiments, the scoring may be based on a scale of 1 to 100, where a score of 100 is the highest score and a score of 1 is the lowest score. One of ordinary skill in the art can understand that many different scoring algorithms and models may be used to calculate the matching score. In some embodiments, if the content element includes a notification, the matching score calculated based on the attributes may indicate the priority level of the content element that needs to be notified, as opposed to matching with a specific surface. For example, when the content element is a notification from a social media application, the score may be based on the priority level of the notification defined by the user within the user's social media account, as opposed to the matching score based on the attributes of the social media content and the attributes of the surface.

[0187] In some embodiments, when the user is relatively stationary within their environment, the list of surfaces may not change much. However, when the user is in motion, the list of surfaces can change very rapidly depending on the speed at which the user is moving. In dynamic situations, a low matching score may be calculated if it is determined that the user cannot be stationary long enough to fully view the content. This determination of whether the user has sufficient time to view the entire content may be an attribute defined by the content designer.

[0188] At 1040, the content element with the highest matching score is selected. When there are competing content elements (e.g., advertisements) that are likely to be displayed to the user, there may be requirements for sorting through the competing content and picking out the preferred content elements. Here, as an example for selecting the preferred content element, one option is to base the competition on the degree to which the attributes of the content element match the attributes of the surface. As another example, the winner may be selected, at least in part, based on the amount of money that the content element provider may desire to be paid to display the push - delivered content. In some embodiments, the preferred content element may be selected based on the content type (e.g., 3D content or a notification from a social media contact).

[0189] At 1050, the matching / mapping of the preferred content to the corresponding surface may be stored in the cache memory or persistent memory. Storing the matching may be important as it may be important to maintain a certain history of the user's environment when the user returns, in situations where the user is in motion and the environment is changing. The matching / mapping may be stored in a table such as the table disclosed in FIG. 18. At 1060, the content is rendered on the corresponding surface. The matching may be a one - to - one or one - to - many matching / mapping of the content element and the surface.

[0190] Disclosed are the present system and method for decomposing content for display within an environment. Additionally, the present system and method may also push-distribute the content to the surface of a user of a virtual reality or augmented reality system.

Example

[0191] Web page Referring to FIG. 11, environment 1100 represents a physical environment and system for implementing the processes described herein (e.g., matching content elements from within a web page to be displayed on a surface within a user's physical environment 1105). Representative physical environments and systems of environment 1100 include a user's physical environment 1105 as viewed by user 1108 through a head-mounted system 1160. Representative systems of environment 1100 further include accessing content (e.g., a web page) via a web browser 1110 operably coupled to a network 1120. In some embodiments, access to the content may be via an application (not shown) such as a video streaming application, and the video stream may be the content being accessed. In some embodiments, the video streaming application may be a sports organization, and the content being streamed may be an actual live game, recap, summary / highlight, box score, play-by-play, team statistics, player statistics, related videos, news feed, product information, and the like.

[0192] Network 1120 may be the Internet, an internal network, a private cloud network, a public cloud network, or the like. Web browser 1110 is also operably coupled to processor 1170 via network 1120. Although processor 1170 is shown as a separate and isolated component from head-mounted system 1160, in alternative embodiments, processor 1170 may be integrated with one or more components of head-mounted system 1160 and / or integrated within other system components within environment 1100, such as, for example, network 1120, and may access computing network 1125 and memory device 1130. Processor 1170 may be configured with software 1150 for receiving and processing information such as video, audio, and content received from head-mounted system 1160, local memory device 1140, web browser 1110, computing network 1125, and memory device 1130. Software 1150 may communicate with computing network 1125 and memory device 1130 via network 1120. Software 1150 may be installed on processor 1170, or in another embodiment, the features and functionality of the software may be integrated within processor 1170. Processor 1170 may also be configured with local memory device 1140 to store information used by processor 1170 for quick access without relying on information remotely stored on an external memory device in the vicinity of user 1108. In other embodiments, processor 1170 may be integrated with head-mounted system 1160.

[0193] The user's physical environment 1105 is the physical surroundings of user 1108 as the user moves around and views the user's physical environment 1105 through the head-mounted system 1160. For example, referring to FIG. 1, the user's physical environment 1105 shows a room with two walls (e.g., main wall 1180 and side wall 1184, where the main wall and the side wall are with respect to the user's view), and a table 1188. On the main wall 1180, there is a rectangular surface 1182 depicted by a black solid line, which represents a physical surface with a physical boundary that can be a candidate surface for projecting certain content (e.g., a painting or the like hanging from or attached to the wall or window). On the side wall 1184, there is a second rectangular surface 1186 depicted by a black solid line, which represents a physical surface with a physical boundary (e.g., a painting or the like hanging from or attached to the wall or window). Different objects may exist on the table 1188. 1) A virtual Rolodex 1190 on which certain content can be stored and displayed, 2) A horizontal surface 1192 depicted by a black solid line, which represents a physical surface with a physical boundary for projecting certain content, and 3) A plurality of stacks of virtual square surfaces 1194 depicted by black dotted lines, which represent, for example, stacked virtual newspapers on which certain content can be stored and displayed. Those skilled in the art will understand that the physical boundaries described above are useful for installing content elements because they already divide the surface into discrete viewing sections and can themselves be surface attributes, but are not necessary for recognizing eligible surfaces.

[0194] The web browser 1110 may also display blog pages from the Internet or within an intranet / private network. Additionally, the web browser 1110 may also be any technology for displaying digital content. The digital content may include, for example, web pages, blogs, digital photos, videos, news articles, newsletters, or music. The content may be stored in the memory device 1130 and be accessible by the user 1108 via the network 1120. In some embodiments, the content may also be streaming content, such as a live video feed or a live audio feed. The memory device 1130 may include, for example, a database, a file system, a persistent memory device, a flash drive, a cache, etc. In some embodiments, the web browser 1110 containing the content (e.g., a web page) is displayed via the computing network 1125.

[0195] The computing network 1125 accesses the memory device 1130, reads and stores the content in order to display the web page on the web browser 1110. In some embodiments, the local memory device 1140 may provide content of interest to the user 1108. The local memory device 1140 may include, for example, a flash drive, a cache, a hard drive, a database, a file system, etc. The information stored in the local memory device 1140 may include the most recently accessed content or the content most recently displayed within the 3D space. The local memory device 1140 enables performance improvement for the system of the environment 1100 by locally providing certain content to the software 1150 to help decompose the content for display on a 3D space environment (e.g., a 3D surface within the physical environment 1105 of the user).

[0196] Software 1150 includes a software program stored in a non-transitory computer-readable medium and implements a function to decompose content for display within the physical environment 1105 of a user. Software 1150 may be launched on processor 1170, which may be locally attached to user 1108, or in some other embodiments, software 1150 and processor 1170 may be included within a head-mounted system 1160. In some embodiments, some of the features and functions of software 1150 may be stored and executed on a computing network 1125 remote from user 1108. For example, in some embodiments, the step of decomposing content may occur on computing network 1125, the result of the decomposition may be stored in storage device 1130, the inventorying of the surface of the user's local environment for presenting the decomposed content may occur within processor 1170, and the surface inventory and matching / mapping may be stored in local storage device 1140. In one embodiment, the processes of decomposing content, inventorying the local surface, matching / mapping the elements of the content to the local surface, and displaying the elements of the content may all occur locally within processor 1170 and software 1150.

[0197] The head-mounted system 1160 may be a virtual reality (VR) or augmented reality (AR) head-mounted system (e.g., a mixed reality device) that includes a user interface, a user sensing system, an environment sensing system, a processor, etc. (not all shown). The head-mounted system 1160 presents an interface for the user 1108 to interact with and experience the digital world. Such interactions may involve the user and the digital world, one or more other users interfacing with the environment 1100, and objects within the digital and physical worlds.

[0198] The user interface may include the step of receiving content and the step of selecting an element within the content by user input through the user interface. The user interface may be at least one or a combination of a tactile interface device, a keyboard, a mouse, a joystick, a motion capture controller, an optical tracking device, and an audio input device. The tactile interface device is a device that enables a human to interact with a computer through physical sensations and movements. Tactility refers to a type of human-computer interaction technology that includes tactile feedback or other physical sensations for performing an action or is processed on a computing device.

[0199] The user perception system may include one or more sensors 1162 operable to detect certain features, characteristics, or information related to the user 1108 wearing the head-mounted system 1160. For example, in some embodiments, the sensor 1162 may include a camera or an optical detection / scanning circuit capable of detecting real-time optical characteristics / measurements of the user 1108, such as, for example, one or more of the following, namely, miosis / mydriasis, angular measurement / positioning of each pupil, sphericity, eye shape (as the eye shape changes over time), and other anatomical data. This data may be used by the head-mounted system 1160 to improve the user's visual experience, provide information (e.g., the user's visual focus point), or be used to calculate it.

[0200] The environmental perception system may include one or more sensors 1164 for obtaining data from the user's physical environment 1105. The objects or information detected by the sensors 1164 may be provided as input to the head-mounted system 1160. In some embodiments, this input may represent user interaction with the virtual world. For example, a user (e.g., user 1108) viewing a virtual keyboard on a desk (e.g., table 1188) may gesture with their finger as if the user is typing on the virtual keyboard. The motion of the finger movement may be captured by the sensors 1164 and provided as input to the head-mounted system 1160, and the input may be used to change the virtual world or create new virtual objects.

[0201] The sensors 1164 may include, for example, a camera or scanner oriented substantially outwardly, for interpreting scene information through, for example, continuously and / or intermittently projected infrared structured light. The environmental perception system may be used to match / map one or more elements of the user's physical environment 1105 around the user 1108 by detecting and registering the local environment, including static objects, dynamic objects, people, gestures, and various lighting, atmospheric, and acoustic conditions. Thus, in some embodiments, the environmental perception system may include image-based 3D reconstruction software built into a local computing system (e.g., processor 1170) and operable to digitally reconstruct one or more objects or information detected by the sensors 1164.

[0202] In one exemplary embodiment, the environmental sensing system provides one or more of the following, namely, motion capture data (including gesture recognition), depth sensing, face recognition, object recognition, unique object feature recognition, voice / audio recognition and processing, acoustic source localization, noise reduction, infrared or similar laser projection, and monochrome and / or color CMOS sensors (or other similar sensors), field of view sensors, and various other optical enhancement sensors. It should be understood that the environmental sensing system may include other components other than those discussed above.

[0203] As described above, in some embodiments, the processor 1170 may be integrated with other components of the system of the environment 1100, which may also be integrated with other components of the head-mounted system 1160, or may be a separate device (separable from the wearable or the user 1108) as shown in FIG. 1. The processor 1170 may be connected to various components of the head-mounted system 1160 through a physical wired connection or through a wireless connection such as, for example, a mobile network connection (including cellular phones and data networks), Wi-Fi, Bluetooth®, or any other wireless connection protocol. The processor 1170 may include a memory module, an integrated and / or additional graphics processing unit, wireless and / or wired Internet connectivity, and a codec and / or firmware capable of converting data from sources (such as the computing network 1125 and the user sensing system and environmental sensing system from the head-mounted system 1160) into image and audio data, and the images / videos and audio may be presented to the user 1108 via a user interface (not shown).

[0204] Processor 1170 handles data processing for various components of the head-mounted system 1160 and data exchange between the head-mounted system 1160 and content from web pages displayed or accessed by the web browser 1110 and the computing network 1125. For example, processor 1170 may buffer and process data streaming between user 1108 and computing network 1125 and thereby be used to enable a smooth, continuous, and high-fidelity user experience.

[0205] The step of decomposing content from a web page into content elements and matching / mapping the elements to be displayed on a surface within a 3D environment may be performed in an intelligent and logical manner. For example, content parser 115 may be a document object model (DOM) parser that receives an input (e.g., an entire HTML page) such that the elements of the content are accessible and easy to manipulate / extract programmatically, decomposes various content elements within the input, and stores the decomposed content elements within a logical structure. A set of predetermined rules may be available to recommend, suggest, or prescribe where a certain type of element / content identified within a web page should be placed. For example, one type of content element may need to be matched / mapped to a physical or virtual object surface that is easy to handle for storing and displaying one or more elements, while another type of content element may be a single object such as the main video or main article within a web page, in which case the single object may be matched / mapped to the surface that makes the most sense for displaying the single object to the user. In some embodiments, the single object may be a video streamed from a video application such that the single content object can be displayed on a surface (e.g., a virtual surface or a physical surface) within the user's environment.

[0206] The environment 1200 of FIG. 12 depicts content (e.g., a web page) that is displayed or accessed by the web browser 1110 and the user's physical environment 1105. The dotted line with arrowheads depicts elements (e.g., a particular type of content) from the content (e.g., a web page) that are matched / mapped to and displayed on the user's physical environment 1105. Some elements from the content are matched / mapped to some physical or virtual object within the user's physical environment 1105 based on either web designer hints or predefined browser rules.

[0207] As an example, the content accessed or displayed by the web browser 1110 may be a web page having a plurality of tabs, with the current active tab 1260 being displayed and the secondary tabs 1250 being hidden until currently selected according to the display on the web browser 1110. What is displayed within the active tab 1260 is typically a web page. In this particular example, the active tab 1260 displays a YOUTUBE (registered trademark) page that includes a main video 1220, user comments 1230, and proposed videos 1240. As depicted in the exemplary FIG. 12, the main video 1220 may be matched / mapped to be displayed on the vertical surface 1182, the user comments 1230 may be matched / mapped to be displayed on the horizontal surface 1192, and the proposed videos 1240 may be matched / mapped to be displayed on a vertical surface 1186 different from the vertical surface 1182. Additionally, the secondary tabs 1250 may be matched / mapped to be displayed on or as the virtual Rolodex 1190 and / or on the multi-stack virtual object 1194. In some embodiments, the specific content within the secondary tabs 1250 may be stored within the multi-stack virtual object 1194. In other embodiments, the entire content resident within the secondary tabs 1250 may be stored and / or displayed on the multi-stack virtual object 1194. Similarly, the virtual Rolodex 1190 may contain the specific content from the secondary tabs 1250, or the virtual Rolodex 1190 may contain the entire content resident within the secondary tabs 1250.

[0208] In some embodiments, the content elements of the web browser 1110 (e.g., the content elements of the web page within the secondary tab 1250) may be displayed on a two-sided planar window virtual object (not shown) within the user's physical environment 1105. For example, the primary content of the web page may be displayed on the first side (e.g., the front side) of the planar window virtual object, and additional information such as surplus content related to the primary content may be displayed on the second side (e.g., the back side) of the planar window virtual object. As an example, a retail store web page (e.g., BESTBUY) may be displayed on the first side, and a set of coupons and discounts may be displayed on the second side. The discount information is an update related to the second side and may reflect the current context of the object being browsed on the first side (e.g., on the second side, only laptops or home appliances are discounted).

[0209] Some web pages can span multiple pages when viewed within web browser 1110. When viewed within web browser 1110, such web pages can be viewed by scrolling within web browser 1110 or by navigating through multiple pages within web browser 1110. When matching / mapping such a web page from web browser 1110 to the user's physical environment 1105, such a web page may be matched / mapped as a double-sided web page. FIGS. 13A-13B illustrate exemplary double-sided web pages according to some embodiments. FIG. 13A shows a smoothie drink, while FIG. 13B illustrates an exemplary back side / second side of the smoothie drink that includes raw materials and instructions for making the smoothie. In some embodiments, the front side of main wall 1180 may include the first side of the double-sided web page, and the back side of main wall 1180 may include the second side of the double-sided web page. In this example, user 1108 would need to walk around the perimeter of main wall 1180 to view both sides of the double-sided web page. In some embodiments, the front side of main wall 1180 may include both sides of the double-sided web page. In this example, user 1108 may toggle between the two sides of the double-sided web page via user input. The double-sided web page may appear to flip from the first side to the second side in response to user input. The double-sided web page is described as being generated from a web page that spans multiple pages when viewed within web browser 1110, but the double-sided web page may be generated from any web page or a portion or multiple portions thereof. A VR and / or AR system may add to existing content (e.g., secondary tab 1250 or web page) and make available a set of HTML properties that are easy to use for a rendering module to render the content on a double-sided 2D browser plane window virtual object. The examples describe a double-sided plane window virtual object, but the virtual object may have any number of sides (N sides).The embodiments describe the step of displaying content on a two-sided planar window virtual object, although the content elements may be on multiple surfaces of a physical object (e.g., the front side of a door and the back side of the door).

[0210] The vertical surface 1182 can be any type of structure that may already be on the main wall 1180 of a room (depicted as the user's physical environment 1105), such as a window glass or a photo frame. In some embodiments, the vertical surface 1182 can be an open wall where the head-mounted system 1160 determines the optimal size of the frame of the vertical surface 1182 appropriate for the user 1108 to view the main video 1220. This determination of the size of the vertical surface 1182 may be based at least in part on the distance of the user 1108 from the main wall 1180, the size and dimensions of the main video 1220, the quality of the main video 1220, the amount of uncoated wall space, and / or the user's posture when looking at the main wall 1180. For example, if the quality of the main video 1220 is high-definition, the size of the vertical surface 1182 may be larger because the quality of the main video 1220 will not be adversely affected by the vertical surface 1182. However, if the video quality of the main video 1220 is of poor quality, having a large vertical surface 1182 can significantly degrade the video quality. In that case, the methods and systems of the present disclosure may resize / redefine the display method of the content displayed within the vertical surface 1182 to be smaller in order to minimize the poor video quality from pixelation.

[0211] The vertical surface 1186 is a vertical surface on an adjacent wall (e.g., side wall 1184) within the user's physical environment 1105, like the vertical surface 1182. In some embodiments, based on the orientation of the user 1108, the side wall 1184 and the vertical surface 1186 may appear to be surfaces tilted on an upward slope. A surface tilted on an upward slope may be an orientation of a certain type of surface in addition to vertical and horizontal surfaces. The proposed video 1240 from a YOUTUBE (registered trademark) web page is installed on the vertical surface 1186 on the side wall 1184, enabling the user 1108 to view the proposed video simply by moving their head slightly to the right in this example.

[0212] The virtual Rolodex 1190 is a virtual object created by the head-mounted system 1160 and displayed to the user 1108. The virtual Rolodex 1190 may have the ability for the user 1108 to cycle bidirectionally through a set of virtual pages. The virtual Rolodex 1190 may contain an entire web page or may contain individual articles or videos or audio. As shown in this example, the virtual Rolodex 1190 may contain a portion of the content from the secondary tab 1250, or in some embodiments, the virtual Rolodex 1190 may contain the entire page of the secondary tab 1250. The user 1108 may cycle bidirectionally through the content within the virtual Rolodex 1190 simply by focusing on a particular tab within the virtual Rolodex 1190, and one or more sensors (e.g., sensor 1162) within the head-mounted system 1160 will detect the user 1108's eye focus and cycle through the tabs within the virtual Rolodex 1190 as appropriate to obtain relevant information for the user 1108. In some embodiments, the user 1108 may select relevant information from the virtual Rolodex 1190 and instruct the head-mounted system 1160 to display it on either a surrounding surface where the relevant information is available or on yet another virtual object such as a virtual display (not shown) that approaches the user 1108.

[0213] Similar to the virtual Rolodex 1190, the multi-stack virtual object 1194 may contain content that encompasses the entire content from one or more tabs, various web pages, or specific content from tabs that the user 1108 has bookmarked, saved for future viewing, or has open (i.e., inactive) tabs. The multi-stack virtual object 1194 also resembles a real-world stack of newspapers. Each stack within the multi-stack virtual object 1194 may be associated with a specific newspaper article, page, magazine issue, recipe, etc. One of ordinary skill in the art can understand that there may be multiple types of virtual objects that serve the same purpose of providing a surface for installing content elements or content from a content source.

[0214] One of ordinary skill in the art can understand that the content accessed or displayed by the web browser 1110 may be more than just a simple web page. In some embodiments, the content may be photos from a photo album, videos from a movie, TV shows, YOUTUBE (registered trademark) videos, two-way forms, etc. In still other embodiments, the content may be an e-book or any electronic means for displaying a book. Finally, in other embodiments, the content may be other types of content that have not yet been described, since generally, the content is presented in the way information is currently presented. If an electronic device can consume the content, the content can be used by the head-mounted system 1160 to disassemble the content and display it within a 3D setting (e.g., AR).

[0215] In some embodiments, the step of matching / mapping the accessed content may include extracting the content (e.g., from a browser) and posting it on a surface (such that the content is no longer within the browser and is only on the surface), and in some embodiments, the step of matching / mapping may include copying the content (e.g., from a browser) and posting it on a surface (such that the content is on both the browser and the surface).

[0216] The step of disassembling content is a technical problem that exists within the area of Internet and computer-related technologies. Digital content such as web pages is constructed using a certain type of programming language such as HTML, which instructs computer processors and technical components where and how to display elements within the web page on a screen for the user. As discussed above, web designers typically work within the confines of a 2D canvas (e.g., a screen) and place and display elements (e.g., content) within the 2D canvas. HTML tags are used to determine how an HTML document or a portion within the HTML document is formatted. In some embodiments, the (extracted or copied) content can maintain HTML tag references, and in some embodiments, the HTML tag references may be redefined.

[0217] Referring briefly to FIG. 4 with respect to this embodiment, the step of receiving content at 410 may involve searching for digital content with the use of the head-mounted system 1160. The step of receiving content at 410 may also include accessing digital content on a server (e.g., the storage device 1130) connected to the network 1120. The step of receiving content at 410 may include browsing the Internet with respect to a web page of interest to the user 1108. In some embodiments, the step of receiving content at 410 may include a voice activation command provided by the user 1108 to search for content on the Internet. For example, the user 1108 may interact with a device (e.g., the head-mounted system 1160), the user 1108 may issue a command to search for a video, and then request the device to search for a specific video on the Internet by stating the name of the video and a brief description of the video. The device may then search the Internet, pull the video onto a 2D browser, and as the video is displayed on the 2D browser of the device, enable the user 1108 to view the video. The user 1108 may then confirm that the video is the video that the user 1108 would desire to view within a spatial 3D environment.

[0218] Once the content is received, the method identifies content elements within the content at 420 and inventories the content elements within the content for display to user 1108. Content elements within the content may include, for example, videos, articles, and newsletters posted on a web page, comments and posts on social media websites, blog posts, photos posted on various websites, audiobooks, and the like. These elements within the content (e.g., a web page) may be distinguishable by HTML tags within a script related to the content and may further comprise HTML tags or HTML-like tags having attributes provided by a content designer to define where a particular element is placed and, in some cases, the time and manner in which the element is to be displayed. In some embodiments, the methods and systems of the present disclosure utilize these HTML tags and attributes as hints and suggestions provided by the content designer to assist in the matching / mapping process at 440 and will determine where and how the elements are to be displayed within the 3D setting. For example, the following is exemplary HTML web page code provided by a content designer (e.g., a web page developer).

[0219] Exemplary HTML web page code provided by the content designer

Chem.

Chem.

[0220] In some embodiments, for example, <ml-container>Tags such as these may enable a content designer to provide specific preference attributes (e.g., hints) regarding where and how a content element is to be displayed within an environment (e.g., a 3D space environment) such that a parser (e.g., parser 115) can interpret the attributes defined within the tag and determine where and how the content element is to be displayed within the 3D space environment. The specific preference attributes may include one or more attributes for defining display preferences regarding the content element. The attributes may include any of the attributes described above.

[0221] One of ordinary skill in the art will appreciate that these suggestions, hints, and / or attributes defined by the content designer may indicate, for example, similar properties for displaying the content element within the 3D space environment. <ml-container>It can be understood that it may be defined within tags such as . In addition, those skilled in the art can also understand that the content designer may define attributes in any combination. The embodiments disclosed herein may interpret the desired display results by using an analyzer (e.g., analyzer 115), or other similar techniques for analyzing the content of a web page and determining the method and location for best displaying content elements within the content.

[0222] Regarding this embodiment, briefly referring to FIG. 5, the step of identifying elements within the content at 510 may be similar to the step of identifying elements within the content at 420 in FIG. 4. The method proceeds to the next step of identifying attributes from tags regarding the location of the content at 520. As discussed above, when designing and configuring a web page, the content designer may associate HTML tags for defining the content elements within the web page and the location and method for displaying each content element. These HTML tags may also include attributes regarding the location of the content elements on a specific portion of the web page. What the head-mounted system 1160 will detect other components of the system, cooperate with them, and use as input regarding the location where a specific element can be displayed are these HTML tags and their attributes. In some embodiments, for example, <ml-container>Tags such as etc. may include attributes defined by the content designer to propose display preference attributes of content elements within a 3D spatial environment, and the tags are associated with the content elements.

[0223] The step of extracting hints or tags from each element is performed at 530. The hints or tags are typically formatting hints or formatting tags provided by the content designer of the web page. As discussed above, the content designer may provide instructions or hints in the form of HTML tags, as shown in, for example, "exemplary HTML web page code provided by the web page developer", and instruct the web browser 1110 to display the content element in a specific part of the page or screen. In some embodiments, the content designer may use additional HTML tag attributes to define additional formatting rules. For example, if the user has a reduced sensitivity regarding a specific color (e.g., red), instead of displaying red, another color may be used, or if a video having a preference for being displayed on a vertical surface cannot be displayed on a vertical surface, alternatively, the video may be displayed on another (physical) surface, or a virtual surface may be created and the video may be displayed on the virtual surface. The following is an exemplary HTML page parser implemented within the browser to parse through the HTML page and extract hints / tags from each element within the HTML page.

[0224] Exemplary HTML page parser implemented within the browser

Chem.

Chem.

Chem.

[0225] The step of looking up / searching for alternative display forms for the content elements is performed at 540. Certain formatting rules may be defined for content elements displayed on a particular viewing device. For example, certain formatting rules may be defined for images on a web page. The system may access the alternative display forms. For example, if web browser 1110 is capable of displaying a 3D version of an image (or more generally, a 3D asset or 3D media), the web page designer may install additional tags or define certain attributes of a particular tag to enable web browser 1110 to recognize that the image may have an alternative version of the image (e.g., a 3D version of the image). Web browser 1110 may then access the alternative version of the image (e.g., a 3D version of the image) for display on a 3D-capable browser.

[0226] In some embodiments, the 3D image within the web page may not be extractable from or copied from the web page for display on the surface within the 3D environment. In these embodiments, the 3D image may be displayed within the user's 3D environment such that the 3D image appears to rotate, shine, etc., and the user can interact with the 3D image, but can only interact within the web page that contains the 3D image. Since the 3D image is not extracted or copied from the web page in these embodiments, the display of the 3D image is within the web page. In this case, the entire web page is extracted and displayed within the user's 3D environment. For example, some content elements within the web page, such as 3D images, may not be extracted or copied from the web page, but can appear in 3D relative to the rest of the web page and be interactable within the web page.

[0227] In some embodiments, the 3D image within the web page may be copied from the web page, but may not be extracted. In these embodiments, the 3D image may be displayed within the user's 3D environment such that the 3D image appears to rotate, shine, etc., and the user can interact with the 3D image not only within the web page that contains the 3D image, but also within the 3D environment outside the web page that contains a copy of the 3D image. The web page appears the same as the 3D image, and outside the web page, there is a copy of the 3D image.

[0228] In some embodiments, 3D images within a web page can be extracted from the web page. In these embodiments, the 3D images may be displayed within the user's 3D environment such that the 3D images appear to rotate, shine, etc., and the user can interact with the 3D images as they are extracted from the web page, but can only interact outside of the web page. Since the 3D images are extracted from the web page, the 3D images are displayed in the 3D environment only, without the web page. In these embodiments, the web page may be reconfigured after the 3D images are extracted from the web page. For example, a version of the web page that includes a blank section within the web page where the 3D image existed prior to being extracted may be presented to the user.

[0229] The foregoing embodiments and examples are described with respect to 3D images within a web page, but one of ordinary skill in the art can understand that the description can be equally applicable to any content element.

[0230] The step of storing the identified content element is performed at 550. The method is used in the synthesis process 140 and may store the identified element in a non-transitory storage medium to match the content element to the surface. The non-transitory storage medium may include a data storage device such as the storage device 1130 or the local storage device 1140. The content element may be stored in a specific table such as the table disclosed in FIG. 14A described below. In some embodiments, the content element may be stored within a hierarchical structure represented as a tree structure as disclosed in FIG. 14B described below, for example. In some embodiments, the content element may be stored in a transitory storage medium.

[0231] Figures 14A-14B show examples of different structures for storing content elements decomposed from content according to some embodiments. In Figure 14A, the element table 1400 is an exemplary table that can store in a database the results of the step of identifying content elements within the content at 510 of Figure 5. The element table 1400 includes, for example, an element identifier (ID) 1410, a preference attribute indicator 1420 for the content element (e.g., priority attribute, orientation attribute, position type attribute, content type attribute, surface type attribute, and equivalents, or some combination thereof), a parent element ID 1430 when a particular content element is included within a parent content element, a child content element ID 1440 when the content element can contain child content elements, and a multiple entity indicator 1450 for indicating whether the content element contains multiple entities, which can ensure compatibility with displaying multiple versions of the content element on the surface or virtual object used to display the content element. The parent content element is a content element / object within the content that can contain sub-content elements (e.g., child content elements). For example, an element ID having a value of 1220 (e.g., main video 1220) has a parent element ID value of 1260 (e.g., active tab 1260), which indicates that the main video 1220 is a child content element of the active tab 1260. Or, in other words, the main video 1220 is included within the active tab 1260. Continuing with the same example, the main video 1220 has a child element ID 1230 (e.g., user comment 1230), which indicates that the user comment 1230 is associated with the main video 1220. One skilled in the art can understand that the element table 1400 may be a table within a relational database or any type of database. Additionally, the element table 1400 may be an array within a computer memory (e.g., cache) that contains the results of the step of identifying content elements within the content at 510 of Figure 5.

[0232] Each row 1460 within the element table 1400 corresponds to a content element from within a web page. The element ID 1410 is a column that contains a unique identifier for each content element (e.g., the element ID). In some embodiments, the uniqueness of a content element may be defined as a combination of the element ID 1410 column and another column within the table (e.g., the preference attribute 1420 column if there is more than one preference attribute identified by the content designer). The preference attribute 1420 is a column whose value can be determined based on tags and attributes that are at least partially defined therein by the content designer and are identified by the system and method as disclosed in the step of extracting hints or tags at 530 of FIG. 5 from each content element. In other embodiments, the preference attribute 1420 column may be determined at least partially based on predetermined rules that may define where a certain type of content element should be displayed within the environment. These predetermined rules may provide the system and method with suggestions for determining the best location to place the content element within the environment.

[0233] The parent element ID 1430 is a column that contains the element ID of the parent content element in which the current in-line specific content element is displayed or to which it is related. The specific content element may be built-in, placed within another content element on the page, or related to another content element on the web page. For example, in this embodiment, the first entry in the element ID 1410 column stores the value of the element ID 1220 corresponding to the main video 1220 in FIG. 12. The value in the preference attribute 1420 column corresponding to the main video 1220 is determined based on tags and / or attributes, and as shown, it means that this content element should be placed in the "main" location of the user's physical environment 1105. Depending on the current location of the user 1108, the main location may be the wall in the living room where the user 1108 is currently looking or the hood above the stove in the kitchen, or if it exists in an open space, it may be a virtual object projected in front of the user 1108's line of sight where the main video 1220 can be projected. Further information regarding how the content element is displayed to the user 1108 will be disclosed elsewhere in the mode for carrying out the invention. Continuing with this example, the parent element ID 1430 column stores the value of the element ID 1260 corresponding to the active tab 1260 in FIG. 12. Therefore, the main video 1220 is a child of the active tab 1260.

[0234] The child element ID 1440 is a column that contains the element ID of the child content element in which the current in-line specific content element is displayed or to which it is related. A specific content element within the web page may be built-in, placed within another content element, or related to another content element. Continuing with this example, the child element ID 1440 column stores the value of the element ID 1230 corresponding to the user comment 1230 in FIG. 12.

[0235] The multiple entity indicator 1450 may ensure the need to be compatible with displaying multiple versions of content elements on the surface or virtual object used to display the element. It is a column indicating whether the content element contains multiple entities (for example, the content element may be user comment 1230, and there may be more than one comment available regarding the main video 1220). Continuing with this example, the multiple entity indicator 1450 column stores the value of "N", indicating that the main video 1220 does not have or does not correspond to multiple main videos within the active tab 1260 (for example, multiple versions of the main video 1220 "do not exist").

[0236] Continuing with this example, the second entry in the element ID 1410 column stores the value of the element ID 1230 corresponding to the user comment 1230 in FIG. 12. The value in the preference attribute 1420 column corresponding to the user comment 1230 indicates a "horizontal" preference, indicating that the user comment 1230 should be placed on a horizontal surface at any location within the user's physical environment 1105. As discussed above, the horizontal surface will be determined based on the available horizontal surfaces within the user's physical environment 1105. In some embodiments, the user's physical environment 1105 may not have a horizontal surface, in which case the system and method of the present disclosure may identify / create a virtual object with a horizontal surface and display the user comment 1230. Continuing with this example, the parent element ID 1430 column stores the value element ID 1220 corresponding to the main video 1220 in FIG. 12, and the multiple entity indicator 1450 column stores the value of "Y", indicating that the user comment 1230 may contain more than one value (for example, more than one user comment).

[0237] The remaining rows in the element table 1400 contain information about the remaining content elements that the user 1108 is interested in. One skilled in the art can understand that once this analysis is performed on the content, remembering the result of the step of identifying the content elements within the content at 510 can be reserved by the present system and method for future analysis of the content when another user is interested in the same content, thus improving the function of the computer itself. The present system and method for decomposing this specific content may be avoided because it has already been completed previously.

[0238] In some embodiments, the element table 1400 may be stored within the storage device 1130. In other embodiments, the element table 1400 may be stored within the local storage device 1140 for quick access to the most recently viewed content or for a potential revisit to the most recently viewed content. In still other embodiments, the element table 1400 may be stored in both the storage device 1130 located remotely from the user 1108 and the local storage device 1140 located locally to the user 1108.

[0239] In FIG. 14B, the tree structure 1405 is an exemplary logical structure that can be used to store in a database the results of steps for identifying elements within the content at 510 in FIG. 5. Storing content elements within a tree structure can be advantageous when various content has a hierarchical relationship with each other. The tree structure 1405 includes a parent node - web page main tab node 1415, a first child node - main video node 1425, and a second child node - proposed video node 1445. The first child node - main video node 1425 includes a child node - user comment node 1435. The user comment node 1435 is a grandchild of the web page main tab node 1415. As an example, referring to FIG. 12, the web page main tab node 1415 may be the web page main tab 1260, the main video node 1425 may be the main video 1220, the user comment node 1435 may be the user comment 1230, and the proposed video node 1445 may be the proposed video 1240. Here, the tree structuring of content elements indicates the hierarchical relationship between various content elements. It can be advantageous to organize and store content elements within a tree structure type of logical structure. For example, when the main video 1220 is displayed on a particular surface, it may be useful for the system to understand that the user comment 1230 is child content of the main video 1220, and it may be beneficial to display the user comment 1230 relatively close to the main video 1220 and / or on a surface in the vicinity of the main video 1220 so that the user can easily see and understand the relationship between the user comment 1230 and the main video 1220. In some embodiments, it may be beneficial that when the user decides to hide or close the main video 1220, it is possible to hide or close the user comment 1230. In some embodiments, it may be beneficial that when the user decides to move the main video 1220 to a different surface, it is possible to move the user comment 1230 to another surface.When the user moves the main video 1220, the system may move the user comment 1230 by moving both the parent node - main video node 1425 and the child node - user comment node 1435 simultaneously.

[0240] Returning to FIG. 4, the method continues, at 430, with the step of determining the surface. The user 1108 views the user's physical environment 1105 through the head - mounted system 1160, and the head - mounted system 1160 may be enabled to capture and identify surrounding surfaces such as walls, tables, paintings, window frames, stoves, refrigerators, TVs, etc. The head - mounted system 1160 perceives real objects within the user's physical environment 1105 for sensors and cameras on the head - mounted system 1160 or using any other type of similar device. In some embodiments, the head - mounted system 1160 matches real objects observed within the user's physical environment 1105 with virtual objects stored in the memory device 1130 or the local memory device 1140 and identifies the surfaces available with such virtual objects. A real object is an object identified within the user's physical environment 1105. A virtual object is an object that does not physically exist within the user's physical environment but can be presented to the user as if the virtual object exists within the user's physical environment. For example, the head - mounted system 1160 may detect an image of a table within the user's physical environment 1105. The table image may be reduced to a 3D point cloud object for comparison and matching in the memory device 1130 or the local memory device 1140. If a match of the real object (e.g., the table) and the 3D point cloud object is detected, the system and method will identify the table as having a horizontal surface since the 3D point cloud object representing the table is defined to have a horizontal surface.

[0241] In some embodiments, the virtual object may be an extracted object. The extracted object is identified within the user's physical environment 1105, but additional processing and associations that could not be performed on the physical object itself (e.g., changing the color of the physical object, highlighting certain features of the physical object, etc.) can be performed on the extracted object so that it is presented to the user as a virtual object within the location of the physical object. Additionally, the extracted object may be a virtual object that is extracted from content (e.g., a web page from a browser) and presented to the user 1108. For example, the user 1108 may select an object such as a bench displayed on a web page for presentation within the user's physical environment 1105. The system may recognize the selected object (e.g., the bench) and present the extracted object (e.g., the bench) to the user 1108 as if the extracted object (e.g., the bench) physically exists within the user's physical environment 1105. Additionally, the virtual object may also include an object having a surface for presenting certain content to the user (e.g., a transparent display screen that approaches the user for viewing certain content), which may be an ideal display surface for presenting certain content from the perspective of presenting the content, although it does not physically exist within the user's physical environment 1105.

[0242] Referring briefly to FIG. 6, the method begins at 610 with the step of determining a surface. The method proceeds to the next step of determining the user's pose at 620, which may include the step of determining a head pose vector. The step of determining the user's pose at 620 is an important step for identifying the user's current surroundings because the user's pose will provide a viewpoint for the user 1108 in relation to the objects within the user's physical environment 1105. For example, referring back to FIG. 11, the user 1108 is observing the user's physical environment 1105 using a head-mounted system 1160. The step of determining the user's pose (i.e., the head pose vector and / or origin position information relative to the world) at 620 will help the head-mounted system 1160 to understand, for example, (1) the height at which the user 1108 is present relative to the ground, (2) the angle at which the user 1108 needs to rotate their head in order to move around the room and capture its image, and (3) the distances between the user 1108 and the table 1188, the main wall 1180, and the side wall 1184. Additionally, the pose of the user 1108 is also useful in determining the angle of the head-mounted system 1160 when observing other surfaces within the user's physical environment 1105, along with the vertical surfaces 1182 and 186.

[0243] At 630, the method determines the attributes of the surfaces. Each surface within the user's physical environment 1105 is tagged and categorized with corresponding attributes. In some embodiments, each surface within the user's physical environment 1105 is also tagged and categorized with corresponding dimension and / or orientation attributes. This information will be useful in matching content elements to the surfaces, at least in part, based on the dimension attributes of the surfaces, the orientation attributes of the surfaces, the distance the user 1108 is away from a particular surface, and the type of information that needs to be displayed with respect to the content elements. For example, a video may be shown away from a blog or article, contain rich information that may not be viewable by the user if the text size of the article is too small when displayed on a distant wall with small dimensions. In some embodiments, the sensor 162 from FIG. 1B provides the raw data to the CVPU 164 for processing, and the CVPU 164 provides the processed data to the perception framework 166 to prepare data for the environment analyzer 168. The environment analyzer 168 analyzes the environmental data from the perception framework 166 and determines the surfaces within the environment.

[0244] At 640, the method stores an inventory of the surfaces in a non-transitory storage medium for use by a composition process / matching / mapping routine to match / map the extracted elements to a particular surface. The non-transitory storage medium may include a data storage device such as the storage device 1130 or the local storage device 1140. The identified surfaces may be stored in a particular table such as the table disclosed in FIG. 15 described below. In some embodiments, the identified surfaces may be stored in a transitory storage medium.

[0245] FIG. 15 shows an example of a table that stores an inventory of surfaces identified from a user's local environment according to some embodiments. The surface table 1500 is an exemplary table that may store the results of the step of identifying the surrounding surfaces and the attribute process in a database. The surface table 1500 includes, for example, information about surfaces in the user's physical environment 1105 having a data column including a surface ID 1510, a width 1520, a height 1530, an orientation 1540, a real or virtual indicator 1550, a multiplicity 1560, a position 1570, and a dot product relative surface orientation for the user 1580. The surface table 1500 may have additional columns representing other attributes of each surface. One of ordinary skill in the art may understand that the surface table 1500 may be a table in a relational database or any type of database. Additionally, the surface table 1500 may be an array in a computer memory (e.g., a cache) that stores the results of the step of determining the surface at 430 in FIG. 4.

[0246] Each row of the rows 1590 in the surface table 1500 may correspond to a surface from the user's physical environment 1105 or a virtual surface that may be presented to the user 1108 within the user's physical environment 1105. The surface ID 1510 is a column that contains a unique identifier and uniquely identifies a particular surface (e.g., the surface ID). The dimensions of a particular surface are stored in the width 1520 and height 1530 columns.

[0247] Orientation 1540 is a column that indicates the orientation of the surface with respect to the user 1108 (e.g., vertical, horizontal, etc.). Real / Virtual 1550 is a column that indicates whether a particular surface is located on a real surface / object within the user's physical environment 1105 such that it is perceived by the user 1108 using the head-mounted system 1160, or whether a particular surface is located on a virtual surface / object that would be generated by the head-mounted system 1160 and displayed within the user's physical environment 1105. The head-mounted system 1160 may need to generate virtual surfaces / objects for situations where the user's physical environment 1105 does not contain sufficient surfaces, cannot contain sufficiently appropriate surfaces based on matching score analysis, or where the head-mounted system 1160 cannot detect sufficient surfaces to display the amount of content that the user 1108 desires to display. In these embodiments, the head-mounted system 1160 may search a database of existing virtual object data that may have appropriate surface dimensions for displaying a certain type of element identified for display. The database may be from the storage device 1130 or the local storage device 1140. In some embodiments, the virtual surface is created substantially in front of the user or is offset from the front vector of the head-mounted system 1160 so as not to occlude the user's and / or the device's primary field of view of the real world.

[0248] Multiplicity 1560 is a column that indicates whether a surface / object is compatible with displaying multiple versions of an element (e.g., the element may be the secondary tab 1250 of FIG. 12, and there may be more than one secondary (i.e., non-active) tab (e.g., one web page per tab) for a particular web browser 1110. When the multiplicity 1560 column has a value of "multiplicity", as in the case of the fourth entry of the surface ID column that stores the value of 1190 corresponding to the virtual Rolodex 1190 of FIG. 12 and the fifth entry of the surface ID column that stores the value of 1194 corresponding to the multi-stack virtual object 1194 of FIG. 12, the system and method will understand that when there are elements that can have multiple versions, such as in the case of non-active tabs, these are of a type of surface that can accommodate multiple versions.

[0249] Position 1570 is a column that indicates the position of a physical surface relative to a frame of reference or reference point. The position of the physical surface may be pre-determined to be the center of the surface, as shown in the column header of position 1570 in FIG. 15. In other embodiments, the position may be pre-determined to be another reference point of the surface (e.g., the front, back, top, or bottom of the surface). The position information can be represented as a vector from the center of the physical surface relative to a frame of reference or reference point and / or as position information. There may be several ways to represent the position within the surface table 1500. For example, the value of the position for surface ID 1194 within the surface table 1500 is represented in the summary to illustrate vector information and reference frame information (e.g., subscript "frame"). x, y, z are 3D coordinates within each spatial dimension, and frame indicates the reference frame to which the 3D coordinates are relative.

[0250] For example, surface ID 1186 indicates that the position of the center of surface 1186 is (1.3, 2.3, 1.3) relative to the real-world origin. As another example, surface ID 1192 indicates that the position of the center of surface 1192 is (x, y, z) relative to the user reference frame, and surface ID 1190 indicates that the position of the center of surface 1190 is (x, y, z) relative to another surface 1182. The reference frame is important to clarify the currently used reference frame. In the case of the real-world origin as the reference frame, this is generally a static reference frame. However, in other embodiments where the reference frame is the user reference frame, the user may move the reference frame, in which case the plane (or vector information) can move and change with the user as the user moves, and the user reference frame is used as the reference frame. In some embodiments, the reference frame for each surface may be the same (e.g., the user reference frame). In other embodiments, the reference frame for the surfaces stored in surface table 1500 may vary depending on the surface (e.g., user reference frame, world reference frame, another surface or object in the room, etc.).

[0251] In this embodiment, the values stored within the surface table 1500 include physical surfaces (e.g., vertical surfaces 1182 and 1186 and horizontal surface 1192) and virtual surfaces (e.g., virtual Rolodex 1190 and multi-stack virtual object 1194) identified within the user's physical environment 1105 of FIG. 12. For example, in this embodiment, the first entry in the surface ID 1510 column stores the value of the surface ID 1182 corresponding to the vertical surface 1182 of FIG. 12. The width value within the width 1520 column and the height value within the height 1530 column correspond to the width and height of the vertical surface 1182, indicating that the vertical surface 1182 has dimensions of 48 inches (W) × 36 inches (H), respectively. Similarly, the orientation value within the orientation 1540 column indicates that the vertical surface 1182 has an "upright" orientation. Additionally, the real / virtual value within the real / virtual 1550 column indicates that the vertical surface 1182 is an "R" (e.g., real) surface. The multiplicity value within the multiplicity 1560 column indicates that the vertical surface 1182 is "single" (e.g., can hold only a single piece of content). Finally, the position within the 1570 column indicates the position of the vertical surface 1182 relative to the user 1108 using the vector information of (2.5, 2.3, 1.2) user to indicate the position of the vertical surface 1182 relative to the user 1108.

[0252] The remaining rows within the surface table 1500 contain information regarding the remaining surfaces within the user's physical environment 1105. One skilled in the art can understand that storing the results of the step of determining the surfaces at 430 of FIG. 4 can be retained by the head-mounted system 1160 for future analysis of the user's surrounding surfaces when another user or the same user 1108 exists within the same physical environment 1105 but is interested in different content, as the computer's own functionality can be improved. The processing steps for determining the surfaces at 430 may be avoided as these processing steps have already been completed previously. The only difference, at least in part, may include the step of identifying available additional or different virtual objects based on the element table 1400 that identifies elements with different content.

[0253] In some embodiments, the surface table 1500 is stored within the memory device 1130. In other embodiments, the surface table 1500 is stored within the local memory device 1140 of the user 1108 for quick access to the most recently viewed content or for potential revisit to the most recently viewed content. In still other embodiments, the surface table 1500 may be stored in both the memory device 1130 located remotely from the user 1108 and the local memory device 1140 located locally to the user 1108.

[0254] Returning to FIG. 4, in some embodiments, the method continues, at 440, to match content elements to a surface using the combination of content elements identified at 420 in identifying content elements within content and surfaces determined at 430, and in some embodiments, using virtual objects as additional surfaces. The step of matching content elements to a surface may involve multiple factors, some of which may include analyzing hints provided by the content designer via HTML tag elements defined by the content designer, such as by using an exemplary HTML page parser as discussed above. Other factors may include selecting from a set of predefined rules of how and where to match / map to certain content as provided by an AR browser, an AR interface, and / or a cloud storage device.

[0255] Referring briefly to FIG. 7A, this depicts a flow diagram illustrating a method for matching content elements to a surface, according to some embodiments. At 710, the method determines whether the identified content element contains a hint provided by a content designer. The content designer may provide a hint regarding the location for optimally displaying the content element. For example, the main video 1220 of FIG. 12 may be a video displayed on a web page within the active tab 1260. The content designer may provide a hint indicating that the main video 1220 is optimally displayed on a flat vertical surface within the direct view of the user 1108.

[0256] In some embodiments, the 3D preview for a web link can be represented as a set of new HTML tags and properties associated with the web page. FIG. 16 shows an exemplary 3D preview for a web link according to some embodiments. A content designer may define a web link having an associated 3D preview for rendering using the new HTML properties. Optionally, the content designer / web developer may define a 3D model for use in rendering the 3D web preview. If the content designer / web developer defines a 3D model for use in rendering the web preview, the web content image may be used as a texture for the 3D model. The web page may be received. If there are preview properties defined for a link tag, a first-level web page may be read, and based on the preview properties, a 3D preview may be generated and loaded onto a 3D model (e.g., sphere 1610) defined by the content designer or a default 3D model. Although the 3D preview is described with respect to web links, the 3D preview may be used for other content types. One of ordinary skill in the art will understand that there are many other ways in which a content designer can provide hints as to where a particular content element should be placed within a 3D environment other than those disclosed herein, and these are some examples of different ways in which a content designer can provide hints for displaying some or all of the content elements of a web page.

[0257] In another embodiment, the tag specification (e.g., HTML tag specification) may include the creation of new tags (e.g., HTML tags) or a similar markup language for providing hints within, for example, an exemplary web page provided by the content designer discussed above. If the tag specification includes these types of additional tags, some embodiments of the method and system will utilize these tags and further provide a matching / mapping of the identified content elements to the identified surfaces.

[0258] For example, a set of web components may be exposed as new HTML tags for use by content designers / web developers to create elements of a web page, which may appear as 3D volumes protruding from a 2D web page or 3D volumes etched within a 2D web page. FIG. 17 shows an example of a web page having a 3D volume etched within a web page (e.g., 1710). These 3D volumes may include web controls (e.g., buttons, handles, joysticks), which are placed on the web page and will enable a user to manipulate the web controls and thereby manipulate the content displayed within the web page. Those skilled in the art will appreciate that there are many other languages other than HTML that can be modified or adopted to further provide hints regarding how content elements should best be displayed within a 3D environment, and that the new HTML tag standard is simply one way to achieve such a goal.

[0259] At 720, the method determines whether to use hints provided by the content designer or a set of predefined rules for matching / mapping content elements to surfaces. At 730, if it is determined that using the hints provided by the content designer is the way to proceed, the system and method analyze the hints and search for a logical structure that includes identified surrounding surfaces that can be used, at least in part, to display specific content elements based on the hints (e.g., querying the surface table 1500 of FIG. 15).

[0260] At 740, the system and method activate a best-fit algorithm and select a best-fit surface for a particular content element based on the hints provided. The best-fit algorithm, for example, may attempt to obtain hints regarding a particular content element and identify a surface that is frontal and central to user 1108 within the environment. For example, the main video 1220 of FIG. 12 has the preference value of "main" in the preference attribute 1420 column of the element table 1400 of FIG. 14A within the active tab 1260, and the vertical surface 1182 is a surface within the direct view of user 1108 and has an optimal size dimension for displaying the main video 1220, and thus is matched / mapped to the vertical surface 1182.

[0261] At 750, the system and method store a matching result having a matching of a content element and a surface. The table may be stored in a non-transitory storage medium for use by a display algorithm to display the content element on its respective matched / mapped surface. The non-transitory storage medium may include a data storage device such as the storage device 1130 or the local storage device 1140. The matching result may be stored in a particular table such as the table disclosed in FIG. 18 below.

[0262] FIG. 18 shows an example of a table for storing the matching of content elements with surfaces according to some embodiments. The matching / mapping table 1800 is an exemplary table that stores the results of the process of matching content elements with surfaces in a database. The matching / mapping table 1800 includes information about, for example, content elements (e.g., element IDs) and the surfaces (e.g., surface IDs) to which the content elements are matched / mapped. One of ordinary skill in the art can understand that the matching / mapping table 1800 may be a table stored in a relational database or any type of database or storage medium. Additionally, the matching / mapping table 1800 may be an array in a computer memory (e.g., cache) that contains the results of the step of matching content elements with surfaces at 440 in FIG. 4.

[0263] Each row of the matching / mapping table 1800 corresponds to a content element that is matched to one or more surfaces within the user's physical environment 1105 or a virtual surface / object displayed to the user 1108, and the virtual surface / object appears to be a surface / object within the user's physical environment 1105. For example, in this embodiment, the first entry in the element ID column stores the value of the element ID 1220 corresponding to the main video 1220. The surface ID value within the surface ID column corresponding to the main video 1220 is 1182, which corresponds to the vertical surface 1182. Thus, the main video 1220 is matched / mapped to the vertical surface 1182. Similarly, the user comment 1230 is matched / mapped to the horizontal surface 1192, the proposed video 1240 is matched / mapped to the vertical surface 1186, and the secondary tab 1250 is matched / mapped to the virtual Rolodex 1190. The element IDs within the matching / mapping table 1800 may be associated with the element IDs stored in the element table 1400 of FIG. 14A. The surface IDs within the matching / mapping table 1800 may be associated with the surface IDs stored in the surface table 1500 of FIG. 15.

[0264] Returning to FIG. 7A, assume that at 760 it is determined that a particular rule should be used. The method then queries a database containing content element to surface matching / mapping rules to determine the type of surface that should be considered for matching / mapping a content element with respect to a particular content element within a web page. For example, the rule returned for main video 1220 from FIG. 12 may indicate that main video 1220 should be matched / mapped to a vertical surface. Thus, after querying surface table 1500, a plurality of candidate surfaces are revealed (e.g., vertical surfaces 1182 and 1186 and virtual Rolodex 1190). At 770, a pre-defined set of rules activates an optimal fit algorithm to select, from the available candidate surfaces, the surface that is the optimal fit for this main video 1220. Based at least in part on the optimal fit algorithm, of all the candidate surfaces, vertical surface 1182 is determined to be the surface within the direct line of sight of user 1108 and having the optimal dimensions for displaying the video, such that main video 1220 should be matched / mapped to vertical surface 1182. Once the matching / mapping of one or more elements is determined at 750, the method stores, within a non-transitory storage medium, the matching / mapping results for the content elements within the element to surface table matching / mapping as described above.

[0265] Returning to FIG. 4, the method continues, at 450, to render content elements as virtual content on a surface that has been matched. The head-mounted system 1160 may include within the head-mounted system 1160 one or more display devices, such as a mini-projector (not shown) for displaying information. One or more elements are displayed on an individual matched surface as was matched at 440. Using the head-mounted system 1160, user 1108 will view the content on the individual matched / mapped surface. One of ordinary skill in the art will appreciate that the content elements are displayed such that they appear to be physically attached on various surfaces (physical or virtual), but in reality, the content elements are actually projected onto a physical surface as perceived by user 1108, and in the case of virtual objects, the virtual objects are displayed such that they appear to be attached on an individual surface of the virtual object. One of ordinary skill in the art will appreciate that when user 1108 turns their head or looks up and down, the display device within the head-mounted system 1160 continues to keep the content elements attached to their individual surfaces and further provides user 1108 with the perception that the content is attached to the matched / mapped surface. In other embodiments, user 1108 may change the content of user's physical environment 1105 by movements made by user 1108's head, hand, eye, or voice.

[0266] Application FIG. 19 shows an example of an environment 1900 that includes content elements matched to a surface, according to some embodiments.

[0267] Referring briefly to FIG. 4 with respect to this embodiment, the parser 115 receives the content 110 from the application at 410. The parser 115 identifies content elements within the content 110 at 420. In this embodiment, the parser 115 identifies a video panel 1902, a highlight panel 1904, a replay 1906, a graphic statistic 1908, a text statistic 1910, and a social media news feed 1912.

[0268] The environment parser 168 determines the surfaces within the environment at 430. In this embodiment, the environment parser 168 determines a first vertical surface 1932, a second vertical surface 1934, a first ottoman top 1936, a second ottoman top 1938, and a second ottoman front 1940. The environment parser 168 may determine additional surfaces within the environment. However, in this embodiment, the additional surfaces are not labeled. In some embodiments, the environment parser 168 continuously determines the surfaces within the environment at 430. In some embodiments, the environment parser 168 determines the surfaces within the environment at 430 as the parser 115 receives the content 110 at 410 and / or identifies content elements within the content 110 at 420.

[0269] The matching module 142 matches the content elements to the surfaces based on the attributes of the content elements and the attributes of the surfaces at 440. In this embodiment, the matching module 142 matches the video panel 1902 to the first vertical surface 1932, the highlight panel 1904 to the second vertical surface 1934, the replay 1906 to the first ottoman top 1936, the graphic statistic 1908 to the second ottoman top 1938, and the text statistic 1910 to the second ottoman front 1940.

[0270] The optional virtual object creation module 144 may create a virtual object for displaying a content element. During the matching process of the matching module 142, it may be determined that the virtual surface may be an optional surface for displaying a certain content element. In this embodiment, the optional virtual object creation module 144 creates a virtual surface 1942. The social media news feed 1912 is matched to the virtual surface 1942. The rendering module 146 renders the content element on its matched surface 450. The resulting FIG. 19 illustrates what the user of the head-mounted display device launching the application will see after the rendering module 146 renders the content element on its matched surface 450.

[0271] Dynamic environment In some embodiments, the environment 1900 is dynamic. That is, the environment itself is changing, objects are moving in and out of the user's and / or device's field of view, creating new surfaces, or the user is moving to a new environment and receiving content elements such that a previously matched surface is no longer eligible under the results of the previous composition process 140. For example, as in FIG. 19, while watching a basketball game within the environment 1900, the user may walk into the kitchen.

[0272] Figures 20A - 20E depict the change of the environment as a function of user movement, and those skilled in the art will understand that the following techniques are also applicable to environments that change around a static user. In FIG. 20A, the user is appreciating the spatialized display of content after the composition process 140 as described throughout this disclosure. FIG. 20B illustrates a larger environment in which the user can immerse and additional surfaces can become available to the user.

[0273] As the user moves from one room to another as depicted in FIG. 20C, it becomes readily apparent that the content initially rendered for display in FIG. 20A no longer meets the matching of the composition process 140. In some embodiments, the sensor 162 prompts the system of a change in the user's environment. The change in the environment may be a change in depth sensor data (the room on the left side of FIG. 20C produces a new virtual mesh structure different from the room on the right side of FIG. 20C where the content was initially rendered for display), a change in head pose data (the IMU produces a motion change exceeding a threshold for the current environment, or a camera on the head-mounted system starts capturing a new object within the field of view of the user and / or the device). In some embodiments, the change in the environment starts a new composition process 140 and finds a new surface for the content elements previously matched and / or currently rendered and displayed. In some embodiments, a change in the environment exceeding a time threshold starts a new composition process 140 and finds a new surface for the content elements previously matched and / or currently rendered and displayed. The time threshold eliminates wasteful computing cycles for minor interruptions to the environmental data (such as simply turning the head, talking to another user, or a short exit from the environment where the user will quickly return).

[0274] In some embodiments, as the user enters room 2002 in FIG. 20D, the composition process 140 matches the active content 2002 that was matched to room 2014 with a new surface. In some embodiments, the active content 2002 is currently active in both rooms 2012 and 2014 but is only displayed in room 2012 (the appearance of the active content 2002 within room 2014 in FIG. 20D depicts that the active content 2002 is still being rendered but not being displayed to the user).

[0275] In this way, the user can walk between rooms 2012 and 2014, but the synthesis process 140 does not need to continuously repeat the matching protocol. In some embodiments, the active content 2002 in room 2014 is set to an idle or sleep state while the user is located in room 2012. Similarly, when the user returns to room 2014, the active content in room 2012 is put into an idle or sleep state. Thus, as the user dynamically changes their environment, they can automatically continue consuming content.

[0276] In some embodiments, the user can pause the active content 2002 in room 2014, enter room 2012, and resume the same content at the same interaction point where it was paused in room 2014. Thus, as the user dynamically changes their environment, they can automatically resume consuming content.

[0277] The idle or sleep state can be characterized in terms of the degree of output that the content element performs. Active content can have the full capacity of the rendered content element, such that the frames of the content are continuously updated with respect to its matched surface, the audio output continues with virtual speakers associated with the location of the matched surface, etc. The idle or sleep state can reduce some of this functionality. In some embodiments, the audio output in the idle or sleep state reduces the volume or enters a mute state. In some embodiments, the rendering cycle is slowed down such that fewer frames are generated. Such a lower frame rate overall conserves computing power, but can introduce a slight wait time when resuming content element consumption when the idle or sleep state returns to the active state (such as when the user returns to the room where the idle or sleep state content element was operating).

[0278] Figure 20E depicts the cessation of rendering content elements in different environments without simply changing to an idle or sleep state. In Figure 20E, the active content within room 2014 has been aborted. In some embodiments, the trigger for the cessation of rendering is changing the content element from one source to another, such as changing the channel of a video stream from a basketball game to a movie. In some embodiments, the active content immediately aborts rendering once sensor 162 detects a new environment and new synthesis process 140 begins.

[0279] In some embodiments, the content rendered and displayed on the first surface at the first location may be paused, for example, at least in part, based on the user's movement from the first location to the second location, and then subsequently resumed on the second surface at the second location. For example, a user viewing content displayed on the first surface at the first location (e.g., a living room) may physically move from the first location to the second location (e.g., a kitchen). The rendering and / or display of the content on the first surface at the first location may be paused in response to a determination that the user has physically moved from the first location to the second location (e.g., based on sensor 162). Once the user moves into the second location, the sensors of the AR system (e.g., sensor 162) may detect that the user has moved into the new environment / location, and the environment analyzer 168 may begin to identify the new surface at the second location and then resume displaying the content on the second surface at the second location. In some embodiments, the content may continue to be rendered on the first surface at the first location while the user moves from the first location to the second location. Once the user has been present in the second location, for example, for a threshold time period (e.g., 30 seconds), the content may stop being rendered on the first surface at the first location and may be rendered on the second surface at the second location. In some embodiments, the content may be rendered on both the first surface at the first location and the second surface at the second location.

[0280] In some embodiments, the step of pausing the rendering and / or display of content on the first surface at the first location may be automatically performed in response to the user physically moving from the first location to the second location. Detection of the user's physical movement may trigger the automatic pause of the content, and the trigger for the user's physical movement may be based at least in part on an inertial measurement unit (IMU) exceeding a threshold, or on a position indication (e.g., GPS) that the user has moved out of or is moving within a predetermined area that may be associated with the first location. Once the second surface is identified, e.g., by the environmental analyzer 168, and matched to the content, the content may automatically resume rendering and / or display on the second surface at the second location. In some embodiments, the content may resume rendering and / or display on the second surface based at least in part on the user's selection of the second surface. In some embodiments, the environmental analyzer 168 may refresh within a particular time frame (e.g., every 10 seconds) to determine whether the surface within the field of view of the user and / or device has changed and / or whether the physical location of the user has changed. If it is determined that the user has moved to a new location (e.g., the user has moved from the first location to the second location), the environmental analyzer 168 may begin identifying a new surface within the second location to resume rendering and / or display of the content on the second surface. In some embodiments, the content rendered and / or displayed on the first surface may not be automatically paused immediately simply because the user may change the field of view (e.g., the user may briefly look at another person at the first location, e.g., to have a conversation). In some embodiments, the rendering and / or display of the content may be automatically paused if the user's changed field of view exceeds a threshold. For example, if the user changes the head pose, and thus the corresponding field of view, over a time period exceeding a threshold, the display of the content may be automatically paused.In some embodiments, the content may automatically pause rendering and / or displaying the content on the first surface at the first location in response to the user leaving the first location, and the content may automatically resume rendering and / or displaying on the first surface at the first location in response to the user physically (re)entering the first location.

[0281] In some embodiments, as the user's and / or the field of view of the user's head-mounted device changes, the content on a particular surface may slowly follow the change in the user's field of view. For example, the content may be within the user's direct field of view. When the user changes the field of view, the content may change position and follow the change in the field of view. In some embodiments, the content may not be immediately displayed on the surface within the direct field of view of the changed field of view. Instead, there may be a slight delay in the change of the content in response to the change in the field of view, and the change in the content location may appear to slowly follow the change in the field of view.

[0282] Figures 20F - 20I illustrate an example in which the content displayed on a specific surface can slowly follow the change in the field of view of the user currently viewing the content. In Figure 20F, user 1108 is located at a seated position on a bench in the room and is viewing a spatial display of content. The seated position has, for example, a first head posture of the user facing towards the main wall 1180 and / or the head - mounted device of the user. As shown in Figure 20F, the spatial display of content is displayed on a first location (e.g., rectangular surface 1182) of the main wall 1180 via the first head posture. Figure 20G illustrates the situation where user 1108 changes the position on the bench from the seated position to the lying position. The lying position has, for example, a second head posture facing towards the side wall 1184 instead of the main wall 1180. The content displayed on the rectangular surface 1182 may continue to be rendered / displayed on the rectangular surface 1182 until a time threshold and / or a head - posture change threshold is met / exceeded. Figure 20H illustrates how the content can slowly follow the user. The movement involves discrete incremental positions little by little to a new location corresponding to the second head posture facing towards the side wall 1185, in contrast to a single update. For example, after a certain time point (e.g., after a certain time threshold) after user 1108 changes from the seated position to the lying position, it appears to be displayed on the first display option / surface 2020. The first display option / surface 2020 may be a virtual display screen / surface within the field of view corresponding to the second head posture because there is no optimal surface available within the direct field of view of user 1108. Figure 20I illustrates how the content can also be displayed on a second display option / surface on the rectangular surface 1186 of the side wall 1184. As disclosed above, in some embodiments, user 1108 may be provided with display options for selecting a display option (e.g., the first display option 2020 or the second rectangular surface 1186) for displaying the content based on the change in the field of view of user 1108 and / or the device.

[0283] In some embodiments, for example, the user may view the content displayed directly in front of the user within the first field of view. The user may turn their head 90 degrees to the left and maintain the second field of view for about 30 seconds. The content displayed directly in front of the user within the first field of view is, for the first time, moved 30 degrees towards the second field of view with respect to the first field of view, and after a certain time threshold (e.g., 5 seconds) has elapsed, it may slowly follow the user and slowly follow the user to the second field of view. The AR system may, for the second time, further move the content 30 degrees so that the content can be displayed 30 degrees behind from the second field of view here, and follow the user to the second field of view.

[0284] Figures 20J - 20N illustrate content that slowly follows the user from a first field of view to a second field of view of the user and / or the user's device, according to some embodiments. Figure 20J illustrates a top view of a user 2030 viewing content 2034 displayed on a surface (e.g., a virtual surface or an actual surface within a physical environment). The entire content 2034 is displayed directly in front of the user 2030 and is completely within the first field of view 2038 of the user 2030 and / or the device at the first head pose position of the user and / or the user's device. The user 2030 views the content 2034. Figure 20K illustrates, as an example, a top view of the user 2030 rotated approximately 45 degrees to the right (e.g., in the clockwise direction) with respect to the first head pose position illustrated in Figure 20J. A portion of the content 2034 (e.g., as depicted by the dashed line) is no longer within the field of view of the user 2030 and / or the device, while a portion of the content 2034 (e.g., as depicted by the solid line) is still being rendered / displayed to the user 2030.

[0285] FIG. 20L illustrates a top view of user 2030 at the completion of a 90-degree rotation of a second head pose position to the right (e.g., clockwise direction) relative to the first head pose position illustrated in FIG. 20J. Content 2034 is no longer visible to user 2030 (e.g., as depicted by the dashed lines around content 2034) because content 2034 is completely outside the field of view 2038 of user 2030 and / or the device. Note that content 2034 has also moved slowly. The slow movement correlates with a waiting time with respect to the time and amount by which content 2034 can move from its original position illustrated in FIGS. 20J / 20K and catch up to the second head pose position.

[0286] FIG. 20M illustrates the content 2034 slowly moving and being displayed in a new position such that a portion of the content 2034 is within the field of view 2038 of the user 2030 and / or the device at the second head pose position (e.g., as depicted by the solid line encompassing a portion of the content 2034). FIG. 20N illustrates the content 2034 having completed its slow movement and having fully caught up to the user's second head pose position. The content 2034 is fully within the field of view 2038 of the user 2030 and / or the device, as shown by the solid line encompassing the entire content 2034. One of ordinary skill in the art can understand that the user may change the field of view from a first field of view to a second field of view where the content is no longer visible, but the user may not desire for the content to be directly displayed within the second field of view. Instead, the user may desire, for example, for the system to slowly follow the content to the second field of view (e.g., the new field of view) without being displayed directly in front of the user until the system prompts the user to select whether the user desires for the content to be displayed directly in front of the user with respect to the user's second field of view, or the user may simply desire for the content to remain visibly displayed peripherally until the user interacts with the content that is visibly displayed peripherally. In other words, in some embodiments, the step of displaying the content / element on one or more surfaces may be moved in response to a change in the user's field of view from a first field of view to a second field of view, and the content / element slowly follows the change in the user's field of view from the first field of view to the second field of view. Further, in some embodiments, the content may move to the front of the second field of view only in response to confirmation received from the user to move the content to the front of the second field of view.

[0287] In some embodiments, the user may view (a) the extracted content elements presented to the user via the AR system and (b) interact with the extracted content elements. In some embodiments, the user may interact with the extracted content by purchasing items / services presented within the extracted content. In some embodiments, similar to online purchases made by a user interacting with a 2D web page, the AR system enables the user to interact with the extracted content presented on a surface and / or virtual object (e.g., a prism or virtual display screen) within the AR system and, as an example, may enable the electronic purchase of items and / or services presented within the extracted content presented on a surface and / or virtual object of the AR system.

[0288] In some embodiments, the user may interact with the extracted content element by further selecting an item within the displayed content element and placing the selected item on different surfaces and / or different virtual objects (e.g., prisms) within the user's physical environment. For example, as an example, the user may (a) use a totem to target a content element within a gallery, (b) press a trigger on the totem to select the content element and hold it for a time period (e.g., about 1 second), (c) move the totem to a desired location within the user's physical environment, and (d) press the trigger on the totem to place the content element at the desired location, and a copy of the content element may be extracted from the gallery by loading and displaying the copy of the content element at the desired location. In some embodiments, a preview of the content element is created and displayed as visual feedback as a result of the user selecting the content element and holding the trigger for a time period. The preview of the content may be created because creating a full-resolution version of the content element for use at the location of the content element may be resource-intensive. In some embodiments, the content element is copied / extracted and displayed for visual feedback as a whole as the user places the extracted content element at a desired location within the user's physical environment.

[0289] Figure 20O illustrates an example of a user who visually perceives the extracted content and interacts with the extracted content 2050 and 2054. Since the user 2040 was unable to detect a display surface suitable for displaying the content extracted by the sensor 1162 (for example, for a bookshelf), the extracted content 2044a-2044d may be visually perceived on a virtual display surface. Instead, the extracted content 2044a-d is displayed on a plurality of virtual display surfaces / screens. The extracted content 2044a is an online web site that sells audio headphones. The extracted content 2044b is an online web site that sells sports shoes. The extracted content 2044c / 2044d is an online furniture web site that sells furniture. The extracted content 2044d may include a detailed view of a specific item (for example, chair 2054) displayed from the extracted content 2044c. The user 2040 may interact with the extracted content by selecting a specific item from the extracted content being displayed and placing the extracted item within the user's physical environment (for example, chair 2054). In some embodiments, the user 2040 may interact with the extracted content by purchasing a specific item (for example, sports shoes 2050) displayed in the extracted content.

[0290] FIG. 21 illustrates an audio transition during such an environmental change. The active content within room 2014 may have, for example, virtual speaker 2122 that delivers spatial audio to the user from a location associated with a content element within room 2014. As the user transitions to room 2012, the virtual speaker may follow the user by positioning and directing the audio to virtual speaker 2124 at the center of the user's head (in a manner substantially the same as conventional headphones), and ceasing the audio playback from virtual speaker 2122. As synthesis process 140 matches the content element to the surface within room 2012, the audio output may shift from virtual speaker 2124 to virtual speaker 2126. The audio output, in this case, maintains a constant consumption of the content element, at least the audio output component, during the environmental transition. In some embodiments, the audio component is always a virtual speaker at the center of the user's head, eliminating the need to adjust the position of the spatial audio virtual speaker.

[0291] System Architecture Overview FIG. 22 is a block diagram of an exemplary computing system 2200 suitable for implementing an embodiment of the present disclosure. Computing system 2200 includes a bus 2206 or other communication mechanism for interconnecting subsystems and devices such as processor 2207, system memory 2208 (e.g., RAM), static storage device 2209 (e.g., ROM), disk drive 2210 (e.g., magnetic or optical), communication interface 2214 (e.g., modem or Ethernet® card), display 2211 (e.g., CRT or LCD), input device 2212 (e.g., keyboard and mouse), etc. for communicating information.

[0292] According to one embodiment of the present disclosure, computing system 2200 performs specific operations by a processor 2207 executing one or more sequences of one or more instructions contained within a system memory 2208. Such instructions may be read into the system memory 2208 from another computer-readable / usable medium such as a static storage device 2209 or a disk drive 2210. In an alternative embodiment, wired circuitry may be used in place of or in combination with software instructions to implement the present disclosure. Accordingly, embodiments of the present disclosure are not limited to any specific combination of hardware circuitry and / or software. In one embodiment, the term "logic" shall mean any combination of software or hardware used to implement all or a portion of the present disclosure.

[0293] The term "computer-readable medium" or "computer-usable medium" as used herein refers to any medium involved in providing instructions for execution to a processor 2207. Such a medium may take many forms including, but not limited to, non-volatile media and volatile media. Non-volatile media includes, for example, optical or magnetic disks such as disk drive 2210. Volatile media includes dynamic memory such as system memory 2208.

[0294] Common forms of computer-readable media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, flash-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read.

[0295] In certain embodiments of the present disclosure, the execution of a sequence of instructions for practicing the present disclosure is performed by a single computing system 2200. According to other embodiments of the present disclosure, two or more computing systems 2200 coupled by a communication link 2215 (e.g., a LAN, PTSN, or wireless network) may cooperate with each other to perform a sequence of instructions required to practice the present disclosure.

[0296] The computing system 2200 may transmit and receive messages, data, and instructions, including a program (i.e., application code), through the communication link 2215 and the communication interface 2214. The received program code may be executed by the processor 2207 as it is received and / or stored in the disk drive 2210 or other non-volatile storage device for later execution. The computing system 2200 may communicate with a database 2232 on an external storage device 2231 through the data interface 2233.

[0297] In the foregoing specification, the present disclosure has been described with reference to its specific embodiments. However, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the present disclosure. For example, the process flows described above are described with reference to a particular order of process actions. However, many of the orders of the described process actions may be changed without affecting the scope or operation of the present disclosure. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.

Claims

1. A method of matching a plurality of content elements of content to a spatial three-dimensional (3D) environment, the method comprising: a content structuring process including decomposing the content by analyzing the content, thereby identifying a plurality of content elements from the content; an environment structuring process including analyzing the environment, thereby identifying a plurality of surfaces from the environment; a synthesis process including matching the plurality of content elements and the plurality of surfaces respectively, and rendering and displaying the plurality of content elements on the matched plurality of surfaces respectively; comprising; wherein the content is a web page, the web page is a double-sided web page, the plurality of surfaces include opposite surfaces of a physical structure, two of the plurality of content elements are arranged on opposite sides of the double-sided web page, the two of the plurality of content elements are matched with the opposite surfaces of the physical structure and are rendered on the opposite surfaces of the physical structure.

2. The method according to claim 1, wherein analyzing the content includes identifying an attribute of the content and storing the attribute of the content in a logical structure for each of the plurality of content elements.

3. The method according to claim 2, wherein analyzing the environment includes identifying an attribute of the plurality of surfaces and storing the attribute of the plurality of surfaces in the logical structure for each of the plurality of surfaces.

4. The method according to claim 3, wherein matching the plurality of content elements and the plurality of surfaces includes comparing the attributes of the plurality of content elements and the attributes of the plurality of surfaces respectively.

5. The method according to claim 4, wherein the plurality of content elements are respectively matched to the plurality of surfaces based on the plurality of content elements and the plurality of surfaces sharing similar or opposite attributes.

6. The method according to claim 2, wherein the logical structure includes one of an ordered array, a hierarchical table, a tree structure, and a logical graph structure.

7. A method for matching a plurality of content elements of content to a spatial three-dimensional (3D) environment, the method comprising: A content structuring process including decomposing the content by analyzing the content, thereby identifying a plurality of content elements from the content; An environment structuring process including analyzing the environment, thereby identifying at least one surface from the environment; A synthesis process including matching the plurality of content elements with a plurality of surfaces including the at least one surface, and rendering and displaying the plurality of content elements on the matched plurality of surfaces; comprising: Analyzing the content includes identifying attributes of the content and storing the attributes of the content in a logical structure for each of the plurality of content elements; The logical structure includes one of an ordered array, a hierarchical table, a tree structure, and a logical graph structure; The logical structure includes a tree structure, the tree structure having a parent node to which a parent content element is assigned, and at least one child node linked to the parent node and to which at least another child content element is assigned, the parent content element and the at least one child content element being matched to the plurality of surfaces based on the proximity of the plurality of surfaces to each other. Claim 8 The method according to claim 1, wherein the plurality of matched surfaces further includes at least one virtual surface. Claim 9 The method according to claim 1, wherein rendering the plurality of content elements on the plurality of surfaces respectively includes scaling the content elements to the respective matched surfaces. Claim 10 The method according to claim 1, wherein at least some of the plurality of content elements matched to the plurality of surfaces in a first room are matched to a plurality of surfaces in a second room when the user exits the first room and enters the second room. Claim 11 The synthesis process according to claim 1, wherein the synthesis process determines that one of the plurality of surfaces is not suitable for one of the plurality of content elements, creates a virtual surface, matches the one content element with the virtual surface, and renders the one content element on the virtual surface.

12. The method according to claim 1, wherein the synthesis process further includes displaying a plurality of candidate surface options to a user for displaying one of the plurality of content elements, and matching the one content element with the surface corresponding to one candidate screen selected by the user among the plurality of displayed candidate surfaces.

13. The method according to claim 12, wherein the plurality of candidate surface options are arranged within the user's field of view, and the synthesis process further includes, when the user's field of view changes to a new field of view where additional candidate surface options are arranged, displaying the additional candidate surface options to the user for displaying the one content element.

14. A method of matching a plurality of content elements of content to a spatial three-dimensional (3D) environment, the method comprising: A content structuring process including decomposing content by analyzing the content, thereby identifying a plurality of content elements from the content; An environment structuring process including analyzing the environment, thereby identifying at least one surface from the environment; A synthesis process including matching each of the plurality of content elements with a plurality of surfaces including the at least one surface, and rendering and displaying each of the plurality of content elements on the matched plurality of surfaces; comprising; The synthesis process further includes displaying a plurality of candidate surface options to a user for displaying one of the plurality of content elements, and matching the one content element with the surface corresponding to one candidate screen selected by the user among the plurality of displayed candidate surfaces. The plurality of candidate surface options are arranged within the user's field of view, and the composition process further includes displaying the additional candidate surface options to the user in order to display the one content element when the user's field of view changes to a new field of view in which the additional candidate surface options are arranged. The method further includes sensing a change in head pose between the field of view and the new field of view, and the additional candidate surface options are only displayed to the user when the sensed change in head pose exceeds a head pose threshold. **Claim 15** The method according to claim 12, wherein the plurality of candidate surfaces are displayed to the user simultaneously. **Claim 16** The method according to claim 12, wherein the plurality of candidate surfaces are displayed to the user in sequence. **Claim 17** The method according to claim 12, wherein the composition process further includes displaying candidate views of the content element on the plurality of candidate surface options. **Claim 18** The method according to claim 12, wherein the composition process further includes storing the candidate surface option selected by the user to display the one content element as a preferred candidate surface option, and subsequently prioritizing the preferred candidate surface option when matching a plurality of content elements with a plurality of surfaces.

Citation Information

Patent Citations

  • User interface for augmented reality enabled devices

    JP2016508257A

  • Assisted Viewing Of Web-Based Resources

    US20150331240A1

  • User-based context sensitive hologram reaction

    US20160266386A1