Matching content to a spatial 3D environment
By identifying and matching content elements with surface attributes, the problem of poor display effects in spatial 3D environments using traditional methods has been solved, achieving more adaptive and dynamic content display and improving the user experience.
Patent Information
- Application Number
- CN202310220705.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-16
- Filing Date
- 2018-05-01
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2038-05-01
AI Technical Summary
When displaying content in a 3D spatial environment, traditional methods cannot effectively utilize the characteristics of the 3D spatial environment, resulting in poor display effects.
By receiving content, identifying elements within the content, determining surrounding surfaces, and matching elements to surfaces for display in a virtual reality or augmented reality system, the process includes identifying element attributes, determining surface attributes, creating virtual surfaces, calculating matching scores, and displaying matches based on user preferences.
It enables more effective content display in a spatial 3D environment, improves display adaptability and user experience, and can dynamically adjust the displayed content according to changes in the user's field of vision and environment.
Smart Images

Figure CN116203731B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 201880043910.4, “Matching of Content to Spatial 3D Environment” (filing date May 01, 2018). TECHNICAL FIELD
[0002] The present disclosure relates to systems and methods for displaying content in a spatial 3D environment. BACKGROUND
[0003] A typical way of viewing content is to open an application that will display the content on a display screen of a display device (e.g., monitor of a computer, smart phone, tablet, etc.). The user will navigate the application to view the content. Typically, when the user views the display screen of the display, there is a fixed format regarding how the content is displayed within the application and on the display screen of the display device.
[0004] With virtual reality (VR), augmented reality (AR), and / or mixed reality (MR) systems (hereinafter collectively referred to as “mixed reality” systems), the application will display the content in a spatial three-dimensional (3D) environment. When used in a spatial 3D environment, the conventional methods of displaying content on a display screen do not perform well. One reason is because, using conventional methods, the display area of the display device is a 2D medium that is limited to the screen area of the display screen on which the content is displayed. As a result, conventional methods are configured to only know how to organize and display content within that screen area of the display screen. In contrast, a spatial 3D environment is not limited to the strict boundaries of the screen area of the display screen. Thus, conventional methods can perform poorly when used in a spatial 3D environment because conventional methods do not necessarily have the functionality or ability to take advantage of a spatial 3D environment to display content.
[0005] Therefore, there is a need for an improved method of displaying content in a spatial 3D environment.
[0006] Nothing in this Background section should be taken as an acknowledgement or admission that anything in the Background section is prior art, merely because it is discussed in the Background section. Similarly, any problems or understandings of the reasons for problems discussed in the Background section should not be taken as prior art merely by virtue of their discussion in the Background section. The subject matter of the Background section can only be taken as indicating different approaches that are themselves also amenable to disclosure. SUMMARY
[0007] Embodiments of the present disclosure provide improved systems and methods to display information in a spatially organized 3D environment. The method includes receiving content, identifying elements in the content, determining surrounding surfaces, matching the identified elements to the surrounding surfaces, and displaying the elements as virtual content on the surrounding surfaces. Additional embodiments of the present disclosure provide improved systems and methods to push content to a user of a virtual reality or augmented reality system.
[0008] In one embodiment, a method includes receiving content. The method also includes identifying one or more elements in the content. The method further includes determining one or more surfaces. Additionally, the method includes matching the one or more elements to the one or more surfaces. Moreover, the method includes displaying the one or more elements as virtual content on the one or more surfaces.
[0009] In one or more embodiments, the content includes at least one of pull content or push content. Identifying the one or more elements can include determining one or more attributes of each of the one or more elements. The one or more attributes include at least one of a priority attribute, an orientation attribute, an aspect ratio attribute, a size attribute, an area attribute, a relative viewing position attribute, a color attribute, a contrast attribute, a location type attribute, an edge attribute, a content type attribute, a focus attribute, a readability index attribute, or a placement surface type attribute. Determining the one or more attributes of each of the one or more elements is based on an explicit indication in the content.
[0010] In one or more embodiments, determining the one or more attributes of each of the one or more elements is based on a placement of the one or more elements within the content. The method further includes storing the one or more elements into one or more logical structures. The one or more logical structures include at least one of an ordered array, a hierarchical table, a tree structure, or a logical graph structure. The one or more surfaces include at least one of a physical surface or a virtual surface. Determining the one or more surfaces includes parsing an environment to determine at least one of the one or more surfaces.
[0011] In one or more embodiments, determining the one or more surfaces includes receiving raw sensor data, simplifying the raw sensor data to produce simplified data, and creating one or more virtual surfaces based on the simplified data. The one or more surfaces include the one or more virtual surfaces. Simplifying the raw sensor data includes filtering the raw sensor data to produce filtered data, and grouping the filtered data into one or more groups by point cloud points. The simplified data includes the one or more groups. Creating the one or more virtual surfaces includes traversing each of the one or more groups to determine one or more real world surfaces, and creating the one or more virtual surfaces based on the one or more real world surfaces.
[0012] In one or more embodiments, determining the one or more surfaces includes determining one or more properties of each of the one or more surfaces. The one or more properties include at least one of a priority property, an orientation property, an aspect ratio property, a size property, an area property, a relative viewing position property, a color property, a contrast property, a location type property, an edge property, a content type property, a focal point property, a legibility indicator property, or a placement surface type property. The method further includes storing the one or more surfaces into one or more logical structures. Matching the one or more elements to the one or more surfaces includes: prioritizing the one or more elements; for each of the one or more elements, comparing one or more properties of the element to one or more properties of each of the one or more surfaces; calculating a match score based on the one or more properties of the element and the one or more properties of each of the one or more surfaces; and identifying a best matching surface having a highest match score. Additionally, for each of the one or more elements, storing an association between the element and the best matching surface.
[0013] In one or more embodiments, an element is matched to one or more surfaces. Further, each of the one or more surfaces is displayed to a user. Additionally, a user selection is received that indicates a winning surface of the displayed one or more surfaces. Further, surface properties of the winning surface are saved in a user preference data structure in accordance with the user selection. The content is data streamed from a content provider. The one or more elements are displayed to the user by a mixed reality device.
[0014] In one or more embodiments, the method further includes displaying one or more additional surface options for displaying the one or more elements based at least in part on a change in field of view of the user. The display of the one or more additional surface options is based at least in part on a time threshold corresponding to the change in field of view. The display of the one or more additional surface options is based at least in part on a head pose change threshold.
[0015] In one or more embodiments, the method further includes overriding a display of the one or more elements on the matched one or more surfaces. The overriding of the display of the one or more elements on the one or more surfaces is based at least in part on a surface that is historically frequently used. The method even further includes moving the one or more elements displayed on the one or more surfaces to a different surface based at least in part on a user selection of moving a particular element displayed at the one or more surfaces to the different surface. The particular element is moved to the different surface that is at least viewable by the user.
[0016] In one or more embodiments, the method additionally includes, in response to a change in the field of view of the user from the first field of view to the second field of view, slowly moving the display of the one or more elements to the new surface to follow the change in the field of view of the user to the second field of view. The one or more elements can only move to be directly in front of the second field of view of the user upon receiving confirmation from the user to move the content to be directly in front of the second field of view of the user. The method includes pausing the display of the one or more elements on the one or more surfaces at the first location and resuming the display of the one or more elements on one or more other surfaces at the second location is based at least in part on the user moving from the first location to the second location. The display of the one or more elements is automatically paused based at least in part on a determination that the user is or has moved from the first location to the second location. The display of the one or more elements is automatically resumed based at least in part on the identification and matching of the one or more other surfaces at the second location to the one or more elements.
[0017] In one or more embodiments, determining the one or more surfaces includes identifying one or more virtual objects for displaying the one or more elements. Identifying the one or more virtual objects is based at least in part on data received from the one or more sensors indicating a lack of a suitable surface. An element of the one or more elements is a television channel. The user interacts with an element of the one or more elements displayed by purchasing one or more items or services displayed to the user.
[0018] In one or more embodiments, the method further includes detecting a change in the environment from the first location to the second location, determining one or more additional surfaces at the second location, matching the one or more elements currently displayed at the first location to the one or more additional surfaces, and displaying the one or more elements as virtual content on the one or more additional surfaces at the second location. The determination of the one or more additional surfaces begins after the change in the environment exceeds a time threshold. The user pauses active content displayed at the first location and resumes active content to be displayed at the second location, the active content resuming at the same interaction point at which the user paused the active content at the first location.
[0019] In one or more embodiments, the method further includes, when the user leaves the first location, transferring spatialized audio delivered to the user from a location associated with the display content at the first location to an audio virtual speaker directed at the center of the user's head, and from the audio virtual speaker directed at the center of the user's head to spatialized audio delivered to the user from the one or more additional surfaces displaying the one or more elements at the second location.
[0020] In another embodiment, a method for pushing content to a user of a mixed reality system includes receiving one or more available surfaces from an environment of the user. The method also includes identifying one or more contents that match a size of one of the one or more available surfaces. The method further includes calculating a score based on comparing one or more constraints of the one or more contents to one or more surface constraints of the one available surface. Moreover, the method includes selecting a content from the one or more contents that has a highest score. In addition, the method includes storing a one-to-one match of the selected content to the one available surface. Furthermore, the method includes displaying the selected content to the user on the available surface.
[0021] In one or more embodiments, the environment of the user is a personal residence of the user. The one or more available surfaces from the environment of the user are peripheral to a focal viewing area of the user. The one or more contents are advertisements. The advertisements are targeted to a specific group of users located in a specific environment. The one or more contents are notifications from an application. The application is a social media application. One of the one or more constraints of the one or more contents is orientation. The selected content is 3D content.
[0022] In another embodiment, an augmented reality (AR) display system includes a head mounted system including one or more sensors and one or more cameras including an outward facing camera. The system also includes a processor executing a set of program code instructions. In addition, the system includes a memory for storing the set of program code instructions, wherein the set of program code instructions includes program code for receiving content. The program code also executes identifying one or more elements in the content. Moreover, the program code also executes determining one or more surfaces. Additionally, the program code also executes matching the one or more elements to the one or more surfaces. Further, the program code also executes displaying the one or more elements as virtual content on the one or more surfaces.
[0023] In one or more embodiments, the content includes at least one of a pull content or a push content. Identifying the one or more elements includes parsing the content to identify the one or more elements. Identifying the one or more elements includes determining one or more attributes of each of the one or more elements. Additionally, the program code also executes storing the one or more elements into one or more logical structures. The one or more surfaces include at least one of a physical surface or a virtual surface. Determining the one or more surfaces includes parsing an environment to determine at least one of the one or more surfaces.
[0024] In one or more embodiments, determining the one or more surfaces includes receiving raw sensor data, simplifying the raw sensor data to produce simplified data, and creating the one or more virtual surfaces based on the simplified data, wherein the one or more surfaces include the one or more virtual surfaces. Determining the one or more surfaces includes determining one or more properties of each of the one or more surfaces. Additionally, the program code further performs storing the one or more surfaces into one or more logical structures.
[0025] In one or more embodiments, matching the one or more elements to the one or more surfaces includes prioritizing the one or more elements, comparing one or more properties of each of the one or more elements to one or more properties of each of the one or more surfaces, calculating a match score based on the one or more properties of the element and the one or more properties of each of the one or more surfaces, and identifying a best matching surface having a highest match score. An element is matched to the one or more surfaces. The content is data streamed from a content provider.
[0026] In one or more embodiments, the program code further performs displaying one or more surface options for displaying the one or more elements based at least in part on a change in a field of view of the user. The program code further performs overlaying a display of the one or more elements on the matched one or more surfaces. The program code further performs moving one or more elements displayed on the one or more surfaces to a different surface based at least in part on a user selection to move a particular element displayed at the one or more surfaces to the different surface. The program code further performs slowly moving a display of the one or more elements onto a new surface to follow a change in the field of view of the user from a first field of view to a second field of view in response to the change in the field of view of the user from the first field of view to the second field of view.
[0027] In one or more embodiments, the program code further performs pausing a display of the one or more elements on the one or more surfaces at a first location and resuming a display of the one or more elements on one or more other surfaces at a second location is based at least in part on a movement of the user from the first location to the second location. Determining the one or more surfaces includes identifying one or more virtual objects for displaying the one or more elements. An element of the one or more elements is a television channel.
[0028] In one or more embodiments, a user interacts with an element of one or more elements displayed to the user by purchasing one or more items or services displayed to the user. The program code further performs: detecting a change in the environment from a first location to a second location; determining one or more additional surfaces at the second location; matching one or more elements currently displayed at the first location to the one or more additional surfaces; and displaying the one or more elements as virtual content on the one or more additional surfaces at the second location.
[0029] In another embodiment, an augmented reality (AR) display system includes a head mounted system including one or more sensors and one or more cameras including an outward facing camera. The system further includes a processor executing a set of program code instructions. The system further includes a memory for storing the set of program code instructions, wherein the set of program code instructions includes program code for receiving one or more available surfaces from an environment of a user. The program code further performs identifying one or more contents that match a size of one available surface of the one or more available surfaces. The program code further performs computing a score based on one or more constraints of the one or more contents compared to one or more surface constraints of the one available surface. The program code further performs selecting a content from the one or more contents having a highest score. In addition, the program code performs storing the selected content to a one-to-one match to the one available surface. The program code further performs displaying the selected content to the user on the available surface.
[0030] In one or more embodiments, the environment of the user is a personal residence of the user. The one or more available surfaces from the environment of the user are peripheral to a focal viewing area of the user. The one or more contents are advertisements. The one or more contents are notifications from an application. One of the one or more constraints of the one or more contents is orientation. The selected content is 3D content.
[0031] In another embodiment, a computer-implemented method for deconstructing 2D content includes identifying one or more elements in the content. The method further includes identifying one or more surrounding surfaces. The method further includes mapping the one or more elements to the one or more surrounding surfaces. In addition, the method includes displaying the one or more elements as virtual content on the one or more surfaces.
[0032] In one or more embodiments, the content is a webpage. An element of the one or more elements is a video. The one or more surrounding surfaces include a physical surface within the physical environment or a virtual object not physically located within the physical environment. The virtual object is a plurality of stacked virtual objects. The first set of results of the identified one or more elements and the second set of results of the identified one or more surrounding surfaces are stored in a database table within a storage device. The storage device is a local storage device. The database table storing the results of the identified one or more surrounding surfaces includes: a surface ID, a width dimension, a height dimension, an orientation description, and a location relative to a reference frame.
[0033] In one or more embodiments, identifying the one or more elements in the content includes: identifying attributes from tags corresponding to placement of the elements; extracting hints from tags for the one or more elements; and storing the one or more elements. Identifying the one or more surrounding surfaces includes: identifying surrounding surfaces of the user currently; determining a pose of the user; identifying dimensions of the surrounding surfaces; and storing the one or more surrounding surfaces. Mapping the one or more elements to the one or more surrounding surfaces includes: looking up predefined rules for identifying candidate surrounding surfaces for mapping; and selecting a best fit surface for each of the one or more elements. Displaying the one or more elements on the one or more surrounding surfaces is performed through an augmented reality device.
[0034] In another embodiment, a method of matching content elements of content to a spatial three-dimensional (3D) environment includes a content structuring process, an environment structuring process, and a composition process.
[0035] In one or more embodiments, the content structuring process reads the content and organizes and / or stores the content into a logical / hierarchical structure for accessibility. The content structuring process includes a parser for receiving the content. The parser parses the received content to identify content elements from the received content. The parser identifies / determines attributes and stores the attributes into the logical / hierarchical structure for each content element.
[0036] In one or more embodiments, an environment structuring process parses data related to an environment to identify surfaces. The environment structuring process includes one or more sensors, a computer vision processing unit (CVPU), a perception framework, and an environment parser. The one or more sensors provide raw data related to real-world surfaces (e.g., point clouds of objects and structures from the environment) to the CVPU. The CVPU simplifies and / or filters the raw data. The CVPU changes the remaining data into group point cloud points by distance and planarity to extract / identify / determine surfaces by downstream processes. The perception framework receives the group point cloud points from the CVPU and prepares environment data for the environment parser. The perception framework creates / determines structures / surfaces / planes and populates one or more data storage devices. The environment parser parses the environment data from the perception framework to determine surfaces in the environment. The environment parser using object recognition identifies objects based on the environment data received from the perception framework.
[0037] In one or more embodiments, a synthesis process matches content elements (e.g., a table of content elements stored in a logical structure) from the parser with surfaces (e.g., a table of surfaces stored in a logical structure) of the environment from the environment parser to determine which content element should be rendered / mapped / displayed on which surface of the environment. The synthesis process includes a matching module, a rendering module, a create virtual object module, a display module, and a receiving module.
[0038] In one or more embodiments, the matching module pairs / matches content elements stored in a logical structure with surfaces stored in a logical structure. The matching module compares attributes of content elements with attributes of surfaces. The matching module matches content elements with surfaces based on the content elements and surfaces sharing similar and / or opposite attributes. The matching module can access one or more preference data structures, such as user preferences, system preferences, and / or passable preferences, and can use the one or more preference data structures in the matching process. The matching module matches one content element with one or more surfaces based at least in part on at least one of a content vector (e.g., an orientation attribute), a head pose vector (e.g., an attribute of a VR / AR device, not a surface), or a surface normal vector of one or more surfaces. The results can be stored in a cache memory or a permanent storage device for further processing. The results can be organized and stored in a table for inventory matching.
[0039] In one or more embodiments, the optional create virtual object module creates a virtual object for displaying a content element based on a determination that creating a virtual object for displaying the content element is the most optimal choice, where the virtual object is a virtual planar surface. The virtual object for displaying the content element can be created based on data received from one or more particular sensors of the plurality of sensors, or due to a lack of sensor input from one or more particular sensors. Data received from environment-centric sensors of the plurality of sensors, such as cameras or depth sensors, indicates a lack of a suitable surface based on the user's current physical environment, or such sensors are simply unable to discern the presence of a surface.
[0040] In one or more embodiments, the render module renders the content elements to the respective matching surfaces, which include real surfaces and / or virtual surfaces. The render module renders the content elements for scaling to fit the matching surfaces. Even when the user moves from a first room to a second room, the content elements that match surfaces in the first room (real and / or virtual) remain matched to the surfaces in the first room. The content elements that match surfaces in the first room are not mapped to surfaces in the second room. When the user returns to the first room, the content elements rendered to the surfaces in the first room will resume display, with other features such as audio playback and / or playback time seamlessly resuming playback as if the user had never left the room.
[0041] In one or more embodiments, when the user leaves a first room and enters a second room, content elements that match surfaces in the first room match surfaces in the second room. A first group of content elements that match surfaces in the first room remain matched to the surfaces in the first room, while a second group of content elements that match surfaces in the first room can move with the device implementing the AR system to the second room. The second group of content elements moves with the device as the device enters the second room from the first room. The content elements are determined to be in the first group or the second group based at least in part on at least one of an attribute of the content element, an attribute of one or more surfaces in the first room that the content element matches, a user preference, a system preference, and / or a world preference that is navigable. A content element can match a surface but not be rendered to the surface when the user is not in proximity to the surface or when the user is not in line of sight of the surface.
[0042] In some embodiments, the content element is displayed on all of the first three surfaces at once. The user can select a surface from the first three surfaces as a preferred surface. The content element is displayed on only one of the first three surfaces at a time, and the user is indicated that the content element can be displayed on the other two surfaces. The user can then browse the other surface options, and as each surface option is activated by the user, the content element can be displayed on the activated surface. The user can then select a surface from the surface options as a preferred surface.
[0043] In some embodiments, a user extracts channels of a television by aligning a totem to a channel, pressing a trigger on the totem to select the channel and holding the trigger for a period of time (e.g., about 1 second), moving the totem around to identify a desired location in an environment for displaying the extracted television channel, and pressing the trigger on the totem to place the extracted television channel at the desired location in the environment. The desired location is a surface suitable for displaying the television channel. A Prism is created at the desired location and the selected channel content is loaded and displayed in the Prism. While moving the totem around to identify the desired location in the environment for displaying the extracted television channel, the user is shown visual content. The visual content can be at least one of a single image illustrating the channel, one or more images illustrating a preview of the channel, or a video stream illustrating current content of the channel.
[0044] In some embodiments, a method includes identifying a first field of view of a user, generating one or more surface options for displaying content, changing from the first field of view of the user to a second field of view, generating one or more additional surface options for displaying content corresponding to the second field of view, presenting the one or more surface options corresponding to the first field of view and the one or more additional surface options corresponding to the second field of view, receiving a selection from the user to display content on a surface corresponding to the first field of view while the user is viewing in the second field of view, and displaying an indication to the user in the direction of the first field of view, the indication indicating to the user that the user should navigate back in the indicated direction to the first field of view to view the selected surface option.
[0045] In some embodiments, a method includes displaying content on a first surface in a first field of view of a user, the first field of view corresponding to a first head pose. The method also includes determining that a time period of a change from the first field of view to a second field of view exceeds a time threshold, displaying an option for the user to change a display location of the content from the first surface in the first field of view to one or more surface options in the second field of view. The second field of view corresponds to a second head pose. In some embodiments, the system displays the option for the user to change the display location of the content immediately upon the field of view of the user changing from the first field of view to the second field of view. The first head pose and the second head pose have a position change that is greater than a head pose change threshold.
[0046] In some embodiments, a method includes rendering and displaying content on one or more first surfaces, where a user viewing the content has a first head pose. The method also includes, in response to the user changing from the first head pose to a second head pose, rendering the content on one or more second surfaces, where the user viewing the content has the second head pose. The method also includes providing the user with an option to change a display location of the content from the one or more first surfaces to the one or more second surfaces. The option is provided to the user to change the display location when the head pose change is greater than a corresponding head pose change threshold. The head pose change threshold is greater than 90 degrees. A time that the head pose change is maintained is greater than a threshold time period. The head pose change is less than the head pose change threshold, where the option to change the display location of the content is not provided.
[0047] In some embodiments, a method includes evaluating a list of surfaces visible to a user as the user moves from a first location to a second location, the list of surfaces being suitable for displaying certain types of content that can be pushed into an environment of the user without requiring the user to search or select the content. The method also includes determining a preference attribute of the user, the preference attribute indicating when and where certain types of pushed content can be displayed. The method also includes displaying the pushed content on one or more surfaces based on the preference attribute.
[0048] In another embodiment, a method for pushing content to a user of an augmented reality system includes determining one or more surfaces and corresponding attributes of the one or more surfaces. The method also includes receiving one or more content elements that match the one or more surfaces based at least in part on a surface attribute. The method also includes calculating a match score based at least in part on how well an attribute of the content element matches the attributes of the one or more surfaces, where the match score is based on a scale of 1-100, where a score of 100 is the highest score and a score of 1 is the lowest score. The method also includes selecting a content element from the one or more content elements having the highest match score. Further, the method includes storing a match of the content element to the surface. Additionally, the method includes rendering the content element onto the matched surface.
[0049] In one or more embodiments, the content element includes a notification and the match score is calculated based on an attribute that can indicate a priority of the content element that needs to be notified rather than a match to a particular surface. The preferred content element is selected based on a competition as to how well the attribute of the content element matches the attribute of the surface. The content element is selected based on a content type, where the content type is 3D content and / or a notification from a social media contact.
[0050] In some embodiments, a method for generating a 3D preview for a web link includes representing a 3D preview of a web link as a new set of HTML tags and properties associated with a web page. The method further includes specifying a 3D model as an object and / or surface to render the 3D preview. The method further includes generating the 3D preview and loading the 3D preview onto the 3D model. The 3D model is a 3D volume etched into a 2D web page.
[0051] Every individual embodiment described and illustrated herein has discrete components and features which can be readily separated from or combined with the components and features of any of the other several embodiments without departing from the scope of the present disclosure.
[0052] Further details of the features, objects and advantages of the present disclosure are described in the following detailed description, appended claims, and drawings. The foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to be restrictive of the scope of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0053] The accompanying drawings illustrate the design and utility of various embodiments of the present disclosure. It should be noted that the drawings are not necessarily drawn to scale and that in the figures, like reference numerals designate corresponding parts throughout the several views. For better understanding of how the above-summarized and other advantages and objects of the present disclosure can be obtained, a more complete understanding can be obtained by reference to the detailed description below in connection with the associated drawings in which certain representative embodiments of the present disclosure are shown. Understanding that these drawings depict only typical embodiments of the present disclosure and are not therefore to be considered limiting of its scope, the present disclosure will be described and explained with additional specificity and detail by the use of the accompanying drawings in which:
[0054] Figures 1A-1B An example system and computer-implemented method of matching content elements of content to a spatial three-dimensional (3D) environment is shown in accordance with some embodiments.
[0055] Figures 2A-2E An example of matching content elements to surfaces in a spatial three-dimensional (3D) environment is shown in accordance with some embodiments.
[0056] Figures 3A-3B An example of web content adjusted to lighting and color conditions is shown in accordance with some embodiments.
[0057] Figure 4 is a flowchart showing a method for matching content elements to surfaces to be displayed in a 3D environment in accordance with some embodiments.
[0058] Figure 5 is a flowchart showing a method for identifying elements in content in accordance with some embodiments.
[0059] Figure 6 is a flowchart illustrating a method for determining surfaces from a user's environment according to some embodiments.
[0060] Figures 7A-7B is a flowchart illustrating various methods for matching elements from content to surfaces according to some embodiments.
[0061] Figure 7C An example of a user moving content to a work area in which the content is subsequently displayed in a display surface according to some embodiments is shown.
[0062] Figure 8 A matching score method according to some embodiments is shown.
[0063] Figure 9 An example of a world location environment API providing location specific environments according to some embodiments is shown.
[0064] Figure 10 is a flowchart illustrating a method for pushing content to a user of a VR / AR system according to some embodiments.
[0065] Figure 11 An augmented reality environment for matching / displaying content elements to surfaces according to some embodiments is shown.
[0066] Figure 12 An augmented reality environment for matching / displaying content elements to surfaces according to some embodiments is shown.
[0067] Figures 13A-13B An example two-sided web page according to some embodiments is shown.
[0068] Figures 14A-14B An example of different structures for storing content elements from content according to some embodiments is shown.
[0069] Figure 15 An example of a table for storing a list of surfaces identified from a user's local environment according to some embodiments is shown.
[0070] Figure 16 An example 3D preview for a web link according to some embodiments is shown.
[0071] Figure 17 An example of a web page with a 3D volume etched into the web page according to some embodiments is shown.
[0072] Figure 18 An example of a table for storing matching / mapping of content elements to surfaces according to some embodiments is shown.
[0073] Figure 19 An example of an environment including content elements that match a surface is shown in accordance with some embodiments.
[0074] Figures 20A-20O An example of a dynamic environment matching protocol for content elements is shown in accordance with some embodiments.
[0075] Figure 21 Audio transfer during environment changes is shown in accordance with some embodiments.
[0076] Figure 22 is a block diagram of an illustrative computing system suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0077] Various embodiments will now be described in detail with reference to the accompanying drawings. The embodiments are provided as illustrative examples of the disclosure so as to enable those skilled in the art to practice the disclosure. It is to be noted that the following figures and examples are not meant to limit the scope of the disclosure. Where certain elements of the disclosure can be implemented in hardware or software (or a combination of both), only those elements that are necessary for an understanding of the present disclosure will be described in detail, and other elements will be omitted so as to not obscure the disclosure. Moreover, various embodiments encompass known equivalents of the components referred to herein by way of example.
[0078] Embodiments of the present disclosure display content or content elements in a spatially organized 3D environment. For example, the content or content elements can include pushed content, pulled content, first-party content, and third-party content. Pushed content is content that a server (e.g., a content designer) sends to a client (e.g., a user), where the initial request originates from the server. Examples of pushed content can include: (a) notifications from various applications, such as stock notifications, news feeds; (b) premium content, such as, for example, updates and notifications from social media applications, email updates, etc.; and / or (c) advertisements targeted to a broad target group and / or a specific target group, etc. Pulled content is content that a client (e.g., a user) requests from a server (e.g., a content designer), where the initial request originates from the client. Examples of pulled content can include: (a) web pages requested by a user using, for example, a browser; (b) streaming data requested by a user from a content provider using, for example, a streaming application, such as a video and / or audio streaming application; (c) and / or any digital format data that a user can request / access / query. First-party content is content that is generated by a client (e.g., a user) on any device owned / used by the client (e.g., a client device such as a mobile device, a tablet, a camera, a head-mounted display device, etc.). Examples of first-party content include photos, videos, etc. Third-party content is content that is generated by a party other than the client (e.g., a television network, a movie streaming service provider, a web page developed by someone other than the user, and / or any data not generated by the user). Examples of third-party content can include web pages generated by someone other than the user, data / audio / video streams and associated content received from one or more sources, any data generated by someone other than the user, etc.
[0079] The content can originate from a web page and / or an application on a head-mounted system, a mobile device (e.g., a cellular phone), a tablet, a television, a server, etc. In some embodiments, the content can be received from another application or device, such as a laptop, a desktop computer, an email application with a link to the content, an electronic message referencing or including a link to the content, etc. The following detailed description includes web pages as examples of content. However, the content can be any content, and the principles disclosed herein will apply.
[0080] Block Diagram
[0081] Figure 1A An example system and computer-implemented method of matching content elements of content to a spatial three-dimensional (3D) environment, in accordance with some embodiments, is shown. The system 100 includes a content structuring process 120, an environment structuring process 160, and a composition process 140. The system 100, or portions thereof, can be implemented on a device, such as a head-mounted display device.
[0082] The content structuring process 120 is a process that reads the content 110 and organizes / stores the content 110 into a logical structure so that the content 110 can be accessed and more easily programmatically extract content elements from the content 110. The content structuring process 120 includes a parser 115. The parser 115 receives the content 110. For example, the parser 115 receives the content 110 from an entity (e.g., a content designer). The entity can be, for example, an application. The entity can be external to the system 100. As described above, the content 110 can be, for example, pushed content, pulled content, first-party content, and / or third-party content. When the content 110 is requested, an external web server can serve the content 110. The parser 115 parses the content 110 to identify content elements of the content 110. The parser 115 can identify and subsequently organize and store the content elements in a logical structure, such as a content table that inventories the content 110. The content table can be, for example, a tree structure (such as a document tree or graph) and / or a database table (such as a relational database table).
[0083] The parser 115 can identify / determine and store attributes of each content element. The attributes of each content element can be explicitly indicated by a content designer of the content 110 or can be determined or inferred by the parser 115, for example, based on placement of the content elements within the content 110. For example, the parser 115 can determine or infer attributes of each content element based on placement of the content elements within the content 110 relative to each other. The attributes of the content elements are described in further detail below. The parser 115 can generate a list of all content elements parsed from the content 110 along with the respective attributes. After parsing and storing the content elements, the parser 115 can order the content elements based on associated priorities (e.g., from highest to lowest).
[0084] Some benefits of organizing and storing the content elements in a logical structure is that once the content elements are organized and stored into a logical structure, the system 100 can query and manipulate the content elements. For example, in a hierarchical / logical structure represented as a tree structure with nodes, if a node is deleted, then all content below the deleted node can also be deleted. Likewise, if a node is moved, then all content below the node can move with it.
[0085] The environment structuring process 160 is a process that parses data related to the environment to identify surfaces. The environment structuring process 160 can include sensors 162, a computer vision processing unit (CVPU) 164, a perception framework 166, and an environment parser 168. The sensors 162 provide raw data about real-world surfaces (e.g., point clouds from objects and structures in the environment) to the CVPU 164 for processing. Examples of the sensors 162 can include a global positioning system (GPS), a wireless signal sensor (Wi-Fi, Bluetooth, etc.), a camera, a depth sensor, an inertial measurement unit (IMU) including a triad of accelerometers and a triad of angular rate sensors, a magnetometer, a radar, a barometer, an altimeter, an accelerometer, a photometer, a gyroscope, etc.
[0086] The CVPU 164 simplifies or filters the raw data. In some embodiments, the CVPU 164 can filter out noise from the raw data to produce simplified raw data. In some embodiments, the CVPU 164 can filter out data that can not be used and / or can not be relevant to the current environment scanning task from the raw data and / or the simplified raw data to produce filtered data. The CVPU 164 can change the remaining data into grouped point cloud points by distance and planarity, making it easier for downstream to extract / identify / determine surfaces. The CVPU 164 provides the processed environment data to the perception framework 166 for further processing.
[0087] The perception framework 166 receives the grouped point cloud points from the CVPU 164 and prepares the environment data for the environment parser 168. The perception framework 166 creates / determines structures / surfaces / planes (e.g., a list of surfaces) and populates one or more data storage devices (such as, for example, an external database, a local database, a dedicated local storage device, local memory, etc.). For example, the perception framework 166 iterates through all the grouped point cloud points received from the CVPU 164 and creates / determines virtual structures / surfaces / planes that correspond to real-world surfaces. The virtual planes can be four vertices (picked from the grouped point cloud points) that create a virtual constructed rectangle (e.g., split into two triangles in a rendering pipeline). The structures / surfaces / planes created / determined by the perception framework 166 are referred to as environment data. When rendered and overlaid on the real-world surfaces, the virtual surfaces are essentially placed on their corresponding one or more real-world surfaces. In some embodiments, the virtual surfaces are perfectly placed on their corresponding one or more real-world surfaces. The perception framework 286 can maintain a one-to-one or one-to-many match / mapping of the virtual surfaces to the corresponding real-world surfaces. The one-to-one or one-to-many match / mapping can be used for queries. The perception framework 286 can update the one-to-one or one-to-many match / mapping when the environment changes.
[0088] The environment parser 168 parses the environment data from the perception framework 166 to determine surfaces in the environment. The environment parser 168 can identify objects using object recognition based on the environment data received from the perception framework 166. More details on object recognition are described in U.S. Patent No. 9,671,566 entitled “PLANAR WAVEGUIDE APPARATUS WITH DIFFRACTION ELEMENT(S) AND SYSTEM EMPLOYING SAME” and U.S. Patent No. 9,761,055 entitled “USING OBJECT RECOGNIZERS IN AN AUGMENTED OR VIRUTAL REALITY SYSTEM,” which are incorporated by reference herein. The environment parser 168 can organize and store surfaces in a logical structure, such as a surface table for inventorying surfaces. The surface table can be, for example, an ordered array, a hierarchical table, a tree structure, a logical graph structure, etc. In one example, an ordered array can be iterated linearly until a good fit surface is determined. In one example, for a tree structure ordered by a particular parameter (e.g., maximum surface area), the best fit surface can be determined by successively comparing whether each surface in the tree is smaller or larger than a requested area. In one example, in a logical graph data structure, searching for the best fit surface can be based on a relevant adjacency parameter (e.g., distance from the observer), or can have a table for quick search for a particular surface request.
[0089] The data structures described above can be where the environment parser 168 stores data corresponding to determined surfaces at runtime (and updates the data based on environment changes, if needed) to process surface matching and run any other algorithms on. In one embodiment, the data structures described above with respect to the environment parser 168 can not be where the data is more persistently stored. The data can be more persistently stored by the perception framework 166, which can be a runtime memory RAM, an external database, a local database, etc. when receiving and processing the data. Prior to processing surfaces, the environment parser 168 can receive surface data from the persistent storage device and populate the logical data structures from them, and then run the matching algorithms on the logical data structures.
[0090] The environment parser 168 can determine and store properties of each surface. The properties of each surface can be meaningful with respect to the properties of the content elements in the content table from the parser 115. The properties of a surface are described in further detail below. The environment parser 168 can generate a list of all surfaces parsed from the environment and the respective properties. After parsing and storing the surfaces, the environment parser 168 can order the surfaces based on the associated priority (e.g., from highest to lowest). The associated priority of a surface can be established when the environment parser 168 receives the surface data from the persistent storage device and populates the logical data structure with them. For example, if the logical data structure comprises a binary search tree, for each surface (received in a regular enumerated list) from the storage device, the environment parser 168 can first compute the priority (e.g., based on one or more properties of the surface) and then insert the surface in the logical data structure at the appropriate location. The environment parser 168 can parse point clouds and extract surfaces and / or planes based on proximity of points / relationships in space. For example, the environment parser 168 can extract horizontal and vertical planes and associate dimensions with the planes.
[0091] The content structuring process 120 parses the content 110 and organizes the content elements into a logical structure. The environment structuring process 160 parses data from the sensors 162 and organizes the surfaces from the environment into a logical structure. The logical structure comprising content elements and the logical structure comprising surfaces are used for matching and manipulation. The logical structure comprising content elements can be different (in type) from the logical structure comprising surfaces.
[0092] The synthesis process 140 is a process that matches content elements from the parser 115 (e.g., a table of content elements stored in a logical structure) with surfaces from the environment of the environment parser 168 (e.g., a table of surfaces stored in a logical structure) to determine which content element to render / map / display on which surface of the environment. In some embodiments, as shown in FIG. 1, the synthesis process 140 can include a matching module 142, a rendering module 146, and an optional virtual object creation module 144. In some embodiments, as shown in FIG. 1, the synthesis process 140 can also include a display module 148 and a receiving module 150. Figure 1A Figure 1B
[0093] The matching module 142 pairs / matches content elements stored in the logical structure with surfaces stored in the logical structure. The matching can be a one-to-one or one-to-many match of content elements to surfaces (e.g., one content element to one surface, one content element to two or more surfaces, two or more content elements to one surface, etc.). In some embodiments, the matching module 142 can pair / match content elements with a portion of a surface. In some embodiments, the matching module 142 can pair / match one or more content elements with one surface. The matching module 142 compares attributes of content elements to attributes of surfaces. The matching module 142 matches content elements to surfaces based on the content elements and surfaces sharing similar and / or opposite attributes. This organized infrastructure of content elements stored in the logical structure and surfaces stored in the logical structure allows for easy creation, updating, and implementation of matching rules, policies, and constraints to support and improve the matching process performed by the matching module 142.
[0094] The matching module 142 can access one or more preference data structures, such as user preferences, system preferences, and / or traversable preferences, and can use the one or more preference data structures in the matching process. The user preferences can be based on, for example, a model of overall preferences based on past actions, and can be specific to a particular content element type. For one content element, the system preferences can include a top two or more surfaces, where the user can have the ability to navigate among the two or more surfaces to select a preferred surface. The top two or more surfaces can be based on user preferences and / or traversable preferences. The traversable preferences can be retrieved from a cloud database, where the traversable preferences can be based on, for example, a model of groupings of other users, similar users, all users, similar environments, content element types, etc. The traversable preferences database can be pre-populated with consumer data (e.g., overall consumer data, consumer test data, etc.) to provide reasonable matching even before a large dataset (e.g., a dataset of users) is accumulated.
[0095] The matching module 142 matches one content element to one or more surfaces based at least in part on a content vector (e.g., an orientation attribute), a head pose vector (e.g., an attribute of a VR / AR device, not a surface), and a surface normal vector of the one or more surfaces. The content vector, the head pose vector, and the surface normal vector are described in detail below.
[0096] The matching module 142 generates matching results having at least one-to-one or one-to-many (e.g., one content element to one surface, one content element to two or more surfaces, two or more content elements to one surface, etc.) matches / mappings of content elements to surfaces. The results can be stored in cache memory or permanent storage for further processing. The results can be organized and stored in a table to inventory the matches.
[0097] In some embodiments, the matching module 142 can generate matching results in which one content element can be matched / mapped to multiple surfaces such that the content element can be rendered and displayed on any of the multiple surfaces. For example, a content element can be matched / mapped to five surfaces. The user can then select one of the five surfaces as a preferred surface on which the content element should then be displayed. In some embodiments, the matching module 142 can generate matching results in which one content element can be matched / mapped to the top three surfaces of multiple surfaces.
[0098] In some embodiments, when the user picks or selects a preferred surface, the selection made by the user can update the user preferences such that the system 100 can more accurately and precisely recommend content elements to surfaces.
[0099] If the matching module 142 matches all content elements to at least one surface, or discards content elements (e.g., for mapping to other surfaces, or no suitable match found), the composition process 140 can proceed to the rendering module 146. In some embodiments, for content elements that do not have a matching surface, the matching module 142 can create a match / mapping of the content element to a virtual surface. In some embodiments, the matching module 142 can eliminate content elements that do not have a matching surface.
[0100] Optional virtual object creation module 144 can create virtual objects, such as virtual planes, for displaying content elements. During the matching process by matching module 142, it can be determined that a virtual surface can be an optional surface for displaying certain content elements thereon. The determination can be based on texture properties, occupied properties, and / or other properties of the surface determined by environment parser 168, and / or properties of the content elements determined by parser 115. Texture properties and occupied properties of a surface are described in detail below. For example, matching module 142 can determine that texture properties and / or occupied properties can be disqualifying properties of a potential surface. Matching module 142 can determine that content elements can be displayed on a virtual surface instead based at least on texture properties and / or occupied properties. The location of the virtual surface can be relative to the location of one or more (real) surfaces. For example, the location of the virtual surface can be a distance from the location of one or more (real) surfaces. In some embodiments, matching module 142 can determine that there is no suitable (real) surface, or that sensor 162 can not detect any surface at all, and thus, virtual object creation module 144 can create a virtual surface to display content elements thereon.
[0101] In some embodiments, a virtual object for displaying content elements can be created based on data received from one or more particular sensors of sensor 162, or due to lack of sensor input from one or more particular sensors. Data received from environment-centric sensors of sensor 162, such as cameras or depth sensors, can indicate a lack of suitable surfaces based on the user’s current physical environment, or such sensors can not be able to discern the presence of surfaces at all (e.g., depending on the quality of the depth sensor, highly absorptive surfaces can make surface recognition difficult, or lack of connectivity such that access to certain shareable maps that can provide surface information is impeded).
[0102] In some embodiments, environment parser 168 can passively determine that there is no suitable surface if it does not receive data from sensor 162 or perception framework 166 within a particular timeframe. In some embodiments, sensor 162 can actively confirm that an environment-centric sensor is unable to determine a surface, and can pass such a determination to environment parser 168 or rendering module 146. In some embodiments, if environment structured 160 has no surfaces to provide to composition 140, or through passive determination by environment parser 168 or through active confirmation by sensor 162, composition process 140 can create a virtual surface or access a stored or registered surface such as from storage module 152. In some embodiments, environment parser 168 can receive surface data directly, such as from a hotspot or third-party perception framework or storage module, without input from the device’s own sensor 162.
[0103] In some embodiments, certain sensors, such as GPS, can determine that the location in which the user is located does not have a suitable surface on which to display a content element (e.g., an empty park or beach), or the only sensor providing data is a sensor that does not provide map information but rather orientation information (e.g., a magnetometer). In some embodiments, a certain type of display content element can require a display surface that can not be available or detectable in the user’s physical environment. For example, a user can want to view a map that displays a walking route from the user’s hotel room to a certain location. In order for the user to be able to maintain a view of the walking map as the user navigates to the location, the AR system can need to consider creating a virtual object (such as a virtual surface or screen) to display the walking map, as the environment resolver 168 can not have a sufficient available surface or detectable surface based on the data received (or not received) from the sensors 162 that would enable the user to continuously view the walking map from the user’s room in the hotel to the destination location on the walking map. For example, the user can have to enter an elevator that can limit or prevent network connectivity, exit the hotel, cross an open area such as a park where there can not be an available surface to display a content element, or there can be too much noise for the sensors to accurately detect the required surface. In this example, the AR system can determine that based on the content to be displayed and potential issues that can include a lack of network connectivity or a lack of a suitable display surface (e.g., based on GPS data of the user’s current location), the AR system can determine that it is best to create a virtual object to display the content element, as opposed to relying on the environment resolver 168 to find a suitable display surface using information received from the sensors 162. In some embodiments, the virtual object created to display the content element can be a prism. Further details regarding prisms are described in commonly owned U.S. Provisional Patent Application No. 62 / 610,101, entitled “METHODS AND SYSTEM FOR MANAGING AND DISPLAYING VIRTUAL CONTENT IN A MIXED REALITY SYSTEM,” filed December 22, 2017, which is incorporated by reference in its entirety. One of ordinary skill in the art can appreciate many examples in which it can be more beneficial to create a virtual surface on which to display a content element than to display the content element on a (real) surface.
[0104] The rendering module 146 renders the content elements to their matching surfaces. The matching surfaces can include real surfaces and / or virtual surfaces. In some embodiments, although a match is made between a content element and a surface, the match can not be a perfect match. For example, a content element can require a 2D area of 1000x500. However, the best matching surface can have dimensions of 900x450. In one example, the rendering module 146 can render the 1000x500 content element to fit best on the 900x450 surface, which can include, for example, scaling the content element while keeping the aspect ratio constant. In another example, the rendering module 146 can crop the 1000x500 content element to fit within the 900x450 surface.
[0105] In some embodiments, the device implementing the system 100 can move. For example, the device implementing the system 100 can move from a first room to a second room.
[0106] In some embodiments, content elements that matched surfaces (real and / or virtual) in a first room can remain matched to surfaces in the first room. For example, the device implementing the system 100 can move from a first room to a second room, and content elements that matched surfaces in the first room will not match surfaces in the second room and thus will not be rendered on them. If the device then moves from the second room to the first room, the content elements that matched surfaces in the first room will be rendered to / cast on the corresponding surfaces in the first room. In some embodiments, the content will continue to be rendered in the first room, although not displayed due to being outside the field of view of the device, but certain features will continue to run, such as audio playback or play time, so that when the device returns to having the matching content in the field of view, the rendering will seamlessly resume (a similar effect as if the user left the room and a movie was playing on a traditional television).
[0107] In some embodiments, content elements that matched surfaces in a first room can match surfaces in a second room. For example, the device implementing the system 100 can move from a first room to a second room, and after the device is in the second room, the environment structuring process 160 and the compositing process 140 can occur / run / execute, and content elements can match surfaces (real and / or virtual) in the second room.
[0108] In some embodiments, some content elements that match surfaces in the first room can remain in the first room, while other content elements that match surfaces in the first room can move to the second room. For example, a first set of content elements that match surfaces in the first room can remain matching surfaces in the first room, while a second set of content elements that match surfaces in the first room can move with the device implementing the system 100 to the second room. The second set of content elements can move with the device as the device enters the second room from the first room. Whether a content element is in the first set or the second set can be determined based on attributes of the content element, attributes of one or more surfaces in the first room that the content element matches, user preferences, system preferences, and / or passable world preferences. The basis for these various cases is that matching and rendering can be exclusive; content can match a surface but not be rendered. Since the user device does not need to constantly match surfaces, this can save computational cycles and power, and selective rendering can reduce latency in re-viewing content at a matched surface.
[0109] Figure 1B An example system and computer-implemented method of matching content elements of content to a spatial 3D environment is shown in accordance with some embodiments. The system 105 includes a content structuring process 120, an environment structuring process 160, and a synthesis process 140 similar to Figure 1A Figure 1B The synthesis process 140 includes additional modules including a display module 148 and a receive module 150.
[0110] As described above, the matching module 142 can generate matching results, wherein a content element can be matched / mapped to multiple surfaces, such that the content element can be rendered and displayed on any one of the multiple surfaces. The display module 148 displays the content element, its outline, or a reduced-resolution version thereof (all referred to herein as “candidate views”) on multiple surfaces or in multiple portions of a single surface. In some embodiments, the multiple surfaces are displayed sequentially, such that the user sees only a single candidate view at a time and can cycle through or scroll through additional candidate view options one after another. In some embodiments, all candidate views are displayed simultaneously, and the user selects a single candidate view (e.g., via voice command, input to a hardware interface, eye tracking, etc.). The receiving module 150 receives from the user the selection of a candidate view on one of the multiple surfaces. The selected candidate view may be referred to as the preferred surface. Preference surfaces can be stored in storage module 152 as user preferences or passable preferences, so that future matches when matching content elements with surfaces can benefit from these preferences, as indicated by information flow 156 from receiving module 150 to matching module 142 or information flow 154 from storage module 152 to matching model 142. Information flow 156 can be an iterative process, such that, according to some embodiments, after several iterations, user preferences can begin to take precedence over system and / or passable preferences. In contrast, information flow 154 can be a fixed output, such that matching priority is always given to the immediate user or other users who are entering or wishing to display content in the same environment. System and / or passable preferences can take precedence over user preferences, but as more information flows 156 continue to influence user preferences... Figure 1A System 100 or Figure 1B With the use of system 105, user preferences may begin to be system-preferred through a natural learning process algorithm. Therefore, in some embodiments, regardless of the availability of other surfaces or the environmental input that would otherwise cause matching module 142 to place content elements elsewhere, content elements will be rendered / displayed on the preferred surface. Similarly, information flow 154 can dictate the matching of a second user to the preferred surface, which is never in the environment and has not yet been established for the first user's preferred surface.
[0111] Advantageously, since the sensor data and virtual model are typically stored in short term computer memory, if there is a device shutdown between content placement sessions, the persistent storage module of the preferred surface can speed up the loop synthesis process 140. For example, if the sensors 162 collect depth information to create a virtual mesh reconstruction through environment structuring 160 to match content in a first session, and the system is shut down emptying the random access memory storing that environmental data, the system will have to repeat the environment structuring pipeline upon restart for the next matching session. However, by storing module 152 updating the match 142 with the preferred surface information saves computational resources without the need for a full iteration of the environment structuring process 160.
[0112] In some embodiments, the content element can be displayed on all three of the front surfaces at once. The user can then select one of the three front surfaces as the preferred surface. In some embodiments, the content element can be displayed on only one of the three front surfaces at a time, and the user can be indicated that the content element can be displayed on two other surfaces. The user can then browse the other surface options, and as the user activates each of the surface options, the content element can be displayed on the activated surface. The user can then select one of the surface options as the preferred surface.
[0113] Figures 2A-2E The content element is depicted as matching with three possible locations within the user's physical environment 1105 (e.g., a computer's monitor, a smartphone, a tablet, a television, a web browser, a screen, etc.). Figure 2A The content element is shown as matched / mapped to the three possible locations as indicated by the view location suggestions 214. The three white dots displayed on the left-hand side of the view location suggestions 214 indicate that there can be three display locations. The fourth white dot with an "x" can be a close button to close the view location suggestions 214 and indicate a selection of a preferred display location based on the selected / highlighted display location when the user selects the "x." The display location 212a is a first option for displaying the content element as indicated by the first white dot highlighted among the three white dots on the left-hand side. Figure 2B The same user environment 1105 is shown, where the display location 212b is a second option for displaying the content element as indicated by the second white dot highlighted among the three white dots on the left-hand side. Figure 2C The same user environment 1105 is shown, where the display location 212c is a third option for displaying the content element as indicated by the third white dot highlighted among the three white dots on the left-hand side. Those of ordinary skill in the art can appreciate that there can be other ways to show the display options for the user to select, and Figures 2A-2CThe example shown in FIG. 12 is just one example. For example, another approach can be to display all display options at once and let the user select the preferred option using the VR / AR device (e.g., a controller, by gaze, etc.).
[0114] It will be appreciated that AR systems have certain fields of view in which virtual content can be projected, and that such fields of view are typically less than the full field of view potential of a human. Humans can typically have a natural field of view between 110 and 120 degrees, and in some embodiments, as shown in FIG. 12, the display field of view of the AR system 224 is less than this potential, meaning that surface candidate 212c can be within the natural field of view of the user, but outside the field of view of the device (e.g., the system is capable of rendering content on this surface, but in practice will not display the content). In some embodiments, a field attribute (attribute described further below) is assigned to surfaces to indicate whether the surface is capable of supporting content displayed to the device's field of display. In some embodiments, as noted above, surfaces outside the field of display are not presented to the user for display options. Figure 2D
[0115] In some embodiments, as shown in FIG. 13, a default location 212 is shown as a virtual surface in front of the user at a prescribed distance, such as at the specific focal length specification of the device's display system. The user can then adjust the default location 212 to a desired location in the environment, such as a registered location from storage device 285 or a matching surface from synthesis 140, such as through head pose or hand gestures measured by sensors 162 or other input means. In some embodiments, the default location can remain fixed relative to the user, such that as the user moves in the environment, the default location 212 remains in substantially the same portion of the user's field of view (same as the device's field of display in this embodiment). Figure 2E
[0116] Figure 2E A virtual television (TV) at a default location 212 is also shown by way of example, which default location has three TV application previews (e.g., TV App 1, TV App 2, TV App 3) associated with the virtual TV. The three TV applications can correspond to different TV channels or different TV applications correspond to different TV channels / TV content providers. A user can extract a single channel for TV play by selecting the corresponding TV application / channel displayed below the virtual TV. The user can extract a TV channel by: (a) using totem to align to the channel; (b) pressing a trigger on the totem to select the channel and holding the trigger for a period of time (e.g., about 1 second); (c) moving the totem around to identify a desired location in the environment for displaying the extracted TV channel; and (d) pressing the trigger on the totem to place the extracted TV channel at the desired location in the environment. Selecting virtual content is further described in U.S. Patent Application 15 / 296,869, which claims priority to October 20, 2015, entitled "SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE," each of which is hereby incorporated by reference herein.
[0117] The desired location can be a surface suitable for displaying a TV channel, or other surface identified in accordance with the teachings of the present disclosure. In some embodiments, a new prism can be created at the desired location, with the selected channel content loaded and displayed within the new prism. Further details regarding totems are described in U.S. Patent No. 9,671,566, entitled "PLANAR WAVEGUIDE APPARATUS WITH DIFFRACTION ELEMENT(S) AND SYSTEM EMPLOYING SAME," which is hereby incorporated by reference herein in its entirety. In some embodiments, the three TV applications can be "channel previews," which are peepholes for viewing content that is playing on individual channels by displaying a dynamic or static depiction of the channel content. In some embodiments, (c) while moving the totem to identify a desired location in the environment for displaying the extracted TV channel, a visual can be displayed to the user. The visual can be, for example, a single image illustrating the channel, one or more images illustrating a preview of the channel, a video stream illustrating current content of the channel, etc. The video stream can be, for example, low resolution or high resolution, and can vary as a function of available resources and / or bandwidth (in resolution, frame rate, etc.).
[0118] Figures 2A-2EDifferent display options are shown to display content (e.g., as elements of virtual content) within the original field of view of the user and / or device (e.g., based on a particular head pose). In some embodiments, the field of view of the user and / or device can change (e.g., the user moves their head from one field of view to another). As a result of changing the field of view, additional surface options for displaying content can be made available to the user based at least in part on the change in the field of view of the user (e.g., the change in head pose). The additional surface options for displaying content can also be made available based at least in part on other surfaces that were not initially available in the original field of view of the user and / or device but are now visible to the user based on the change in the field of view of the user. Thus, Figures 2A-2D The view position option 214 can also depict additional options for displaying content. For example, Figures 2A-2D Three display options are depicted. As the field of view of the user changes, more display options can be available, which can result in the view position option 214 displaying more dots to indicate the additional display options. Likewise, if the new field of view has fewer surface options, the view position option 214 can display fewer than 3 dots to indicate the multiple display options available for displaying content in the new field of view. Thus, one or more additional surface options for displaying content can be displayed to the user for selection based on the changing field of view of the user, which corresponds to a change in head pose of the user and / or device.
[0119] In some embodiments, a user and / or device can have a first field of view. The first field of view can be used to generate surface options for displaying content thereon. For example, three surface options in the first field of view can be available for displaying content elements thereon. The user can then change their field of view from the first field of view to a second field of view. The second field of view can then be used to generate additional surface options for displaying content thereon. For example, two surface options in the second field of view can be available for displaying content elements thereon. Between the surfaces in the first field of view and the surfaces in the second field of view, there can be a total of five surface options. The five surface options can be displayed to the user as view position suggestions. If the user is viewing in the second field of view and selects a view position suggestion in the first field of view, the user can receive an indication (e.g., an arrow, a glow, etc.) in the direction of the first field of view that indicates to the user that they should navigate back in the indicated direction to the first field of view to view the selected surface option / view position.
[0120] In some embodiments, a user can be viewing content displayed on a first surface in a first field of view. The first field of view can have an associated first head pose. If the user changes their field of view from the first field of view to a second field of view, after a period of time, the system can provide the user with an option to change the display location of the content from the first surface in the first field of view to one or more surface options in the second field of view. The second field of view can have an associated second head pose. In some embodiments, once the user’s field of view changes from the first field of view and thus the first head pose of the user and / or device to the second field of view and thus the second head pose of the user and / or device, the system can immediately provide the user with the option to move the content, the change in position of the first head pose and the second head pose being greater than a head pose change threshold. In some embodiments, a time threshold (e.g., 5 seconds) that the user remains with the second field of view and thus the second head pose can determine whether the system provides the user with the option to change the display location of the content. In some embodiments, the change in field of view can be a slight change, such as less than a corresponding head pose change threshold (e.g., less than 90 degrees in any direction relative to the first field of view and thus the first head pose) to trigger the system to provide the option to change the display location of the content. In some embodiments, the change in head pose can be greater than a head pose change threshold (e.g., greater than 90 degrees in any direction) before the system provides the user with the option to change the display location of the content. Thus, the one or more additional surface options for displaying the content based on the change in field of view can be displayed based at least in part on a time threshold corresponding to the change in field of view. In some embodiments, the one or more additional surface options for displaying the content based on the user’s change in field of view can be displayed based at least in part on a head pose change threshold.
[0121] In some embodiments, the system can render / display content on one or more first surfaces with a user viewing the content having a first head pose. The user viewing the content can change the head pose of themselves and / or the device from the first head pose to a second head pose. In response to the change in head pose, the system can render / display the content on one or more second surfaces with the user viewing the content having the second head pose. In some embodiments, the system can provide the user with an option to change the rendering / display location of the content from the one or more first surfaces to the one or more second surfaces. In some embodiments, the system can provide the user with the option to move the content as soon as the head pose of the user has changed from the first head pose to the second head pose. In some embodiments, the system can provide the user with the option to change the rendering / display location of the content if the change in head pose is greater than a corresponding head pose change threshold (e.g., 90 degrees). In some embodiments, the system can provide the user with the option to change the rendering / display location of the content if the change in head pose remains for a threshold period of time (e.g., 5 seconds). In some embodiments, the change in head pose can be a slight change, such as less than a corresponding head pose change threshold (e.g., less than 90 degrees), to trigger the system to provide the option to change the rendering / display location of the content.
[0122] Attribute
[0123] General Properties
[0124] As described above, the parser 115 can identify / determine and store properties of each content element, while the environment parser 168 can determine and store properties of each surface. The properties of the content elements can be explicitly indicated by a content designer of the content 110, or can be determined or otherwise inferred by the parser 115. The properties of the surfaces can be determined by the environment parser 168.
[0125] The properties that both content elements and surfaces can have include, for example, orientation, aspect ratio, size, area (e.g., magnitude), relative viewing position, color, contrast, legibility indicator, and / or time. Further details regarding these properties are provided below. One of ordinary skill in the art can appreciate that content elements and surfaces can have additional properties.
[0126] For content elements and surfaces, the orientation property indicates an orientation. The orientation value can include vertical, horizontal, and / or a specific angle (e.g., horizontal at 0 degrees, vertical at 90 degrees, or any position between 0-90 degrees for an angled orientation). The specific angle orientation property can be specified / determined in degrees or radiant, or can be specified / determined relative to the x-axis or y-axis. In some embodiments, a tilted surface can be defined for displaying, for example, a water flow of content flowing at a tilt angle, to display different art pieces. In some embodiments, for content elements, the applied navigation bar can be defined as a horizontal orientation, but tilted at a specific angle.
[0127] For content elements and surfaces, the aspect ratio property indicates an aspect ratio. The aspect ratio property can be specified as, for example, a 4:3 or 16:9 ratio. The content element can be scaled based on the aspect ratio property of the content element and the corresponding one or more surfaces. In some embodiments, the system can determine the aspect ratio of a content element (e.g., a video) based on other properties of the content element (e.g., size and / or area), and scale the content element based on the determined aspect ratio. In some embodiments, the system can determine the aspect ratio of a surface based on other properties of the surface.
[0128] Within the aspect ratio property, a content designer of a content element can use a specific characteristic to recommend a specific aspect ratio for maintaining or changing the content element. In one example, if the specific characteristic is set to “maintain” or a similar keyword or phrase, the aspect ratio of the content element will be maintained (i.e., not changed). In one example, if the specific property is set to, for example, “free” or a similar keyword or phrase, the aspect ratio of the content element can be changed (e.g., scaled or otherwise) to, for example, match the aspect ratio of the one or more surfaces to which the content element is matched. The default value of the aspect ratio property can maintain the original aspect ratio of the content element, and the default value of the aspect ratio property can be overwritten if the content designer specifies some other value or keyword for the aspect ratio property of the content element, and / or if the system determines that the aspect ratio property should be overwritten to better match the content element to the one or more surfaces.
[0129] For content elements and surfaces, the size property indicates a size. The size property of a content element can indicate the size of the content element as a function of pixels (e.g., 800 pixels by 600 pixels). The size property of a surface can indicate the size of the surface as a function of meters or any other unit of measurement (e.g., 0.8 meters by 0.6 meters). The size property of a surface can indicate a measurable range of the surface, where the measurable range can include length, width, depth, and / or height. For content elements, a content designer can specify the size property to suggest a certain shape and outer dimensions of a surface on which the content element is to be displayed.
[0130] For content elements and surfaces, the area attribute indicates an area or size. The area attribute of a content element can indicate the area of the content element as a function of pixels (e.g., 480,000 square pixels). The area attribute of a surface can indicate the area of the surface as a function of meters or any other unit of measurement (e.g.,.48 square meters). For surfaces, the area can be a perceived area as perceived by a user or an absolute area. The perceived area can be defined by increasing the absolute area as well as the angle and distance of the displayed content element from the user such that when the content element is displayed closer to the user, the content element is perceived as a smaller size, and when the content element is farther from the user, the content element can be scaled up accordingly so that the user still perceives it as having the same particular size, and vice versa when the content element is closer to the user. The absolute area can simply be defined by, for example, square meters, regardless of the distance of the displayed content element in the environment.
[0131] For content elements and surfaces, the relative viewing position attribute relates to a position relative to a head pose vector of the user. The head pose vector can be a combination of the position and orientation of a head-mounted device worn by the user. The position can be a fixed point of the device worn on the user’s head that is tracked in a real-world coordinate system using information received from an environment and / or user sensing system. The orientation component of the head pose vector of the user can be defined by a relationship between a three-dimensional device coordinate system local to the head-mounted device and a three-dimensional real-world coordinate system. The device coordinate system can be defined by three orthogonal directions: a forward viewing direction that approximates the forward line of sight of the user through the device, an upright direction of the device, and a right direction of the device. Other reference directions can also be chosen. Information obtained by sensors in the environment and / or user sensing system can be used to determine the orientation of the local coordinate system relative to the real-world coordinate system.
[0132] To further illustrate the device coordinate system, if a user is wearing the device and is upside down, the upright direction of the user and the device is actually pointing in the direction of the ground (e.g., the downward direction of gravity). However, from the perspective of the user, the relative upright direction of the device is still aligned with the upright direction of the user; for example, if the user is reading a book in the typical top-to-bottom, left-to-right manner while upside down, another user who is standing normally but not upside down would see the user holding the book upside down in the real-world coordinate system, but relative to the local device coordinate system that approximates the angle of the user, the book is oriented upright.
[0133] For a content element, the relative viewing position attribute can indicate at what position relative to the head pose vector the content element should be displayed. For a surface, the relative viewing position attribute can indicate the position of the surface in the environment relative to the head pose vector of the user. It will be appreciated that component vectors of the head pose vector, such as the forward viewing direction vector, can also be used as a criterion for determining the relative viewing position attribute of a surface and / or for determining the relative viewing position attribute of a content element. For example, a content designer can indicate that a content element, such as a search bar, should always be within 30 degrees of the left or right side of the user’s head pose vector, and if the user moves more than 30 degrees to the left or right, the search bar should be adjusted so that it is still within 30 degrees of the left or right side of the user’s head pose vector. In some embodiments, the content is adjusted on the fly. In some embodiments, the content is adjusted once a time threshold has been met. For example, the user moves 30 degrees to the left or right and the search bar should be adjusted after a 5 second time threshold has elapsed.
[0134] A relative viewing angle can be specified to maintain an angle or range of angles relative to the head pose vector of the user. For example, a content element such as a video can have a relative viewing position attribute that indicates that the video should be displayed on a surface that is approximately orthogonal to the forward viewing vector of the user. If the user stands in front of a surface, such as a wall, looking directly forward, the relative viewing position attribute of the wall relative to the forward viewing vector of the user can satisfy the relative viewing position attribute requirement of the content element. However, if the user looks down at the floor, the relative viewing position attribute of the wall changes and the relative viewing position attribute of the floor better satisfies the relative viewing position attribute requirement of the content element. In this case, the content element can be moved so that it is projected on the floor instead of the wall. In some embodiments, the relative viewing position attribute can be a depth or distance from the user. In some embodiments, the relative viewing position attribute can be a relative position relative to the current viewing position of the user.
[0135] For content elements and surfaces, a color attribute indicates a color. For a content element, the color attribute can indicate one or more colors, whether the color can be changed, opacity, etc. For a surface, the color attribute can indicate one or more colors, a color gradient, etc. The color attribute can be associated with the readability of the content element and / or how the content element is perceived on the surface. In some embodiments, a content designer can define the color of a content element to be, for example, white or light. In some embodiments, a content designer can not want the system to change the color of a content element (e.g., a company logo). In these embodiments, the system can change the background of one or more surfaces on which the content element is displayed to create the necessary contrast for readability.
[0136] For content elements and surfaces, a contrast property indicates contrast. For content elements, the contrast property can indicate the current contrast, whether the contrast can be changed, instructions about how the contrast can be changed, etc. For surfaces, the contrast property can indicate the current contrast. A contrast preference property can be associated with the readability and / or how a content element will be perceived on a surface. In some embodiments, a content designer can want to display a content element with high contrast relative to the background of a surface. For example, a version of a content element can be rendered as white text on a black background in a webpage on a monitor of a computer, smartphone, tablet, etc. A white wall can be matched to display a text content element that is also white. In some embodiments, the system can change the text content element to a darker color (e.g., black) to provide contrast that satisfies the contrast property.
[0137] In some embodiments, the system can change the background color of a surface without changing a content element to provide a color and / or contrast that satisfies the color and / or contrast properties. The system can change the background color of a (real) surface by creating a virtual surface at the location of the (real) surface, where the color of the virtual surface is the desired background color. For example, if the color of a logo should not be changed, the system can provide sufficient contrast by changing the background color of the surface to provide a color contrast that satisfies the color and / or contrast preference properties while preserving the logo.
[0138] For content elements and surfaces, a readability indicator property can indicate a readability metric. For content elements, the readability indicator property indicates a readability metric that should be maintained for the content element. For content elements, the system can use the readability indicator property to determine the priority of other properties. For example, if the readability indicator for a content element is“high,” the system can set the priority of these properties to“high.” In some examples, even if a content element is in focus and there is sufficient contrast, the system can scale the content element based on the readability indicator property to ensure that the readability metric is maintained. In some embodiments, if the priority of a particular content element is set to“high,” a high readability indicator property value for the particular content element can take precedence or be prioritized over other explicit properties for other content elements. For surfaces, the readability indicator property can indicate how a user will perceive a content element containing text if displayed on the surface.
[0139] Text legibility is a difficult problem for pure VR environments. The problem becomes more complex in AR environments because real-world colors, brightness, lighting, reflections, and other capabilities directly impact a user's ability to read text rendered by an AR device. For example, web content rendered by a web browser can be primarily text driven. As an example, a set of JavaScript APIs (e.g., through new extensions of the current W3C Camera API) can provide content designers with the current world palette and a contrast alternative palette for font and background colors. A set of JavaScript APIs can provide content designers with the unique ability to adjust web content color palettes according to real-world color patterns to improve content contrast and text legibility (e.g., readability). Content designers can use this information by setting font colors to provide better legibility for web content. These APIs can be used to track this information in real-time, so a web page can adjust its contrast and color palettes according to the ambient light changes. For example, Figure 3A shows web content 313 adjusted to be legible with respect to the light and color conditions of a dark real-world environment by at least adjusting the text of web content 313 that is to be displayed in a light color mode. Figure 3A As shown, the text in web content 313 has a light color, while the background in web content 313 has a dark color. Figure 3B shows web content 315 adjusted to be legible with respect to the light and color conditions of a bright real-world environment by at least adjusting the text of web content 315 that is to be displayed in a dark color mode. Figure 3B As shown, the text in web content 315 has a dark color, while the background in web content 313 has a light color. One of ordinary skill in the art can appreciate that other factors can also be adjusted, such as the background color of web content 313 (e.g., a darker background and lighter text) or web content 315 (e.g., a lighter background and darker text) to provide contrast in colors such that the text is more legible based at least in part on the light and color conditions of the real-world environment.
[0140] For a content element, the time attribute indicates how long the content element should be displayed. The time attribute can be short (e.g., less than 5 seconds), medium (e.g., between 5 and 30 seconds), long (e.g., greater than 30 seconds). In some embodiments, the time attribute can be infinite. If the time attribute is infinite, the content element can remain until it is removed and / or another content element is loaded. In some embodiments, the time attribute can be a function of input. In one example, if the content element is an article, the time attribute can be a function of input that indicates that the user has reached the end of the article and stayed there for a threshold period of time. In one example, if the content element is a video, the time attribute can be a function of input that indicates that the user has reached the end of the video.
[0141] For a surface, the time attribute indicates how long the surface will be available. The time attribute can be short (e.g., less than 5 seconds), medium (e.g., between 5 and 30 seconds), long (e.g., greater than 30 seconds). In some embodiments, the time attribute can be infinite. In some embodiments, the time attribute can be a function of sensor input, e.g., from sensors 162. Sensor input from sensors 162 (e.g., from IMUs, accelerometers, gyroscopes, etc.) can be used to predict the availability of a surface relative to the field of view of the device. In one example, if the user is walking, surfaces near the user can have a short time attribute, surfaces somewhat farther from the user can have a medium time attribute, and surfaces farther away can have a long time attribute. In one example, if the user is sitting idly on a couch, the wall in front of the user can have an infinite time attribute until a change in data greater than a threshold is received from sensors 162, after which the time attribute of the wall in front of the user can change from infinite to another value.
[0142] Content element attributes
[0143] A content element can have attributes specific to the content element, such as a priority, surface type, location type, edge, content type, and / or focus attribute. Further details regarding these attributes are provided below. One of ordinary skill in the art can appreciate that a content element can have additional attributes.
[0144] The priority attribute indicates a priority value for a content element (e.g., a video, a picture, or text). The priority value can include a high, medium, or low priority, a numerical range from 0 to 100, and / or an indicator of required or not required. In some embodiments, a priority value can be specified for the content element itself. In some embodiments, a priority value can be specified for a particular attribute. For example, a legibility indicator attribute of a content element can be set to high, indicating that the content designer has placed an emphasis on the legibility of the content element.
[0145] The attribute of surface type or "surface type" attribute indicates the type of surface that the content element should match. The surface type attribute can be based on semantics, such as whether certain content elements should be placed in certain locations and / or on certain surfaces. In some examples, a content designer can suggest that a particular content element not be displayed on a window or a painting. In some examples, a content designer can suggest that a particular content element always be displayed substantially on the largest vertical surface in front of the user.
[0146] The position type attribute indicates the position of the content element. The position type attribute can be dynamic or fixed. A dynamic position type can assume, for example, that the content element is fixed on the user's hand such that when the user's hand moves, the content element moves dynamically with the user's hand. A fixed position type assumes, for example, that the content element is fixed relative to a surface, environment, or a particular position in the virtual world relative to the user's body or head / view position, examples of which are described in detail below.
[0147] The term "fixed" can also have different levels, such as: (a) world fixed, (b) object / surface fixed, (c) body fixed, and (d) head fixed. For (a) world fixed, the content element is fixed relative to the world. For example, if the user walks around in the world, the content element does not move, but is fixed at a position relative to the world. For (b) object / surface fixed, the content element is fixed to an object or surface such that if the object or surface moves, the content element moves with the object or surface. For example, the content element can be fixed to a notepad held by the user. In this case, the content is fixed to the object of the notepad surface and moves accordingly with the notepad. For (c) body fixed, the content element is fixed relative to the user's body. If the user moves their body, the content element will move with the user to maintain the fixed position relative to the user's body. For (d) head fixed, the content element is fixed relative to the user's head or pose. If the user rotates their head, the content element will move relative to the user's head motion. Likewise, if the user walks, the content element will also move relative to the user's head.
[0148] An edge (or padding) attribute indicates an edge around a content element. The edge attribute is a layout attribute that describes the placement of a content element relative to other content elements. For example, the edge attribute represents a distance from a content element's boundary to the nearest permissible boundary of another content element. In some embodiments, the distance is an edge based on x, y, z coordinates and can be measured from a vertex of the content element's boundary or other specified location; in some embodiments, the distance is an edge based on polar coordinates and can be measured from the center of the content element or other specified location, such as a vertex of the content element. In some embodiments, the edge attribute defines a distance from the content element to the actual content inside the content element. In some embodiments, such as for exploded content elements, the edge attribute represents how much of an edge to maintain relative to the surface that the exploded content element is matched to, such that the edge acts as an offset between the content element and the matched surface. In some embodiments, the edge attribute can be extracted from the content element itself.
[0149] A content type attribute or "content type" attribute indicates a type of the content element. The content type can include a reference and / or link to the corresponding media. For example, the content type attribute can designate the content element as an image, a video, a music file, a text file, a video image, a 3D image, a 3D model, container content (e.g., can be any content wrapped inside a container), an advertisement, and / or a rendering canvas defined by a content designer (e.g., a 2D canvas or a 3D canvas). The rendering canvas defined by a content designer can include, for example, a game, a rendering, a map, a data visualization, etc. The advertisement content type can include attributes that define what sound or advertisement should be presented to the user when the user focuses on or is located near a particular content element. The advertisement can be: (a) a sound, such as a ding, (b) a visual, such as a video / image / text, and / or (c) a haptic indicator, such as a vibration in a user controller or headphones, etc.
[0150] A focus attribute indicates whether a content element should be in focus. In some embodiments, the focus can be a function of the distance from the user to the surface on which the content element is displayed. If the focus attribute of a content element is set to always be in focus, then the system will put the content element in focus no matter how far the user is from the content element. If the focus attribute of a content element is not specified, then the system can put the content out of focus when the user is a certain distance from the content element. This can depend on other attributes of the content element, such as the size attribute, the area attribute, the relative viewing position attribute, etc.
[0151] A surface attribute
[0152] A surface can have surface-specific properties, e.g., surface profile, texture, and / or occupancy properties. Further details regarding these properties are provided below. One of ordinary skill in the art can appreciate that a surface can have additional properties.
[0153] In some embodiments, the environment parser 168 can determine surface profile properties (and associated properties), such as a surface normal vector, an orientation vector, and / or an upright vector of one and / or all surfaces. In a 3D case, the surface normal or just normal of a surface at a point P is a vector that is perpendicular to the tangent plane of the surface at point P. The term "normal" can also be used as an adjective; a line normal to a plane, the normal component of a force, a normal vector, etc.
[0154] At least one component of the surface normal vector and the head pose vector of the surface of the environment surrounding the user discussed above can be important to the matching module 142, as although certain properties of a surface (e.g., size, texture, aspect ratio, etc.) can be ideal for displaying certain content elements (e.g., video, three-dimensional models, text, etc.), the positioning of such surface relative to the corresponding surface normal of the user's line of sight approximated by at least one component vector of the user's head pose vector can be poor. By comparing the surface normal vector to the user's head pose vector, surfaces that can otherwise be suitable for the displayed content can be disqualified or filtered.
[0155] For example, a surface's surface normal vector can be in substantially the same direction as the user's head pose vector. This means that the user and the surface are facing in the same direction, rather than facing each other. For example, if the user's forward direction is north, a surface with a normal vector pointing north either faces the user's back or the user faces the back of the surface. If the user cannot see the surface because it is facing away from the user, that particular surface would not be the best surface to display content on, despite potentially having additional beneficial properties values of the surface.
[0156] A comparison between the device's forward viewing vector (approximating the user's forward viewing direction) and the surface normal vector can provide a numerical value. For example, a dot product function can be used to compare the two vectors and determine a numerical relationship that describes the relative angle between the two vectors. Such a calculation can result in a number between 1 and -1, where more negative values correspond to a more favorable relative angle for viewing because the surface is closer to being orthogonal to the user's forward viewing direction, enabling the user to comfortably view virtual content placed on the surface. Thus, based on the identified surface normal vector, a feature for good surface selection can relate to the user's head pose vector or components thereof, such that the content should be displayed on a surface that faces the user's forward viewing vector. It will be appreciated that constraints can be placed on the acceptable relationship between the head pose vector components and the surface normal components. For example, all surfaces that will result in a negative dot product with the user's forward viewing vector can be selected for content display. Depending on the content, content provider or algorithm or user preferences that affect the acceptable range can be considered. In the case of a video that needs to be displayed substantially perpendicular to the user's forward direction, a smaller range of dot product outputs can be allowed. Those skilled in the art will appreciate that many design options are possible depending on other surface properties, user preferences, content properties, etc.
[0157] In some embodiments, a surface can be very suitable according to size and location and head pose perspective, but the surface can not be a good option for selection because the surface can include attributes such as texture attributes and / or occupied attributes. Texture attributes can include materials and / or designs that can change a clean and clear surface for a simple appearance for presentation to an undesirable cluttered surface for presentation. For example, a brick wall can have a large blank area that is ideal for displaying content. However, due to a red stacked brick in the brick wall, the system can consider the brick wall undesirable for displaying content directly on it. This is because the texture of the surface has a roughness variation between the bricks and mortar and has a non-neutral red color that can cause a contrast enhancement of the content. Another undesirable texture example can include a surface with a wallpaper design that not only has a background design pattern and color, but also includes imperfections such as bubbles or uneven coating, resulting in surface roughness variations. In some embodiments, the wallpaper design can include many patterns and / or colors such that displaying content directly on the wallpaper can not display the content in a favorable view. The occupied attribute can indicate that the surface is currently occupied by another content, such that displaying additional content on a particular surface with a value indicating that the surface is occupied can cause the new content to not be displayed on top of the occupied content, and vice versa. In some embodiments, the occupied attribute represents the presence of a small real-world defect or an object occupying the surface. Such occupying real-world objects can include items such as cracks or nails that can be indistinguishable to a depth sensor within the sensor suite 162 but are perceptible to a camera within the sensors 162 as a negligible surface area that is different from the surface. Other occupying real-world objects can include a picture or a poster with a surface disposed thereon hanging from a wall with low texture variation, and certain sensors 162 cannot distinguish it as different from the surface, but a camera of 162 can recognize and the occupied attribute updates the surface accordingly to exclude the system from making a determination that the surface is an "empty canvas."
[0158] In some embodiments, content can be displayed on a virtual surface whose relative position is related to the (real) surface. For example, if the texture attribute indicates that the surface is not simple / clean and / or the occupied attribute indicates that the surface is occupied, the content can be displayed on a virtual surface in front of the (real) surface, e.g., within the edge attribute tolerance range. In some embodiments, the edge attribute of a content element is a function of the texture attribute and / or the occupied attribute of the surface.
[0159] Flow
[0160] Matching content elements to surfaces (advanced)
[0161] Figure 4This is a flowchart illustrating a method for matching content elements to surfaces according to some embodiments. The method includes: receiving content at 410; identifying content elements in the content at 420; determining a surface at 430; matching the content elements to the surface at 440; and rendering the content elements as virtual content onto the matched surface at 450. A parser 115 receives content 110 at 410. The parser 115 identifies content elements in content 110 at 420. The parser 115 may identify / determine and store attributes of each content element. An environment parser 168 determines surfaces in the environment at 430. The environment parser 168 may determine and store attributes of each surface. In some embodiments, the environment parser 168 continuously determines surfaces in the environment at 430. In some embodiments, the environment parser 168 determines surfaces in the environment at 430 while the parser 115 receives content 110 at 410 and / or identifies content elements in content 110 at 420. A matching module 142 matches content elements to surfaces at 440 based on the attributes of the content elements and the attributes of the surfaces. Rendering module 146 renders the content element 450 onto its matching surface. Storage module 152 registers the surface for future use, such as by user specification, so that content elements can be placed on that surface in the future. In some embodiments, storage module 152 may be within a perceptual frame 166.
[0162] Identify content elements in the content
[0163] Figure 5 This is a flowchart illustrating a method for identifying content elements in content according to some embodiments. Figure 5 This is a disclosure based on some embodiments. Figure 4 The detailed process for identifying elements within the content at position 420. This method includes identifying content elements within the content at position 510, similar to... Figure 4identifies elements in the content at 420. The method proceeds to the next step 520 of identifying / determining attributes. For example, attributes can be identified / determined from tags associated with the placement of the content. For example, a content designer, when designing and configuring the content, can use attributes (as described above) to define where and how to display content elements. The attributes can be related to placement of the content elements in a particular location relative to each other. In some embodiments, the step 520 of identifying / determining attributes can include inferring attributes. For example, attributes of each content element can be determined or inferred based on placement of the content elements relative to each other within the content. At 530, extracting hints / tags from each content element is performed. The hints or tags can be formatting hints or formatting tags provided by a content designer of the content. At 540, finding / searching for alternative display forms of the content elements is performed. Certain formatting rules can be specified for content elements displayed on a particular viewing device. For example, certain formatting rules can be specified for an image on a webpage. The system can access the alternative display forms. At 550, storing the identified content elements is performed. The method can store the identified elements into a non-transitory storage medium for use in the composition process 140 to match content elements to surfaces. In some embodiments, the content elements can be stored in a transitory storage medium.
[0164] Determining surfaces in an environment
[0165] Figure 6 is a flowchart showing a method for determining surfaces from a user's environment according to some embodiments. Figure 6 is an example detailed flow of determining surfaces at 430. Figure 4 Figure 6 Determination of surfaces at 610. Determination of surfaces at 610 can include collection of depth information of the environment from the depth sensor of sensor 162 and performing reconstruction and / or surface analysis. In some embodiments, sensor 162 provides a map of points and system 100 reconstructs a series of connected vertices between the points to create a virtual grid representative of the environment. In some embodiments, plane extraction or analysis is performed to determine grid characteristics that indicate what a common surface or interpreted surface can be (e.g., a wall, a ceiling, etc.). The method proceeds to the next step 620 of determining the pose of the user, which can include determining a head pose vector from sensor 162. In some embodiments, sensor 162 collects inertial measurement unit (IMU) data to determine the rotation of the device on the user; in some embodiments, sensor 162 collects camera images to determine the position of the device on the user relative to the real world. In some embodiments, the head pose vector is derived from one or both of the IMU and camera image data. Determination of the pose of the user at 620 is an important step in identifying surfaces, as the pose of the user will provide the user’s perspective relative to the surfaces. At 630, the method determines the properties of the surfaces. Each surface is labeled and categorized with a corresponding property. This information will be used when matching content elements to surfaces. In some embodiments, sensor 162 provides raw data from FIG. 1 to CVPU 164 for processing, and CVPU 164 provides processed data to perception framework 166 to prepare data for environment resolver 168. Environment resolver 168 resolves the environment data from perception framework 166 to determine the surfaces in the environment and corresponding properties. At 640, the method stores a list of the surfaces to a non-transitory storage medium for use by the synthesis process / matching / mapping routine to match / map extracted elements to particular surfaces. The non-transitory storage medium can include a data storage device. The determined surfaces can be stored in a particular table, such as the table disclosed in Figure 15 FIG. 1. In some embodiments, the identified surfaces can be stored in a transitory storage medium. In some embodiments, the storing at 640 includes designating the surfaces as preferred surfaces for future matching of content elements.
[0166] Matching content elements to surfaces (particular)
[0167] Figures 7A-7B is a flowchart showing various methods for matching content elements to surfaces.
[0168] Figure 7A depicts a flowchart showing a method for matching content elements to surfaces according to some embodiments. Figure 7A is a detailed flow of matching content elements to surfaces at 440. Figure 4
[0169] At 710, the method determines whether the identified content element contains a hint provided by the content designer. The content designer can provide a hint as to where the content element is best displayed.
[0170] In some embodiments, this can be accomplished by using existing tag elements (e.g., HTML tag elements) to further define how the content element is displayed if a 3D environment is available. As another example, the content designer can provide a hint that a 3D image, rather than a 2D image, is available as a resource for a particular content element. For example, in the case of a 2D image, in addition to providing the basic tag identifying the resource for the content element, the content designer can provide other, less commonly used tags to identify resources that include a 3D image corresponding to the 2D image, and in addition, provide a hint to highlight it in front of the user's view if the 3D image is used. In some embodiments, the content designer can provide this additional "hint" to the resource for the 2D image in the event that the display device rendering the content can not have 3D display capabilities to take advantage of the 3D image.
[0171] At 720, the method determines whether to use the hint provided by the content designer or to use a predefined set of rules to match / map the content element to a surface. In some embodiments, in the event that the content designer does not provide a hint for a particular content element, the system and method can use a predefined set of rules to determine the best way to match / map the particular content element to a surface. In some embodiments, even when there can be a hint provided by the content designer for a content element, the system and method can determine that it is better to use a predefined set of rules to match / map the content element to a surface. For example, if the content provider provides a hint to display video content on a horizontal surface, but the system has a predefined rule to display video content on a vertical surface, the predefined rule can override the hint. In some embodiments, the system and method can determine that the hint provided by the content designer is sufficient, and thus use the hints to match / map the content element to a surface. Finally, the determination of whether to use the hint provided by the content designer or to use a predefined set of rules to match / map the content element to a surface is a final decision of the system.
[0172] At 730, if the system utilizes the hint provided by the content designer, the system and method analyzes the hint and searches for a logical structure that includes the identified surrounding surface available to display the particular content element based at least in part on the hint.
[0173] At 740, the system and method run a best-fit algorithm to select the best-fit surface for a particular content element based on the provided hints. For example, the best-fit algorithm can hint a particular content element to suggest a direct view and attempt to identify a surface that is in front of and centered relative to the user's and / or device's current field of view.
[0174] At 750, the system and method stores the matching results having the matching of content elements to surfaces. The table can be stored in a non-transitory storage medium for use by a display algorithm to display content elements on their respectively matched / mapped surfaces.
[0175] Figure 7B A flowchart showing a method for matching / mapping elements from content elements to surfaces is depicted in accordance with some embodiments. Figure 7B is shown with reference to FIG. 1, as in Figure 4 At step 440 of FIG. 4, a flow of matching / mapping content elements stored in a logical structure to surfaces stored in the logical structure is disclosed.
[0176] At 715, the content elements stored in the logical structure resulting from the content structuring process 120 of FIG. 1 are ordered based on the associated priority. In some embodiments, a content designer can define a priority attribute for each content element. It can be beneficial for a content designer to set a priority for each content element to ensure that certain content elements are highlighted within the environment. In some embodiments, the content structuring process 120 can determine the priority of a content element, for example, if the content designer does not define the priority of the content element. In some embodiments, if no content element does not have a developer-provided priority attribute, the system will make the dot product relationship of the surface orientation the default priority attribute.
[0177] At 725, the attributes of the content elements are compared to the attributes of the surfaces to identify if there are surfaces that match the content elements and to determine the best matching surface. For example, starting with the content element having the highest associated priority (e.g., the "master" or parent element ID described in further detail below), the system can compare the attributes of the content element to the attributes of the surfaces to identify the best matching surface, and then continue processing the content element having the second highest associated priority, and so on, to continuously iterate through the logical structure including the content elements. Figure 14A
[0178] At 735, a matching score is computed based on how well the attributes of the content element match the attributes of the corresponding best matching surface. Those of ordinary skill in the art can appreciate that many different scoring algorithms and models can be used to compute the matching score. For example, in some embodiments, the score is a simple sum of the attribute values of the content element and the attribute values of the surface.Figure 8 Various matching score methods are shown.
[0179] Figure 8 Three hypothetical content elements with attributes that can be in a logical structure and three hypothetical surfaces are depicted, which are described in further detail below Figures 14A-14B A person skilled in the art will understand that the values in the content element structure can reflect the content itself (e.g., element C has a high color weight) or reflect the desired surface attributes (e.g., element B is biased toward a smoother surface to be rendered). Further, although depicted as numerical values, other attribute values are of course possible, such as explicit colors in a color gamut or precise dimensions or locations in a room to a user.
[0180] At 745, the surface with the highest matching score is identified. Returning to Figure 8 In the illustrated summation example, element A scores highest on surfaces A and C, element B scores highest on surface B, and element C scores highest on surface C. In such an illustrative example, the system can render element A to surface A, element B to surface B, and element C to surface C. Element A scores equally on surfaces A and C, but the highest score of element C on surface C causes element A to be assigned to surface A. In other words, the system can iterate the second summation of matching scores to determine the combination of content elements and surfaces that produces the highest total matching score. It should be noted that, Figure 8 The sample numbers of the dot products reflect attribute values rather than objective measurement values; for example, a -1 dot product result is a good mathematical relationship, but to avoid introducing negative numbers into the equation, the surface attribute scores the -1 dot product relationship as a positive 1 of the surface attribute.
[0181] In some embodiments, the highest score is identified at 745 either by marking the surface with the highest matching score when evaluating the list of surfaces and unmarking previously marked surfaces, or by tracking the highest matching score and a link to the surface with the highest matching score, or by tracking the highest matching score of all content elements that match a surface. In one embodiment, once a surface is identified as having a sufficient matching score, it can be removed from the list of surfaces, thereby excluding it from further processing. In one embodiment, once a surface is identified as having the highest matching score, it can remain in the list of surfaces with an indication that it has already been matched with a content element. In that embodiment, several surfaces can be matched with a single content element, and the matching score of each can be stored.
[0182] In some embodiments, when each surface in the list of surrounding surfaces is evaluated, a match score is calculated one at a time, a match is determined (e.g., 80% or more of the listed attributes of the content element supported by the surface constitutes a match), and if so, the respective surface is marked as the best match, then proceeding to the next surface, and if the next surface is more of a match, that surface is marked as the best match. Once all surfaces are evaluated for a content element, the surface that is still marked as the best match is the best match for the given surface. In some embodiments, the highest match score can need to be greater than a predefined threshold to qualify as a good match. For example, if it is determined that the best match surface is only 40% of a match (whether in terms of the number of attributes supported or the percentage of the target match score), and the threshold to qualify as a good match is higher than 75%, it is better to create a virtual object to display the content element on than to rely on a surface from the user's environment. This is especially true when the user's environment is a beach, for example, with no identifiable surfaces outside of the beach, the ocean, and the sky. Those of ordinary skill in the art can appreciate that many different match / mapping algorithms can be defined for this process, and this is merely one example of many different types of algorithms.
[0183] At 750, the match / mapping results are stored as described above. In one embodiment, if a surface is removed from the list of surfaces at 745, the stored matches can be considered final. In one embodiment, if a surface remains in the list of surfaces at 745 and several surfaces match to a single content element, an algorithm can be run on the conflicting content elements and surfaces to resolve the conflict and have a one-to-one match rather than a one-to-many match or a many-to-one match. If a high priority content element does not match to a surface, the high priority content element can be matched / mapped to a virtual surface. If a low priority content element does not match to a surface, the rendering module 146 can choose not to render the low priority content element. The match results can be stored in a particular table, such as the table disclosed in Figure 18 Match Table, described below.
[0184] Referring back to Figure 7AAt 760, assuming the method determines to use the predefined rules, the method queries a database containing matching rules for content elements to surfaces and determines that a particular content element should consider the type of surface to which the content element should be matched. At 770, the predefined set of rules can run a best-fit algorithm to select one or more surfaces from the available candidate surfaces that best fit the content element. Based at least in part on the best-fit algorithm, it is determined that the content element should be matched / mapped to a particular surface because that particular surface is the surface in all of the candidate surfaces whose attributes best match the attributes of the content element. Once the match of the content element to the surface is determined, at 750, the method stores the match of the content element to the surface in a table in the non-transitory storage medium as described above.
[0185] In some embodiments, a user can override the matched surface. For example, even when a surface is determined to be the best surface for content by a matching algorithm, a user can select where to overlay the surface to display the content. In some embodiments, a user can select a surface from one or more surface options provided by the system, where the one or more surface options can include a surface that is smaller than the best surface. The system can present one or more display surface options to the user, where the display surface options can include a physical surface within the user’s physical environment, a virtual surface for displaying content in the user’s physical environment, and / or a virtual screen. In some embodiments, a user can select a stored screen (e.g., a virtual screen) to display content. For example, for a particular physical environment in which the user is currently located, the user can prefer to display certain types of content (e.g., videos) on certain types of surfaces (e.g., stored screens having a default screen size, a distance from the user’s location, etc.). The stored screens can be historically frequently used surfaces, or the stored screens can be stored screens identified in the user’s profile or preference settings for displaying certain types of content. Accordingly, the display of the one or more elements on the one or more surfaces can be based at least in part on historically frequently used surfaces and / or stored screens.
[0186] Figure 7C An example is shown in which a user can move content 780 from a first surface to any surface available to the user. For example, a user can be able to move content 780 from a first surface to a second surface (i.e., a vertical wall 795). The vertical wall 795 can have a work area 784. The work area 784 of the vertical wall 795 can be determined by, for example, the environment resolver 168. The work area 784 can have a display surface 782 that can display the content 780, for example, that is not blocked by other content / objects. The display surface 782 can be determined by, for example, the environment resolver 168. In Figure 7CIn the illustrated example, the work area 784 includes a photo frame and a light, which can make the display surface of the work area 784 smaller than the entire work area 784 as illustrated by display surface 782. Movement of the content to the vertical wall 795 (e.g., movements 786a-786c) can not necessarily be a perfect placement of the content 780 in the display surface 782, the center of the work area 784, and / or to the vertical wall 795. Rather, the content can be moved within at least a portion of the peripheral workspace of the vertical wall 795 (e.g., work area 784 and / or display surface 782) based on a user’s gesture to move the content to the vertical wall. As long as the content 780 falls within the vertical wall 795, the work area 784, and / or the display surface 782, the system displays the content 780 in the display surface 782.
[0187] In some embodiments, the peripheral workspace is an abstracted boundary that surrounds a target display surface (e.g., display surface 782). In some embodiments, a user’s gesture can be a selection through the totem / controller 790 to select the content 780 at the first surface and move the content 780 so that at least a portion is within the peripheral workspace of the display surface 782. The content 780 can then be aligned with the profile and orientation of the display surface 782. Selecting virtual content is further described in U.S. Patent Application 15 / 296,869, which claims priority to October 20, 2015, entitled “SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE,” and aligning content with a selected surface is further described in U.S. Patent Application 15 / 673,135, which claims priority to August 11, 2016, entitled “AUTOMATIC PLACEMENT OF A VIRTUAL OBJECT IN A THREE-DIMENSIONAL SPACE,” the contents of each of which are incorporated by reference herein.
[0188] In some embodiments, the user's gesture can be a hand pose, which can include the following indications: (a) selecting content from the first surface, (b) moving the content from the first surface to the second surface, and (c) placing the content at the second surface. In some embodiments, the content is moved from the first surface to a particular portion of the second surface. In some embodiments, the second surface is fitted / filled (e.g., scaled to fit, filled, etc.) when the content is placed at the second surface. In some embodiments, the content placed at the second surface retains the size it had when at the first surface. In these embodiments, the second surface can be larger than the first surface and / or the second surface can be larger than the size required to display the content. The AR system can display the content at the second surface at or near the location indicated by the user to display the content. In other words, the movement of the content from the first surface to the second surface can not require the system to place the content perfectly into the entire workable space of the second surface. The content can only have to end up at least in a first peripheral region of the second surface that is at least visible to the user.
[0189] Environment-driven content
[0190] To date, content has been disclosed that drives where in an environment a content element is displayed. In other words, the user can be selecting various content (e.g., content pulled from a webpage) to be displayed into the user's environment. However, in some embodiments, the environment can drive what content is displayed to the user based at least in part on the user's environment and / or surfaces in the environment. For example, the environment resolver 168 continually evaluates a list of surfaces based on data from the sensors 162. Because the environment resolver 168 continually evaluates the list of surfaces as the user moves from one environment to another or around within an environment, new / additional surfaces can become available that can be suitable for displaying certain types of content (e.g., pushed content) that can be pushed into the user's environment without the user having to search for or select the content and that can not originate from a webpage. For example, certain types of pushed content can include: (a) notifications from various applications, such as stock notifications, news summaries, (b) priority content, such as updates and notifications from social media applications, email updates, etc., and / or (c) advertisements targeted to a broad target group and / or a specific target group, etc. Each of these types of pushed content can have associated attributes, such as size, dimensions, orientation, etc., in order to display the advertisements in the most effective form. Depending on the environment, certain surfaces can provide an opportunity to display these environment-driven content (e.g., pushed content). In some embodiments, pulled content can be matched / mapped to surfaces in the environment first, and pushed content can be matched / mapped to any surfaces in the environment that have not already had pulled content matched / mapped to them.
[0191] Taking advertisements as an example of a type of push content, consider a scenario where a user is in an environment, where there can be many surfaces of various sizes and orientations in the environment. Certain advertisements can be best displayed on surfaces of certain sizes and orientations and in certain locations (e.g., geographic locations such as at home, at work, at a baseball field, at a grocery store, etc., and item locations such as in front of certain physical products in the environment). In these cases, the system can search a database of push content to determine which push content best matches the surfaces of the environment. If a match is found, the content can be displayed on the matching surface in the particular location. In some embodiments, the system provides a list of surfaces to an advertisement server, which uses built-in logic to determine the content to push.
[0192] Unlike traditional online advertising, which relies on the layout of the web page being viewed by the user to determine which portions of the web page window have available space for displaying online advertisements, the present disclosure includes an environment parser 168 that identifies surfaces within the environment and determines candidate surfaces for certain push content (e.g., advertisements).
[0193] In some embodiments, a user can specify preferences for when and where certain types of push content can be displayed. For example, a user can indicate that preference attributes for high-priority content from certain people or organizations are highlighted on surfaces in front of the user, while other types of push content, such as advertisements, are displayed on smaller surfaces on the periphery of the user's primary focal view area, where the user's primary focal view area is the view area generally toward the direction the user is looking, rather than the peripheral view to the side of the user's primary focal view area. In some embodiments, user-selected high-priority content elements (e.g., pulled content, rather than pushed content) are displayed on the most prominent surfaces in the user's environment (e.g., within the user's focal view area), while other unmatched / unmapped surfaces on the periphery of the user's focal view area can be available for push content.
[0194] In some embodiments, a world location environment API can be provided to content designers / web developers / advertisers to create location-aware content. The world location environment API can provide a set of functionality that describes the local environment specific to a particular location where the user is currently located. The world location environment API can provide location environment information that can include identification of particular types of rooms (e.g., living room, gym, office, kitchen), particular queries performed by the user from various locations (e.g., user tends to search for movies from living room, music from gym, recipes from kitchen, etc.), and particular services and applications used by the user at various locations (e.g., Mail client from office, Netflix from living room). The content designers can associate certain actions about the world location environment as attributes of content elements.
[0195] The content providers can use this information along with search history, object recognizer, and application data to provide location-specific content. For example, if the user is searching in the kitchen, the ads and search results will be primarily related to food, as the search engine will know that the user is searching from the user's kitchen. The information provided by the world location environment API can provide accurate information for each room or each location, which makes it more accurate than geo-location, and the environment awareness is higher than geo-fencing. Figure 9 is an example of how the world location environment API can be used to provide location-specific environment. As an example, the user's kitchen 905 can include location-specific content 915a, 915b, and / or 915c displayed on certain surfaces within the user's current physical location. The content 915a can be a recipe for a particular meal, the content 915b can be an advertisement for the meal, and / or 915c can be a suggestion for the meal to prepare in the kitchen.
[0196] Figure 10 is an example of a method 1000 for pushing content to a user of a VR / AR system. At 1010, one or more surfaces and their attributes are determined. The one or more surfaces can be determined from the environment structuring process 160, where the environment parser 168 parses the environment data to determine surfaces in the environment and organize and store the surfaces in a logical structure. The user's environment can carry location attributes of surfaces, such as the user's personal residence, a particular room within the residence, the user's work location, etc. The one or more surfaces can be peripheral to the user's focal viewing area. In some embodiments, the one or more surfaces can be within the user's focal viewing area depending on the push content that the user can wish to be notified of (e.g., emergency notifications from an authority, whitelisted applications / notifications, etc.). The size of the surfaces can be 2D and / or 3D dimensions.
[0197] At 1020, one or more content elements that match one or more surfaces are received. In some embodiments, a content element or individual content elements are received based on at least one surface attribute. For example, a location attribute of "kitchen" can prompt a content element corresponding to a food item to be pushed. In another example, a user can be viewing a first content element on a first surface, and the content element has a sub-content element that is only displayed on a second surface when the surface has certain surface attributes.
[0198] At 1030, a match score is computed based on how well the attributes of the content element match the attributes of the surface. In some embodiments, the score can be based on a scale of 1-100, where a score of 100 is the highest score and a score of 1 is the lowest score. Those of ordinary skill in the art can appreciate that many different scoring algorithms and models can be used to compute the match score. In some embodiments where the content element comprises a notification, the match score computed based on the attributes can indicate a priority of the content element that needs to be notified, rather than a match to a particular surface. For example, when the content element is a notification from a social media application, the score can be based on a priority of the notification defined by the user in the user's social media account, rather than a match score based on attributes of the social media content and attributes of the surface.
[0199] In some embodiments, when a user is relatively stationary in their environment, the list of surfaces can not change much. However, when the user moves, depending on the speed at which the user is traveling, the list of surfaces can change rapidly. In dynamic situations, if it is determined that the user's period of stationarity can not be long enough to fully view the content, a lower match score can be computed. The determination of whether the user has enough time to view the entire content can be an attribute defined by the content designer.
[0200] At 1040, the content element with the highest match score is selected. When there are competing content elements (e.g., advertisements) that want to be displayed to the user, it can be necessary to rank the competing content and pick the preferred content element. Here, as an example, one option to select the preferred content element is to have the competition based on how well the attributes of the content element match the attributes of the surface. As another example, the winner can be selected based at least in part on an amount that the content element provider can be willing to pay for the display of the pushed content. In some embodiments, the preferred content element can be selected based on the type of content (e.g., 3D content or a notification from a social media contact).
[0201] At 1050, the match / mapping of the preferred content to the corresponding surface can be stored in a cache memory or a permanent memory. The storage of the match can be important because it can be important to be able to keep some history of the user's environment when the user moves and the environment changes. The match / mapping can be stored in a database such as a relational database, a NoSQL database, or a graph database.Figure 18 The table such as the table disclosed in the middle. At 1060, the content is rendered on the corresponding surface. The matching can be one-to-one or one-to-many matching / mapping of content elements to surfaces.
[0202] A system and method for deconstructing content for display in an environment has been disclosed. Additionally, the system and method can also push content to surfaces of a user of a virtual reality or augmented reality system.
[0203] Example
[0204] web page
[0205] Reference Figure 11 The environment 1100 represents the physical environment and systems for implementing the processes described herein (e.g., matching content elements from content in a web page to display on surfaces in a user's physical environment 1105). The representative physical environment and systems of the environment 1100 include the user's physical environment 1105 as observed by the user 1108 through the head mounted system 1160. The representative systems of the environment 1100 also include accessing content (e.g., a web page) via a web browser 1110 operably coupled to a network 1120. In some embodiments, the content can be accessed via an application (not shown) such as a video streaming application, where the video stream can be the content being accessed. In some embodiments, the video streaming application can be a sports organization and the content being streamed can be an actual live game, a recap, a brief news / hit, a technical statistic, a live report, a team data, a player statistic, a related video, a news source, a product information, etc.
[0206] The network 1120 can be the Internet, an intranet, a private cloud network, a public cloud network, etc. The web browser 1110 is further operatively coupled to a processor 1170 via the network 1120. Although shown as a separate component from the head mounted system 1160, in alternative embodiments, the processor 1170 can be integrated with one or more components of the head mounted system 1160, and / or can be integrated into other system components within the environment 1100, such as into the network 1120 to access the computing network 1125 and storage device 1130. The processor 1170 can be configured with software 1150 for receiving and processing information, such as video, audio, and content received from the head mounted system 1160, the local storage device 1140, the web browser 1110, the computing network 1125, and the storage device 1130. The software 1150 can communicate with the computing network 1125 and the storage device 1130 via the network 1120. The software 1150 can be installed on the processor 1170 or, in another embodiment; features and functionality of the software can be integrated into the processor 1170. The processor 1170 can be further configured with the local storage device 1140 for storing information used by the processor 1170 for fast access, without having to rely on information stored remotely on an external storage device in the vicinity of the user 1108. In other embodiments, the processor 1170 can be integrated within the head mounted system 1160.
[0207] When the user moves around the user's physical environment 1105 and looks through the head-mounted system 1160, the user's physical environment 1105 is the physical environment of the user 1108. For example, with reference to FIG. 1, the user's physical environment 1105 shows a room with two walls (e.g., a main wall 1180 and a side wall 1184, which are relative to the user's view) and a table 1188. On the main wall 1180, there is a rectangular surface 1182 delineated by a solid black line to display a physical surface with a physical boundary (e.g., a painting hung or affixed on a wall or window, etc.) that can be a candidate surface to project certain content onto. On the side wall 1184, there is a second rectangular surface 1186 delineated by a solid black line to display a physical surface with a physical boundary (e.g., a painting hung or affixed on a wall or window, etc.). On the table 1188, there can be different objects. 1) a virtual Rolodex 1190 in which certain content can be stored and displayed; 2) a horizontal surface 1192 delineated by a solid black line to represent a physical surface with a physical boundary to project certain content onto; and 3) a plurality of stacked virtual square surfaces 1194 delineated by black dashed lines to represent, for example, stacked virtual newspapers that can store and display certain content. Those skilled in the art will appreciate that the above-mentioned physical boundaries, while helpful for placing content elements as they have broken the surface into discrete viewing portions and can be the surface properties themselves, are not necessary for identifying eligible surfaces.
[0208] The web browser 1110 can also display blog pages from the Internet or an intranet / extranet network. Additionally, the web browser 1110 can also be any technology that displays digital content. The digital content can include, for example, web pages, blogs, digital pictures, videos, news articles, newsletters, or music. The content can be stored in a storage device 1130 accessible by the user 1108 via the network 1120. In some embodiments, the content can also be streaming content, for example, a live video feed or a live audio feed. The storage device 1130 can include, for example, a database, a file system, a persistent memory device, a flash drive, a cache, etc. In some embodiments, the web browser 1110 containing the content (e.g., a web page) is displayed via a computing network 1125.
[0209] The computing network 1125 accesses the storage device 1130 to retrieve and store content for display in a webpage on the web browser 1110. In some embodiments, the local storage device 1140 can provide content of interest to the user 1108. The local storage device 1140 can include, for example, a flash drive, a cache, a hard drive, a database, a file system, etc. Information stored in the local storage device 1140 can include recently accessed content or recently displayed content in a 3D space. The local storage device 1140 allows for improved performance of the system of the environment 1100 by locally providing certain content to the software 1150 to assist in deconstructing content for display of the content on a 3D surface environment (e.g., a 3D surface in the physical environment 1105 of the user).
[0210] The software 1150 includes software programs stored in non-transitory computer-readable media to perform the functions of deconstructing content to be displayed within the physical environment 1105 of the user. The software 1150 can run on a processor 1170, which can be locally attached to the user 1108, or in some other embodiments, the software 1150 and the processor 1170 can be included within the head-mounted system 1160. In some embodiments, portions of the features and functions of the software 1150 can be stored and executed on the computing network 1125, remote from the user 1108. For example, in some embodiments, the deconstruction of content can be performed on the computing network 1125, and the results of the deconstruction can be stored in the storage device 1130, where the surfaces of the local environment of the user on which to present the deconstructed content can be inventoried within the processor 1170, where the inventory and mapping / matching of the surfaces are stored within the local storage device 1140. In one embodiment, the processes of deconstructing content, inventorying local surfaces, mapping / matching elements of content to local surfaces, and displaying elements of content can be performed locally within the processor 1170 and the software 1150.
[0211] The head-mounted system 1160 can be a virtual reality (VR) or augmented reality (AR) head-mounted system (e.g., a mixed reality device) that includes a user interface, a user sensing system, an environment sensing system, and a processor (all not shown). The head-mounted system 1160 presents an interface to the user 1108 for interacting with and experiencing a digital world. Such interactions can involve the user and the digital world, one or more other users in communication with the environment 1100, and objects within the digital and physical worlds.
[0212] The user interface can include receiving content and selecting elements within the content by user input through the user interface. The user interface can be at least one or a combination of a haptic interface device, a keyboard, a mouse, a joystick, a motion capture controller, an optical tracking device, and an audio input device. A haptic interface device is a device that allows a human to interact with a computer through the senses and motions of the body. Haptics refers to a human-computer interaction technique that incorporates haptic feedback or other body sensations to perform actions or processes on a computing device.
[0213] The user sensing system can include one or more sensors 1162 operable to detect certain features, characteristics, or information related to the user 1108 wearing the head-mounted system 1160. For example, in some embodiments, the sensors 1162 can include a camera or optical detection / scanning circuitry capable of detecting real-time optical characteristics / measures of the user 1108, such as one or more of: pupil constriction / dilation, angular measurement / positioning of each pupil, spherical degree, eye shape (as eye shape changes over time), and other structural data. This data can provide or be used to calculate information that can be used by the head-mounted system 1160 to augment the user's viewing experience (e.g., the user's visual focal point).
[0214] The environment sensing system can include one or more sensors 1164 for acquiring data from the user's physical environment 1105. Objects or information detected by the sensors 1164 can be provided as input to the head-mounted system 1160. In some embodiments, this input can represent user interaction with the virtual world. For example, a user (e.g., the user 1108) viewing a virtual keyboard on a table (e.g., the table 1188) can gesture with their fingers as if the user were typing on the virtual keyboard. The motion of the fingers can be captured by the sensors 1164 and provided as input to the head-mounted system 1160, where the input can be used to change the virtual world or create new virtual objects.
[0215] The sensors 1164 can include, for example, a camera or scanner generally facing outward, for interpreting scene information, for example, by projected infrared structured light, continuously and / or intermittently. By detecting and registering the local environment, the environment sensing system can be used to match / map one or more elements of the user's physical environment 1105 around the user 1108, including static objects, dynamic objects, people, gestures, and various lighting, atmospheric, and acoustic conditions. Thus, in some embodiments, the environment sensing system can include image-based 3D reconstruction software embedded in the local computing system (e.g., the processor 1170) and operable to digitally reconstruct one or more objects or information detected by the sensors 1164.
[0216] In one example embodiment, the environmental sensing system provides one or more of: movement capture data (including gesture recognition), depth sensing, facial recognition, object recognition, unique object feature recognition, speech / audio recognition and processing, sound source localization, noise reduction, infrared or similar laser projection, and monochrome and / or color CMOS sensors (or other similar sensors), field of view sensors, and various other optical enhanced sensors. It should be appreciated that the environmental sensing system can include other components in addition to those discussed above.
[0217] As discussed above, in some embodiments, the processor 1170 can be integrated with other components of the head mounted system 1160, integrated with other components of the system of the environment 1100, or can be a separate device (wearable or separate from the user 1108) as shown in FIG. 1. The processor 1170 can be connected to the various components of the head mounted system 1160 through physical, wired connections or through wireless connections (e.g., mobile network connections (including cellular phone and data networks), Wi-Fi, Bluetooth, or any other wireless connection protocol). The processor 1170 can include memory modules, integrated and / or additional graphics processing units, wireless and / or wired internet connections, and codecs and / or firmware capable of converting data from sources (e.g., the computing network 1125 and from the user sensing system and the environmental sensing system of the head mounted system 1160) into image and audio data, where the image / video and audio can be presented to the user 1108 through a user interface (not shown).
[0218] The processor 1170 handles data processing for the various components of the head mounted system 1160 and data exchange between the head mounted system 1160 and content from web pages displayed or accessed by the web browser 1110 and the computing network 1125. For example, the processor 1170 can be used to buffer and process data streams between the user 1108 and the computing network 1125, enabling a smooth, continuous, and high-fidelity user experience.
[0219] The deconstruction of content from a web page into content elements and the matching / mapping of elements to be displayed on surfaces in a 3D environment can be done in an intelligent and logical manner. For example, the content parser 115 can be a document object model (DOM) parser that receives input (e.g., an entire HTML page), deconstructs various content elements within the input and stores the deconstructed content elements in a logical structure such that the elements of content are accessible and easily programmatically manipulated / extracted. A predetermined set of rules can be used to recommend, suggest or dictate placement locations for certain types of elements / content identified within, for example, a web page. For example, certain types of content elements can have one or more content elements that can need to be matched / mapped to a physical or virtual object surface that is suitable for storing and displaying the one or more elements, while other types of content elements can be single objects, such as a primary video or a primary article within a web page, in which case the single object can be matched / mapped to the most meaningful surface to display the single object to the user. In some embodiments, the single object can be a video streamed from a video application such that the single content object can be displayed on a surface (e.g., a virtual surface or a physical surface) within the user's environment.
[0220] Figure 12 In the example of FIG. 12, the environment 1200 depicts content (e.g., a web page) displayed or accessed by the web browser 1110 and the user's physical environment 1105. The dashed arrowed lines depict elements (e.g., certain types of content) from the content (e.g., the web page) being matched or mapped to and displayed on the user's physical environment 1105. Certain elements from the content are matched / mapped to certain physical or virtual objects in the user's physical environment 1105 based on either web page designer cues or predefined browser rules.
[0221] As an example, the content accessed or displayed by the web browser 1110 can be a web page having multiple tabs, where the currently active tab 1260 is displayed and a second tab 1250 is currently hidden until selected to be displayed on the web browser 1110. Typically displayed within the active tab 1260 is a web page. In this particular example, the active tab 1260 is displaying a YOUTUBE page that includes a primary video 1220, user comments 1230, and suggested videos 1240. As this example illustrates, the primary video 1220 is a single object that can be matched / mapped to a surface in the user's physical environment 1105 that is suitable for displaying the primary video 1220. In this example, the primary video 1220 is matched / mapped to a virtual surface 1215 in the user's physical environment 1105 that is suitable for displaying the primary video 1220. The user comments 1230 and suggested videos 1240 are multiple objects that can be matched / mapped to a physical or virtual object surface that is suitable for storing and displaying the multiple objects. In this example, the user comments 1230 and suggested videos 1240 are matched / mapped to a physical surface 1210 in the user's physical environment 1105 that is suitable for storing and displaying the user comments 1230 and suggested videos 1240. Figure 12As shown, the primary video 1220 can be matched / mapped for display on the vertical surface 1182, the user comments 1230 can be matched / mapped for display on the horizontal surface 1192, and the suggested videos 1240 can be matched / mapped for display on a different vertical surface 1186 than the vertical surface 1182. In addition, the second tab 1250 can be matched / mapped for display on and / or as the virtual Rolodex 1190 and / or on the plurality of stacked virtual objects 1194. In some embodiments, particular content within the second tab 1250 can be stored in the plurality of stacked virtual objects 1194. In other embodiments, all content located within the second tab 1250 can be stored and / or displayed on the plurality of stacked virtual objects 1194. Similarly, the virtual Rolodex 1190 can contain particular content from the second tab 1250, or the virtual Rolodex 1190 can contain all content located within the second tab 1250.
[0222] In some embodiments, content elements of the web browser 1110 (e.g., content elements of a web page in the second tab 1250) can be displayed on a two-sided planar window virtual object (not shown) in the user's physical environment 1105. For example, the primary content of a web page can be displayed on a first side (e.g., front side) of the planar window virtual object, while additional information such as extra content related to the primary content can be displayed on a second side (e.g., back side) of the planar window virtual object. As an example, a merchant web page (e.g., BEST BUY) can be displayed on the first side, and a set of coupons and discounts can be displayed on the second side. The discount information can be updated on the second side to reflect the current context of the content the user is viewing on the first side (e.g., only laptop or home appliance discounts on the second side).
[0223] Some web pages can span multiple web pages when viewed in the web browser 1110. Such web pages can be viewed by scrolling in the web browser 1110 or navigating multiple pages in the web browser 1110 when viewed in the web browser 1110. When such web pages from the web browser 1110 are matched / mapped to the user's physical environment 1105, such web pages can be matched / mapped as two-sided web pages. Figures 13A-13B An example two-sided web page is shown in accordance with some embodiments. Figure 13A A smoothie drink is shown, while Figure 13BAn exemplary back / second side of a smoothie drink is shown, which includes ingredients and instructions for making the smoothie. In some embodiments, the front side of the main wall 1180 can include a first side of a double-sided web page, while the back side of the main wall 1180 can include a second side of the double-sided web page. In this example, the user 1108 would have to walk around the main wall 1180 to see both sides of the double-sided web page. In some embodiments, the front side of the main wall 1180 can include both sides of the double-sided web page. In this example, the user 1108 can switch between the two sides of the double-sided web page via user input. In response to the user input, the double-sided web page can appear to flip from the first side to the second side. Although the double-sided web page is described as being generated from a web page when viewed across multiple pages in the web browser 1110, the double-sided web page can be generated from any web page or portion thereof or multiple thereof. The VR and / or AR system can provide a set of easy-to-use HTML features that can be added to existing content (e.g., the second tab 1250 or web page) to make it available to the rendering module to render the content onto a double-sided 2D browser plane window virtual object. Although this example describes a double-sided plane window virtual object, the virtual object can have any number of sides (N-sided). Although this example describes displaying content on a double-sided plane window virtual object, content elements can be on multiple surfaces of a real object (e.g., the front of a door and the back of the door).
[0224] The vertical surface 1182 can be any type of structure that is already on the main wall 1180 of the room (depicted as the physical environment 1105 of the user) such as a windowpane or a picture frame. In some embodiments, the vertical surface 1182 can be an empty wall, where the head-mounted system 1160 determines the optimal size of the frame of the vertical surface 1182 that is suitable for the user 1108 to view the main video 1220. This determination of the size of the vertical surface 1182 can be based at least in part on the distance of the user 1108 from the main wall 1180, the size and dimensions of the main video 1220, the quality of the main video 1220, the amount of uncovered wall space, and / or the pose of the user when viewing the main wall 1180. For example, if the quality of the main video 1220 is high definition, the size of the vertical surface 1182 can be larger because the quality of the main video 1220 is not adversely affected by the vertical surface 1182. However, if the video quality of the main video 1220 is of lower quality, having a larger vertical surface 1182 can greatly impede the video quality, in which case the methods and systems of the present disclosure can resize / redefine how the content is displayed within the vertical surface 1182 to be smaller to minimize the poor video quality from pixilation.
[0225] Vertical surface 1186, like vertical surface 1182, is a vertical surface on an adjacent wall (e.g., side wall 1184) in the user's physical environment 1105. In some embodiments, based on the orientation of user 1108, side wall 1184 and vertical surface 1186 can appear to be a sloped surface on a slope. In addition to vertical surfaces and horizontal surfaces, sloped surfaces on a slope can also be a type of orientation of a surface. In this example, suggested video 1240 from the YOUTUBE webpage can be placed on vertical surface 1186 on side wall 1184 to allow user 1108 to view the suggested video by simply moving his head slightly to the right.
[0226] Virtual Rolodex 1190 is a virtual object created and displayed to user 1108 by head mounted system 1160. Virtual Rolodex 1190 can enable user 1108 to cycle through a set of virtual pages in both directions. Virtual Rolodex 1190 can contain an entire webpage, or it can contain individual articles or videos or audio. As shown in this example, virtual Rolodex 1190 can contain a portion of the content from second tab 1250, or in some embodiments, virtual Rolodex 1190 can contain the entire page of second tab 1250. User 1108 can cycle through the content within virtual Rolodex 1190 by simply focusing on a particular tab within virtual Rolodex 1190, and one or more sensors within head mounted system 1160 (e.g., sensors 1162) will detect the eye focus of user 1108 and accordingly cycle through the tabs within virtual Rolodex 1190 to obtain relevant information for user 1108. In some embodiments, user 1108 can select relevant information from virtual Rolodex 1190 and instruct head mounted system 1160 to display the relevant information onto an available surrounding surface or another virtual object (such as a virtual display proximate to user 1108) (not shown).
[0227] Similar to virtual Rolodex 1190, multiple stacked virtual objects 1194 can contain the following: complete content from one or more tabs or specific content from various webpages or tabs that user 1108 has bookmarked for future viewing or has opened (i.e., inactive tabs). Multiple stacked virtual objects 1194 are also similar to a stack of real-world newspapers. Each stack within multiple stacked virtual objects 1194 can belong to a particular newspaper article, page, magazine issue, recipe, etc. Those of ordinary skill in the art will appreciate that there can be multiple types of virtual objects to achieve the same purpose of this provided surface to place content elements or content from a content source.
[0228] As can be appreciated by one of ordinary skill in the art, the content accessed or displayed by the web browser 1110 can not be just a web page. In some embodiments, the content can be a picture from a photo album, a video from a movie, a television show, a YOUTUBE video, an interactive form, etc. However, in other embodiments, the content can be an e-book or any electronic way of displaying a book. Finally, in other embodiments, the content can be other types of content not yet described, as content is generally a presentation of current information. If the electronic device can consume the content, the head mounted system 1160 can use the content to deconstruct and display the content in a 3D setting (e.g., AR).
[0229] In some embodiments, matching / mapping the accessed content can include extracting the content (e.g., from the browser) and placing it on the surface (such that the content is no longer in the browser and only on the surface), and in some embodiments, matching / mapping can include copying the content (e.g., from the browser) and placing it on the surface (such that the content is in both the browser and on the surface).
[0230] Deconstructing content is a technical problem that exists in the field of internet and computer related technologies. Digital content such as web pages are constructed using certain types of programming languages such as HTML to instruct computer processors and technical components where and how to display elements within the web page on a screen for a user. As described above, web designers typically work within the confines of a 2D canvas (e.g., a screen) to place and display elements (e.g., content) within the 2D canvas. HTML tags are used to determine how to format an HTML document or portions of an HTML document. In some embodiments, the (extracted or copied) content can maintain HTML tag references, and in some embodiments, the HTML tag references can be redefined.
[0231] With respect to this example, briefly refer to Figure 4Receiving content at 410 can involve using the head-mounted system 1160 to search for digital content. Receiving content at 410 can also include accessing digital content on a server (e.g., storage device 1130) connected to the network 1120. Receiving content at 410 can include browsing the Internet for web pages of interest to the user 1108. In some embodiments, receiving content at 410 can include a voice-activated command given by the user 1108 to search for content on the Internet. For example, the user 1108 can be interacting with a device (e.g., head-mounted system 1160) where the user 1108 requests the device to search for a particular video by saying a command to search for a video and then saying the name of the video and a short description of the video. The device can then search the Internet and pull up the video on a 2D browser to allow the user 1108 to see the video displayed on the 2D browser of the device. The user 1108 can then confirm that the video is the video that the user 1108 wants to watch in the spatial 3D environment.
[0232] Once the content is received, the method identifies content elements in the content at 420 to inventory the content elements within the content for display to the user 1108. The content elements within the content can include, for example, videos, articles and newsletters posted on web pages, comments and posts on social media websites, blogs, pictures posted on various websites, audiobooks, etc. These elements within the content (e.g., web pages) can be discernible in the content script by HTML tags, and can further include HTML tags or similar HTML tags that have attributes provided by the content designer to define where to place the particular element and in some cases when and how to display the element. In some embodiments, the methods and systems of the present disclosure will utilize these HTML tags and attributes as hints and suggestions provided by the content designer to assist in the matching / mapping process at 440 to determine where and how to display the elements in the 3D setting. For example, the following is an example HTML web page code provided by a content designer (e.g., web page developer).
[0233] Example HTML web page code provided by a content designer
[0234]
[0235]
[0236]
[0237] The example HTML webpage code provided by the content designer includes preferences about how to display the primary video on the webpage, as well as preferences about how to display the recommendations (or suggested videos). The preferences can be conveyed as one or more attributes in the tag. Example attributes for the content element are described above and below. The attributes can be determined or inferred as described above. In particular, this HTML webpage code uses a "style" tag to specify how to display the primary video using a "vertical" type value to specify a vertical surface to display the video. Additionally, within the "style" tag, additional hints provided by the content designer can include a "priority" preference attribute for the matching algorithm to use for prioritizing which HTML element / content in the webpage (e.g., the primary video) should be matched / mapped to which potential surface area. In the example HTML webpage code, the priority is set to a value of 100 for videos that have a vertical plane layout, where in this example, a higher priority value indicates a higher priority. Additionally, in this example, the content designer indicates a preference attribute to place the suggested videos in a stack layout with a type value of "horizontal" in the stack, where the distance between objects of the stack (e.g., in this case, a suggested video with respect to another suggested video) should be 20 centimeters.
[0238] In some embodiments, such as <ml-container>The tags can allow content designers to provide specific preference attributes (e.g., hints) about where and how content elements should be displayed in an environment (e.g., a 3D space environment) such that a resolver (e.g., resolver 115) can be able to interpret the attributes specified within the tags to determine where and how content elements should be displayed in the 3D space environment. The specific preference attributes can include one or more attributes that define display preferences for the content elements. The attributes can include any of the attributes described above.
[0239] As can be appreciated by one of ordinary skill in the art, these suggestions, hints, and / or attributes defined by the content designer can be in a variety of formats, such as <ml-container>defined in the tags of the content, which can indicate similar characteristics for displaying the content elements in the 3D spatial environment. Additionally, one of ordinary skill in the art can also appreciate that the content designer can specify the attributes in any combination. The embodiments disclosed herein can interpret the desired display results by analyzing the content of the web page using a parser (e.g., parser 115) or other similar techniques to determine how and where to best display the content elements within the content.
[0240] Briefly referring to Figure 5 , with respect to this example, identifying elements within the content at 510 can be similar to identifying elements in the content at 420 of Figure 4 . The method proceeds to the next step of identifying attributes from the tags about placement of the content at 520. As described above, the content designer, when designing and configuring the web page, can associate the content elements within the web page with HTML tags to define where and how to display each of the content elements. These HTML tags can also include attributes about placing the content elements on specific portions of the web page. The head mounted system 1160 will detect these HTML tags and their attributes and coordinate with other components of the system to be used as input about where a particular element can be displayed. In some embodiments, such as <ml-container>The tags can include properties specified by a content designer to suggest display preference properties for content elements in a 3D spatial environment, where the tags are associated with the content elements.
[0241] At 530, extracting the hints or tags from each element is performed. The hints or tags are typically formatting hints or formatting tags provided by a content designer of the web page. As described above, the content designer can provide instructions or hints, for example, in the form of HTML tags as shown in the "example HTML web page code provided by web developer" to instruct the web browser 1110 to display the content element in a particular portion of the page or screen. In some embodiments, the content designer can use additional HTML tag properties to define additional formatting rules. For example, if the user has reduced sensitivity to a particular color (e.g., red), the red color is not displayed, but another color is used, or if a preference for displaying a video on a vertical surface cannot be displayed on a vertical surface, the video is displayed on another (physical) surface, or a virtual surface is created and the video is displayed on the virtual surface. The following is an example HTML page parser implemented in a browser to parse an HTML page to extract hints / tags from each element in the HTML page.
[0242] Example HTML page parser implemented in a browser
[0243]
[0244]
[0245]
[0246]
[0247] The example HTML page parser shows how an HTML page containing HTML tags for providing display preference properties for particular content elements can be parsed and recognized and / or extracted / copied. As disclosed in the example HTML page parser, the disclosed example code can be used to parse the content elements. The HTML page parser can recognize / extract certain HTML tags (e.g., ML.layout, ML.container, etc.) using various element names and values to determine how a particular element is to be displayed to the user in a 3D environment (e.g., by matching the content element to a particular surface).
[0248] At 540, a look / search for an alternative display form of the content element is performed. Certain formatting rules can be specified for content elements that are displayed on a particular viewing device. For example, certain formatting rules can be specified for images on a webpage. The system can access the alternative display form. For example, if the web browser 1110 is capable of displaying a 3D version of an image (or more generally a 3D asset or 3D media), then a web designer can place additional tags or define certain properties of particular tags to allow the web browser 1110 to recognize that the image can have an alternative version of the image (e.g., a 3D version of the image). The web browser 1110 can then access the alternative version of the image (e.g., a 3D version of the image) that is to be displayed in the 3D-enabled browser.
[0249] In some embodiments, a 3D image within a webpage can not be extractable or copyable from the webpage to be displayed on a surface in a 3D environment. In these embodiments, the 3D image can be displayed within the user's 3D environment, where the 3D image appears to be rotating, glowing, etc., and the user can interact with the 3D image but only within the webpage that includes the 3D image. In these embodiments, the 3D image is displayed within the webpage because the 3D image is not extracted or copied from the webpage. In this case, the entire webpage is extracted and displayed in the user's 3D environment, and some content elements within the webpage, such as the 3D image, appear in 3D relative to the rest of the webpage and are interactive within the webpage, although the 3D image is not extracted or copied from the webpage.
[0250] In some embodiments, a 3D image within a webpage can be copied but not extractable from the webpage. In these embodiments, the 3D image can be displayed within the user's 3D environment, where the 3D image appears to be rotating, glowing, etc., and the user can interact with the 3D image not only within the webpage that includes the 3D image but also in the 3D environment outside of the webpage that includes a copy of the 3D image. The webpage appears the same as the 3D image, and there is a copy of the 3D image outside of the webpage.
[0251] In some embodiments, a 3D image within a webpage can be extractable from the webpage. In these embodiments, the 3D image can be displayed within the user's 3D environment, where the 3D image appears to be rotating, glowing, etc., and the user can interact with the 3D image but only outside of the webpage because the 3D image is extracted from the webpage. Because the 3D image is extracted from the webpage, the 3D image is displayed only in the 3D environment and not in the webpage without. In these embodiments, the webpage can be reconfigured after the 3D image is extracted from the webpage. For example, a version of the webpage can be presented to the user that includes a blank space within the webpage from which the 3D image was previously extracted.
[0252] While previous embodiments and examples are described with respect to 3D images within a webpage, one of ordinary skill in the art will appreciate that the description can similarly apply to any content element.
[0253] At 550, storing the identified content elements is performed. The method can store the identified elements into a non-transitory storage medium for use in the composition process 140 to match content elements to surfaces. The non-transitory storage medium can include a data storage device, such as the storage device 1130 or the local storage device 1140. The content elements can be stored in a particular table, for example, a table as disclosed in Figure 14A some embodiments, the content elements can be stored in a hierarchical structure, for example, represented as a tree structure as disclosed in Figure 14B some embodiments, the content elements can be stored in a non-transitory storage medium.
[0254] Figures 14A-14B Examples of different structures for storing content elements deconstructed from content are shown in accordance with some embodiments. In Figure 14A element table 1400 is an exemplary table that can store Figure 5 the results of identifying content elements within the content at 510. The element table 1400 includes, for example, information about one or more content elements within the content, including an element identification (ID) 1410, a preference attribute indicator 1420 for the content element (e.g., a priority attribute, an orientation attribute, a position type attribute, a content type attribute, a surface type attribute, etc., or some combination thereof), a parent element ID 1430 (if the particular content element is included within a parent content element), a child content element ID 1440 (if the content element can contain sub-content elements), and a multiple entity indicator 1450 indicating whether the content element contains multiple embodiments that can warrant the use of a surface or virtual object for displaying the content element that is compatible with displaying multiple versions of the content element. A parent content element is a content element / object within the content that can contain sub-content elements (e.g., child content elements). For example, an element ID with a value of 1220 (e.g., a main video 1220) has a parent element ID value of 1260 (e.g., an activity tab 1260) indicating that the main video 1220 is a sub-content element of the activity tab 1260. Or stated differently, the main video 1220 is included within the activity tab 1260. Continuing the same example, the main video 1220 has a child element ID 1230 (e.g., a user comment 1230) indicating that the user comment 1230 is associated with the main video 1220. Those of ordinary skill in the art will appreciate that the element table 1400 can be a table in a relational database or any type of database. Additionally, the element table 1400 can be an array in a computer memory (e.g., a cache) that contains Figure 5 the results of identifying content elements within the content at 510.
[0255] Each row 1460 in the element table 1400 corresponds to a content element within the web page. The element ID 1410 is a column that contains a unique identifier (e.g., an element ID) for each content element. In some embodiments, the uniqueness of the content element can be defined as a combination of the element ID 1410 column and another column in the table (e.g., the preference attribute 1420 column if the content designer identified more than one preference attribute). The preference attribute 1420 is a column whose values can be determined based at least in part on the labels and attributes defined by the content designer in the content and in the content design file 1300. The parent element ID 1430 is a column that contains a unique identifier for a parent content element if the particular content element is included within a parent content element. The child content element ID 1440 is a column that contains a unique identifier for a child content element if the content element can contain sub-content elements. The multiple entity indicator 1450 is a column that indicates whether the content element contains multiple embodiments that can warrant the use of a surface or virtual object for displaying the content element that is compatible with displaying multiple versions of the content element. Figure 5 The preference attribute 1420 column is a column identified by the disclosed systems and methods from the hints or labels of each content element at 530. In other embodiments, the preference attribute 1420 column can be determined based at least in part on predefined rules that specify where certain types of content elements should be displayed in the environment. These predefined rules can provide suggestions to the systems and methods to determine where the content elements are best placed in the environment.
[0256] Parent Element ID 1430 is a column containing the element IDs of the parent content element in which or related content elements are displayed in the current row. A specific content element can be embedded within, placed within, or related to another content element on the page. For example, in this embodiment, the first entry in the element ID 1410 column stores information related to... Figure 12 The value of element ID 1220 corresponding to the main video 1220. The value in the preference attribute column 1420 corresponding to the main video 1220 is determined based on tags and / or attributes, and as shown, this content element should be placed in the "main" position of the user's physical environment 1105. Depending on the user 1108's current position, this main position could be a wall in the living room or a stovetop in the kitchen that the user 1108 is currently looking at, or, if in a spacious area, a virtual object projected in front of the user 1108's line of sight, onto which the main video 1220 can be projected. More information about how the content element is displayed to the user 1108 will be disclosed elsewhere in the detailed description. Continuing with the current example, the parent element ID column 1430 stores the value of the parent element ID. Figure 12 The value of element ID 1260 corresponding to activity tag 1260. Therefore, main video 1220 is a child of activity tag 1260.
[0257] The child element ID 1440 column is a list of element IDs containing child content elements that are currently displayed in or related to that content element. A specific content element within a webpage can be embedded, placed within another content element, or related to another content element. Continuing with the current example, the child element ID 1440 column stores... Figure 12 The value of the element ID 1230 corresponding to user comment 1230.
[0258] Multiple Entity Indicator 1450 is a column indicating whether a content element contains multiple entities that can guarantee compatibility between the surface or virtual object required for the display element and the content element displaying multiple versions (e.g., the content element could be user comment 1230, where, for the main video 1220, there may be more than one comment available). Continuing the current example, the Multiple Entity Indicator 1450 column stores the value "N" to indicate that the main video 1220 does not have or correspond to multiple main videos in the activity tag 1260 (e.g., "No" multiple versions of the main video 1220).
[0259] Continuing with the current example, the second entry in column 1410 of element ID is stored with... Figure 12 corresponding to the user comment 1230. The value in the preference attribute 1420 column corresponding to the user comment 1230 displays a "horizontal" preference to indicate that the user comment 1230 should be placed on a horizontal surface somewhere in the user's physical environment 1105. As described above, the horizontal surface will be determined based on available horizontal surfaces in the user's physical environment 1105. In some embodiments, the user's physical environment 1105 can not have a horizontal surface, in which case the system and methods of the present disclosure can identify / create a virtual object with a horizontal surface to display the user comment 1230. Continuing the current example, the parent element ID 1430 column stores the value element ID 1220 corresponding to the main video 1220, and the multiple entity indicator 1450 column stores the value "Y" to indicate that the user comment 1230 can contain more than one value (e.g., more than one user comment). Figure 12
[0260] The remaining rows within the element table 1400 contain information for the remaining content elements of interest to the user 1108. As can be appreciated by one of ordinary skill in the art, storing the results identifying the content elements within the content at 510 improves the functionality of the computer itself, as once that analysis is performed on the content, it can be retained by the system and methods for future analysis of the content if another user is interested in the same content. The system and methods used to deconstruct that particular content can be avoided as it has been done previously.
[0261] In some embodiments, the element table 1400 can be stored in the storage device 1130. In other embodiments, the element table 1400 can be stored in the local storage device 1140 for quick access to recently viewed content or for possible re-access to recently viewed content. In other embodiments, the element table 1400 can be stored at both the storage device 1130 remote from the user 1108 and the local storage device 1140 local to the user 1108.
[0262] In some embodiments, the element table 1400 can be stored in the storage device 1130. In other embodiments, the element table 1400 can be stored in the local storage device 1140 for quick access to recently viewed content or for possible re-access to recently viewed content. In other embodiments, the element table 1400 can be stored at both the storage device 1130 remote from the user 1108 and the local storage device 1140 local to the user 1108. Figure 14B In some embodiments, the element table 1400 can be stored in the storage device 1130. In other embodiments, the element table 1400 can be stored in the local storage device 1140 for quick access to recently viewed content or for possible re-access to recently viewed content. In other embodiments, the element table 1400 can be stored at both the storage device 1130 remote from the user 1108 and the local storage device 1140 local to the user 1108. Figure 5 The exemplary logical structure in which the results of identifying elements within content at 510 are stored into a database. When various content has a hierarchical relationship to one another, it can be advantageous to store the content elements in a tree structure. The tree structure 1405 includes a parent node - the web page main tag node 1415, a first child node - the main video node 1425, and a second child node - the suggested video node 1445. The first child node - the main video node 1425 includes a child node - the user comment node 1435. The user comment node 1435 is a grandchild of the web page main tag node 1415. For example, referring to Figure 12 The web main tag node 1415 can be the web main tag 1260, the main video node 1425 can be the main video 1220, the user comment node 1435 can be the user comment 1230, and the suggested video node 1445 can be the suggested video 1240. Here, the tree structure organization of the content elements shows the hierarchical relationship between the various content elements. It can be advantageous to organize and store content elements in a logical structure of the tree structure type. For example, if the main video 1220 is being displayed on a particular surface, it can be useful for the system to know that the user comment 1230 is a child content of the main video 1220, and it can be beneficial to display the user comment 1230 relatively close to the main video 1220 and / or on a surface near the main video 1220 so that the user can easily view and understand the relationship between the user comment 1230 and the main video 1220. In some embodiments, it can be beneficial to be able to hide or close the user comment 1230 if the user decides to hide or close the main video 1220. In some embodiments, it can be beneficial to be able to move the user comment 1230 to another surface if the user decides to move the main video 1220 to a different surface. When the user moves the main video 1220 by moving both the parent node-main video node 1425 and the child node-user comment node 1435, the system can move the user comment 1230.
[0263] Returning to Figure 4 The method continues with determining surfaces at 430. The user 1108 can view the user's physical environment 1105 through the head-mounted system 1160 to allow the head-mounted system 1160 to capture and recognize the surrounding surfaces, such as walls, tables, paintings, window frames, fireplaces, refrigerators, televisions, etc. The head-mounted system 1160 is aware of real objects in the user's physical environment 1105 due to sensors and cameras on the head-mounted system 1160 or any other type of similar device. In some embodiments, the head-mounted system 1160 can match real objects observed in the user's physical environment 1105 with virtual objects stored in the storage device 1130 or the local storage device 1140 to identify surfaces available for those virtual objects. A real object is an object identified in the user's physical environment 1105. A virtual object is an object that is not physically present in the user's physical environment but can be displayed to the user to make it appear as if the virtual object is present in the user's physical environment. For example, the head-mounted system 1160 can detect an image of a table within the user's physical environment 1105. The table image can be reduced to a 3D point cloud object for comparison and matching at the storage device 1130 or the local storage device 1140. If a match of a real object to a 3D point cloud object (e.g., of a table) is detected, the system and method will identify the table as having a horizontal surface because the 3D point cloud object representing the table is defined as having a horizontal surface.
[0264] In some embodiments, virtual objects can be extracted objects, where an extracted object can be a physical object recognized within the user's physical environment 1105, but displayed to the user as a virtual object in the location of the physical object, such that additional processing and associations can be performed on the extracted object that would not be possible on the physical object itself (e.g., changing the color of the physical object to highlight a particular feature of the physical object, etc.). Additionally, an extracted object can be a virtual object extracted from content (e.g., a webpage from a browser) and displayed to the user 1108. For example, the user 1108 can select an object displayed on a webpage (such as a sofa) to be displayed in the user's physical environment 1105. The system can recognize the selected object (e.g., the sofa) and display the extracted object (e.g., the sofa) to the user 1108 as if the extracted object (e.g., the sofa) physically exists in the user's physical environment 1105. Furthermore, virtual objects can also include objects with surfaces for displaying content (e.g., a transparent display screen positioned very close to the user for viewing certain content) that do not even physically exist in the user's physical environment 1105, but can be ideal display surfaces for presenting certain content to the user from a display content perspective.
[0265] Briefly referring to Figure 6 the method begins at 610 with determining a surface. The method proceeds to 620 with determining the user's pose, which can include determining a head pose vector. Determining the user's pose at 620 is an important step in identifying the user's current surroundings, as the user's pose will provide the user 1108 with a perspective on the objects within the user's physical environment 1105. For example, referring back to Figure 11 , the user 1108 using the head-mounted system 1160 is observing the user's physical environment 1105. Determining the user's pose (i.e., the head pose vector and / or the position information relative to the origin of the world) at 620 will help the head-mounted system 1160 understand, for example: (1) how tall the user 1108 is relative to the ground; (2) the angle at which the user 1108 must rotate his head to move and capture an image of the room; and (3) the distance of the user 1108 to the table 1188, the main wall 1180, and the side wall 1184. Additionally, the user's pose helps determine the angle at which the head-mounted system 1160 views the vertical surfaces 1182 and 186 and other surfaces within the user's physical environment 1105.
[0266] At 630, the method determines the properties of the surfaces. Each surface within the user's physical environment 1105 is labeled and categorized using a corresponding property. In some embodiments, each surface within the user's physical environment 1105 is also labeled and categorized using a corresponding size and / or orientation property. This information will assist in matching content elements to surfaces based at least in part on: the size property of the surface, the orientation property of the surface, the distance of the user 1108 from a particular surface, and the type of information that needs to be displayed for the content element. For example, a video can be displayed farther away than a blog or article that can contain a large amount of information, where the text size of the article can be too small for the user to see if displayed in small size on a distant wall. In some embodiments, raw data from sensors 162 is provided to CVPU 164 for processing, and CVPU 164 provides the processed data to perception framework 166 for preparing data for environment resolver 168. Environment resolver 168 resolves the environment data from perception framework 166 to determine the surfaces in the environment. Figure 1B
[0267] At 640, the method stores the inventory of surfaces to a non-transitory storage medium for use by the synthesis process / matching / mapping routine to match / map the extracted elements to particular surfaces. The non-transitory storage medium can include a data storage device such as storage device 1130 or local storage device 1140. The identified surfaces can be stored in a particular table such as the table disclosed in Figure 15
[0268] Figure 15 An example of a table for storing an inventory of surfaces identified from a user's local environment is shown in accordance with some embodiments. Surface table 1500 is an example table that can store the results of processing identified surrounding surfaces and properties in a database. Surface table 1500 includes, for example, information about the surfaces within the user's physical environment 1105 with data columns including: surface ID 1510, width 1520, height 1530, orientation 1540, real or virtual indicator 1550, multiple 1560, location 1570, and dot product surface orientation relative to user 1580. Surface table 1500 can have additional columns representing other properties of each surface. One of ordinary skill in the art can appreciate that surface table 1500 can be a table in a relational database or any type of database. Additionally, surface table 1500 can be an array in computer memory (e.g., cache) storing the results of determining surfaces at 430. Figure 4
[0269] Each row 1590 in surface table 1500 may correspond to a surface from the user's physical environment 1105 or a virtual surface that can be displayed to the user 1108 within the user's physical environment 1105. Surface ID 1510 is a column containing a unique identifier (e.g., surface ID) used to uniquely identify a specific surface. The dimensions of a specific surface are stored in width 1520 and height 1530 columns.
[0270] Orientation 1540 is a column indicating the orientation (e.g., vertical, horizontal, etc.) of the surface relative to the user 1108. Real / Virtual 1550 is a column indicating whether a particular surface is located on a real surface / object in the user's physical environment 1105, or whether a particular surface is located on a virtual surface / object, as perceived by the user 1108 using the head-mounted system 1160, which will be generated and displayed by the head-mounted system 1160 within the user's physical environment 1105. The head-mounted system 1160 may have to generate virtual surfaces / objects if the user's physical environment 1105 may not contain enough surfaces, if the matching score analysis does not contain enough suitable surfaces, or if the head-mounted system 1160 may not detect enough surfaces to display the amount of content the user 1108 wishes to display. In these embodiments, the head-mounted system 1160 may search a database of existing virtual objects, which may have appropriate surface sizes to display certain types of elements identified for display. The database may come from storage device 1130 or local storage device 1140. In some embodiments, the virtual surface is created substantially in front of the user, or offsets the forward vector of the head-mounted system 1160 so as not to obstruct the user's and / or the device's primary real-world field of view.
[0271] Multiple 1560 columns are columns indicating whether a surface / object is compatible with elements that display multiple versions (e.g., the element could be...). Figure 12 The second tag 1250, where, for a specific web browser 1110, there can be more than one second (i.e., inactive) tag (e.g., one page per tag). If the value of column 1560 is "multiple", such as for the corresponding Figure 12 The fourth entry in the surface ID column of the virtual Roledex 1190, with a stored value of 1190, and corresponding to... Figure 12 In the case of multiple stacked virtual objects 1194 storing a value of 1194 as the fifth entry in the surface ID column, the system and methods will know whether there are elements that may have multiple versions, such as inactive tags, which are surface types that can accommodate multiple versions.
[0272] Position 1570 is a column indicating the position of a physical surface relative to a reference frame or reference point. For example... Figure 15 The position of a physical surface can be predetermined to be the center of the surface, as shown in the column header for position 1570. In other embodiments, the position can be predetermined to be another reference point face of the surface (e.g., the front, back, top, or bottom of the surface). The position information can be expressed as a vector from the center of the physical surface and / or position information relative to a certain frame of reference or reference point. There can be several ways to express the position in the surface table 1500. For example, the value for the position of surface ID 1194 in the surface table 1500 is expressed in abstract form to show vector information and frame of reference information (e.g., "frame" subscript). x, y, z are 3D coordinates in each spatial dimension, and the frame designates which frame of reference the 3D coordinates are relative to.
[0273] For example, surface ID 1186 shows the position of the center of surface 1186 as (1.3, 2.3, 1.3) relative to the real world origin. As another example, surface ID 1192 shows the position of the center of surface 1192 as (x, y, z) relative to the user frame of reference, and surface ID 1190 shows the position of the center of surface 1190 as (x, y, z) relative to another surface 1182. The frame of reference is important to disambiguate which frame of reference is currently being used. In the case of the real world origin as the frame of reference, it is typically a static frame of reference. However, in other embodiments, when the frame of reference is the user frame of reference, the user can be a moving frame of reference, in which case the plane (or vector information) can move and change with the user if the user is moving and the user frame of reference is used as the frame of reference. In some embodiments, the frame of reference for each surface can be the same (e.g., the user frame of reference). In other embodiments, the frame of reference for a surface stored within the surface table 1500 can be different depending on the surface (e.g., the user frame of reference, the world frame of reference, another surface or object in the room, etc.).
[0274] In the current example, the values stored within the surface table 1500 include the physical surfaces (e.g., vertical surfaces 1182 and 1186, and horizontal surface 1192) and virtual surfaces (e.g., virtual Rolodex 1190 and multiple stacked virtual objects 1194) identified in the user's physical environment 1105 of Figure 12 For example, in the current embodiment, the first entry of the surface ID 1510 column stores the surface ID 1182 for the vertical surface 1182. The second entry of the surface ID 1510 column stores the surface ID 1186 for the vertical surface 1186. The third entry of the surface ID 1510 column stores the surface ID 1190 for the virtual Rolodex 1190. The fourth entry of the surface ID 1510 column stores the surface ID 1192 for the horizontal surface 1192. The fifth entry of the surface ID 1510 column stores the surface ID 1194 for the first virtual object 1194. The sixth entry of the surface ID 1510 column stores the surface ID 1194 for the second virtual object 1194. The seventh entry of the surface ID 1510 column stores the surface ID 1194 for the third virtual object 1194. Figure 12 corresponding to the value of the surface ID 1182 of the vertical surface 1182. The width value in the width 1520 column and the height value in the height 1530 column correspond to the width and height, respectively, of the vertical surface 1182, indicating that the vertical surface 1182 has dimensions of 48 inches (wide) by 36 inches (high). Similarly, the orientation value in the orientation 1540 column indicates that the vertical surface 1182 has an "upright" orientation. Additionally, the real / virtual value in the real / virtual 1550 column indicates that the vertical surface 1182 is an "R" (e.g., real) surface. The multiple value in the multiple 1560 column indicates that the vertical surface 1182 is "single" (e.g., can only hold a single piece of content). Finally, the position 1570 column indicates that the vertical surface 1182 has vector information of (2.5, 2.3, 1.2) for its position relative to the user 1108. 用户 .
[0275] The remaining rows in the surface table 1500 contain information for the remaining surfaces in the user's physical environment 1105. One of ordinary skill in the art can understand that, Figure 4 The storing of the results of determining the surfaces at 430 improves the functionality of the computer itself, as once this analysis has been performed on the surrounding surfaces, the analysis can be retained by the head mounted system 1160 for use in the future to analyze the user's surrounding surfaces if another user or the same user 1108 is in the same physical environment 1105 but interested in different content. The processing steps of determining the surfaces at 430 can be avoided as these processing steps have already been completed. The only difference can include identifying additional or different virtual objects to be available based at least in part on the element table 1400 identifying elements having different content.
[0276] In some embodiments, the surface table 1500 is stored in the storage device 1130. In other embodiments, the surface table 1500 is stored in the local storage device 1140 of the user 1108 for quick access to recently viewed content or possible re-visit of recently viewed content. In other embodiments, the surface table 1500 can be stored at both the storage device 1130 remote from the user 1108 and the local storage device 1140 located locally to the user 1108.
[0277] Returning to Figure 4 The method continues with matching the content elements to surfaces using a combination of the identified content elements from step 420 identifying content elements in the content and the determined surfaces from step 430 determining surfaces, and in some embodiments using virtual objects as additional surfaces at 440. Matching the content elements to surfaces can involve a number of factors, where some factors can include analyzing hints provided by the content designer through content designer defined HTML tag elements by using, for example, an HTML page parser such as the example HTML page parser discussed above. Other factors can include selecting from a predefined set of rules provided by the AR browser, AR interface, and / or cloud storage device regarding how and where to match / map certain content.
[0278] Briefly referring to Figure 7A depicts a flowchart showing a method for matching content elements to surfaces according to some embodiments. At 710, the method determines whether the identified content elements contain hints provided by the content designer. The content designer can provide hints regarding where the content elements are best displayed. For example, Figure 12 The main video 1220 of the web page 1200 can be a video displayed on the web page within the active tab 1260. The content designer can provide hints to indicate that the main video 1220 is best displayed on a flat vertical surface in the direct line of sight of the user 1108.
[0279] In some embodiments, the 3D preview for a web link can be represented as a new set of HTML tags and features associated with a web page. Figure 16 An example 3D preview for a web link is shown according to some embodiments. The content designer can use the new HTML feature to specify which web links have associated 3D previews to be rendered for them. Optionally, the content designer / web developer can specify a 3D model to be used to render the 3D web preview on. If the content designer / web developer specifies a 3D model to be used to render the web preview, the web content image can be used as a texture for the 3D model. A web page can be received. If there are preview features specified for certain link tags, the first level web page can be retrieved and based on the preview features, a 3D preview can be generated and loaded to the 3D model specified by the content designer or a default 3D model (e.g., sphere 1610). Although the 3D preview is described for web links, the 3D preview can be used for other content types. Those skilled in the art can appreciate that there can be many other ways for the content designer to provide hints regarding where a particular content element should be placed in a 3D environment in addition to what has been disclosed here, which are some examples of different ways in which the content designer can provide hints to display certain or all content elements of a web page content.
[0280] In another embodiment, the tag standard (e.g., HTML tag standard) can include additional tags (e.g., HTML tags) such as the example web page provided by the content designer discussed above or a creation of a similar markup language for providing hints. If the tag standard includes these types of additional tags, certain embodiments of the method and system will utilize these tags to further provide the matching / mapping of the identified content elements to the identified surfaces.
[0281] For example, a set of web page components can be exposed as new HTML tags for the content designer / web page developer to create elements of the web page that will themselves appear as 3D volumes that protrude from the 2D web page or are etched into the 2D web page. Figure 17 An example of a web page with 3D volumes etched into the web page is shown (e.g., 1710). These 3D volumes can include web page controls (e.g., buttons, handles, joysticks) that will be placed on the web page, allowing the user to manipulate the web page controls to manipulate the content displayed within the web page. Those skilled in the art will appreciate that there are many other languages besides HTML that can be modified or adopted to further provide hints as to how content elements should be best displayed in a 3D environment, and that the new HTML markup standard is just one way to achieve this purpose.
[0282] At 720, the method determines whether to use the hints provided by the content designer or to use a predefined set of rules to match / map the content elements to the surfaces. At 730, if it is determined that the hints provided by the content designer are the way to proceed, the system and method analyzes the hints and searches the logical structure including the identified surrounding surfaces that can be used to display the particular content element based at least in part on the hints (e.g., query Figure 15 of the surface table 1500).
[0283] At 740, the system and method runs a best fit algorithm to select the best fit surface for the particular content element based on the provided hints. For example, the best fit algorithm can hint the particular content element and try to identify a surface in the environment that is front and center with respect to the user 1108. For example, Figure 12 The main video 1220 is matched / mapped to the vertical surface 1182 because the main video 1220 is within the active tag 1260 Figure 14A The preference attribute 1420 column of the element table 1400 has a preference value of "main" for the main video 1220, and the vertical surface 1182 is the surface that is directly in front of the user 1108 and has the best size dimensions to display the main video 1220.
[0284] At 750, the system and method stores the matching results having the content elements matched to surfaces. The table can be stored in a non-transitory storage medium for use by a display algorithm to display the content elements on their respective matched / mapped surfaces. The non-transitory storage medium can include a data storage device, such as storage device 1130 or local storage device 1140. The matching results can be stored in a particular table, such as the table disclosed below in Figure 18 .
[0285] Figure 18 An example of a table for storing content element to surface matches is shown in accordance with some embodiments. The match / mapping table 1800 is an exemplary table that stores the results of matching content elements to surfaces in a database. The match / mapping table 1800 includes information, for example, about the content elements (e.g., element ID) and the surfaces (e.g., surface ID) to which the content elements are matched / mapped. Those of ordinary skill in the art will appreciate that the match / mapping table 1800 can be a table stored in a relational database or any type of database or storage medium. Additionally, the match / mapping table 1800 can be an array in computer memory (e.g., a cache) that contains the results of matching content elements to surfaces at 440 of FIG. 4. Figure 4 .
[0286] Each row of the match / mapping table 1800 corresponds to a content element that is matched to one or more surfaces in the user's physical environment 1105 or a virtual surface / object displayed to the user 1108 that appears to be a surface / object in the user's physical environment 1105. For example, in the current embodiment, the first entry of the element ID column stores a value of element ID 1220 corresponding to the main video 1220. The surface ID value in the surface ID column corresponding to the main video 1220 is 1182 corresponding to the vertical surface 1182. In this way, the main video 1220 is matched / mapped to the vertical surface 1182. Similarly, the user comment 1230 is matched / mapped to the horizontal surface 1192, the suggested video 1240 is matched / mapped to the vertical surface 1186, and the second tab 1250 is matched / mapped to the virtual Rolodex 1190. The element IDs in the match / mapping table 1800 can be associated with the element IDs stored in the element table 1400 of FIG. 4. The surface IDs in the match / mapping table 1800 can be associated with the surface IDs stored in the surface table 1500 of FIG. 5. Figure 14A . Figure 15 .
[0287] Returning to Figure 7A At point 760, assuming that using predefined rules is the preferred method, this method queries a database containing matching / mapping rules between content elements and surfaces, and determines which type of surface should be considered for matching / mapping the content element for a specific content element within the webpage. For example, for a content element from... Figure 12 The rules for the returned main video 1220 can indicate that the main video 1220 should be matched / mapped to a vertical surface, and thus, after searching the surface table 1500, multiple candidate surfaces are revealed (e.g., vertical surfaces 1182 and 1186, and the virtual Roledex 1190). At 770, a predefined set of rules can be used to run a best-fit algorithm to select which surface from the available candidate surfaces is best suited for the main video 1220. Based at least in part on the best-fit algorithm, it is determined that the main video 1220 should be matched / mapped to the vertical surface 1182, because among all candidate surfaces, the vertical surface 1182 is the surface within the direct line of sight of the user 1108, and the vertical surface 1182 has the optimal size for displaying the video. Once the matching / mapping of one or more elements is determined, at 750, the method, as described above, stores the matching / mapping results of the content elements in an element-to-surface matching / mapping table in a non-transitory storage medium.
[0288] return Figure 4 The method continues at 450 to render content elements as virtual content onto the matching surface. The head-mounted system 1160 may include one or more display devices (such as a microprojector (not shown)) within the head-mounted system 1160 to display information. One or more elements are displayed on the corresponding matching surface as matched at 440. Using the head-mounted system 1160, the user 1108 will see content on the corresponding matching / mapped surface. Those skilled in the art will understand that while content elements are displayed as if physically attached to various surfaces (physical or virtual), in reality, the content elements are actually projected onto the physical surface perceived by the user 1108, and in the case of virtual objects, the virtual objects are displayed as appearing attached to the various surfaces of the virtual objects. Those skilled in the art will understand that as the user 1108 turns their head or looks up or down, the display devices within the head-mounted system 1160 may continue to hold the content elements fixed to their respective surfaces to further provide the user 1108 with the perception that content is fixed to the matching / mapped surface. In other embodiments, user 1108 can alter the content of the user's physical environment 1105 through movements of the user's head, hands, eyes, or voice.
[0289] application
[0290] Figure 19 An example of an environment 1900 including content elements that match a surface, according to some embodiments, is shown.
[0291] With respect to this example, briefly refer to Figure 4 The parser 115 receives 410 the content 110 from the application. The parser 115 identifies 420 content elements in the content 110. In this example, the parser 115 identifies the video panel 1902, the highlight panel 1904, the replay 1906, the graphical statistics 1908, the textual statistics 1910, and the social media news feed 1912.
[0292] The environment parser 168 determines 430 surfaces in the environment. In this example, the environment parser 168 determines a first vertical surface 1932, a second vertical surface 1934, a top of a first ottoman 1936, a top of a second ottoman 1938, and a front of the second ottoman 1940. The environment parser 168 can determine additional surfaces in the environment; however, in this example, no additional surfaces are labeled. In some embodiments, the environment parser 168 continuously determines 430 surfaces in the environment. In some embodiments, the environment parser 168 determines 430 surfaces in the environment as the parser 115 receives 410 the content 110 and / or identifies 420 content elements in the content 110.
[0293] The matching module 142 matches 440 content elements to surfaces based on attributes of the content elements and attributes of the surfaces. In this example, the matching module 142 matches the video panel 1902 to the first vertical surface 1932, matches the highlight panel 1904 to the second vertical surface 1934, matches the replay 1906 to the top of the first ottoman 1936, matches the graphical statistics 1908 to the top of the second ottoman 1938, and matches the textual statistics 1910 to the front of the second ottoman 1940.
[0294] The optional virtual object creation module 144 can create virtual surfaces for displaying content elements. During the matching process of the matching module 142, it can be determined that a virtual surface can be an optional surface for displaying certain content elements. In this example, the optional virtual object creation module 144 creates a virtual surface 1942. The social media news feed 1912 is matched to the virtual surface 1942. The rendering module 146 renders 450 the content elements to their matched surfaces. The resulting Figure 19 This shows what the user of the head-mounted display device running the application would see after the rendering module 146 renders 450 the content elements to their matched surfaces.
[0295] Dynamic Environment
[0296] In some embodiments, the environment 1900 is dynamic: the environment itself is changing and objects move in / out of the user's and / or device's field of view to create new surfaces, or the user moves to a new environment and simultaneously receives content elements such that the previously matched surfaces no longer conform to the previous composition process 140 results. For example, while watching a basketball game in the environment 1900 as shown, Figure 19 the user can walk into the kitchen.
[0297] While one skilled in the art will appreciate that the following techniques will apply to changing environments with respect to static users, the following Figures 20A-20E depicts changes in the environment according to the user's movement. In Figure 20A the user is watching spatialized display of content after the composition process 140 described throughout this disclosure. Figure 20B a larger environment in which the user can be immersed is shown, as well as additional surfaces eligible for the user.
[0298] As shown in Figure 20C when the user moves from one room to another, it becomes apparent that the content originally rendered for display in Figure 20A no longer satisfies the match of the composition process 140. In some embodiments, the sensors 162 cue the system to a change in the user's environment. The change in the environment can be a change in depth sensor data (a new virtual grid structure is generated by the room on the Figure 20C left than the room on the Figure 20C right in which the content was originally rendered for display), a change in head pose data (movement produced by the IMU changes beyond a threshold for the current environment, or a camera on the headset begins to capture new objects in its and / or the device's field of view). In some embodiments, the change in the environment initiates a new composition process 140 to find new surfaces for the content elements that were previously matched and / or currently rendered and displayed. In some embodiments, a change in the environment that exceeds a time threshold initiates a new composition process 140 to find new surfaces for the content elements that were previously matched and / or currently rendered and displayed. The time threshold can avoid wasted computational cycles for minor disruptions to the environment data (e.g., simply turning one's head to talk to another user, or a brief exit from the environment that the user returns from shortly).
[0299] In some embodiments, when the user enters the room 2002, Figure 20D the composition process 140 matches the active content 2002 with the room 2014 with new surfaces. In some embodiments, the active content 2002 is now active content in both rooms 2012 and 2014, although only displayed in room 2012 ( Figure 20D the appearance of the active content 2002 in room 2014 in depicts that the active content 2002 is still rendered, although not displayed to the user).
[0300] In this manner, the user can walk between rooms 2012 and 2014 and the synthesis process 140 need not continuously repeat the matching protocol. In some embodiments, when the user is in room 2012, the active content 2002 in room 2014 is set to an idle or sleep state, and similarly, if the user returns to room 2014, the active content in room 2012 is put into an idle or sleep state. Thus, the user can automatically continue consuming content as they dynamically change their environment.
[0301] In some embodiments, the user can pause active content 2002 in room 2014 and enter room 2012, and resume the same content at the same interaction point as where they paused in room 2014. Thus, the user can automatically continue content consumption as they dynamically change their environment.
[0302] The idle or sleep state can be characterized by the degree of output that the content element performs. Active content can have the full capability of rendered content elements, such that frames of content continue to update their matching surfaces, audio output continues to the virtual speakers associated with the matching surface locations, and so on. The idle or sleep state can reduce some of this functionality; in some embodiments, the audio output of the idle or sleep state is reduced in volume or enters a mute state; in some embodiments, the rendering cycle is slowed such that fewer frames are generated. Such slower frame rates can conserve overall computing power, but introduce a small delay when resuming content element consumption if the idle or sleep state returns to an active state, such as the user returning to the room in which the idle or sleep state content element was running.
[0303] Figure 20E It is depicted that rendering of content elements is stopped in different environments, not just changed to an idle or sleep state. In Figure 20E In some embodiments, the trigger for stopping rendering is changing the content element from one source to another, such as changing the channel of a video stream from a basketball game to a movie; in some embodiments, active content is immediately stopped from rendering once the sensors 162 detect the new environment and the new synthesis process 140 begins.
[0304] In some embodiments, for example, based at least in part on a user moving from a first location to a second location, content rendered and displayed on a first surface in the first location can be paused and then resumed on a second surface in the second location. For example, a user viewing content displayed on a first surface in a first location (e.g., a living room) can physically move from the first location to a second location (e.g., a kitchen). Upon determining (e.g., based on sensors 162) that the user has physically moved from the first location to the second location, the rendering and / or display of the content on the first surface in the first location can be paused. Once the user moves to the second location, sensors of the AR system (e.g., sensors 162) can detect that the user has moved to the new environment / location, and the environment resolver 168 can begin to identify a new surface in the second location, and then can resume displaying the content on a second surface in the second location. In some embodiments, the content can continue to be rendered on the first surface in the first location as the user moves from the first location to the second location. Once the user is in the second location, for example, after a threshold period of time (e.g., 30 seconds), the content can stop being rendered on the first surface in the first location and can be rendered on the second surface in the second location. In some embodiments, the content can be rendered on both the first surface in the first location and the second surface in the second location.
[0305] In some embodiments, the rendering and / or display of content at a first surface in a first location can be automatically paused in response to a user physically moving from the first location to a second location. Detection of the user's physical movement can trigger the automatic pausing of the content, where the triggering of the user's physical movement can be based at least in part on an inertial measurement unit (IMU) exceeding a threshold or a location indication indicating that the user has moved out of or is moving out of a predetermined area (e.g., GPS) that can be associated with the first location. Once a second surface is identified, e.g., by the environment parser 168, and matches the content, the content can automatically resume rendering and / or display on the second surface in the second location. In some embodiments, the content can resume rendering and / or display on the second surface based at least in part on a user selection of the second surface. In some embodiments, the environment parser 168 can refresh within a certain schedule (e.g., every 10 seconds) to determine whether the surface within the user's and / or device's field of view has changed and / or the user's physical location has changed. If it is determined that the user has moved to a new location (e.g., the user moved from the first location to the second location), the environment parser 168 can begin identifying new surfaces within the second location for resuming the rendering and / or display of the content on the second surface. In some embodiments, the content rendered and / or displayed on the first surface can not automatically pause immediately due to the user changing the field of view (e.g., the user briefly looks at another person in the first location to, e.g., have a conversation). In some embodiments, the rendering and / or display of the content can be automatically paused if the user's changing the field of view exceeds a threshold. For example, if the user changes the head pose and thus the corresponding field of view exceeds a threshold for a period of time, the display of the content can be automatically paused. In some embodiments, the content can automatically pause rendering and / or display the content on the first surface in the first location in response to the user leaving the first location, and the content can automatically resume rendering and / or display on the first surface in the first location in response to the user physically (re)entering the first location.
[0306] In some embodiments, the content on a particular surface can slowly follow the changes in the user's field of view as the user and / or the user's head mounted device changes the field of view. For example, the content can be within the user's direct field of view. If the user changes the field of view, the content can change position to follow the change in the field of view. In some embodiments, the content can not immediately be displayed on the surface that is within the changed field of view. Rather, there can be a slight delay in the change of the content relative to the change in the field of view, where the change in the content position can appear to slowly follow the change in the field of view.
[0307] Figures 20F-20I An example is shown in which content displayed on a particular surface can slowly follow changes in the field of view of a user currently viewing the content. In Figure 20F In this example, user 1108 is viewing a spatialized display of content on a couch in a room in a seated position that has a first head pose in which the user and / or the user’s head-mounted device is facing, for example, a main wall 1180. As shown, the spatialized display of content is displayed at a first location (e.g., rectangular surface 1182) of main wall 1180 by first head pose. Figure 20F Figure 20G As shown, user 1108 changes from a seated position to a reclined position on the couch, the reclined position having a second head pose in which the user is facing, for example, a side wall 1184 instead of main wall 1180. The content displayed on rectangular surface 1182 can continue to be rendered / displayed at rectangular surface 1182 until a time threshold and / or a head pose change threshold has been reached / exceeded. Figure 20H As shown, the content can slowly follow the user, i.e., move in small discrete incremental positions to a new location corresponding to the second head pose facing side wall 1185 instead of a single update, and appear to be displayed at a first display option / surface 2020, for example, after some point in time after user 1108 has changed from a seated position to a reclined position (e.g., after some time threshold). First display option / surface 2020 can be a virtual display screen / surface within the field of view corresponding to the second head pose, as there is no optimal surface available within the direct field of view of user 1108. Figure 20I As shown, the content can also be displayed at a second display option / surface at rectangular surface 1186 on side wall 1184. As described above, in some embodiments, user 1108 can be provided with display options to select which display options to display content (e.g., first display option 2020 or second rectangular surface 1186) based on changes in the field of view of user 1108 and / or the device.
[0308] In some embodiments, for example, a user can be viewing content that is displayed in a first field of view directly in front of the user. The user can rotate their head 90 degrees to the left and hold the second field of view for about 30 seconds. The content that is displayed in the first field of view directly in front of the user can slowly follow the user to the second field of view by moving 30 degrees to the second field of view relative to the first field of view the first time to slowly follow the user after a certain time threshold (e.g., 5 seconds) has elapsed. The AR system can move the content an additional 30 degrees to follow the user to the second field of view the second time, such that the content is now displayed 30 degrees behind the second field of view.
[0309] Figures 20J-20N As shown, content slowly follows a user from a first field of view to a second field of view of the user and / or the user’s device according to some embodiments. Figure 20J A top-down view showing user 2030 viewing content 2034 displayed on a surface (e.g., a virtual surface in a physical environment or an actual surface). User 2030 is viewing content 2034 such that the entire content 2034 is displayed directly in front of user 2030 and is entirely within a first field of view 2038 of user 2030 and / or the device at a first head pose position of the user and / or the user’s device. Figure 20K A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown. Figure 20J A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown.
[0310] Figure 20L A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown. Figure 20J A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown. Figure 20J A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown.
[0311] Figure 20M A top-down view showing user 2030 as an example rotated about 45 degrees to the right (e.g., in a clockwise direction) from the first head pose position shown. Figure 20N It is shown that the content 2034 has completed its slow movement to fully follow the user's second head pose position. The content 2034 is fully within the field of view 2038 of the user 2030 and / or device, as shown by the solid line encompassing the entire content 2034. As can be appreciated by one of ordinary skill in the art, although the user can have changed the field of view from the first field of view to a second field of view in which the content is no longer visible, the user can not wish to have the content displayed directly in the second field of view. Instead, the user can wish for the content to slowly follow the user to the second field of view (e.g., new field of view) without being displayed directly in front of the user until, for example, the system prompts the user to select whether the user wishes to have the content displayed directly in front of the user with respect to the second field of view of the user, or just have the user be able to view the displayed content peripherally until the user re-engages with the displayed content that the user can view peripherally. In other words, in some embodiments, the display of the content / element on the surface(s) can be moved in response to the change in the field of view of the user from the first field of view to the second field of view, where the content / element slowly follows the change in the field of view of the user from the first field of view to the second field of view. Further, in some embodiments, the content can be moved to be directly in front of the second field of view only upon receiving confirmation from the user to move the content to be directly in front of the second field of view.
[0312] In some embodiments, the user can (a) view the extracted content element displayed to the user via the AR system, and (b) interact with the extracted content element. In some embodiments, the user can interact with the extracted content by purchasing an item / service displayed within the extracted content. In some embodiments, similar to online purchases made by users interacting with 2D webpages, the AR system can allow the user to interact with the extracted content displayed on a surface and / or virtual object (e.g., a prism or virtual display screen) within the AR system to, for example, electronically purchase items and / or services presented within the extracted content displayed on the surface and / or virtual object of the AR system.
[0313] In some embodiments, a user can interact with the extracted content elements by further selecting items within the displayed content elements and placing the selected items on different surfaces and / or different virtual objects (e.g., prisms) within the user’s physical environment. For example, a user can extract a content element such as an image, a video, and / or a model from a gallery by, for example, (a) aiming a totem at a content element in the gallery, (b) pressing a trigger on the totem to select the content element and holding for a period of time (e.g., about 1 second), (c) moving the totem around to a desired location in the user’s physical environment, and (d) pressing the trigger on the totem to place the content element at the desired location, where a copy of the content element is loaded and displayed at the desired location. In some embodiments, as a result of the user selecting the content element and holding the trigger for a period of time, a preview of the content element is created and displayed as visual feedback, as creating a full resolution version of the content element to place the content element can consume greater resources. In some embodiments, when the user places the extracted content element in a desired location in the user’s physical environment, the entire content element is copied / extracted and displayed for visual feedback.
[0314] Figure 20O An example is shown in which a user views extracted content and interacts with extracted content 2050 and 2054. User 2040 can be viewing extracted content 2044a-2044d on a virtual display surface as sensors 1162 are unable to detect a suitable display surface for displaying the extracted content (e.g., due to a bookshelf). Instead, extracted content 2044a-d is displayed on multiple virtual display surfaces / screens. Extracted content 2044a is an online website selling audio headphones. Extracted content 2044b is an online website selling athletic shoes. Extracted content 2044c / 2044d is an online furniture website selling furniture. Extracted content 2044d can include a detailed view of a particular item (e.g., chair 2054) displayed from extracted content 2044c. User 2040 can interact with the extracted content by selecting a particular item from the displayed extracted content and placing the extracted item in the user’s physical environment (e.g., chair 2054). In some embodiments, user 2040 can interact with the extracted content by purchasing a particular item (e.g., athletic shoes 2050) displayed in the extracted content.
[0315] Figure 21 Audio transfer during such environmental changes is shown. Active content in room 2014 can have a virtual speaker 2122 that delivers spatialized audio to the user, for example, from a location associated with the content element in room 2014. As the user moves to room 2012, the virtual speaker can follow the user by positioning and directing audio to a virtual speaker 2124 in the center of the user's head (much like the way a traditional headset would) and stopping audio playback from virtual speaker 2122. When the synthesis process 140 matches the content element to a surface in room 2012, the audio output can be transferred from virtual speaker 2124 to virtual speaker 2126. In this case, the audio output maintains a constant consumption of the content element, at least the audio output component, during the environmental transfer. In some embodiments, the audio component is always a virtual speaker in the center of the user's head, thereby eliminating the need to adjust the location of the spatialized audio virtual speaker.
[0316] System Architecture Overview
[0317] Figure 22 is a block diagram of an illustrative computing system 2200 suitable for implementing embodiments of the present disclosure. The computing system 2200 includes a bus 2206 or other communication mechanism for communicating information, and a processor 2207 coupled with the bus 2206 for processing information. The computing system 2200 also includes a system memory 2208, such as one or more of random access memory (RAM) 2210, and a read only memory (ROM) 2212 for storing instructions for the processor 2207. A basic input / output system (BIOS) 2214, containing the basic routines that help to transfer information between elements within the computing system 2200, such as during startup, is typically stored in the ROM 2212. The computing system 2200 further includes a magnetic hard disk drive 2216 for storing information, such as an operating system 2218, and one or more application programs 2220, according to an embodiment of the present disclosure. Also
[0318] According to an embodiment of the present disclosure, the computing system 2200 performs a particular operation involving the execution of one or more sequences of instructions contained in system memory 2208 by processor 2207. Such instructions can be read into system memory 2208 from another computer readable medium, such as the static storage device 2209 or the disk drive 2210. In alternative embodiments, hard-wired circuitry can be used in place of or in combination with software instructions to implement the present disclosure. Thus, embodiments of the present disclosure are not limited to any specific combination of hardware circuitry and / or software. In one embodiment, the term "logic" shall mean any combination of software or hardware that is used to implement all or part of the present disclosure.
[0319] The term "computer readable media" or "computer usable medium" as used herein refers to any media that participates in providing instructions to processor 2207 for execution. Such media can take many forms, including but not limited to, nonvolatile media and volatile media. Nonvolatile media includes, for example, optical or magnetic disks, such as disk drive 2210. Volatile media includes dynamic memory, such as system memory 2208.
[0320] Common forms of computer readable media include, for example, floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, a FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read.
[0321] In embodiments of the present disclosure, execution of sequences of instructions for practicing the present disclosure is performed by a single computing system 2200. According to other embodiments of the present disclosure, two or more computing systems 2200 coupled by communication link 2215 (e.g., LAN, PTSN, or wireless network) can perform in conjunction with one another to execute the sequences of instructions required to practice the present disclosure.
[0322] Computing system 2200 can send and receive messages, data, and instructions, including program (i.e., application code) over communications link 2215 and communication interface 2214. Received program code can be executed by processor 2207 as it is received, and / or stored in disk drive 2210, or other non-volatile storage for later execution. Computing system 2200 can communicate with a database 2232 on external storage device 2231 through data interface 2233.
[0323] In the foregoing specification, the disclosure has been described with reference to specific embodiments thereof. It will, however, be evident that various modifications and changes can be made thereto without departing from the broader spirit and scope of the disclosure. For example, the order in which steps are described in the above processes flow can be changed. In addition, the above-described processes flow can be performed by more than one entity, and performed in a distributed manner. Thus, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method for matching content to a plurality of surfaces of an environment of a user, the method comprising: identifying a content element having a plurality of different attributes corresponding to a plurality of different attributes of each of a plurality of surfaces; determining a plurality of different attributes of each of the plurality of surfaces corresponding respectively to the plurality of different attributes of the content element; comparing the plurality of different attributes of the content element to the plurality of different attributes of each of the plurality of surfaces respectively; based on respective comparisons, computing a plurality of scores of respective plurality of surfaces; selecting a surface having a highest score from the plurality of surfaces; storing a mapping of the content element to the selected surface; and displaying the content element to the user on the selected surface. The identified content element is a 3D content element.
2. The method of claim 1, wherein, The plurality of different attributes of the content element are weighted differently.
3. The method of claim 1, wherein, The plurality of different attributes of the content element include dot product orientation surface relationships, texture, and color.
4. The method of claim 1, wherein, The surface on which the content element is displayed to the user is the selected surface.
5. The method of claim 1, wherein, The highest score is compared to a threshold score, based on the comparison, displaying the content element on the selected surface or on a virtual surface.
6. The method of claim 1, further comprising: The content element is displayed on the selected surface if the threshold score is greater than the threshold score and on the virtual surface if the threshold score is less than the threshold score.
7. The method of claim 6, wherein, Overriding the selected surface and selecting another surface, wherein the surface on which the content element is displayed to the user is the other surface.
8. The method of claim 1, further comprising: Moving the displayed content element from the surface to another surface.
9. The method of claim 1, further comprising: Moving the displayed content element from the surface to another surface by a gesture of the user.
10. The method of claim 1, wherein, 11. An augmented reality (AR) display system comprising: a head mounted system comprising: one or more sensors, and one or more cameras including an outward facing camera; a processor executing a set of program code instructions; and a memory storing the set of program code instructions, wherein the set of program code instructions includes program code that: identifies a content element having a plurality of different attributes corresponding to a plurality of different attributes of each of a plurality of surfaces; determines a plurality of different attributes of each of the plurality of surfaces corresponding respectively to the plurality of different attributes of the content element; compares the plurality of different attributes of the content element to the plurality of different attributes of each of the plurality of surfaces respectively; based on respective comparisons, computes a plurality of scores of respective plurality of surfaces; selects a surface having a highest score from the plurality of surfaces; stores a mapping of the content element to the selected surface; and displays the content element to a user on the selected surface. The identified content element is a 3D content element.
12. The system of claim 11, wherein, The plurality of different attributes of the content element are weighted differently.
13. The system of claim 11, wherein, The plurality of different attributes of the content element include dot product orientation surface relationships, texture, and color.
14. The system of claim 11, wherein, The surface on which the content element is displayed to the user is the selected surface.
15. The system of claim 11, wherein, 16. The system of claim 11, wherein, The program code also performs comparing the highest score to a threshold score, displaying the content element on a selected surface or on a virtual surface based on the comparison.
17. The system of claim 16, wherein, The content element is displayed on the selected surface if the threshold score is greater than the threshold score and on the virtual surface if the threshold score is less than the threshold score.
18. The system of claim 11, wherein, The program code also performs covering the selected surface and selecting another surface, wherein the surface on which the content element is displayed to the user is the other surface.
19. The system of claim 11, wherein, The program code also performs moving the displayed content element from the surface to another surface.
20. The system of claim 11, wherein, The program code also allows moving the displayed content element from the surface to another surface by a gesture of the user.
Citation Information
Patent Citations
Selecting virtual objects in a three-dimensional space
US10521025B2
Automatic placement of a virtual object in a three-dimensional space
US20180045963A1
Planar waveguide apparatus with diffraction element(s) and system employing same
US9671566B2
Using object recognizers in an augmented or virtual reality system
US9761055B2
User Interface for Augmented Reality Enabled Devices
US20140168262A1