Intelligent content rendering method and device on augmented reality system
By integrating camera devices and processors in augmented reality devices, tracking and object recognition of real-world environments is solved, and the problem in the prior art is difficult to distinguish between the situation where the user interacts with the physical world and the situation where the user interacts with the virtual world, and the adaptability of user experience and content rendering is improved.
Patent Information
- Application Number
- CN202380076824.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-13
AI Technical Summary
When existing augmented reality systems provide context analysis and intelligent content rendering, it is difficult to effectively distinguish the user's interaction with the physical world and the user's interaction with the virtual world, resulting in a reduced user experience.
By integrating camera devices and processors in an augmented reality device, tracking and object recognition of real-world environments is achieved, areas of interest in the scene are determined, and locations of content provided on the display are adaptively changed according to environmental interactions.
Effectively distinguish the interaction between users and the physical world and the interaction between users and the virtual world, improve user experience, and reduce the interference of AR content on real-world interaction.
Smart Images

Figure CN120153338A_ABST
Abstract
Description
Technical Field
[0001] The various examples of the present disclosure generally may relate to methods, apparatuses, and computer program products for providing context analysis and intelligent content rendering on an augmented reality system. Background Art
[0002] Augmented reality is a form of reality that has been adjusted in some way before being presented to a user, and this form of reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. AR devices, VR devices, MR devices, and hybrid reality devices typically provide content through a visual device (e.g., through a headset, such as glasses).
[0003] Many augmented reality devices utilize a camera to present information for rendering added information and / or content on the physical world and can perform various AR operations and simulations. For example, an augmented reality device can display a hologram overlaid on a screen or a display.
[0004] However, real content and augmented content may interfere with each other and / or make it difficult for a user to focus on the content of interest. For example, the overlaid holographic content may obscure the view on a mobile phone screen or a wearable display, resulting in a degraded user experience. In another example, a user may attempt to read a book, watch TV, talk to someone, or otherwise view and / or interact with one or more objects in the physical world, but the AR content may distract from and interfere with the real-world interaction. Such distractions can be problematic and may cause a person to become disengaged from their real-world surroundings. Therefore, there is a need to assist in distinguishing between situations where an AR user intends to interact with the physical world or with the virtual world and to determine, in context, for example, when an interaction (e.g., a physical interaction versus a virtual interaction) may be appropriate or desired. Summary of the Invention
[0005] To address the described challenges, the present disclosure provides systems and methods for operating an augmented reality device.
[0006] According to one aspect, a device is provided that includes: a display configured to provide content in a visible area of the display; a camera device configured to track a scene of a real-world environment captured in a field of view of the camera device; one or more processors and a non-transitory memory including computer-executable instructions that, when executed, cause the device to at least perform the following operations: determine a region of interest in the scene; perform object recognition on the scene tracked by the camera device; determine an environmental interaction based on the object recognition and the region of interest; and adaptively change a position of the content provided by the display based on the environmental interaction.
[0007] In some embodiments, the environmental interaction includes at least one of the following: an approaching object; an approaching person; a departing object; a departing person; an interaction with one or more objects; an interaction with one or more persons; or a gesture.
[0008] In some embodiments, when the one or more processors further execute the instructions, the device is configured to: determine that the region of interest transitions from the display to the scene, or the region of interest transitions from the scene to the display; in an instance where the region of interest transitions from the display to the scene, minimize or reduce a size of the content; and in an instance where the region of interest transitions from the scene to the display, maximize or increase a size of the content.
[0009] In some embodiments, the device includes an augmented reality device. The augmented reality device can be a head-mounted device.
[0010] In some embodiments, the device further includes: a second camera device configured to track a gaze of at least one eye, wherein the second camera device includes a left-eye tracking camera and a right-eye tracking camera.
[0011] In some embodiments, the camera device includes at least one outward-facing camera for tracking the scene. In some embodiments, the display is configured to provide holographic content including augmented reality content.
[0012] In some embodiments, the scene captured in the field of view is associated with a gaze of at least one eye tracked by the second camera device.
[0013] In some embodiments, when the one or more processors further execute the instructions, the device is configured to: determine the region of interest by determining a gaze depth of at least one eye captured by the second camera device or a gaze direction of the at least one eye.
[0014] In some embodiments, the object recognition identifies a person or an object in the scene.
[0015] According to another aspect, a method is provided that includes: performing object recognition on a scene of a real-world environment captured in the field of view of a camera device; determining a region of interest in the scene; determining an environmental interaction based on the object recognition and the region of interest; and adaptively changing the position of content provided by a display based on the environmental interaction.
[0016] In some embodiments, the method further includes: determining a region of interest based on the gaze of at least one eye tracked by a second camera device; determining a transition of the region of interest between the display and the scene based on the gaze; minimizing or reducing the size of the content in an instance where the region of interest transitions from the display to the scene; and maximizing or increasing the size of the content in an instance where the region of interest transitions from the scene to the display.
[0017] In some embodiments, the method further includes: associating a second region of interest with the environmental interaction; associating a third region of interest with the content; and moving the position of the content to reduce interference between the third region of interest associated with the content and the region of interest associated with the environmental interaction.
[0018] In some embodiments, the method further includes: applying one or more machine learning techniques to determine the environmental interaction based on training data associating one or more tracked gazes with one or more scenes.
[0019] In some embodiments, the method further includes: adaptively changing the position of the content by the device.
[0020] According to yet another aspect, a computer-readable medium is provided that stores instructions that, when executed, cause: performing object recognition on a scene tracked by a camera device; determining a region of interest in the scene; determining an environmental interaction based on the object recognition and the region of interest; and adaptively changing the position of content provided on a display based on the environmental interaction.
[0021] In some embodiments, the object recognition is performed continuously in real time.
[0022] In some embodiments, the instructions, when executed, further cause: determining a region of interest corresponding to the gaze of at least one eye tracked by a second camera device, wherein the region of interest is based on the gaze depth or the gaze direction of the at least one eye.
[0023] In some embodiments, the instructions, when executed, further cause: moving the position of the content from the display, reducing the content from the display, or minimizing the content from the display in an instance where an interesting object approaches at a predetermined threshold speed.
[0024] Some examples can include a gaze tracking camera device configured to track gaze.
[0025] The computer-executable instructions can also determine that a region of interest transitions from a display to a scene, or from a scene to a display, based on the gaze; in an instance where the region of interest transitions from the display to the scene, minimize or reduce the size of the content; and in an instance where the region of interest transitions from the scene to the display, maximize or increase the size of the content. In other examples, in an instance where an object of interest approaches at a predetermined threshold speed, the position of the content on the display can be moved, reduced, and / or minimized.
[0026] In some examples of the present disclosure, one or more camera devices and a display can be mounted on an augmented reality device (e.g., a head-mounted device). In various examples of the present disclosure, the augmented reality device can also include glasses (e.g., smart glasses), a head-mounted viewer, a display, a microphone, a speaker, and any combination of various peripheral devices, as well as a computing system. At least one camera device can include a left-eye tracking camera and a right-eye tracking camera. At least one camera device can also include at least one outward-facing camera configured to track a scene captured in the field of view of the camera device. The display can provide holographic content, and in some examples, the field of view can correspond to the gaze of at least one eye of a user captured by the camera device. In some examples, determining a region of interest can be associated with determining the gaze and / or gaze direction of at least one eye.
[0027] Examples of the present disclosure can include one or more machine learning modules and techniques for determining environmental interactions. The training data can include the association between the tracked gaze and the scene. Machine learning algorithms can be used to adaptively change the position of the content. Object recognition can also be performed continuously in real time. Additional advantages will be partially set forth in the description that follows, or can be learned by practice. These advantages will be realized and obtained by the elements and combinations particularly pointed out in the appended claims. It will be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive, as claimed.
[0028] In another example of the present disclosure, a system is provided. The system can include: one or more processors and a memory including computer program code instructions. The system can also include a computer-readable medium storing instructions that, when executed, cause: performing object recognition on a scene tracked by a camera device; determining a region of interest in the scene; determining an environmental interaction based on the object recognition and the region of interest; and adaptively changing the position of the content provided on a display based on the environmental interaction.
[0029] It will be recognized that any feature described herein as being suitable for incorporation into one or more aspects or embodiments of the present disclosure is intended to be generalizable to any and all aspects and embodiments of the present disclosure. Those skilled in the art can understand other aspects of the present disclosure based on the specification, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are merely exemplary and explanatory and do not limit the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The invention content and the following detailed description can be further understood when read in conjunction with the accompanying drawings. For the purpose of illustrating the disclosed subject matter, examples of the present disclosure are shown in the drawings; however, the disclosed subject matter is not limited to the specific methods, compositions, and devices disclosed. Additionally, the drawings are not necessarily to scale. In the drawings:
[0031] Figure 1 A side view of an augmented reality system according to an example of the present disclosure is shown.
[0032] Figure 2 An internal view and an external view of an augmented reality system according to an example of the present disclosure are shown.
[0033] Figure 3 A flowchart of performing context analysis according to an example of the present disclosure is shown.
[0034] Figure 4 A flowchart of adaptively changing content according to an example of the present disclosure is shown.
[0035] Figure 5 Another flowchart of adaptively changing content according to an example of the present disclosure is shown.
[0036] Figure 6 An augmented reality system including a head-mounted viewer according to an example of the present disclosure is shown.
[0037] Figure 7 A block diagram of an example device according to an example of the present disclosure is shown.
[0038] Figure 8 A block diagram of an example computing system according to an example of the present disclosure is shown.
[0039] Figure 9 A machine learning and training model according to an example of the present disclosure is shown.
[0040] Figure 10 A computing system according to an example of the present disclosure is shown.
[0041] These figures depict various embodiments for illustrative purposes only. Those skilled in the art will readily recognize, based on the following discussion, that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles described herein. Detailed Description
[0042] Some embodiments of the present invention will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the invention are shown. In fact, the various embodiments of the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Like reference numerals refer to like elements throughout. As used herein, the terms "data", "content", "information" and similar terms may be used interchangeably to refer to data capable of being sent, received, and / or stored according to embodiments of the present invention. Additionally, as used herein, the term "exemplary" is not used to convey any qualitative assessment but is used merely to convey an illustration of an example.
[0043] Furthermore, as used in the specification including the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an" and "the" include the plural, and references to a particular numerical value include at least that particular value. As used herein, the term "plurality" means more than one. When expressing a range of values, another example includes starting from a particular value and / or ending at another particular value. Similarly, when a value is expressed as an approximation by use of the antecedent "about", it will be understood that the particular value forms another example. All ranges are inclusive of the end points and combinable. It will be understood that the terms used herein are for the purpose of describing particular aspects only and are not intended to be limiting.
[0044] As defined herein, a "computer-readable storage medium" refers to a non-transitory physical or tangible storage medium (e.g., a volatile or non-volatile storage device), which can be distinguished from a "computer-readable transmission medium" that refers to an electromagnetic signal.
[0045] As mentioned herein, holographic content may represent two-dimensional or three-dimensional virtual objects and / or interactive applications.
[0046] As mentioned herein, the Metaverse can represent an immersive virtual space or immersive virtual world in which devices can be used in a network where there can, but does not have to, exist one or more social connections among users in the network or with the environment in the virtual space or virtual world. The Metaverse or Metaverse network can be associated with three-dimensional (3D) virtual worlds, online games (e.g., video games), one or more content items (e.g., images, videos, non-fungible tokens (NFTs)), and wherein the content items can, for example, be purchased using digital currency (e.g., cryptocurrency) and other suitable currency. In some examples, the Metaverse or Metaverse network can implement the generation and provision of an immersive virtual space in which remote users can (including by using augmented reality / virtual reality / mixed reality) socialize, collaborate, learn, shop, and / or participate in various other activities within the virtual space.
[0047] It will be recognized that certain features of the disclosed subject matter described herein in the context of multiple separate examples for clarity can also be provided combinatorially in a single embodiment. Conversely, various features of the disclosed subject matter described in the context of a single example embodiment for brevity can also be provided separately or in any sub-combination. Additionally, any reference to a value stated in a range includes each value within that range.
[0048] In various aspects, systems and methods can implement context analysis and understanding to enhance the user experience and interaction with AR technology. Examples can combine gaze tracking technology with the tracked scene to identify at least one region of interest, determine environmental interactions, and adaptively change the position of content on a display based on the environmental interactions.
[0049] Examples can utilize gaze tracking technology and object recognition technology to detect whether a user is viewing one or more objects in the physical world (e.g., a phone, a book, or a newspaper) and / or attempting to interact with the one or more objects. In instances where such viewing and / or interaction with the physical world is determined, content on a display associated with an AR device (e.g., AR glasses or other head-mounted AR device) can be moved, minimized, phased out, and / or adaptively adjusted to facilitate the intended user interaction with the physical world.
[0050] In various examples, aspects discussed herein may also be enhanced and applied to smartphones, gaming devices, and other handheld technologies. For example, context analysis can determine whether a user is actively interacting with a phone (e.g., tapping or scrolling on a screen), and such context clues can cause the AR device to pause, move, and / or progressively eliminate content that may obscure the user's view of interacting with the phone. In some examples, the AR content can be moved to a different location that does not obscure the user's line of sight / field of view of the phone.
[0051] Thus, by performing context analysis on what the user is looking at and interacting with, the system, method, and device can intelligently determine how and when to render AR content and enable the user to stay and interact in the physical world.
[0052] Figure 1 A side view of an example AR system in accordance with aspects discussed herein is shown. Device 110 may be an AR device (e.g., Figure 6 augmented reality system 600 in), such as AR glasses or other head-mounted AR devices. Device 110 may be configured to utilize virtual reality, augmented reality, mixed reality, hybrid reality, virtual reality reality, or some combination thereof. Device 110 may include a display on which a user can view content. The content may include holographic content, a display of the physical world (e.g., a real-world environment), or a combination of both.
[0053] In some examples, AR device 110 may include one or more camera devices 112, 114, 116, 120, 122. An eye-tracking camera device including one or more of the one or more cameras 120, 122 may include at least one camera that tracks the eye 115 to capture the user's gaze. The scene-tracking camera devices 112, 114, 116 may include at least one camera that tracks the scene 170 in the physical world. One or more of the cameras in the eye-tracking camera devices 120, 122 may point inward, while one or more of the cameras in the scene-tracking camera devices 112, 114, 116 may point outward to capture the scene.
[0054] As Figure 1As shown, the individual cameras 112, 114, 116, 120, 122 can capture different fields of view. For example, in the outward-facing scene-tracking camera devices 112, 114, 116, one or more of the cameras 112, 114, 116 can each capture a corresponding field of view 140, 150, 160 of their own scene in the physical world. The fields of view 140, 150, 160 can be combined to form the captured scene 170. In some examples, the captured scene can be displayed on the display 105 to allow the user to view the physical world.
[0055] The inward-facing eye-tracking camera devices similarly include cameras 120, 122 that have corresponding fields of view 125, 127 and an overlapping field of view 135 that includes the eyes 115. The corresponding fields of view 125, 127 can capture images of one or both of the user's eyes and track eye movement to determine and track the gaze of one or both eyes. Each of the corresponding fields of view 125, 127 of the cameras 120, 122 can provide information that can be combined to more accurately evaluate eye movement and track the gaze of one or both eyes.
[0056] Gaze tracking can analyze various positions and characteristics of one or both eyes to, for example, determine that the user is focusing in a particular direction and / or at a particular depth (e.g., gaze direction and / or gaze depth). According to one aspect, at least one machine learning module can be applied / implemented to analyze eye movement patterns and determine the gaze associated with a behavior. For example, a gaze that includes repeated left-eye and right-eye movements can indicate that the user is reading. A long fixation on a particular location can indicate that the user is viewing an object. Eye movement and characteristics (e.g., pupil size) can indicate whether the user is looking at something near or far. A decrease in pupil size can indicate that one or both eyes are focusing on a nearby object or area, while an increase in pupil size can indicate that one or both eyes are focusing on a more distant object or area.
[0057] In various examples, an eye-tracking camera device including one or more cameras 120, 122 and a scene-tracking camera device including one or more cameras 112, 114, 116 may apply various configurations and arrangements of the one or more individual cameras. The eye-tracking camera device including one or more cameras 120, 122 may include one or more cameras 112, 114, 116, 120, 122 mounted on the AR device 110. One or more cameras of the eye-tracking camera device 120, 122 may capture one or both eyes of a user. For example, the respective separate cameras 112, 114, 116, 120, 122 may be positioned to focus on the left eye or the right eye, e.g., as a left-eye tracking camera and a right-eye tracking camera. Some cameras may capture both eyes, while other cameras may capture one eye. The scene-tracking camera device including one or more cameras 112, 114, 116 may include one or more outward-facing cameras 112, 114, 116 for tracking a scene 170. The tracked scene may correspond to the gaze of one or both eyes. In some examples, the gaze of one or both eyes may correspond to a region of the scene (e.g., in the physical world), and vice versa.
[0058] In some AR devices 110, the display 105 may be transparent, allowing the eyes 115 to directly view / watch the scene and any holographic content that may be provided on the display 105. As discussed herein, an AR glasses device (e.g., AR device 110) may allow a user to view the physical world (e.g., the real-world environment), while AR content may be overlaid via the display 105. In other examples, the display 105 and the AR device 110 may not be transparent, and the scene may be reproduced on the display 105 based on regions of the fields of view 140, 150, 160 captured by one or more cameras 112, 114, 116 of the scene-tracking camera device. In such examples, one or more computing devices and / or processors may be used to combine the regions associated with the fields of view 140, 150, 160 to reproduce the scene 170 on the display 105 and render content. As discussed herein, holographic content and / or other AR content may also be provided on the display 105.
[0059] Figure 2 An internal view 210 and an external view 220 of an augmented reality system according to various examples discussed herein are shown. As discussed herein, an AR device (e.g., AR device 110) may be a head-mounted device. The internal view 210 may correspond to the side (e.g., the inside) that a user of the AR device directly views when wearing or using the AR device. The opposite side of the internal view 210 may correspond to the external view 220 of the AR device.
[0060] The inner side 210 may include a display 270 and a camera device that includes one or more cameras 230a, 230b, 230c, 230d (also referred to herein as cameras 230a through 230d). In some examples, the cameras may be positioned along the edge of the display 270. The cameras (e.g., cameras 230a through 230d) may capture the movement of one or both eyes, and the images captured from cameras 230a through 230d may be used to track the gaze of one or both eyes, which includes at least one of the gaze depth and / or gaze direction of one or both eyes. In one example, cameras 230a and 230d may track the user's left eye, while cameras 230b, 230c may track the user's right eye.
[0061] The display 270 may provide content corresponding to real-world content 260a, 260b, 260c (also referred to herein as real-world content 260a through 260d) and AR content 250 (e.g., holographic content). As discussed herein, the real-world content 260a through 260d may be a rendering of content captured from one or more of the outward-facing cameras 240a, 240b, 240c, 240d, 240e, 240f (also referred to herein as outward-facing cameras 240a through 240f). In other examples, similar to a user viewing the real world through glasses, the real-world content (e.g., real-world content 260a through 260c) may be viewed through a transparent display and / or through a lens.
[0062] An external view 220 of the augmented reality system shows the positioning of one or more outward-facing cameras 240a, 240b, 240c, 240d, 240e, 240f (also referred to herein as outward-facing cameras 240a through 240f). The cameras 240a through 240f may be positioned, for example, along the edge of the outward-facing side of the AR device. The positioning of the cameras may be changed to capture different fields of view. The cameras may be embedded in the AR device (e.g., AR device 110) to reduce or minimize the visibility of each camera on the AR device. The cameras (e.g., cameras 240a through 240f) may also be colored and / or positioned to blend in with the outer surface of the AR device.
[0063] The outward-facing cameras 240a through 240f may be used to track scenes in the real physical world and optionally provide information for rendering the scene on the display 270. The information from the outward-facing cameras 240a through 240f may be combined with the information from the inward-facing eye-tracking cameras 230a through 230d to provide context analysis and adaptively adjust the positioning of the AR content 250 on the display 270.
[0064] As discussed herein, object recognition techniques can be applied to the tracked scene. Object recognition techniques can assist in determining environmental events and interactions that can indicate what the user is focusing on and / or interacting with. The depicted display 270 can show multiple example regions of interest that contain one or more objects, and these regions of interest can indicate environmental interactions when combined with eye tracking. Based on the gaze of one or both tracked eyes, regions of interest 260a, 260b, 260c (also referred to herein as regions of interest 260a to 260c, or regions 260a to 260c) can be initially recognized. For example, based on the gaze direction and / or gaze depth of one or both eyes, the AR device can infer / determine whether the user is viewing a distant region (e.g., region 260a) or a closer region (e.g., regions 260b, 260c). The object recognition techniques applied by the AR device to the tracked scene can identify various objects within the scene, and these objects can be objects of interest. For example, for purposes of illustration and not limitation, the object recognition techniques can identify one or more objects such as a car or a mobile phone, or a person.
[0065] The characteristics of the object can indicate specific environmental actions that the user is engaged in or likely to be engaged in. For example, the eye tracking cameras 230a to 230d can identify that the user is generally viewing region 260b. The object recognition performed by the AR device on the scene can indicate an object that indicates a mobile phone within region 260b. The characteristics of the mobile phone or the tracked eye movements can indicate that the user is likely interacting with the mobile phone in region 260b, such characteristics as the size of the mobile phone relative to distant objects (e.g., the car and the person in region 260a) (e.g., a larger size relative to the sizes of other objects in this example), and such eye movements as indicating viewing a nearby object, reading, and / or scrolling, etc. Therefore, the AR content 250 can be positioned by the AR device so that it does not overlap with region 260b and / or one or more objects with which the user is interacting.
[0066] In another example, the tracked eye movements can indicate that the user is viewing the general areas of regions 260a and 260c. A long-term distant gaze of one or both eyes at region 260a can indicate that the user is viewing the person in region 260a. Therefore, the content 250 can be positioned so that it does not overlap with region 260a.
[0067] In another example, the tracked eye movement can indicate that the user has shifted their gaze (e.g., the gaze of one or both eyes) to region 260c. The AR device can then update the positioning of the AR content 250 such that the AR content does not overlap with region 260c. In some instances, based on the tracked gaze of one or both eyes, it may not be clear whether the user is viewing region 260a or 260c. In such a scenario, object recognition can further assist the AR device's determination by identifying the characteristics of the object and correlating those characteristics with potential environmental interactions within the region of interest. The size of the person in region 260c may be increasing, indicating that the object is approaching. As the scene is tracked, a wave may be recognized in the scene. In some examples, a gesture can indicate a specific environmental interaction. The combination of the wave and the approaching person can indicate to the AR device that the user is focusing on region 260c and intends to interact with the person identified in region 260c. Thus, the AR content 250 can be kept away from region 260c.
[0068] In an example where the people in region 260c are getting smaller, this can enable the AR device to determine that those people are moving away from or leaving the user. This environmental interaction utilized by the AR device can indicate that the objects within region 260c are no longer of interest, and thus, the AR content 250 can optionally be positioned over region 260c.
[0069] In yet another example, the tracked eye movement can enable the AR device to determine that the user is viewing the AR content 250. In this example, the AR device can optionally zoom in, move, and / or maximize the AR content 250 to enhance the viewing experience of the AR content 250.
[0070] In some examples, although the user may be focused on the AR content 250, one or more environmental interactions and events may cause the content AR 250 to be shifted, minimized, or otherwise moved. The object recognition technology and scene tracking technology utilized / implemented by the AR device can indicate that the person in region 260c is approaching. The approaching person can indicate a potential interaction with the user (e.g., of the AR device). In such a case, the AR content 250 can be positioned such that it does not obscure or generally interfere with any user interactions within the region of interest (e.g., region 260c).
[0071] Thus, an approaching object and a receding object can be environmental interactions that cause the AR content 250 to shift such that the user can perceive the approaching object or, in this case, a potential interaction with the person in region 260c. Similarly, an interaction or gesture with an object or person (such as pointing, waving, hugging, etc.) can indicate an environmental interaction and can cause the AR device to adaptively reposition the AR content 250 such that the AR content 250 does not overlap with the interaction with any object of interest.
[0072] Figure 3 A flowchart showing execution of context analysis according to various examples discussed herein is shown. At block 310, a device (e.g., AR device 110) can identify a region of interest corresponding to a gaze (e.g., the gaze of an eye) tracked by an eye-tracking camera device. As discussed herein, such an eye-tracking camera device (e.g., eye-tracking camera devices 120, 122) can include one or more inward-facing cameras configured to capture the movement of one or both of the user's eyes.
[0073] At block 320, a device (e.g., AR device 110) can perform object recognition on a scene (e.g., scene 170) tracked by a scene-tracking camera device. As discussed herein, such a scene-tracking camera device (e.g., the scene-tracking camera device including one or more cameras 112, 114, 116, for example) can include one or more outward-facing cameras configured to capture a scene from the physical world (e.g., a real-world environment). Object recognition techniques can identify one or more objects or one or more persons, such as an object or person with which the user is interacting. Example objects that can be recognized can include a phone, a book, a newspaper, or other devices / items.
[0074] At block 330, a device (e.g., AR device 110) can determine an environmental interaction based on the object recognition and the region of interest. As discussed herein, an environmental interaction can, for example, include one or more of the following: one or more approaching objects or persons; one or more receding objects or persons; an interaction with one or more objects or persons; one or more gestures, etc. An environmental interaction can also indicate a possible interaction between the user and one or more objects or persons.
[0075] At block 340, a device (e.g., AR device 110) can adaptively change the position of content (e.g., AR content 250) provided on a display (e.g., display 105, display 270) based on environmental interactions. Certain types of environmental interactions can cause the positioning of the content to change. For example, a person or object approaching rapidly can indicate that something the user might want to perceive is approaching. In some examples, an approaching object (e.g., a curb on a sidewalk) can represent a danger or an object the user might want to see. Similarly, a person approaching, standing near the user, or otherwise appearing to make contact with or interact with the user can be an environmental interaction the user also wishes to engage in. Thus, the content (e.g., AR content) can be adaptively changed by the device (e.g., AR device 110) so that the content does not overlap with an object or person. In some cases, the content positioning can be changed so that the content does not overlap with, obstruct, occlude, or otherwise distract the user from contact with an object or person.
[0076] In various examples, the adaptive content positioning can be adjusted based on user preferences. In other examples, a machine learning module can be used to adjust the positioning to identify specific object interactions of the user and / or associated gazes and / or interactions.
[0077] In some examples, as shown in block 350, one or more machine learning modules can assist with any of the operations associated with blocks 310, 320, 330, 340. For example, a machine learning module can assist in correlating the tracked eye movements with the gaze of one or both eyes. Other modules can assist in adaptively learning the directionality associated with the tracked gaze of one or both eyes, and / or correlating the tracked gaze with information about the scene (e.g., scene 170) obtained from a scene tracking camera device (e.g., a scene tracking camera device including one or more cameras 112, 114, 116). Such adaptive modules can also improve the understanding and / or rendering of the scene obtained by multiple cameras. In some examples, the above-mentioned machine learning modules, other modules, and adaptive modules can be implemented by the device (e.g., AR device 110).
[0078] As described above, environmental interactions can be determined based on prior training data associating the tracked objects with one or more user interactions. Specific user behaviors, such as looking up a book, newspaper, or phone, can be identified based on object recognition and specific behaviors within the region of interest corresponding to the tracked gaze of one or both eyes. The content positioning can also be adaptively learned, for example, based on one or more user preferences, manual adjustments, and / or prior positioning.
[0079] Figure 4A flowchart showing the adaptive change of content is presented. At block 410, a device (e.g., AR device 110) can determine the transition of the region of interest between the display and the scene (e.g., scene 170) based on the gaze (e.g., the gaze of one or both eyes). As discussed herein, such a transition can be identified based on the gaze direction and / or gaze depth of one or both eyes. For example, in an instance where the user is viewing an area where there may be no displayed content, the transition from the display (e.g., display 105, display 270) to the scene can be identified by the gaze direction of one or both eyes. The gaze depth, which can be identified based on the pupil size or other eye characteristics, can indicate that the user has shifted from viewing something nearby to viewing something far away. The transition from the scene to the display can be the reverse. The gaze direction can indicate the area where the user is viewing the displayed content. The gaze depth can indicate that the user has shifted from viewing something far away to viewing something nearby.
[0080] At block 420, in an instance where the region of interest transitions from the display to the scene, the device (e.g., AR device 110) can minimize or reduce the size of the content (e.g., AR content 250). In other examples, the content can be gradually eliminated. In various examples, the type of transition can be set manually (e.g., by the user). Based on the scene context and what the user is viewing, the transition style can also vary. For example, in an instance where the user quickly looks up and scans the scene to check the surroundings, the content can become transparent briefly (e.g., on the display). In an instance where the user transitions from the display to the scene and a person or object is approaching, the content can be minimized or reduced because the approaching object can indicate a possible interaction between the user and the object in the scene. In another example, the transition can cause the device to shift the content to an area that may not interfere with the region of interest the user is viewing.
[0081] At block 430, in an instance where the region of interest transitions from the scene to the display, the device (e.g., AR device 110) can maximize or increase the size of the content. In other examples, the content can be phased in (e.g., by the device). Similar to the above examples, based on the type of environmental interaction, for the user, the positioning of the objects in the scene and / or real-world objects with respect to the displayed content can become more visible and easier to view.
[0082] Figure 5Shows another flowchart of adaptively changing content according to various examples of the present disclosure. At block 510, a device (e.g., AR device 110) may associate a second region of interest with an environmental interaction. The second region of interest may correspond to a region on a display (e.g., display 270, display 105). At block 520, the device (e.g., AR device 110) may associate a third region of interest with content (e.g., AR content 250). The third region of interest may correspond to a region on the display that is covered by the content. The second and third regions of interest may be adaptively (e.g., in size) changed by the device based on any corresponding movement and / or change of the environmental interaction and the content.
[0083] At block 520, the device (e.g., AR device 110) may continuously move the position of the content (e.g., AR content 250) to reduce interference between the third region of interest associated with the content and the second region of interest associated with the environmental interaction. Accordingly, user interaction with the environment may be improved and interference that obscures (e.g., the user's) field of view associated with the content may be reduced.
[0084] Figure 6 Shows an example augmented reality system 600. The augmented reality system 600 may include a head-mounted display (HMD) 610 (e.g., smart glasses), the head-mounted display including a frame 612, one or more displays 614, and a computing device 608 (also referred to herein as computer 608). The display 614 may be transparent or translucent, thereby allowing a user wearing the HMD 610 to view the real world (e.g., real-world environment) through the display 614 and simultaneously display visual augmented reality content to the user. The HMD 610 may include an audio device 606 (e.g., speaker / microphone), which may provide audio augmented reality content to the user. The HMD 610 may include one or more cameras 616, 618 that may capture images and / or video of the environment. In one example, the HMD 610 may include one or more cameras 618 that may be rear-facing cameras that track the movement and / or gaze of the user's eyes.
[0085] One of the cameras 616 can be a front-facing camera that captures images and / or videos of the environment that a user wearing the HMD 610 can view. The HMD 610 can include an eye-tracking system that is configured to track the vergence movement of a user wearing the HMD 610. In one example, one or more cameras 618 can be the eye-tracking system. The HMD 610 can include a microphone of the audio device 606 to capture voice input from the user. The augmented reality system 600 can also include a controller 604 that includes a touchpad and one or more buttons. The controller 604 can receive input from the user and forward the input to the computing device 608. The controller can also provide haptic feedback to one or more users. The computing device 608 can be connected to the HMD 610 and the controller via a wired connection or a wireless connection. The computing device 608 can control the HMD 610 and the controller to provide augmented reality content to one or more users and receive input from one or more users. In some examples, the controller 604 can be a stand-alone controller or integrated within the HMD 610. The computing device 608 can be a stand-alone host computer device, an on-board computer device integrated with the HMD 610, a mobile device, or any other hardware platform capable of providing augmented reality content to a user and receiving input from the user. In some examples, the HMD 610 can include an augmented reality system / virtual reality system.
[0086] Figure 7 A block diagram of an exemplary hardware / software architecture of the UE 30 is shown. As Figure 7 shown, the UE 30 (also referred to herein as node 30) can include: a processor 32; a non-removable memory 44; a removable memory 46; a speaker / microphone 38; a keypad 40; a display, touchpad, and / or indicator 42; a power supply 48; a global positioning system (GPS) chipset 50; and other peripheral devices 52. The UE 30 can also include a camera 54. In one example, the camera 54 is an intelligent camera configured to sense images that appear within one or more bounding boxes. The UE 30 can also include communication circuitry, such as a transceiver 34 and a transmit / receive element 36. It will be appreciated that the UE 30 can include any sub-combination of the foregoing elements while remaining consistent with the various examples discussed herein.
[0087] The processor 32 can be a dedicated processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. Generally speaking, the processor 32 can execute computer-executable instructions stored in the memory of the node 30 (e.g., the memory 44 and / or the memory 46) to perform various required functions of the node. For example, the processor 32 can execute signal encoding, data processing, power control, input / output processing, and / or any other function that enables the node 30 to operate in a wireless environment or a wired environment. The processor 32 can run application layer programs (e.g., a browser) and / or radio access-layer (RAN) programs and / or other communication programs. The processor 32 can also perform security operations, such as security operations such as authentication, security key negotiation, and / or operations related to passwords in the access layer and / or the application layer.
[0088] The processor 32 is coupled to its communication circuit (e.g., the transceiver 34 and the transmit / receive element 36). The processor 32 can control the communication circuit by executing computer-executable instructions to enable the node 30 to communicate with other nodes via the network to which it is connected.
[0089] The transmit / receive element 36 can be configured to send signals to other nodes or network devices or receive signals from other nodes or network devices. For example, the transmit / receive element 36 can be an antenna configured to transmit and / or receive radio frequency (RF) signals. The transmit / receive element 36 can support various networks and air interfaces, such as wireless local area network (WLAN), wireless personal area network (WPAN), and cellular, etc. In another example, the transmit / receive element 36 can be configured to send and receive both RF signals and optical signals. It will be appreciated that the transmit / receive element 36 can be configured to send and / or receive any combination of wireless signals or wired signals.
[0090] The transceiver 34 can be configured to modulate signals to be transmitted by the transmit / receive element 36 and to demodulate signals received by the transmit / receive element 36. As described above, the node 30 can have multi-mode capabilities. Thus, the transceiver 34 can include multiple transceivers to enable the node 30 to communicate via multiple radio access technologies (RATs), such as universal terrestrial radio access (UTRA) and Institute of Electrical and Electronics Engineers (IEEE) 802.11.
[0091] The processor 32 can access information from and store data in any type of suitable memory, such as non-removable memory 44 and / or removable memory 46. For example, as described above, the processor 32 can store session context in its memory. The non-removable memory 44 can include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 46 can include a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card, among others. In other examples, the processor 32 can access information from and store data in a memory that is not physically located on the node 30 (e.g., located on a server or a home computer).
[0092] The processor 32 can receive power from the power supply 48 and can be configured to distribute power to other components in the node 30 and / or control the power to other components in the device 30. The power supply 48 can be any suitable device for powering the node 30. For example, the power supply 48 can include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells, among others.
[0093] The processor 32 can also be coupled to a GPS chipset 50, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the node 30. It will be appreciated that the node 30 can obtain location information by any suitable location determination method while remaining consistent with the examples of the present disclosure.
[0094] Figure 8is a block diagram of an exemplary computing system 800, which can also be used to implement components of a system or as part of the UE 30. The computing system 800 can include a computer or a server and can be mainly controlled by computer-readable instructions, which can be in the form of software, regardless of where or by what means such software is stored or accessed. Such computer-readable instructions can be executed within a processor (e.g., a central processing unit (CPU) 91) to cause the computing system 200 to operate. In many workstations, servers, and personal computers, the central processing unit 91 can be implemented by a single-chip CPU called a microprocessor. In other machines, the central processing unit 91 can include multiple processors. A coprocessor 81 can be an optional processor that is different from the main CPU 91 and performs additional functions or assists the CPU 91.
[0095] In operation, the CPU 91 fetches, decodes, and executes instructions and transfers information to and from other resources via the computer's main data transfer path, the system bus 80. Such a system bus connects multiple components in the computing system 200 and defines the medium for data exchange. The system bus 80 typically includes data lines for sending data, address lines for sending addresses, and control lines for sending interrupts and for the operating system bus. An example of such a system bus 80 is the Peripheral Component Interconnect (PCI) bus.
[0096] The memory coupled to the system bus 80 includes RAM 82 and ROM 93. Such a memory can include circuitry that allows for the storage and retrieval of information. The ROM 93 typically contains stored data that is not easily modified. The data stored in the RAM 82 can be read or changed by the CPU 91 or other hardware devices. Access to the RAM 82 and / or ROM 93 can be controlled by a memory controller 92. The memory controller 92 can provide an address translation function that translates virtual addresses into physical addresses when instructions are executed. The memory controller 92 can also provide a memory protection function that isolates processes within the system and isolates system processes from user processes. Thus, a program running in a first mode can only access the memory mapped by its own process virtual address space; the program cannot access the memory within another process's virtual address space unless memory sharing has been set up between the processes.
[0097] In addition, the computing system 200 may include a peripheral device controller 83, which is responsible for transmitting instructions from the CPU 91 to peripheral devices such as a printer 94, a keyboard 84, a mouse 95, and a disk drive 85.
[0098] A display 86 controlled by a display controller 96 is used to display the visual output generated by the computing system 200. Such visual output may include text, graphics, animated graphics, and video. The display 86 may be implemented using a cathode-ray tube (CRT)-based video display, a liquid-crystal display (LCD)-based flat panel display, a gas plasma-based flat panel display, or a touchpad. The display controller 96 includes the electronic components required to generate the video signal sent to the display 86.
[0099] In addition, the computing system 800 may include communication circuitry, such as a network adapter 97, which may be used to connect the computing system 200 to an external communication network (e.g., Figure 7 the network 12 in), so that the computing system 200 can communicate with other nodes of the network (e.g., the UE 30).
[0100] Figure 9 A framework 900 adopted by a software application (e.g., an algorithm) for evaluating gesture attributes is shown. The framework 900 may be remotely hosted. Alternatively, the framework 900 may reside within Figure 7 the UE 30 shown and / or be processed by Figure 8 the computing system 800 shown. The machine learning model 910 is operatively coupled to training data 920 stored in a database. In some examples, the machine learning model 910 may be associated with Figure 3 the operation of block 350 in. In some other examples, the machine learning model 910 may be associated with other operations. For example, in some examples, the machine learning model 910 may be associated with Figure 3 operations 310, 320, 330, 340 in. In other examples, the machine learning model 910 may be associated with Figure 4 operations 410, 420, 430 in and / or Figure 5 operations 510, 520, 530 in. The machine learning model 910 may be implemented by one or more machine learning modules and / or another device (e.g., the AR device 110).
[0101] In one example, the training data 920 can include the attributes of thousands of objects. For example, the objects can be smartphones, people, books, newspapers, signs, cars, etc. The attributes can include, but are not limited to, the size, shape, orientation, location, etc. of the objects. The training data 920 used by the machine learning model 910 can be fixed or updated periodically. Alternatively, the training data 920 can be updated in real time based on the evaluation performed by the machine learning model 910 in the non-training mode. This is shown by the bi-directional arrow connecting the machine learning model 910 and the stored training data 920.
[0102] In operation, the machine learning model 910 can evaluate the attributes of the images / videos acquired by the hardware (e.g., of the AR device 110, UE 30, etc.). For example, the eye-tracking camera devices 120, 122 and / or the scene-tracking camera devices 112, 114, 116 of the AR device 110 and / or Figure 7 The camera 54 of the UE 30 shown can sense and collect the images / videos that appear in or around the bounding box of the software application, such as an object approaching or moving away, object interactions, gestures, and / or other objects. Then, the attributes of the collected images (e.g., the attributes of the image of the collected object or person) can be compared with the corresponding attributes of the stored training data 920 (e.g., pre-stored objects). A likelihood of similarity between each of the obtained attributes (e.g., the images of one or more collected objects) and the stored training data 920 (e.g., pre-stored objects) can be given a determined confidence score. In one example, in instances where the confidence score exceeds a predetermined threshold, one or more attributes can be included in the image description that can ultimately be transmitted to the user via the user interface of the computing device (e.g., UE 30, computing system 800). In another example, the description can include a specific number / amount of attributes that may have exceeded a predetermined threshold for sharing with the user. The sensitivity of sharing more or fewer attributes can be customized based on the needs of a specific user.
[0103] Figure 10FIG. 1000 illustrates an example computer system 1000. In various examples, one or more computer systems 1000 may perform one or more steps of one or more of the methods described or illustrated herein. In a particular example, one or more computer systems 1000 provide the functionality described or illustrated herein. In some examples, software running on one or more computer systems 1000 performs one or more steps of one or more of the methods described or illustrated herein, or provides the functionality described or illustrated herein. Other examples may include one or more portions of one or more computer systems 1000. In this document, references to computer systems may include computing devices, as appropriate, and vice versa. Additionally, references to computer systems may include one or more computer systems, as appropriate.
[0104] The present disclosure contemplates any suitable number of computer systems 1000. The present disclosure contemplates computer systems 1000 in any suitable physical form. By way of example and not limitation, computer system 1000 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these systems. As appropriate, computer system 1000 may include one or more computer systems 1000; be single or distributed; span multiple locations; span multiple machines; span multiple data centers; or be located in the cloud (which may include one or more cloud components in one or more networks). As appropriate, one or more computer systems 1000 may perform one or more steps of one or more of the methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 1000 may perform one or more steps of one or more of the methods described or illustrated herein in real time or in batch processing mode. As appropriate, one or more computer systems 1000 may perform one or more steps of one or more of the methods described or illustrated herein at different times or at different locations.
[0105] In each example, computer system 1000 includes a processor 1002, a memory 1004, a storage 1006, an input / output (I / O) interface 1008, a communication interface 1010, and a bus 1012. Although this disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0106] In some examples, the processor 1002 includes hardware for executing a plurality of instructions, such as those that make up a computer program. By way of example and not limitation, to execute the plurality of instructions, the processor 1002 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1004, or storage 1006; decode and execute the instructions; and then write one or more results to an internal register, an internal cache, memory 1004, or storage 1006. In a particular example, the processor 1002 may include one or more internal caches for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 1002 that includes any suitable number of any suitable internal caches. By way of example and not limitation, the processor 1002 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The instructions in the instruction cache may be copies of the instructions in memory 1004 or storage 1006, and the instruction cache may accelerate the retrieval of these instructions by the processor 1002. The data in the data cache may be a copy of data in memory 1004 or storage 1006 for use by instructions executed at the processor 1002; may be the result of a previous instruction executed at the processor 1002 for access by a subsequent instruction executed at the processor 1002 or for writing to memory 1004 or storage 1006; or may be other suitable data. The data cache may accelerate read or write operations of the processor 1002. The TLB may accelerate virtual address translation by the processor 1002. In a particular example, the processor 1002 may include one or more internal registers for data, instructions, or addresses. In appropriate instances, the present disclosure contemplates a processor 1002 that includes any suitable number of any suitable internal registers. In appropriate instances, the processor 1002 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 1002. Although the present disclosure describes and shows particular processors, the present disclosure contemplates any suitable processor.
[0107] In some examples, memory 1004 includes a main memory that stores instructions for execution by processor 1002 or data for operation by processor 1002. By way of example and not limitation, computer system 1000 can load instructions from memory 1006 or another source (e.g., another computer system 1000) into memory 1004. Then, processor 1002 can load these instructions from memory 1004 into internal registers or an internal cache. To execute these instructions, processor 1002 can retrieve these instructions from the internal registers or internal cache and decode them. During or after execution of these instructions, processor 1002 can write one or more results (which can be intermediate or final results) to the internal registers or internal cache. Then, processor 1002 can write one or more of these results to memory 1004. In a particular example, processor 1002 only executes instructions in one or more internal registers or in one or more internal caches, or in memory 1004 (different from or other than memory 1006), and only operates on data in one or more internal registers or internal caches, or in memory 1004 (different from or other than memory 1006). One or more memory buses (each memory bus can include an address bus and a data bus) can couple processor 1002 to memory 1004. As described below, bus 1012 can include one or more memory buses. In some examples, one or more memory management units (MMUs) are located between processor 1002 and memory 1004 and facilitate access to memory 1004 requested by processor 1002. In a particular example, memory 1004 includes random access memory (RAM). In appropriate circumstances, the RAM can be volatile memory. In appropriate circumstances, the RAM can be dynamic RAM (DRAM) or static RAM (SRAM). Additionally, in appropriate circumstances, the RAM can be single-port RAM or multi-port RAM. The present disclosure contemplates any suitable RAM. In appropriate circumstances, memory 1004 can include one or more memories 1004. Although the present disclosure describes and illustrates particular memories, the present disclosure contemplates any suitable memory.
[0108] In some examples, the memory 1006 includes a mass storage for data or instructions. By way of example and not limitation, the memory 1006 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, a tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these memories. In appropriate cases, the memory 1006 can include removable media or non-removable (or fixed) media. In appropriate cases, the memory 1006 can be internal or external to the computer system 1000. In certain examples, the memory 1006 is a non-volatile solid-state memory. In a particular example, the memory 1006 includes a read-only memory (ROM). In appropriate cases, the ROM can be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. The present disclosure contemplates mass storage 1006 in any suitable physical form. In appropriate cases, the memory 1006 can include one or more memory control units that facilitate communication between the processor 1002 and the memory 1006. In appropriate cases, the memory 1006 can include one or more memories 1006. Although the present disclosure describes and illustrates particular memories, the present disclosure contemplates any suitable memory.
[0109] In some examples, I / O interface 1008 includes hardware, software, or both that provide one or more interfaces for communication between computer system 1000 and one or more I / O devices. In appropriate instances, computer system 1000 may include one or more of these I / O devices. One or more of these I / O devices may enable communication between a person and computer system 1000. By way of example, and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, camera, another suitable I / O device, or a combination of two or more of these I / O devices. An I / O device may include one or more sensors. The present disclosure contemplates any suitable I / O device and any suitable I / O interface 1008 for that I / O device. In appropriate instances, I / O interface 1008 may include one or more device or software drivers that enable processor 1002 to drive one or more of these I / O devices. In appropriate instances, I / O interface 1008 may include one or more I / O interfaces 1008. Although the present disclosure describes and shows particular I / O interfaces, the present disclosure contemplates any suitable I / O interface.
[0110] In some examples, communication interface 1010 includes hardware, software, or both that provide one or more interfaces for communication (e.g., packet-based communication) between computer system 1000 and one or more other computer systems 1000 or one or more networks. By way of example and not limitation, communication interface 1010 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network such as a WI-FI network. The present disclosure contemplates any suitable network and any suitable communication interface 1010 for that network. By way of example and not limitation, computer system 1000 can communicate with a network such as an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these networks. One or more portions of one or more of these networks can be wired or wireless. By way of example, computer system 1000 can communicate with a network such as a wireless PAN (WPAN) (e.g., a Bluetooth WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (e.g., a Global System for Mobile Communications (GSM) network), or other suitable wireless networks, or a combination of two or more of these networks. In appropriate instances, computer system 1000 can include any suitable communication interface 1010 for any of these networks. In appropriate instances, communication interface 1010 can include one or more communication interfaces 1010. Although the present disclosure describes and shows particular communication interfaces, the present disclosure contemplates any suitable communication interface.
[0111] In a particular example, bus 1012 includes hardware, software, or both that couple multiple components of computer system 1000 to each other. By way of example and not limitation, bus 1012 can include: an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these buses. Where appropriate, bus 1012 can include one or more buses 1012. Although the present disclosure describes and shows particular buses, the present disclosure contemplates any suitable bus or interconnect.
[0112] In this document, where appropriate, one or more computer-readable non-transitory storage media may include: one or more semiconductor-based integrated circuits (ICs) or other ICs (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical disc drives, floppy disks, floppy disk drives, magnetic tapes, solid-state drives (SSDs), RAM drives, Secure Digital cards or Secure Digital drives, any other suitable computer-readable non-transitory storage medium, or any suitable combination of two or more of these storage media. Where appropriate, the computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0113] In this document, unless otherwise expressly stated or the context otherwise indicates, "or" is inclusive rather than exclusive. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A or B" means "A, B, or both". Further, unless otherwise expressly stated or the context otherwise indicates, "and" is both joint and several. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A and B" means "A and B, jointly or severally".
[0114] The scope of the present disclosure includes all changes, substitutions, variations, transformations, and modifications to the examples described or shown herein, which would be understood by a person of ordinary skill in the art. The scope of the present disclosure is not limited to the examples described or shown herein. Additionally, although the various examples herein are described and shown as including specific components, elements, features, functions, operations, or steps, any of these examples can include any combination or arrangement of any of the components, elements, features, functions, operations, or steps described or shown anywhere herein that would be understood by a person of ordinary skill in the art. Further, the recitation in the appended claims of a device or system, or a component in a device or system, that is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function includes that device, system, component, whether or not the particular function is activated, turned on, or unlocked, so long as the device, system, or component is so adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to. Additionally, although the present disclosure describes or shows a particular example as providing a particular advantage, the particular example may not provide that advantage, or may provide some or all of that advantage.
[0115] Alternative embodiments
[0116] The foregoing description of the embodiments has been presented for purposes of illustration; the foregoing description is not intended to be exhaustive or to limit the patent rights to the precise forms disclosed. Many modifications and variations are possible in light of the above disclosure, as will be recognized by those of ordinary skill in the relevant art.
[0117] Some portions of this specification describe multiple embodiments in terms of applications and symbolic representations of operations on information. These application descriptions and representations are commonly used by those skilled in the data processing arts to effectively convey the substance of their work to others skilled in the art. Although these operations are described functionally, computationally, or logically, these operations are understood to be implemented by a computer program or equivalent circuitry, or microcode, etc. Additionally, it has sometimes proven convenient, without loss of generality, to refer to these operational arrangements as modules. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combination thereof.
[0118] Any step, operation, or process described herein can be performed or implemented using one or more hardware or software modules, alone or in combination with other devices. In one embodiment, the software modules are implemented using a computer program product that includes a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or processes.
[0119] Each embodiment may also relate to an apparatus for performing the operations herein. The apparatus may be specially constructed for the required purposes, and / or the apparatus may include a computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium, or stored in any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. In addition, any computing system mentioned in this specification may include a single processor, or may be an architecture adopting a multi-processor design for increasing computing power.
[0120] Each embodiment may also relate to a product produced by the computing processes described herein. Such a product may include information produced by the computing process, wherein the information is stored on a non-transitory tangible computer-readable storage medium, and such a product may include any embodiment of the computer program product or other data combinations described herein.
[0121] Finally, the language used in the specification is mainly chosen for readability and guidance purposes, and the language may not be chosen to define or limit the inventive subject matter. Therefore, it is intended that the scope of the patent rights not be limited by this specific implementation, but rather by any claims published based on the application herein. Thus, the disclosure of each embodiment is intended to illustrate rather than limit the scope of the patent rights set forth in the claims.
Claims
1. A device, comprising: a display configured to provide content in a visible area of the display; a camera device configured to track a scene of a real - world environment captured in a field of view of the camera device; one or more processors and a non - transitory memory, the non - transitory memory including computer - executable instructions that, when executed, cause the device to at least perform the following operations: determine a region of interest in the scene; perform object recognition on the scene tracked by the camera device; determine an environmental interaction based on the object recognition and the region of interest; and adaptively change a position of the content provided by the display based on the environmental interaction.
2. The device according to claim 1, wherein, the environmental interaction includes at least one of the following: an approaching object; an approaching person; a departing object; a departing person; an interaction with one or more objects; an interaction with one or more persons; or a gesture.
3. The device according to claim 1 or 2, wherein, when the one or more processors further execute the instructions, the device is configured to: determine that the region of interest changes from the display to the scene, or the region of interest changes from the scene to the display; in an instance where the region of interest changes from the display to the scene, minimize or reduce a size of the content; and in an instance where the region of interest changes from the scene to the display, maximize or increase a size of the content.
4. The device according to any one of the preceding claims, wherein, the device includes an augmented reality device, and optionally, wherein the augmented reality device includes a head - mounted device.
5. The device according to any one of the preceding claims, further comprising: a second camera device configured to track a gaze of at least one eye, wherein the second camera device includes a left - eye tracking camera and a right - eye tracking camera.
6. The device according to any one of the preceding claims, wherein, the camera device includes at least one outward - facing camera for tracking the scene.
7. The device according to any one of the preceding claims, wherein, the display is configured to provide holographic content including augmented reality content.
8. The device according to any one of the preceding claims, wherein, the scene captured in the field of view is associated with a gaze of at least one eye tracked by a second camera device.
9. The device according to any one of the preceding claims, wherein, when the one or more processors further execute the instructions, the device is configured to: determine the region of interest by determining a gaze depth of at least one eye captured by the second camera device or a gaze direction of the at least one eye.
10. The device according to any one of the preceding claims, wherein, the object recognition identifies a person or an object in the scene.
11. A method, comprising: Perform object recognition on a scene of a real-world environment captured in the field of view of a camera device; Determine a region of interest in the scene; Determine an environmental interaction based on the object recognition and the region of interest; And Based on the environmental interaction, adaptively change the position of the content provided by a display.
12. The method according to claim 11, The method further Comprises: Determine the region of interest based on the gaze of at least one eye tracked by a second camera device; Determine a transition of the region of interest between the display and the scene based on the gaze; In an instance where the region of interest transitions from the display to the scene, minimize or reduce the size of the content; And In an instance where the region of interest transitions from the scene to the display, maximize or increase the size of the content, and / or The method further comprises: Associate a second region of interest with the environmental interaction; Associate a third region of interest with the content; and Move the position of the content to reduce interference between the third region of interest associated with the content and the region of interest associated with the environmental interaction.
13. The method according to any one of claims 11 or 12, The method further Comprises: Apply one or more machine learning techniques to determine the environmental interaction based on training data associating one or more tracked gazes with one or more scenes, And / or The method further comprises: The device performs the adaptive change of the position of the content.
14. A computer-readable medium storing instructions that, when executed, cause: Perform object recognition on a scene tracked by a camera device; Determine a region of interest in the scene; Determine an environmental interaction based on the object recognition and the region of interest; And Based on the environmental interaction, adaptively change the position of the content provided on a display.
15. The computer-readable medium according to claim 14, Wherein: The object recognition is performed continuously in real time, And / or The instructions, when executed, further cause: Determine a region of interest corresponding to the gaze of at least one eye tracked by a second camera device, wherein the region of interest is based on the gaze depth of the at least one eye or the gaze direction of the at least one eye, And / or The instructions, when executed, further cause: In an instance where an interesting object approaches at a predetermined threshold speed, move the position of the content from the display, reduce the content from the display, or minimize the content from the display.
Citation Information
Cited By
Mixed Reality Emergency Response Training System
TWI937986B