Managing visual clutter in artificial reality environment

By assessing the clutter levels of virtual and real-world elements through image analysis and machine learning, and combining this with user gaze characteristics to predict reaction time, the challenge of managing visual clutter in artificial real-world environments has been solved, improving user experience and resource utilization efficiency.

CN122055686APending Publication Date: 2026-05-15CTRL-LABS CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CTRL-LABS CORP
Filing Date
2024-09-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to adaptively manage visual clutter in artificial reality environments, potentially overwhelming users with information in mixed reality settings, compromising security, and wasting computing resources.

Method used

By receiving images, using image analysis techniques and machine learning models, the degree of clutter in virtual and real-world elements is assessed, and reaction time is predicted based on the user's gaze characteristics. The overall clutter metric is calculated, and then actions are taken to manage the clutter.

Benefits of technology

Effectively manage visual clutter, improve user experience, optimize computing resource utilization, reduce information interference, and enhance security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122055686A_ABST
    Figure CN122055686A_ABST
Patent Text Reader

Abstract

In a particular embodiment, a computing system may receive an image that includes one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment. The system may determine a first metric indicative of a degree of clutter in the virtual environment and a second metric indicative of a degree of clutter in the real-world environment, respectively. The system may determine a gaze feature associated with the user based on the user activity and predict a response time of the user based on the gaze feature using a machine learning model. The system may determine a third metric indicative of a degree of clutter in the image based on the predicted reaction time. The system may calculate an overall clutter metric based on the first metric, the second metric, and the third metric. The system may perform one or more actions based on the overall clutter metric to manage clutter in the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to computer graphics and clutter management. In particular, this disclosure relates to adaptively managing visual clutter in artificial reality environments, such as mixed reality environments. Background Technology

[0002] Artificial reality is a form of reality that has been modified in some way before being presented to a user. Artificial reality can include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), or some combination and / or derivative thereof. Artificial reality content (e.g., mixed reality images) can include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content can include video, audio, haptic feedback, or some combination thereof, any of which can be presented in single-channel or multi-channel (e.g., stereoscopic video that gives the viewer a three-dimensional effect). Artificial reality can be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or for use in artificial reality (e.g., performing activities in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted devices (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.

[0003] "Passthrough" is a feature that allows users to see their physical environment while wearing an artificial reality system, such as a mixed reality (MR) headset. By having the MR headset display information captured by the headset's externally facing camera, information about the user's physical environment is visually "passed through" to the user.

[0004] The mixed reality imagery or user interface (UI) displayed to a user wearing an artificial reality system (e.g., an MR headset) may include (1) a pass-through image representing the user's physical or real-world environment, and (2) one or more virtual elements (e.g., digital avatars, VR / AR applications, AR objects, etc.) overlaid on the real-world pass-through imagery. When a user interacts with such a mixed reality imagery or UI, it is assumed that the user entrusts all their thoughts to the interface. In some instances, the UI presented to the user may be cluttered and not properly organized. For example, many applications may be open simultaneously, including unnecessary applications, numerous interactive options within applications, several notifications, etc. Such cluttered imagery or UI presented to the user may pose security concerns, especially when the user is wandering in a real-world environment. Furthermore, computational resources may be unnecessarily wasted on displaying elements that the user is not even interested in (e.g., because the user is not paying attention to these elements in their UI).

[0005] Therefore, visual clutter management is required to adaptively manage cluttered UIs or mixed reality images so that users are not overwhelmed by information (e.g., visual content) when engaging in artificial reality environments (e.g., mixed reality environments involving virtual and real-world interactions). Summary of the Invention

[0006] According to a first aspect of this disclosure, a method is provided, the method comprising: receiving an image by a computing system, the image including one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; using one or more image analysis techniques to determine a first metric indicating the degree of clutter in the virtual environment based on the one or more virtual elements, and to determine a second metric indicating the degree of clutter in the real-world environment based on the one or more real-world elements; determining a plurality of gaze features associated with a user based on user activity with respect to the image including the one or more virtual elements and the one or more real-world elements; using a machine learning model to predict the user's reaction time when performing user activities based on the plurality of gaze features; determining a third metric indicating the degree of clutter in the image based on the predicted reaction time; calculating an overall clutter metric based on: the first metric determined based on the one or more virtual elements; the second metric determined based on the one or more real-world elements; and the third metric determined based on the predicted reaction time; and performing one or more actions to manage clutter in the image based on the overall clutter metric.

[0007] In some embodiments, the method further includes training a machine learning model, wherein training the machine learning model includes: accessing multiple training samples obtained based on multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponding to a user trial includes gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; using the machine learning model, predicting reaction times of the multiple users in the multiple user trials based on the gaze characteristics included in the multiple training samples; comparing the predicted reaction times with multiple ground truth reaction times; and updating the machine learning model based on the comparison.

[0008] In some embodiments, the method further includes: comparing an overall clutter metric with a predetermined threshold; and determining that the overall clutter metric is higher than the predetermined threshold; wherein one or more actions for managing clutter are performed in response to determining that the overall clutter metric is higher than the predetermined threshold.

[0009] In some embodiments, one or more actions for managing clutter include: removing one or more virtual elements from an image; modifying information associated with one or more virtual elements; changing the layout or position of one or more virtual elements in an image; or adjusting the size of one or more virtual elements.

[0010] In some embodiments, calculating the overall disorder measure includes: taking a weighted average of the first, second, and third measures based on the weights assigned to each of the first, second, and third measures.

[0011] In some embodiments, multiple gaze features include: gaze or saccade velocity; saccade probability; saccade altitude; fixation point; fixation duration; saccade duration; saccade length; or saccade angular velocity.

[0012] In some embodiments, user activities include: a user searching for a specific virtual element among one or more virtual elements and one or more real-world elements in an image.

[0013] In some embodiments, the method further includes: decomposing an image into a virtual layer and a real-world layer, the virtual layer including one or more virtual elements associated with a virtual environment, and the real-world layer including one or more real-world elements associated with a real-world environment, wherein: a first measure indicating the degree of clutter in the virtual environment is determined by performing one or more image analysis techniques on the virtual layer; and a second measure indicating the degree of clutter in the real-world environment is determined by performing one or more image analysis techniques on the real-world layer.

[0014] In some embodiments, the one or more image analysis techniques include: feature congestion techniques; subband entropy techniques; or edge density techniques.

[0015] In some embodiments, the image is captured by an artificial reality system.

[0016] In some embodiments: the one or more virtual elements include applications installed on an artificial reality system; and the one or more real-world elements include physical objects existing in a real-world environment.

[0017] In some embodiments, the one or more virtual elements are overlaid on the one or more real-world elements.

[0018] According to another aspect of this disclosure, one or more computer-readable non-transitory storage media are provided, the one or more computer-readable non-transitory storage media comprising software, which, when executed, is operable to: receive an image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; using one or more image analysis techniques, determine a first metric indicating the degree of clutter in the virtual environment based on the one or more virtual elements, and determine a second metric indicating the degree of clutter in the real-world environment based on the one or more real-world elements; determine a plurality of gaze features associated with a user based on user activity with respect to an image comprising the one or more virtual elements and the one or more real-world elements; use a machine learning model to predict the user's reaction time when performing user activities based on the plurality of gaze features; determine a third metric indicating the degree of clutter in the image based on the predicted reaction time; calculate an overall clutter metric based on: the first metric determined based on the one or more virtual elements; the second metric determined based on the one or more real-world elements; and the third metric determined based on the predicted reaction time; and perform one or more actions to manage clutter in the image based on the overall clutter metric.

[0019] In some embodiments, the software, when executed, is also operable to train a machine learning model, wherein training the machine learning model includes: accessing multiple training samples obtained based on multiple user trials or studies conducted in different cluttered scenarios, wherein each training sample corresponding to a user trial includes gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; using the machine learning model, based on the gaze characteristics included in the multiple training samples, predicting the reaction times of the multiple users in the multiple user trials; comparing the predicted reaction times with multiple ground truth reaction times; and updating the machine learning model based on the comparison.

[0020] In some embodiments, the software, when executed, is also operable to: compare an overall clutter metric with a predetermined threshold; and determine that the overall clutter metric is higher than the predetermined threshold, wherein one or more actions for managing clutter are performed in response to determining that the overall clutter metric is higher than the predetermined threshold.

[0021] In some embodiments, one or more actions for managing clutter include: removing one or more virtual elements from an image; modifying information associated with one or more virtual elements; changing the layout or position of one or more virtual elements in an image; or adjusting the size of one or more virtual elements.

[0022] According to another aspect of this disclosure, a system is provided, comprising: one or more processors; and one or more computer-readable non-transitory storage media coupled to one or more of the one or more processors, and including instructions that, when executed by one or more of the one or more processors, are operable to cause the system to: receive an image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; and, using one or more image analysis techniques, determine a first measure indicative of the degree of clutter in the virtual environment based on the one or more virtual elements, and, based on the one or more real-world elements... The method involves: determining a second metric indicating the degree of clutter in a real-world environment based on elements; identifying multiple gaze features associated with the user based on user activity with respect to an image including the one or more virtual elements and the one or more real-world elements; using a machine learning model to predict the user's reaction time when performing user activities based on the multiple gaze features; determining a third metric indicating the degree of clutter in the image based on the predicted reaction time; calculating an overall clutter metric based on: a first metric determined based on the one or more virtual elements; a second metric determined based on the one or more real-world elements; and a third metric determined based on the predicted reaction time; and performing one or more actions to manage clutter in the image based on the overall clutter metric.

[0023] In some embodiments, the one or more processors are also operable to cause the system to operate when training a machine learning model, wherein training the machine learning model includes: accessing multiple training samples obtained based on multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponding to a user trial includes gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; using the machine learning model, predicting the reaction times of the multiple users in the multiple user trials based on the gaze characteristics included in the multiple training samples; comparing the predicted reaction times with multiple ground truth reaction times; and updating the machine learning model based on the comparison.

[0024] In some embodiments, the one or more processors are also operable, when executing instructions, to cause the system to: compare an overall messiness metric with a predetermined threshold; and determine that the overall messiness metric is higher than the predetermined threshold, wherein one or more actions for managing messiness are performed in response to determining that the overall messiness metric is higher than the predetermined threshold.

[0025] In some embodiments, one or more actions for managing clutter include: removing one or more virtual elements from an image; modifying information associated with one or more virtual elements; changing the layout or position of one or more virtual elements in an image; or adjusting the size of one or more virtual elements.

[0026] The specific embodiments described herein relate to systems and methods for adaptively managing visual clutter in artificial reality environments (e.g., mixed reality environments). The clutter management discussed herein can be performed on visual scenes presented by an artificial reality system (e.g., a mixed reality head-mounted viewer). This visual scene can be a mixed reality image, which may consist of: (1) a real-world environment including one or more real-world elements (e.g., cars, trees, people, roads, etc.) and (2) a virtual reality environment (or virtual environment) including one or more virtual elements (e.g., applications on the artificial reality system, one or more digital avatars, AR / VR objects, etc.). These one or more virtual elements may be overlaid on top of real-world elements. As discussed elsewhere in this document, if the visual scene presented to the user is too cluttered (e.g., too many virtual elements, too dense / crowded physical environment), this can interfere with the user experience. Simply using existing image analysis methods to manage clutter based on the degree of virtual clutter and / or real-world clutter may not account for the amount of clutter experienced by the user at a given time. For example, different users may have different perceptions of an image. As an example, the first user and the second user may perceive a visual scene differently. This could be because different users react differently to their environment, and therefore have different reaction times to the same information presented to them. Thus, user perception characteristics or user components should also be part of the process when assessing how cluttered a particular environment is, whether there is a real need to manage visual clutter according to user perception, and how much content (e.g., elements in a virtual UI) should be modified to manage visual clutter.

[0027] In certain embodiments, systems and methods for clutter management can use three components to adaptively and intelligently perform clutter assessment and management. These components include a virtual user interface (UI) component, a real-world component, and a user component. A separate clutter metric can be calculated for each of the virtual UI component, the real-world component, and the user component. Each clutter metric can indicate the degree of clutter in its respective component. For example, a first clutter metric can indicate the degree of clutter in the virtual UI (e.g., how cluttered the virtual UI is based on the number of virtual elements present in the virtual UI (e.g., VR / AR applications)). A second clutter metric can indicate the degree of clutter in the real world (e.g., how cluttered the real-world environment is based on the number of real-world elements present in the real-world environment (e.g., physical objects)). In certain embodiments, the first clutter metric indicating the degree of clutter in the virtual UI and the second clutter metric indicating the degree of clutter in the real world can be calculated using one or more image analysis techniques, such as feature congestion techniques, subband entropy techniques, or edge density techniques.

[0028] In certain embodiments, a third clutter metric can be calculated based on a third component that is a user component. As previously mentioned, simply using image analysis methods to manage clutter based on the degree of clutter in a virtual UI and / or in the real world may not account for the amount of clutter experienced by a user at a given time, as well as the user's perceptual characteristics. For example, different users may have different perceptions and reactions to images. In other words, one user may be able to react to elements in an image faster than another user. This user's reaction time may be key to determining or assessing how cluttered a particular environment is. Reaction time is one of the best physiological correlations with visual clutter. Increased visual clutter leads to increased cognitive, visual, memory, and motor loads, or a combination thereof. This results in increased reaction time. In certain embodiments, reaction time can be determined based on multiple gaze characteristics associated with the user when performing a specific task (e.g., saccade speed, saccade duration, saccade length, fixation point, fixation duration, scanpath rate, etc.).

[0029] In a particular embodiment, a machine learning (ML) model can be trained to predict reaction time in real time based on multiple gaze features provided as input to the ML model. This ML model is trained on a wide variety of user trials or studies performed in different environments, including clutter-free, low-clutter, medium-clutter, and high-clutter environments. For each user trial, gaze features are observed while performing a given task, and the user's reaction time for performing the given task is recorded. Based on the gaze features and reaction times associated with different user trials or studies, an ML model is constructed to predict reaction time in a given mixed reality environment in real time.

[0030] During inference, gaze features are determined while the user is viewing a mixed reality image, and a trained ML model is used to predict reaction time based on these gaze features. The reaction time output by the ML model indicates the time a user might take to react to a specific task in a given mixed reality environment (including virtual UI clutter and real-world clutter). Based on this reaction time, a third clutter metric can be calculated. The third clutter metric indicates the degree of clutter in the mixed reality environment or merged image (including virtual UI clutter and real-world clutter). Once the first clutter metric based on virtual UI clutter, the second clutter metric based on real-world clutter, and the third clutter metric based on reaction time are determined, the system discussed in this paper can calculate the overall clutter metric. Depending on the user's perception and whether clutter needs to be managed to improve the user's viewing experience, better manage computational resources, better utilize display space, etc., the overall clutter metric can be a true indicator of visual clutter in the user's environment (e.g., a mixed reality environment).

[0031] The system discussed herein can perform one or more actions to manage clutter in a visual scene or image presented to a user via an artificial reality system, based on an overall clutter metric. In a particular embodiment, the one or more actions can be performed in response to determining that the overall clutter metric is higher than a certain threshold. As an example, the overall clutter metric can be a value between 0 and 1, and the predetermined threshold can be 0.75, and the one or more actions are performed if the overall clutter metric is higher than 0.75. In a particular embodiment, the one or more actions for clutter management can include, for example, but not limited to, closing one or more applications from the virtual UI, changing the layout currently displayed to the user, adjusting (e.g., enlarging, shrinking) the size of one or more elements in the virtual UI, modifying information associated with one or more applications in the virtual UI (e.g., expanding or collapsing information), changing the position of one or more applications in the virtual UI, etc.

[0032] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Specific embodiments may include all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein, or may exclude the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments of the invention are particularly disclosed in the appended claims for methods, storage media, systems, and computer program products, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from the intentional reference to any prior claims (especially multiple dependencies) may also be claimed, such that any combination of multiple claims and their multiple features, regardless of the dependencies chosen in the appended claims, is disclosed and can be claimed. The claimable subject matter includes not only various combinations of the features set forth in the appended claims, but also any other combination of the features in the appended claims, wherein each feature mentioned in the claims may be combined with any other feature in the claims, or with a combination of other features. Furthermore, any embodiment and feature of the multiple embodiments and features described or depicted herein may be claimed in a single claim, and / or any combination of any embodiment and feature of the multiple embodiments and features described or depicted herein with any embodiment or feature described or depicted herein, or any combination of any embodiment and feature of the multiple embodiments and features described or depicted herein with any feature of the appended claims.

[0033] It will be appreciated that any feature described herein that is suitable for incorporation into one or more aspects or embodiments of this disclosure is intended to be generalizable in any and all aspects and embodiments of this disclosure. Other aspects of this disclosure will be understood by those skilled in the art based on the specification, claims, and drawings of this disclosure. The foregoing general description and the following detailed description are exemplary and illustrative only and are not intended to limit the scope of the claims. Attached Figure Description

[0034] Figure 1 A block diagram is shown for calculating an overall measure of clutter used to adaptively manage clutter in artificial reality environments.

[0035] Figure 2 An example training process is shown for training a machine learning model to predict a user's reaction time based on the user's gaze characteristics.

[0036] Figure 3 An example method for adaptively managing visual clutter in an artificial reality environment, according to a specific embodiment, is shown.

[0037] Figure 4 An example of an artificial reality system worn by a user is shown.

[0038] Figure 5 An example network environment associated with an artificial reality system is shown.

[0039] Figure 6 An example computer system is shown. Detailed Implementation

[0040] The mixed reality imagery or user interface (UI) displayed to a user wearing an artificial reality system (e.g., artificial reality system 400) may include (1) a pass-through image representing the user's physical or real-world environment, and (2) one or more virtual elements (e.g., digital avatars, VR / AR applications, AR objects, etc.) overlaid on the real-world pass-through imagery. When a user interacts with such a mixed reality imagery or UI, it is assumed that the user entrusts all their thoughts to the interface. In some instances, the mixed reality imagery may be very cluttered and not properly organized. For example, many applications may be open simultaneously, including unnecessary applications, numerous interactive options within applications, several notifications, etc. Such cluttered imagery or UI presented to the user may pose security concerns, especially when the user is wandering in a real-world environment. Furthermore, computational resources may be unnecessarily wasted on displaying elements that the user is not even interested in (e.g., because the user is not paying attention to these elements in their UI). Thus, clutter management needs to be adaptively implemented to better manage cluttered UIs or mixed reality images so that users are not overwhelmed by information (e.g., visual content) when engaging in artificial reality environments (e.g., mixed reality environments involving virtual and real-world interactions).

[0041] Figure 1 A block diagram 100 is shown for calculating a general measure of clutter for adaptively managing clutter in artificial reality environments. The clutter management discussed herein can be targeted at artificial reality systems (e.g., Figure 4The visual scene presented by the artificial reality system 400 shown is executed. For example, the visual scene can be a mixed reality visual scene, which can consist of a real-world environment including one or more real-world elements (e.g., cars, trees, people, roads, etc.) and a virtual reality environment (or virtual environment) including one or more virtual elements (e.g., applications on the artificial reality system, one or more digital avatars, AR / VR objects, etc.). The one or more virtual elements can be overlaid on top of the real-world elements. As discussed elsewhere in this paper, if the visual scene presented to the user is too cluttered (e.g., too many virtual elements, too dense / crowded physical environment), this can interfere with the user experience. Simply using existing image analysis methods to manage clutter based on the degree of virtual clutter and / or real-world clutter may not account for the amount of clutter a user is experiencing at a given time. For example, different users may have different perceptions of an image. As an example, the first user and the second user may have different perceptions of the visual scene. More specifically, the first user may perceive the visual scene presented to them as cluttered, while the second user may perceive the visual scene as less cluttered. This may be because different users react differently to their environment, and therefore have different reaction times to the same information presented to them. Thus, when assessing how cluttered a particular environment is, the user's perceptual characteristics or user components should also be part of the process.

[0042] In certain embodiments, a system (e.g., computer unit 408 of artificial reality system 400, or computer system 600) may use three components to adaptively and intelligently perform clutter assessment and management. Figure 1 As shown, these components include a virtual user interface (UI) component 102, a real-world component 104, and a user component 106. In a particular embodiment, the virtual UI component 102 and the real-world component 104 can be obtained by decomposing a captured visual scene or image, which may be captured using an artificial reality system. As an example, it can be used... Figure 4 The external cameras 405A and 405B of the artificial reality system 400 shown capture the image. As previously described, the image can be a mixed reality image, which may include one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment, wherein the one or more virtual elements may be overlaid on the one or more real-world elements. This mixed reality image can be decomposed into two separate image layers, including a virtual layer corresponding to the virtual UI component 102 and a real-world layer corresponding to the real-world component 104. One or more image analysis techniques can be used to analyze each of the virtual layer and the real-world layer to assess the degree of clutter in the virtual environment and the real-world environment, respectively, as discussed in further detail below.

[0043] In a particular embodiment, one or more image analysis techniques or methods can be applied to the virtual layer corresponding to the virtual UI component 102 to assess virtual UI clutter 112 in a virtual reality environment or virtual reality space. More specifically, the one or more image analysis techniques can be used to assess the degree of clutter in the UI of the virtual environment presented to the user. The UI of the virtual environment or the virtual UI can consist of one or more virtual elements, such as, but not limited to, applications installed on an artificial reality system, one or more digital avatars, AR / VR objects, etc. In other words, virtual UI clutter 112 can represent all virtual and / or augmented reality applications presented on the user's display screen. As discussed in further detail below, image analysis techniques or methods can be used to analyze the virtual layer or image representing the UI of the virtual environment and determine virtual UI clutter 112 and a first clutter measure 122 indicating the degree of clutter in the virtual environment. More specifically, the image analysis technique can analyze the virtual layer to determine how cluttered the virtual UI is based on the number of virtual elements (e.g., VR / AR applications) present in the virtual UI.

[0044] The image analysis techniques discussed in this article may include, for example, but not limited to, feature congestion techniques, subband entropy techniques, or edge density techniques. Feature congestion techniques are based on various features in an image. Essentially, feature congestion techniques consider all possible features, such as contrast, shape, size, color, etc., and take into account each feature for a measure of clutter, such as virtual UI clutter 112. As an example, if you have a cluttered desktop and you want to search for a specific element on the desktop (e.g., a keychain), how long it takes to find that element is the basis behind feature congestion. Subband entropy techniques are based on the idea of ​​how much an image can be compressed by preserving details at the perceptual level, or in other words, how much you can compress an image. Therefore, if an image is very cluttered, it may not be compressed well because important details will be lost. If the image is less cluttered, it can be compressed well without losing details. Edge density techniques may involve calculating the ratio of edges within a scene to one or more objects. Therefore, the more objects there are, the more edges there will be. Calculating the number of pixels at the edges (e.g., edge pixels) and dividing by the total number of pixels gives the level of clutter, such as clutter 112 for virtual UI.

[0045] Based on the virtual UI clutter 112 determined using one or more image analysis techniques, a first clutter measure 122 can be calculated, indicating the degree of clutter in the virtual environment. In a particular embodiment, the first clutter measure 122 can be a normalized value between 0 and 1, where 0 indicates a clutter-free environment and 1 indicates a highly cluttered environment. By way of example and not limitation, the first clutter measure can be 0 if no applications are present in the virtual UI, 0.5 if two to three applications are present in the virtual UI, and 1 if more than a certain number of applications are present in the virtual UI (e.g., 10 applications).

[0046] In a particular embodiment, similar to how one or more image analysis techniques can be applied to a virtual layer corresponding to virtual UI component 102 to evaluate virtual UI clutter 112 and a first clutter metric 122, one or more image analysis techniques can be applied to a real-world layer corresponding to real-world component 104 to evaluate real-world clutter 114 and a second clutter metric 124. More specifically, these one or more image analysis techniques can be used to evaluate the degree of clutter in a real-world environment presented to a user. A real-world environment can consist of one or more real-world elements that may exist in a user's physical environment, such as, but not limited to, cars, trees, people, buildings, roads, etc. In other words, real-world clutter 114 can represent how dense or crowded a user's physical environment is. As discussed elsewhere in this document, image analysis techniques or methods can be used to analyze a real-world layer or image representing a real-world environment and determine real-world clutter 114 and a second clutter metric 124 indicating the degree of clutter in the real-world environment. More specifically, the image analysis technique can analyze the real-world layer to determine how cluttered the real-world environment is based on the number of real-world elements present in the real-world environment.

[0047] Once the real-world clutter 114 has been analyzed, a second clutter metric 124 can be calculated, indicating the degree of clutter in the real-world environment. In a particular embodiment, similar to the first clutter metric 122, the second clutter metric 124 can be a normalized value between 0 and 1, where 0 indicates a clutter-free environment and 1 indicates a highly cluttered environment. As an example and not a limitation, the first clutter metric could be 0 if a user is staring at a blank wall or a blank screen, and 0.8 if a user is walking in Times Square, New York on New Year's Eve.

[0048] In a particular embodiment, the third clutter measure 126 may be calculated based on a third component, which is the user component 106. As previously mentioned, simply using image analysis methods to manage clutter based on the degree of virtual UI clutter 112 and / or real-world clutter 114 may not account for the amount of clutter experienced by a user at a given time. For example, different users may have different perceptions and reactions to images. In other words, one user may be able to react to elements in an image faster than another user. The user's reaction time to elements may be key to determining or assessing how cluttered a particular environment is. In a particular embodiment, reaction time 120 may be determined based on multiple gaze features (e.g., gaze feature 116) associated with the user when performing a specific task, and a machine learning (ML) model 118 may be trained to predict reaction time 120 based on gaze feature 116 provided as input to the ML model 118.

[0049] To determine gaze feature 116, a mixed reality image is provided to a user wearing an artificial reality system (e.g., artificial reality system 400) that includes one or more virtual UI elements (e.g., AR / VR applications) and one or more real-world elements (e.g., buildings, cars, trees, etc.). While viewing the mixed reality image, the user can perform a specific task. This task may include searching for specific elements within the mixed reality image. As an example, the mixed reality image may include an image of the Golden Gate Bridge and its surrounding environment that the user can view through their mixed reality headset (e.g., artificial reality system 400), and five applications (including a music application, a game application, a video application, a messaging application, and an image capture application) may be overlaid on the image. The user might look for a stop button in the music application to stop the music while viewing the mixed reality image. As another example within the same mixed reality image, the user may receive a notification of a received message, and the user might look for a messaging application within the mixed reality image to read the received message. While a user is performing a task (e.g., searching for a stop button in a music app or a messaging app), one or more eye-tracking sensors in an AI system worn by the user can track the user's eye movements or gaze movements as the user performs that task. Based on the tracked eye movements or gaze movements, multiple gaze features 116 can be determined. In a particular embodiment, multiple gaze features 116 may include, for example, but not limited to, gaze or saccade duration (e.g., how long a user gazes when searching for a specific element), saccade / gaze speed (e.g., how fast a user gazes when moving from one point to another to search for a specific element), fixation duration (e.g., how long a user fixates on a specific point), fixation point (e.g., the location where the user's eyes are fixed), saccade height, saccade angular velocity, saccade probability, saccade length (e.g., the distance the user's eyes travel from one point to another), fixation probability or fixation probability, scan path rate, etc.

[0050] Once the gaze feature 116 is determined, it can be provided as input to the ML model 118, which is trained to predict the user's reaction time 120 for a specific task (e.g., searching for a specific element in a mixed reality image) based on the gaze feature 116. In a particular embodiment, the ML model 118 is trained based on various user trials or studies performed in different environments, including clutter-free, low-clutter, medium-clutter, and high-clutter mixed reality environments. For each user trial, multiple gaze features are observed while performing a given task, and the user's reaction time for performing that given task is recorded. Based on the gaze features and reaction times associated with different user trials or studies, the ML model 118 is constructed to predict reaction times for a given mixed reality environment in real time. Model architectures such as logistic regression or deep learning models (e.g., temporal convolutional networks (TCNs), recurrent neural networks (RNNs)) are good candidates for the ML models discussed herein. See below for reference. Figure 2 The training of ML model 118 is discussed in detail.

[0051] In a particular embodiment, the reaction time 120 output by the ML model 118 can indicate the time a user might take to react to a specific task in a given mixed reality environment (including virtual UI clutter 112 and real-world clutter 114). Based on the reaction time 120, a third clutter metric 126 can be determined. The third clutter metric 126 can indicate the degree of clutter in the mixed reality environment or blended image (including virtual UI clutter 112 and real-world clutter 114). In a particular embodiment, similar to the first clutter metric 122 and the second clutter metric 124, the third clutter metric 126 can also be a normalized value between 0 and 1, where 0 indicates a clutter-free environment and 1 indicates a highly cluttered environment. However, the third clutter metric 126 here is based on the reaction time 120, and the value of the third clutter metric 126 can vary depending on how low or high the reaction time 120 is. In other words, the value of the third clutter metric 126 can be proportional to the reaction time 120. For example, if the response time of 120 is low, the third clutter metric 126 will be low, indicating a low-clutter environment. Conversely, if the response time of 120 is high, the third clutter metric 126 will be high, indicating a high-clutter environment. As an example, and not a limitation, if a user spends 5 seconds in a music app pausing music by selecting the stop button, the third clutter metric 126 could be 0.3, indicating a low-clutter environment. However, if the user spends 30 seconds performing the same action, the third clutter metric 126 could be 0.8, indicating a high-clutter environment.

[0052] Once a first clutter metric 122 based on virtual UI clutter 112, a second clutter metric 124 based on real-world clutter 114, and a third clutter metric 126 based on reaction time 120 are determined, the system (e.g., computer unit 408 of artificial reality system 400, or computer system 600) can calculate a total clutter metric 130. Depending on the user's perception and whether clutter needs to be managed to improve the user's viewing experience, better manage computing resources, better utilize display space, etc., the total clutter metric 130 can be a realistic indication of visual clutter in the user's environment (e.g., a mixed reality environment).

[0053] In a particular embodiment, the overall noise metric 130 can be calculated by weighting the first noise metric 122, the second noise metric 124, and the third noise metric 126, and then taking the weighted sum of these three metrics. In some embodiments, equal weights can be assigned to the three metrics 122, 124, and 126, and the overall noise metric 130 can be calculated by simply taking the average of the three metrics. In other embodiments, different weights can be assigned to the three metrics 122, 124, and 126, and the overall noise metric 130 can be calculated by summing the three metrics according to their weights. For example, equal weights can be applied to the first noise metric 122 and the second noise metric 124, but a relatively higher weight can be applied to the third noise metric 126.

[0054] Once the overall clutter metric 130 is calculated, the computational system discussed herein can perform one or more actions 140 to manage clutter in a visual scene or image presented to a user via an artificial reality system (e.g., artificial reality system 400). In a particular embodiment, one or more actions 140 may be performed in response to determining that the overall clutter metric 130 is higher than a specific threshold. For example, the computational system may compare the overall clutter metric 130 to a predetermined threshold to determine whether the overall clutter metric 130 is higher or lower than the predetermined threshold, and if the overall clutter metric 130 is determined to be higher than the predetermined threshold, then one or more actions 140 are performed. As an example, the overall clutter metric 130 may be a value between 0 and 1, and the predetermined threshold may be 0.75; if the value of the overall clutter metric 130 is higher than 0.75, then one or more actions 140 are performed. In a particular embodiment, one or more actions for clutter management may include, for example, but not limited to, closing one or more applications from the virtual UI, changing the layout currently displayed to the user, adjusting (e.g., enlarging, shrinking) the size of one or more elements in the virtual UI, modifying information associated with one or more applications in the virtual UI (e.g., expanding or collapsing information), changing the position of one or more applications in the virtual UI, etc.

[0055] Figure 2 An example training process 200 is shown for training machine learning model 118 to predict user reaction time based on user gaze characteristics. (See above reference.) Figure 1 The user's reaction time, as discussed, can be used to assess how cluttered the user's displayed image (e.g., a mixed reality image displayed via a mixed reality head-mounted viewer) is, and whether clutter management is truly necessary based on the user's perception. In a particular embodiment, the ML model 118 can be trained to predict reaction time in real time (i.e., at inference time) based on multiple gaze features associated with the user. As shown, the ML model 118 is trained using training data 202, which may include multiple training samples 204a, 204b, 204c, ..., 204n (referred to individually and / or collectively as 204 herein). By way of example and not limitation, the ML model 118 can be trained based on 100 training samples, of which 90 samples can be used to train the ML model 118 and 10 samples (e.g., test sample 205) can be used to test the ML model 118.

[0056] Multiple training samples 204 can be obtained based on multiple user trials or studies in different cluttered scenarios, including no-clutter, low-clutter, medium-clutter, and high-clutter environments. Each training sample 204 corresponds to a user trial and includes gaze characteristics of a specific user observed over a specific duration (e.g., 30 seconds) during which the specific user performs an assigned task in a specific cluttered scenario, and the reaction time of the specific user in performing the assigned task is recorded. The recorded reaction times can be used as ground truth for testing the ML model 118 (e.g., ground truth reaction time 208), as discussed later below.

[0057] As previously described, training samples 204 used to train the ML model 118 may include users' gaze features and recorded reaction times (e.g., ground truth reaction times) of them performing multiple user trials in different cluttered scenarios or environments. For example, training sample 204a includes gaze features corresponding to user trial 1 that can be performed in cluttered scenario 1 (e.g., low visual UI clutter and low real-world clutter), training sample 204b includes gaze features corresponding to user trial 2 that can be performed in cluttered scenario 2 (e.g., low visual UI clutter but high real-world clutter), training sample 204c includes gaze features corresponding to user trial 3 that can be performed in cluttered scenario 3 (e.g., high visual UI clutter but low real-world clutter), and training sample 204n includes gaze features corresponding to user trial N that can be performed in cluttered scenario N.

[0058] The ML model 118 is trained based on a plurality of training samples 204. Once the model 118 has been trained on a set number of training samples, it can be tested. To test the ML model 118, a set of test samples 205 can be accessed. The test samples 205 can be part of the training data 202. These test samples 205 can include gaze features of the user observed in different cluttered scenes. The test samples 205 are provided as input to the ML model 118, which can then use the input gaze features to predict the user's reaction time 206 in different cluttered scenes. The ground truth reaction time 208 of these users (e.g., real or actual reaction time) is accessed, and the predicted reaction time generated by the ML model 118 and the ground truth reaction time 208 are compared, as indicated by reference numeral 210. Based on this comparison, a loss function can be calculated to determine the error rate or metric (e.g., root-mean-square error (RMSE) or root-mean-square deviation (RMSD)) between the value of the reaction time 206 predicted by the ML model and the value of the ground truth reaction time 208. Using the calculated loss function and / or the comparison, the ML model 118 can be updated, as indicated by reference numeral 212. The ML model 118 can be updated to minimize the loss function.

[0059] In a particular embodiment, updating the ML model 118 may include updating one or more parameters or components of the ML model 118. The training process 200 may be repeated until the loss function is minimized (e.g., the RMSE approaches zero or reaches a certain threshold), all training and test samples have been utilized, and / or a predetermined number of training iterations have been performed. Once it is determined that the ML model 118 is sufficiently trained, the trained ML model 118 can be used to predict a user's reaction time at inference time, such as, for example, at least by referencing... Figure 1 As discussed above, the trained ML model 118 can be stored in the memory of an artificial reality device (e.g., artificial reality system 400).

[0060] Figure 3An example method 300 for adaptively managing visual clutter in an artificial reality environment, according to a particular embodiment, is illustrated. Method 300 may begin at step 310, where a computing system (e.g., computer 408) associated with an artificial reality system (e.g., artificial reality system 400) may receive an image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment. The image may be captured using external cameras 405A to 405B of the artificial reality system 400, and the captured image may be a mixed reality image. In a particular embodiment, the image may be decomposed into a virtual layer (or virtual UI component 102) and a real-world layer (or real-world component 104), the virtual layer (or virtual UI component 102) comprising one or more virtual elements associated with a virtual environment, and the real-world layer (or real-world component 104) comprising one or more real-world elements associated with a real-world environment. In some embodiments, the one or more virtual elements may include applications installed on an artificial reality system (e.g., AR / VR applications), and the one or more real-world elements may include physical objects existing in the real environment (e.g., buildings, cars, people, trees, etc.). The one or more virtual elements may be overlaid on the one or more real-world elements.

[0061] In step 320, the computing system (e.g., computer 408 of the artificial reality system 400) may use one or more image analysis techniques to determine a first metric (e.g., first clutter metric 122) indicating the degree of clutter in the virtual environment (e.g., virtual UI clutter 112) based on the one or more virtual elements, and to determine a second metric (e.g., second clutter metric 124) indicating the degree of clutter in the real-world environment (e.g., real-world clutter 114) based on the one or more real-world elements. In a particular embodiment, the first metric (e.g., first clutter metric 122) indicating the degree of clutter in the virtual environment may be determined by performing one or more image analysis techniques or methods on a virtual layer corresponding to a virtual UI component (e.g., virtual UI component 102). The second metric (e.g., second clutter metric 124) indicating the degree of clutter in the real-world environment may be determined by performing one or more image analysis techniques on a real-world layer corresponding to a real-world component (e.g., real-world component 104). The one or more image analysis techniques may include, for example, but not limited to, feature congestion techniques, subband entropy techniques, or edge density techniques.

[0062] In step 330, a computing system (e.g., computer 408 of the artificial reality system 400) may determine multiple gaze features associated with the user based on user activity with respect to an image comprising one or more virtual elements and one or more real-world elements. The determination of these gaze features may be triggered in response to the user performing some activity. In some embodiments, user activity may include the user searching for a specific virtual element among one or more virtual elements and one or more real-world elements in the mixed reality image (e.g., an application). The computing system may determine multiple gaze features based on user activity. These multiple gaze features may be tracked or collected using eye-tracking sensors associated with the artificial reality system. In certain embodiments, the multiple gaze features may include, for example, but not limited to, gaze or saccade velocity, saccade probability, saccade height, fixation point, fixation duration, saccade duration, saccade length, saccade angular velocity, etc.

[0063] In step 340, the computing system (e.g., computer 408 of the artificial reality system 400) can use a machine learning model (e.g., ML model 118) to predict the user's reaction time when performing a user activity based on multiple gaze features. In a particular embodiment, the machine learning model can be trained by: (1) accessing multiple training samples obtained based on multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponds to a user trial and includes gaze features of a particular user observed over a specific duration during which the particular user performs an assigned task in a particular cluttered scenario; (2) using the machine learning model to predict the reaction times of multiple users in the multiple user trials based on the gaze features included in the multiple training samples; (3) comparing these predicted reaction times with multiple ground truth reaction times; and (4) updating the machine learning model based on the comparison.

[0064] In step 350, the computing system (e.g., computer 408 of the artificial reality system 400) may determine a third metric (e.g., third clutter metric 126) indicating the degree of clutter in the image based on reaction time predicted by a machine learning model. In step 360, the computing system may calculate an overall clutter metric (e.g., overall clutter metric 130) based on: a first metric determined based on one or more virtual elements; a second metric determined based on one or more real-world elements; and a third metric determined based on the predicted reaction time. In some embodiments, calculating the overall clutter metric may include taking a weighted average of the first, second, and third metrics according to weights assigned to each of the first, second, and third metrics.

[0065] In step 370, the computing system (e.g., computer 408 of the artificial reality system 400) may perform one or more actions to manage clutter in an image (e.g., a mixed reality image) based on an overall clutter metric. In a particular embodiment, performing one or more actions may be based on determining whether the overall clutter metric is higher or lower than a predetermined threshold. For example, the computing system may compare the overall clutter metric to a predetermined threshold and determine that the overall clutter metric is higher than the predetermined threshold. In response to determining that the overall clutter metric is higher than the predetermined threshold, the computing system may perform one or more actions for clutter management. One or more actions for managing clutter in a visual scene or image may include, for example, but not limited to, removing one or more virtual elements from the image, modifying information associated with the one or more virtual elements, changing the layout or position of the one or more virtual elements in the image, adjusting the size of the one or more virtual elements, etc. If the overall clutter metric is lower than the predetermined threshold, no clutter management actions may be performed, and the content presented to the user (e.g., images or UI) may remain unchanged.

[0066] Where appropriate, certain embodiments may be repeated. Figure 3 One or more steps of the method. Although this disclosure will Figure 3 The specific steps of the method are described and shown to be performed in a particular order, but this disclosure contemplates that they may be performed in any suitable order. Figure 3 Any suitable steps of the method. Furthermore, although this disclosure describes and illustrates example methods for adaptively managing visual clutter in artificial reality environments (including...) Figure 3 This disclosure considers any suitable method (including any suitable steps) for adaptively managing visual clutter in artificial reality environments, where appropriate, the method may include... Figure 3 A subset of the steps of the method. Furthermore, although this disclosure describes and illustrates the execution of... Figure 3 The method may refer to a specific component, device, or system of a particular step, but this disclosure is intended to perform... Figure 3 Any suitable step of the method, any suitable component, device or system, or any suitable combination.

[0067] Figure 4An example of an artificial reality system 400 worn by user 402 is shown. The artificial reality system 400 can be used to implement some of the various embodiments / examples disclosed herein. The artificial reality system 400 can be configured to operate as a virtual reality display, an augmented reality display, and / or a mixed reality display. In a particular embodiment, the artificial reality system 400 may include a head-mounted device (“HMD”) 404, a controller 406, and a computing system 408. The HMD 404 may be worn on the user's eyes and provide visual content to the user 402 via an internal display (not shown). The HMD 404 may have two separate internal displays, one for one of the user 402's eyes. Figure 4 As shown, HMD 404 can completely cover the user's field of view. By acting as the sole provider of visual information to user 402, HMD 404 achieves the goal of providing an immersive artificial reality experience. In a particular embodiment, HMD 404 can be configured to present a view of the user's surroundings or external physical environment as one or more transparent images (e.g., user 402 can still see the external physical environment while wearing HMD 404).

[0068] The HMD 404 can have an outward-facing camera, for example... Figure 4 The two forward-facing cameras 405A and 405B are shown. Although only two forward-facing cameras 405A and 405B are shown, the HMD 404 can have any number of cameras facing any direction (e.g., an upward-facing camera for capturing a ceiling or interior light, a downward-facing camera for capturing a portion of the user's face and / or body, a rearward-facing camera for capturing a portion of an object behind the user, and / or an internal camera for capturing the user's eye gaze for eye-tracking purposes). Outward-facing cameras are configured to capture the physical environment around the user and can do so continuously to generate a series of frames (e.g., as video).

[0069] A 3D representation can be generated based on depth measurements of a physical object observed through cameras 405A and 405B. Depth can be measured in various ways. In a particular embodiment, depth can be calculated based on stereo images. For example, two forward-facing cameras 405A and 405B can share overlapping fields of view and can be configured to capture images simultaneously. Therefore, the same physical object can be captured simultaneously by both cameras 405A and 405B. For example, in the image captured by camera 405A, specific features of the object may appear in a pixel p. A In one location, the same feature can appear in another pixel p in an image captured by camera 405B. BAs long as the depth measurement system knows that two pixels correspond to the same feature, it can use triangulation techniques to calculate the depth of the observed feature. For example, based on the position of camera 405A in 3D space and p A Relative to the pixel position of the field of view of camera 405A, a line can be projected from camera 405A through pixel p. A A similar line can be projected from another camera, the 405B, across the pixel p. B Since it is assumed that the two pixels correspond to the same physical feature, the two lines should intersect. The two intersecting lines, along with an imaginary line drawn between the two cameras 405A and 405B, form a triangle that can be used to calculate the distance of the observed feature from either camera 405A or 405B, or from a point in space where the observed feature is located.

[0070] In certain embodiments, the pose (e.g., position and orientation) of the HMD 404 within the environment may be required. For example, in order to render an appropriate display for the user 402 as the user moves around in the virtual environment, the system 400 will need to determine the user's position and orientation at any given time. The system 400 can also determine the viewpoint of either of the cameras 405A and 405B, or the viewpoint of either of the user's eyes, based on the pose of the HMD. In certain embodiments, the HMD 404 may be equipped with an inertial-measurement unit (IMU). Data generated by the IMU, along with stereo images captured by the outward-facing cameras 405A and 405B, allows the system 400 to calculate the pose of the HMD 404, for example, using simultaneous localization and mapping (SLAM) or other suitable techniques.

[0071] In a particular embodiment, the artificial reality system 400 may also include one or more controllers 406 that enable the user 402 to provide input. The controllers 406 may communicate with the HMD 404 or a separate computing unit 408 via a wireless or wired connection. The controllers 406 may have any number of buttons or other mechanical input mechanisms. Furthermore, the controllers 406 may have an IMU, allowing the position of the controllers 406 to be tracked. The controllers 406 may also be tracked based on a predetermined pattern on the controller. For example, the controllers 406 may have several infrared light-emitting diodes (LEDs) or other known observable features that collectively form the predetermined pattern. The system 400 may be able to use a sensor or camera to capture images of the predetermined pattern on the controller. The system may calculate the position and orientation of the controller relative to the sensor or camera based on the observed orientation of these patterns.

[0072] The artificial reality system 400 may also include a computer unit 408. The computer unit may be a physically separate, independent unit from the HMD 404, or it may be integrated with the HMD 404. In embodiments where the computer 408 is a separate unit, the computer may be communicatively coupled to the HMD 404 via a wireless or wired link. The computer 408 may be a high-performance device (e.g., a desktop or laptop computer) or a resource-constrained device (e.g., a mobile phone). A high-performance device may have a dedicated graphics processing unit (GPU) and a high-capacity or constant-power power supply. On the other hand, a resource-constrained device may not have a GPU and may have a limited battery capacity. Therefore, the algorithms that the artificial reality system 400 can actually use depend on the capabilities of its computer unit 408.

[0073] Figure 5 An example network environment 500 associated with an artificial reality system is shown. Although Figure 5 While illustrated as a virtual reality system, this example network environment 500 may include one or more other artificial reality systems, such as mixed reality systems, augmented reality systems, etc. Network environment 500 includes a user 501 interacting with a client system 530, a social networking system 560, and a third-party system 570, which are interconnected via network 510. Although... Figure 5 A specific arrangement of user 501, client system 530, social networking system 560, third-party system 570, and network 510 is shown; however, this disclosure contemplates any suitable arrangement of user 501, client system 530, social networking system 560, third-party system 570, and network 510. By way of example and not limitation, two or more of user 501, client system 530, social networking system 560, and third-party system 570 may bypass network 510 and be directly connected to each other. As another example, two or more of client system 530, social networking system 560, and third-party system 570 may be physically or logically located entirely or partially in the same location. Furthermore, although... Figure 5 A specific number of users 501, client systems 530, social networking systems 560, third-party systems 570, and networks 510 are shown, but this disclosure contemplates any suitable number of client systems 530, social networking systems 560, third-party systems 570, and networks 510. As an example and not a limitation, network environment 500 may include multiple users 501, multiple client systems 530, multiple social networking systems 560, multiple third-party systems 570, and multiple networks 510.

[0074] This disclosure considers any suitable network 510. As an example and not a limitation, one or more portions of network 510 may include an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular telephone network, or a combination of two or more of these networks. Network 510 may include one or more networks 510.

[0075] Link 550 can connect client system 530, social networking system 560, and third-party system 570 to communication network 510, or can connect client system 530, social networking system 560, and third-party system 570 to each other. This disclosure contemplates any suitable link 550. In a particular embodiment, one or more links 550 include one or more wired (e.g., Digital Subscriber Line (DSL) or Data OverCable Service Interface Specification (DOCSIS)) links, one or more wireless (e.g., Wi-Fi or Worldwide Interoperability for Microwave Access (WiMAX)) links, or one or more optical (e.g., Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)) links. In a particular embodiment, one or more links 550 each include an ad hoc network, intranet, extranet, VPN, LAN, WLAN, WAN, WWAN, MAN, a portion of the Internet, a portion of the PSTN, a cellular-based network, a satellite-based network, another link 550, or a combination of two or more such links 550. Multiple links 550 need not be identical throughout the network environment 500. In one or more aspects, one or more first links 550 may differ from one or more second links 550.

[0076] In a particular embodiment, client system 530 may be an electronic device comprising hardware, software, or embedded logic components, or a combination of two or more such components, and capable of performing appropriate functions implemented or supported by client system 530. By way of example and not limitation, client system 530 may include computer systems such as desktop computers, laptop or notebook computers, netbooks, tablet computers, e-book readers, GPS devices, cameras, personal digital assistants (PDAs), handheld electronic devices, cellular phones, smartphones, virtual reality headsets and controllers, or mixed reality headsets and controllers, other suitable electronic devices, or any suitable combination thereof. This disclosure contemplates any suitable client system 530. Client system 530 enables network users at client system 530 to access network 510. Client system 530 enables its users to communicate with other users at other client systems 530. Client system 530 may generate virtual reality or mixed reality environments for users to interact with content.

[0077] In a particular embodiment, client system 530 may include a virtual reality (or augmented reality or mixed reality) head-mounted viewer 532 and one or more virtual reality input devices 534 (e.g., virtual reality controllers). A user at client system 530 may wear the virtual reality head-mounted viewer 532 and use the one or more virtual reality input devices to interact with the virtual reality environment 536 generated by the virtual reality head-mounted viewer 532. Although not shown, client system 530 may also include a separate processing computer and / or any other components of the virtual reality system. The virtual reality head-mounted viewer 532 may generate the virtual reality environment 536, which may include system content 538 (including but not limited to an operating system), such as software or firmware updates, and the virtual reality environment 536 may also include third-party content 540, such as content from an application or content dynamically downloaded from the Internet (e.g., web page content). The virtual reality head-mounted viewer 532 may include one or more sensors 542 (e.g., accelerometers, gyroscopes, magnetometers) for generating sensor data that tracks the position of the head-mounted viewer device 532. The head-mounted viewer 532 may also include an eye tracker for tracking the position of the user's eyes or the direction of the user's gaze. The client system 530 may use data from the one or more sensors 542 to determine velocity, orientation, and gravity relative to the head-mounted viewer. One or more virtual reality input devices 534 may include one or more sensors 544 (e.g., accelerometers, gyroscopes, magnetometers, and touch sensors) to generate sensor data that tracks the position of the input device 534 and the position of the user's fingers. The client system 530 may use outside-in tracking, in which a tracking camera (not shown) is positioned outside the virtual reality head-mounted viewer 532 and within its line of sight. In outside-in tracking, the tracking camera may track the position of the virtual reality head-mounted viewer 532 (e.g., by tracking one or more infrared LED markers on the virtual reality head-mounted viewer 532). Alternatively or additionally, the client system 530 may utilize inside-out tracking, in which a tracking camera (not shown) may be placed on or inside the virtual reality headset 532. In inside-out tracking, the tracking camera may capture images of its surroundings in the real world, and the changing perspective of the real world may be used to determine the tracking camera's position in space.

[0078] In a particular embodiment, client system 530 (e.g., HMD) may include a pass-through engine 546 for providing the pass-through functionality described herein, and the client system may have one or more add-ons, plugins, or other extensions. A user at client system 530 can connect to a specific server (e.g., server 562 or a server associated with a third-party system 570). The server can accept requests and communicate with client system 530.

[0079] Third-party content 540 may include a web browser and may have one or more add-ons, plugins, or other extensions. A user at client system 530 may enter a Uniform Resource Locator (URL) or other address to direct the web browser to a specific server (e.g., server 562, or a server associated with third-party system 570), and the web browser may generate a Hypertext Transfer Protocol (HTTP) request and send that HTTP request to the server. The server may receive the HTTP request and, in response, send one or more Hypertext Markup Language (HTML) files to client system 530. Client system 530 may render a web interface (e.g., a webpage) based on the HTML files from the server for presentation to the user. This disclosure considers any suitable source file. By way of example and not limitation, the web interface may be rendered from HTML files, Extensible Hypertext Markup Language (XHTML) files, or Extensible Markup Language (XML) files, depending on specific needs. This interface can also execute scripts, such as, but not limited to, combinations of markup languages ​​and scripts. In this document, references to the web interface include one or more corresponding source files (which the browser can use to render the web interface), and vice versa, where appropriate.

[0080] In a particular embodiment, the social networking system 560 may be a network-addressable computing system capable of hosting online social networks. For example, the social networking system 560 may generate, store, receive, and transmit social networking data, such as user profile data, concept profile data, social graph information, or other suitable data related to the online social network. Other components of the network environment 500 may access the social networking system 560 directly or via network 510. By way of example and not limitation, the client system 530 may use a web browser of third-party content 540 or a native application associated with the social networking system 560 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof) to access the social networking system 560 directly or via network 510. In a particular embodiment, the social networking system 560 may include one or more servers 562. Each server 562 may be a single server or a distributed server spanning multiple computers or multiple data centers. Servers 562 may be of various types, such as, but not limited to: web servers, news servers, mail servers, messaging servers, advertising servers, file servers, application servers, exchange servers, database servers, proxy servers, another server suitable for performing the functions or processes described herein, or any combination thereof. In a particular embodiment, each server 562 may include hardware, software, or embedded logic components, or a combination of two or more such components, for performing appropriate functions implemented or supported by the server 562. In a particular embodiment, the social networking system 560 may include one or more data repositories 564. Data repositories 564 may be used to store various types of information. In a particular embodiment, the information stored in the data repository 564 may be organized according to a specific data structure. In a particular embodiment, each data repository 564 may be a relational database, a columnar database, an association database, or other suitable database. Although this disclosure describes or illustrates specific types of databases, this disclosure also contemplates any suitable type of database. Specific embodiments may provide interfaces that enable client system 530, social networking system 560, or third-party system 570 to manage, retrieve, modify, add, or delete information stored in the data repository 564.

[0081] In a particular embodiment, the social network system 560 may store one or more social graphs in one or more data repositories 564. In a particular embodiment, the social graph may include multiple nodes—which may include multiple user nodes (each user node corresponds to a specific user) or multiple concept nodes (each concept node corresponds to a specific concept)—and multiple edges connecting these nodes. The social network system 560 may provide users of the online social network with the ability to communicate and interact with other users. In a particular embodiment, a user can join an online social network via the social network system 560 and then add connections (e.g., relationships) to some other users in the social network system 560 that they wish to connect with. In this document, the term "friend" may refer to any other user with whom a user in the social network system 560 has already formed a connection, association, or relationship through the social network system 560.

[0082] In a particular embodiment, the social networking system 560 may provide users with the ability to take actions on various types of items or objects supported by the social networking system 560. By way of example, and not limitation, these items and objects may include groups or social networks to which the user of the social networking system 560 may belong, events or calendar entries that the user may be interested in, computer-based applications available to the user, transactions involving items that allow the user to buy or sell through services, interactions with user-executable advertisements, or other suitable items or objects. Users may interact with anything that can be represented in the social networking system 560, or with anything that can be represented by an external system 570, separate from and coupled to the social networking system 560 via network 510.

[0083] In a particular embodiment, the social networking system 560 may be able to link various entities. By way of example and not limitation, the social networking system 560 may enable users to interact with each other and receive content from third-party systems 570 or other entities, or allow users to interact with these entities through application programming interfaces (APIs) or other communication channels.

[0084] In certain embodiments, third-party system 570 may include one or more types of servers, one or more data repositories, one or more interfaces (including but not limited to APIs), one or more web services, one or more content sources, one or more networks, or any other suitable component to which the server can communicate. Third-party system 570 may be operated by an entity different from the entity operating social networking system 560. However, in certain embodiments, social networking system 560 and third-party system 570 may operate collaboratively to provide social networking services to users of either social networking system 560 or third-party system 570. In this sense, social networking system 560 may provide a platform or backbone that other systems (e.g., third-party system 570) can use to provide social networking services and functionality to users on the Internet.

[0085] In a particular embodiment, third-party system 570 may include a third-party content object provider. The third-party content object provider may include one or more sources of content objects that can be delivered to client system 530. As an example, and not a limitation, content objects may include information related to things or activities of interest to the user, such as movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. As another example, and not a limitation, content objects may include incentive content objects, such as coupons, discount tickets, gift certificates, or other suitable incentives.

[0086] In a particular embodiment, the social networking system 560 also includes user-generated content objects, which can enhance user interaction with the social networking system 560. User-generated content can include any content that a user can add, upload, send, or "post" to the social networking system 560. As an example, and not a limitation, a user transmits a post from the client system 530 to the social networking system 560. Posts can include data such as status updates or other text data, location information, photos, videos, links, music, or other similar data or media. Content can also be added to the social networking system 560 by third parties via a "communication channel" (e.g., a news feed or stream).

[0087] In certain embodiments, the social networking system 560 may include various servers, subsystems, programs, modules, logs, and data repositories. In certain embodiments, the social networking system 560 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, an action log, a third-party content object public log, an inference module, an authorization / privacy server, a search module, an ad targeting module, a user interface module, a user profile repository, a contact repository, a third-party content repository, or a location repository. The social networking system 560 may also include suitable components such as a web interface, security mechanisms, load balancers, failover servers, management and network operations consoles, other suitable components, or any suitable combination thereof. In certain embodiments, the social networking system 560 may include one or more user profile repositories for storing user profiles. User profiles may include, for example, biometric information, personal background information, behavioral information, social information, or other types of descriptive information (e.g., work experience, educational history, hobbies or preferences, interests, kinship, or location). Interest information may include interests associated with one or more categories. Categories may be general or specific. As an example, and not a limitation, if a user “likes” items related to a shoe brand, the category could be that brand, or it could be a generic category like “shoes” or “clothing.” A contact repository can be used to store contact information about users. Contact information can indicate users who have similar or shared work experience, group memberships, hobbies, education, or are in any way related to or share common attributes. Contact information can also include user-defined connections (both internal and external) between different users and content. A web server can be used to link the social networking system 560 to one or more client systems 530 or one or more third-party systems 570 via network 510. The web server may include a mail server or other messaging functionality for receiving and routing messages between the social networking system 560 and one or more client systems 530. An API request server can allow third-party systems 570 to access information from the social networking system 560 by calling one or more APIs. An action logger can be used to receive information from the web server related to user actions related to logging into or out of the social networking system 560. Combined with the action log, a log of third-party content objects that users interact with can be maintained. The notification controller can provide information about content objects to the client system 530. This information can be pushed to the client system 530 as a notification, or retrieved from the client system 530 in response to a request received from the client system 530. The authorization server can be used to implement one or more privacy settings for users of the social networking system 560.A user's privacy settings determine how specific information associated with that user can be shared. The authorization server can allow users to choose, for example, by setting appropriate privacy settings, whether or not to allow the social network system 560 to record their actions or share their actions with other systems (e.g., third-party system 570). A third-party content object repository can be used to store content objects received from third parties (e.g., third-party system 570). A location repository can be used to store location information received from client systems 530 associated with the user. The advertising pricing module can combine social information, current time, location information, or other suitable information to deliver relevant advertisements to users in the form of notifications.

[0088] Figure 6 An example computer system 600 is illustrated. In a particular embodiment, one or more computer systems 600 perform one or more steps of one or more processes, algorithms, techniques, or methods described or illustrated herein. In a particular embodiment, one or more computer systems 600 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 600 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular embodiments include one or more portions of one or more computer systems 600. Throughout this document, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.

[0089] This disclosure contemplates any suitable number of computer systems 600. This disclosure contemplates computer systems 600 employing any suitable physical form. By way of example and not limitation, computer system 600 may be: an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service terminal (kiosk), a mainframe, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented reality / virtual reality device, or a combination of two or more of these systems. Where appropriate, computer system 600 may include one or more computer systems 600; may be single or distributed; may span multiple locations; may span multiple machines; may span multiple data centers; or may be located in a cloud comprising one or more cloud components in one or more networks. Where appropriate, one or more computer systems 600 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 600 may execute one or more steps of the methods described or shown herein in real time or in batch mode. Where appropriate, one or more computer systems 600 may execute one or more steps of the methods described or shown herein at different times or at different locations.

[0090] In a particular embodiment, computer system 600 includes a processor 602, memory 604, storage device 606, input / output (I / O) interface 608, communication interface 610, and bus 612. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of suitable components in any suitable arrangement.

[0091] In a particular embodiment, processor 602 includes hardware for executing instructions, such as instructions constituting a computer program. By way of example, and not limitation, to execute instructions, processor 602 may retrieve (or read) these instructions from internal registers, internal cache, memory 604, or storage device 606; decode and execute these instructions; and subsequently write one or more results to internal registers, internal cache, memory 604, or storage device 606. In a particular embodiment, processor 602 may include one or more internal caches that can be used for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 602 including any suitable number of suitable internal caches. By way of example, and not limitation, processor 602 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). The multiple instructions in the instruction cache may be copies of multiple instructions in memory 604 or storage device 606, and the instruction cache may accelerate the retrieval of these instructions by processor 602. The data in the data cache may be a copy of the data in memory 604 or storage device 606 for operation by instructions executed at processor 602; the result of a previous instruction executed at processor 602 for access by subsequent instructions executed at processor 602, or for writing to memory 604 or storage device 606; or the data in the data cache may be other suitable data. The data cache can accelerate read or write operations performed by processor 602. The TLB can accelerate virtual address translation by processor 602. In a particular embodiment, processor 602 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 602 including any suitable number of suitable internal registers. Where appropriate, processor 602 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 602. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.

[0092] In a particular embodiment, memory 604 includes main memory for storing instructions to be executed by processor 602 or data to be operated on by processor 602. By way of example and not limitation, computer system 600 may load multiple instructions from storage device 606 or another source (e.g., another computer system 600) into memory 604. Processor 602 may then load these instructions from memory 604 into internal registers or internal cache. To execute the instructions, processor 602 may retrieve the instructions from internal registers or internal cache and decode them. During or after instruction execution, processor 602 may write one or more results (which may be intermediate or final results) into internal registers or internal cache. Processor 602 may then write one or more of these results into memory 604. In a particular embodiment, processor 602 executes only the instructions in one or more internal registers or internal cache or memory 604 (not storage device 606 or elsewhere) and operates only on the data in one or more internal registers or internal cache or memory 604 (not storage device 606 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) couple processor 602 to memory 604. Bus 612 may include one or more memory buses as described below. In a particular embodiment, one or more memory management units (MMUs) are located between processor 602 and memory 604 and facilitate access to memory 604 requested by processor 602. In a particular embodiment, memory 604 includes random access memory (RAM). Where appropriate, the RAM is volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port RAM or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 604 may include one or more memories 604. Although this disclosure describes and illustrates specific memories, this disclosure contemplates any suitable memory.

[0093] In a particular embodiment, storage device 606 includes a mass storage device for data or instructions. By way of example, and not limitation, storage device 606 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these storage devices. Where appropriate, storage device 606 may include removable or non-removable (or fixed) media. Where appropriate, storage device 606 may be located inside or outside of computer system 600. In a particular embodiment, storage device 606 is a non-volatile solid-state memory. In a particular embodiment, storage device 606 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these ROMs. This disclosure contemplates a mass storage device 606 in any suitable physical form. Where appropriate, storage device 606 may include one or more memory control units facilitating communication between processor 602 and storage device 606. Where appropriate, storage device 606 may include one or more storage devices 606. Although this disclosure describes and illustrates specific storage devices, this disclosure contemplates any suitable storage device.

[0094] In a particular embodiment, I / O interface 608 includes hardware, software, or both that provides one or more interfaces for communication between computer system 600 and one or more I / O devices. Where appropriate, computer system 600 may include one or more of these I / O devices. One or more of these I / O devices may enable communication between a person and computer system 600. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, trackball, video camera, another suitable I / O device, or a combination of two or more of these I / O devices. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 608 for such I / O devices. Where appropriate, I / O interface 608 may include one or more device drivers or software drivers that enable processor 602 to drive one or more of these I / O devices. Where appropriate, I / O interface 608 may include one or more I / O interfaces 608. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.

[0095] In a particular embodiment, communication interface 610 includes hardware, software, or both that provides one or more interfaces for communication (e.g., packet-based communication) between computer system 600 and one or more other computer systems 600 or with one or more networks. By way of example and not limitation, communication interface 610 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (e.g., Wi-Fi networks). This disclosure contemplates any suitable network and any suitable communication interface 610 for that network. By way of example and not limitation, computer system 600 may communicate with one or more of the following networks: ad hoc networks; personal area networks (PANs); local area networks (LANs); wide area networks (WANs), metropolitan area networks (MANs), or the Internet, or a combination of two or more of these networks. One or more of these networks may be wired or wireless. As an example, computer system 600 may communicate with: a wireless PAN (WPAN) (e.g., Bluetooth WPAN); a Wi-Fi network; a Wi-Fi Max network; a cellular telephone network (e.g., a Global System for Mobile Communication (GSM) network); or other suitable wireless networks; or a combination of two or more of these. Where appropriate, computer system 600 may include any suitable communication interface 610 for any of these networks. Where appropriate, communication interface 610 may include one or more communication interfaces 610. Although specific communication interfaces are described and shown in this disclosure, any suitable communication interface is contemplated in this disclosure.

[0096] In a particular embodiment, bus 612 includes hardware, software, or both, that couple multiple components of computer system 600 to each other. By way of example and not limitation, bus 612 may include an Accelerated GraphicsPort (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth interconnect, a low-pin-count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these. Where appropriate, bus 612 may include one or more buses 612. Although this disclosure describes and illustrates specific buses, this disclosure contemplates any suitable bus or interconnection.

[0097] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based integrated circuits (ICs) or other ICs (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disc drives (ODDs), magneto-optical disk drives (MODs), floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, security digital cards or security digital drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these storage media. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.

[0098] In this document, unless otherwise expressly stated or the context otherwise requires, "or" is open-ended rather than exclusive. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, "A or B" means "A, B, or both." Furthermore, unless otherwise expressly stated or the context otherwise requires, "and" is both common and separate. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, "A and B" means "A and B, commonly or separately."

[0099] The scope of this disclosure includes all changes, substitutions, variations, transformations, and modifications to the exemplary embodiments described or illustrated herein, which will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although this disclosure describes and illustrates various embodiments including specific components, elements, features, functions, operations, or steps, any embodiment in these embodiments may include any combination or arrangement of any component, element, feature, function, operation, or step described or illustrated anywhere herein as will be understood by those skilled in the art. Furthermore, the device, system, or component mentioned in the appended claims being adapted, arranged, enabled, configured, activated, operable, or operable to perform a specific function includes the device, system, or component, whether or not it or the specific function is activated, turned on, or unlocked, provided that the device, system, or component is so adapted, arranged, enabled, configured, activated, operable, or operable. Moreover, although this disclosure describes or illustrates specific embodiments to provide specific advantages, specific embodiments may not provide these advantages, provide some of these advantages, or provide all of these advantages.

Claims

1. A method comprising a computing system: Receive an image, the image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; Using one or more image analysis techniques, a first metric indicating the degree of clutter in the virtual environment is determined based on the one or more virtual elements, and a second metric indicating the degree of clutter in the real-world environment is determined based on the one or more real-world elements; Based on user activity with respect to the image including the one or more virtual elements and the one or more real-world elements, determine multiple gaze features associated with the user; Using a machine learning model, the user's reaction time when performing the user activity is predicted based on the multiple gaze features; A third measure indicating the degree of clutter in the image is determined based on the predicted reaction time; The overall disorder metric is calculated based on the first metric determined based on the one or more virtual elements; The second metric is determined based on the one or more real-world elements; And the third metric determined based on the predicted reaction time; as well as One or more actions are performed based on the overall clutter metric to manage clutter in the image.

2. The method according to claim 1, further comprising: Training the machine learning model, wherein training the machine learning model includes: Access is based on multiple training samples obtained from multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponding to a user trial includes the gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; Using the machine learning model, based on gaze features included in the plurality of training samples, the reaction time of a plurality of users in the plurality of user trials is predicted. Compare the predicted reaction times with multiple true ground reaction times; and The machine learning model is updated based on the comparison.

3. The method according to claim 1 or 2, further comprising: The overall disorder measure is compared with a predetermined threshold. as well as Determine that the overall disorder metric is higher than the predetermined threshold. The one or more actions for managing the clutter are performed in response to determining that the overall clutter metric is higher than the predetermined threshold.

4. The method according to any one of claims 1 to 3, wherein, The one or more actions used to manage the clutter include: Remove one or more virtual elements from the image; Modify the information associated with the one or more virtual elements; Change the layout or position of one or more virtual elements in the image; or Adjust the size of the one or more virtual elements.

5. The method according to any one of claims 1 to 4, wherein, Calculating the overall disorder measure includes: Calculate the weighted average of the first metric, the second metric, and the third metric based on the weights assigned to each of them.

6. The method according to any one of claims 1 to 5, wherein, The plurality of gaze features include: The speed of staring or saccades; Probability of scanning; Scan height; fixation point; gaze duration; Duration of the scan; The length of the scan; or Angular velocity of the scan.

7. The method according to any one of claims 1 to 6, wherein, The user activities include: The user searches for a specific virtual element among one or more virtual elements and one or more real-world elements in the image.

8. The method according to any one of claims 1 to 7, further comprising: The image is decomposed into a virtual layer and a real-world layer. The virtual layer includes one or more virtual elements associated with the virtual environment, and the real-world layer includes one or more real-world elements associated with the real-world environment, wherein: The first metric indicating the degree of clutter in the virtual environment is determined by performing one or more image analysis techniques on the virtual layer; and The second metric, which indicates the degree of clutter in the real-world environment, is determined by performing one or more image analysis techniques on the real-world layer.

9. The method according to any one of claims 1 to 8, wherein, The one or more image analysis techniques include: Feature-based congestion techniques; Subband entropy technique; or Edge density technology.

10. The method according to any one of claims 1 to 9, wherein, The image was captured by an artificial reality system; Furthermore, optionally The one or more virtual elements include applications installed on the artificial reality system; and The one or more real-world elements include physical objects existing in the real-world environment.

11. The method according to any one of claims 1 to 10, wherein, The one or more virtual elements are overlaid on the one or more real-world elements.

12. One or more computer-readable non-transitory storage media, said one or more computer-readable non-transitory storage media comprising software, said software being operable to: Receive an image, the image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; Using one or more image analysis techniques, a first metric indicating the degree of clutter in the virtual environment is determined based on the one or more virtual elements, and a second metric indicating the degree of clutter in the real-world environment is determined based on the one or more real-world elements; Based on user activity with respect to the image including the one or more virtual elements and the one or more real-world elements, determine multiple gaze features associated with the user; Using a machine learning model, the user's reaction time when performing the user activity is predicted based on the multiple gaze features; A third measure indicating the degree of clutter in the image is determined based on the predicted reaction time; The overall disorder metric is calculated based on the first metric determined based on the one or more virtual elements; The second metric is determined based on the one or more real-world elements; And the third metric determined based on the predicted reaction time; as well as One or more actions are performed based on the overall clutter metric to manage clutter in the image.

13. The medium according to claim 12, further comprising at least one of the following features: (a) Among them, When executed, the software is also capable of training the machine learning model, wherein training the machine learning model includes: Access is based on multiple training samples obtained from multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponding to a user trial includes the gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; Using the machine learning model, based on gaze features included in the plurality of training samples, the reaction time of a plurality of users in the plurality of user trials is predicted. Compare the predicted reaction times with multiple true ground reaction times; and Update the machine learning model based on the comparison; and / or (b) wherein, when the software is executed, it is also capable of: The overall disorder measure is compared with a predetermined threshold; and Determine that the overall disorder metric is higher than the predetermined threshold. Wherein, the one or more actions for managing the clutter are performed in response to determining that the overall clutter metric is higher than the predetermined threshold; and / or (c) wherein the one or more actions for managing the disorder include: Remove one or more virtual elements from the image; Modify the information associated with the one or more virtual elements; Change the layout or position of one or more virtual elements in the image; or Adjust the size of the one or more virtual elements.

14. A system comprising: One or more processors; as well as One or more computer-readable non-transitory storage media, said one or more computer-readable non-transitory storage media being coupled to one or more of said one or more processors and including instructions that, when executed by said one or more of said one or more processors, are operable to cause the system to: Receive an image, the image comprising one or more virtual elements associated with a virtual environment and one or more real-world elements associated with a real-world environment; Using one or more image analysis techniques, a first metric indicating the degree of clutter in the virtual environment is determined based on the one or more virtual elements, and a second metric indicating the degree of clutter in the real-world environment is determined based on the one or more real-world elements; Based on user activity with respect to the image including the one or more virtual elements and the one or more real-world elements, determine multiple gaze features associated with the user; Using a machine learning model, the user's reaction time when performing the user activity is predicted based on the multiple gaze features; A third measure indicating the degree of clutter in the image is determined based on the predicted reaction time; The overall disorder metric is calculated based on the first metric determined based on the one or more virtual elements; The second metric is determined based on the one or more real-world elements; And the third metric determined based on the predicted reaction time; as well as One or more actions are performed based on the overall clutter metric to manage clutter in the image.

15. The system of claim 14, further comprising at least one of the following features: (a) Among them, The one or more processors, when executing the instructions, are also capable of operating to enable the system to train the machine learning model, wherein... Training the machine learning model includes: Access is based on multiple training samples obtained from multiple user trials or studies in different cluttered scenarios, wherein each training sample corresponding to a user trial includes the gaze characteristics of a specific user observed over a specific duration during which the specific user performs an assigned task in a specific cluttered scenario; Using the machine learning model, based on the gaze features included in the plurality of training samples, predict multiple reaction times of multiple users in the plurality of user trials; The predicted reaction times are compared with the true ground reaction times; and Update the machine learning model based on the comparison; and / or (b) wherein the one or more processors, when executing the instructions, are also capable of operating such that the system: The overall disorder measure is compared with a predetermined threshold; and Determine that the overall disorder metric is higher than the predetermined threshold. Wherein, the one or more actions for managing the clutter are performed in response to determining that the overall clutter metric is higher than the predetermined threshold; and / or (c) wherein the one or more actions for managing the disorder include: Remove one or more virtual elements from the image; Modify the information associated with the one or more virtual elements; Change the layout or position of one or more virtual elements in the image; or Adjust the size of the one or more virtual elements.