Multi-modal input for computer-generated reality

Through multimodal input technology, combined with facial expressions, gestures and voice, the problem of insufficient interaction between virtual objects and physical environment in augmented reality technology is solved, and a better user interaction and immersion experience is achieved.

CN120276604APending Publication Date: 2025-07-08APPLE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510489077.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-09
Filing Date
2020-09-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine the physical environment and the virtual environment to provide interaction and immersion. Especially in augmented reality technology, the interaction between virtual objects and the physical environment and the responsiveness of sensory inputs is insufficient.

Method used

Through multimodal inputs, such as facial expressions, gestures, voice and explicit hardware inputs, combined with a computer-generated reality system, track user actions and adjust the virtual environment in real time, providing interactive and immersiveness.

Benefits of technology

It realizes effective interaction between virtual objects and physical environments in an augmented reality environment, enhances the immersive experience and interactivity of users, and improves the responsiveness and real-timeness of computer-generated reality systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276604A_ABST
    Figure CN120276604A_ABST
Patent Text Reader

Abstract

The invention relates to multi-modal input for computer-generated reality. Particular implementations of the subject technology provide for determining an operating mode of an electronic device based at least in part on whether the electronic device is communicatively coupled to an associated infrastructure. Based on the determined mode of operation, the subject technology identifies a set of input modalities for initiating recording of content within the field of view of the electronic device. The subject technology monitors sensor information generated by at least one sensor included in or communicatively coupled to the electronic device. Further, when the monitored sensor information indicates that at least one of the identified set of input modalities has been triggered, the subject technology initiates recording of content within the field of view of the electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application that enters the Chinese national phase from a PCT application with an international filing date of September 9, 2020, a national application number of 202080057569.5, and an invention title of "Multimodal Inputs for Computer-Generated Reality".

[0002] Cross-Reference to Related Applications

[0003] This patent application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 897,909, filed on September 9, 2019, with the title "Multimodal Inputs for Computer-Generated Reality", the disclosure of which is hereby incorporated by reference in its entirety. Technical Field

[0004] This specification generally relates to computer-generated reality environments, including utilizing multimodal inputs in computer-generated reality environments. Background Art

[0005] Augmented reality technology aims to bridge the gap between virtual and physical environments by providing an enhanced physical environment augmented with electronic information. Thus, the electronic information appears to be part of the physical environment perceived by the user. In one example, augmented reality technology further provides a user interface to interact with the electronic information overlaid in the augmented physical environment. Brief Description of the Drawings

[0006] Some features of the subject technology are shown in the appended claims. However, for purposes of explanation, several implementations of the subject technology are set forth in the following drawings.

[0007] Figure 1 An exemplary system architecture including various electronic devices that can implement the subject system according to one or more specific implementations is shown.

[0008] Figure 2 An exemplary software architecture that can be implemented on an electronic device according to one or more specific implementations of the subject technology is shown.

[0009] Figure 3A An example of facial expression tracking to initiate computer-generated reality recording according to a specific implementation of the subject technology is shown.

[0010] Figure 3B and Figure 3C An example of tracking a gaze direction to initiate computer-generated reality recording according to a specific implementation of the subject technology is shown.

[0011] Figure 4A , Figure 4BAnd Figure 4C illustrates an example of determining a region of interest within a computer-generated reality environment and initiating a recording based on the region of interest, in accordance with some specific implementations of the present subject matter technology.

[0012] Figure 5A 、 Figure 5B and Figure 5C illustrates an example of providing annotations to various objects or entities within a computer-generated reality environment, in accordance with some specific implementations of the present subject matter technology.

[0013] Figure 6 illustrates a flowchart of an example process for initiating a recording of content within the field of view of an electronic device.

[0014] Figure 7 illustrates a flowchart of an exemplary process for updating a set of input modalities for initiating a recording of content on an electronic device, in accordance with one or more specific implementations.

[0015] Figure 8 illustrates a flowchart of an exemplary process for determining a quality of service metric associated with an operating mode of an electronic device, in accordance with one or more specific implementations.

[0016] Figure 9 illustrates an electronic system that can implement one or more specific implementations of the present subject matter technology. Detailed Description

[0017] The detailed description presented below is intended as a description of various configurations of the present subject matter technology and is not intended to represent the only configuration in which the present subject matter technology can be practiced. The drawings are incorporated herein and constitute a part of the detailed description. The detailed description includes specific details intended to provide a thorough understanding of the present subject matter technology. However, the present subject matter technology is not limited to the specific details set forth herein and may be practiced using one or more other specific implementations. In one or more specific implementations, structures and components are shown in block diagram form to avoid obscuring the concepts of the present subject matter technology.

[0018] A computer-generated reality (CGR) system enables physical and virtual environments to be combined in varying degrees to facilitate real-time interaction with a user. Thus, as described herein, such CGR systems can include various possible combinations of physical and virtual environments, including augmented reality, which primarily includes physical elements and is closer to the physical environment than a virtual environment (e.g., without physical elements). In this way, the physical environment can be connected to the virtual environment through the CGR system. A user immersed in a CGR environment can navigate within such an environment, and the CGR system can track the user's viewpoint to provide visualization based on how the user is located within the environment.

[0019] The physical environment refers to the physical world that people can sense and / or interact with without the help of an electronic system. Physical environments such as physical parks include physical objects such as physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell.

[0020] In contrast, a computer-generated reality (CGR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In CGR, a subset of a person's physical movements or representations thereof are tracked, and in response, one or more characteristics of one or more virtual objects simulated in the CGR environment are adjusted in a manner that complies with at least one physical law. For example, a CGR system can detect a person's body and / or head turning, and in response, adjust the graphical content and sound field presented to the person in a manner similar to how such views and sounds would change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of the characteristics of virtual objects in a CGR environment can be made in response to a representation of a physical movement (e.g., a voice command).

[0021] A person can use any of their senses to sense and / or interact with CGR objects, including vision, hearing, touch, taste, and smell. For example, a person can sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides the perception of point audio sources in 3D space. As another example, an audio object can enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some CGR environments, a person can sense and / or interact only with audio objects.

[0022] Examples of CGR include virtual reality and mixed reality.

[0023] A virtual reality (VR) environment refers to a simulated environment that is designed to be completely computer-generated sensory input for one or more senses. A VR environment includes multiple virtual objects that a person can sense and / or interact with. For example, computer-generated images of trees, buildings, and avatars representing people are examples of virtual objects. A person can sense and / or interact with the virtual objects in a VR environment through the simulation of the person's presence within the computer-generated environment and / or through the simulation of a subgroup of the person's physical movements within the computer-generated environment.

[0024] Compared with a VR environment designed to be completely based on computer-generated sensory inputs, a mixed reality (MR) environment is a simulated environment designed to include, in addition to computer-generated sensory inputs (e.g., virtual objects), sensory inputs or their representations from the physical environment. On the virtual continuum, an MR environment is any environment between, but not including, a completely physical environment at one end and a virtual reality environment at the other end.

[0025] In some MR environments, the computer-generated sensory inputs can respond to changes in the sensory inputs from the physical environment. Additionally, some electronic systems for presenting an MR environment can track the position and / or orientation relative to the physical environment so that virtual objects can interact with real objects (i.e., physical items from the physical environment or their representations). For example, the system can cause movement such that a virtual tree appears stationary relative to the physical ground.

[0026] An augmented reality (AR) environment is a simulated environment in which one or more virtual objects are superimposed on the physical environment or its representation. For example, an electronic system for presenting an AR environment can have a transparent or translucent display through which a person can directly view the physical environment. The system can be configured to present virtual objects on the transparent or translucent display such that the person perceives the virtual objects superimposed on a portion of the physical environment using the system. Alternatively, the system can have an opaque display and one or more imaging sensors that capture images or videos of the physical environment, which are representations of the physical environment. The system combines the images or videos with the virtual objects and presents the composition on the opaque display. The person indirectly views the physical environment through the images or videos of the physical environment using the system and perceives the virtual objects superimposed on and / or behind a portion of the physical environment. As used herein, a video of the physical environment displayed on an opaque display is referred to as “passthrough video,” meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting the AR environment on the opaque display. Further alternatively, the system can have a projection system that projects virtual objects into the physical environment, such as as a hologram or on a physical surface, such that the person perceives the virtual objects superimposed on the physical environment using the system.

[0027] An augmented reality environment is also a simulated environment in which a representation of the physical environment is transformed by computer-generated sensory information. For example, in providing see-through video, the system can transform one or more sensor images to impose an alternative perspective (e.g., a viewpoint) different from the perspective captured by the imaging sensor. As another example, a representation of the physical environment can be transformed by graphically modifying (e.g., magnifying) portions thereof such that the modified portions can be a representative but not a true version of the originally captured image. As yet another example, a representation of the physical environment can be transformed by graphically removing portions thereof or blurring portions thereof.

[0028] An augmented virtual (AV) environment is a simulated environment in which a virtual or computer-generated environment incorporates one or more sensory inputs from the physical environment. The sensory input can be a representation of one or more characteristics of the physical environment. For example, an AV park can have virtual trees and virtual buildings, but the human face is a realistic reproduction from an image of a physical person. As another example, a virtual object can assume the shape or color of a physical item imaged by one or more imaging sensors. As yet another example, a virtual object can assume a shadow that conforms to the positioning of the sun in the physical environment.

[0029] There are many different types of electronic systems that enable a person to sense and / or interact with various CGR environments. Examples include mobile devices, tablet devices, projection-based systems, head-up displays (HUDs), head-mounted systems, vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eye (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or hand-held controllers with or without haptic feedback), smart phones, tablet or slate devices, and desktop / laptop computers. For example, a head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, the head-mounted system can be configured to receive an external opaque display (e.g., a smart phone). The head-mounted system can incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system can have a transparent or translucent display instead of an opaque display. The transparent or translucent display can have a medium through which light representing an image is directed to a person's eye. The display can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium can be an optical waveguide, holographic medium, optical combiner, optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection techniques that project a graphical image onto a person's retina. The projection system can also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface.

[0030] Specific implementations of the subject technology described herein provide CGR systems that can use different input modalities to enable multimodality for recording content within a CGR environment. Examples of different input modalities include facial expressions, gestures, speech, and / or explicit hardware input, each of which can work alone and / or in combination with one or more of the other input modalities. Thus, the input modalities described herein can function in a complementary manner. Additionally, the subject technology enables selection of regions of interest in a CGR environment and provides annotation of objects and / or events detected within the CGR environment.

[0031] Figure 1Exemplary system architecture 100 is shown that includes various electronic devices that can implement the systems of the subject matter according to one or more specific implementations. However, not all of the depicted components may be used in all specific implementations, and one or more specific implementations may include additional or different components compared to those shown in the figures. Variations in the arrangement and type of these components can be made without departing from the spirit or scope of the claims listed herein. Additional components, different components, or fewer components may be provided.

[0032] System architecture 100 includes electronic device 105, handheld electronic device 104, electronic device 110, electronic device 115, and server 120. For purposes of explanation, system architecture 100 is shown in Figure 1 as including electronic device 105, handheld electronic device 104, electronic device 110, electronic device 115, and server 120; however, system architecture 100 may include any number of electronic devices and any number of servers or a data center including multiple servers.

[0033] Electronic device 105 can be implemented, for example, as a tablet device, a handheld and / or mobile device, or as a head-mounted portable system (e.g., worn by user 101). Electronic device 105 includes a display system capable of presenting a visualization of a computer-generated reality environment to the user. Electronic device 105 can be powered by a battery and / or another power source. In one example, the display system of electronic device 105 provides a stereoscopic presentation of a computer-generated reality environment to the user, enabling a three-dimensional visual display of a particular scene rendering. In one or more specific implementations, instead of using electronic device 105 to access a computer-generated reality environment or in addition thereto, the user can use handheld electronic device 104, such as a tablet computer, a watch, a mobile device, etc.

[0034] The electronic device 105 may include one or more cameras, such as camera 150 (e.g., visible light camera, infrared camera, etc.). Additionally, the electronic device 105 may include various sensors 152, including but not limited to cameras, image sensors, touch sensors, microphones, inertial measurement units (IMUs), heart rate sensors, temperature sensors, depth sensors (e.g., lidar sensors, radar sensors, sonar sensors, time-of-flight sensors, etc.), GPS sensors, Wi-Fi sensors, near field communication sensors, radio frequency sensors, etc. Additionally, the electronic device 105 may include hardware elements that can receive user input, such as hardware buttons or switches. User input detected by such sensors and / or hardware elements corresponds to various input modalities for initiating a co-occurrence session within an application, for example. Such input modalities may include but are not limited to face tracking, eye tracking (e.g., gaze direction), hand tracking, gesture tracking, biometric readings (e.g., heart rate, pulse, pupil dilation, respiration, temperature, electroencephalogram, olfaction), recognition of speech or audio (e.g., specific hot words), and activation of buttons or switches, etc.

[0035] In one or more specific embodiments, the electronic device 105 may be communicatively coupled to a base device, such as electronic device 110 and / or electronic device 115. Generally speaking, compared with the electronic device 105, such base devices may include more computing resources and / or available power. In one example, the electronic device 105 may operate in various modes. For example, the electronic device 105 may operate in an independent mode independent of any base device. When the electronic device 105 operates in the independent mode, the number of input modalities may be constrained by the power and / or processing limitations of the electronic device 105 (such as the available battery power of the device). In response to the power limitation, the electronic device 105 may deactivate certain sensors within the device itself to conserve battery power and / or relieve processing limitations.

[0036] The electronic device 105 may also operate in a wireless tethered mode (e.g., connected to a base device via a wireless connection) to work in conjunction with a given base device. The electronic device 105 may also operate in a connected mode where the electronic device 105 is physically connected to the base device (e.g., via a cable or some other physical connector), and may utilize the power resources provided by the base device (e.g., in the case where the base device charges the electronic device 105 when physically connected).

[0037] When the electronic device 105 operates in a wireless connection mode or a connected mode, at least a part of processing user input and / or rendering a computer-generated reality environment can be offloaded to the base device, thereby reducing the processing burden on the electronic device 105. For example, in a specific implementation, the electronic device 105 works in combination with the electronic device 110 or the electronic device 115 to generate a computer-generated reality environment, and the extended reality environment includes physical objects and / or virtual objects that enable different forms of interaction (e.g., visual, auditory, and / or physical or tactile interaction) between the user and the computer-generated reality environment in real time. In one example, the electronic device 105 provides a rendering of a scene corresponding to the computer-generated reality environment, and the scene can be perceived by the user and interacted with in real time, such as a host environment for a co-presence session with another user. Additionally, as part of presenting the rendered scene, the electronic device 105 can provide sound and / or tactile or haptic feedback to the user. The content of a given rendered scene may depend on available processing power, network availability and capacity, available battery power, and the current system workload.

[0038] The electronic device 105 can also detect an event that has occurred within the scene of the computer-generated reality environment. Examples of such events include detecting the presence of a specific person, entity, or object in the scene. In response to the detected event, the electronic device 105 can provide an annotation (e.g., in the form of metadata) corresponding to the detected event in the computer-generated reality environment.

[0039] The network 106 can communicatively (directly or indirectly) couple, for example, the electronic device 104, the electronic device 105, the electronic device 110, and / or the electronic device 115 to each other and / or to the server 120. In one or more specific implementations, the network 106 can be an interconnected network that may include the Internet or devices communicatively coupled to the Internet.

[0040] The electronic device 110 can include a touch screen and can be, for example, a smart phone including a touch screen, a portable computing device such as a laptop computer including a touch screen, a companion device including a touch screen (e.g., a digital camera, headphones), a tablet device including a touch screen, a wearable device including a touch screen (such as a watch, a wristband, etc.), any other suitable device including, for example, a touch screen, or any electronic device having a touchpad. In one or more specific implementations, the electronic device 110 may not include a touch screen but can support touch screen-like gestures, such as in a computer-generated reality environment. In one or more specific implementations, the electronic device 110 can include a touchpad. In Figure 1In this example, the electronic device 110 is depicted as a mobile smart phone device with a touch screen. In one or more specific embodiments, the electronic device 110, the handheld electronic device 104, and / or the electronic device 105 can be and / or can include all or part of the electronic devices discussed below with respect to the electronic systems discussed below with respect to Figure 9 In one or more specific embodiments, the electronic device 110 can be another device, such as an Internet Protocol (IP) camera, a tablet computer, or an accessory device such as an electronic stylus.

[0041] The electronic device 115 can be, for example, a desktop computer, a portable computing device such as a laptop computer, a smart phone, an accessory device (e.g., a digital camera, headphones), a tablet device, a wearable device such as a watch, a wristband, etc. In Figure 1 In this example, the electronic device 115 is depicted as a desktop computer. The electronic device 115 can be and / or can include all or part of the electronic systems discussed below with respect to Figure 9 discussed.

[0042] The server 120 can form all or part of a computer network or server group 130, such as in a cloud computing or data center implementation. For example, the server 120 stores data and software and includes specific hardware (e.g., processors, graphics processors, and other dedicated or custom processors) for rendering and generating content for a computer-generated reality environment such as graphics, images, videos, audio, and multimedia files. In one specific embodiment, the server 120 can be used as a cloud storage server that stores any of the aforementioned computer-generated reality content generated by the above devices and / or the server 120.

[0043] Figure 2 An exemplary software architecture 200 that can be implemented on the electronic device 105 according to one or more specific embodiments is shown. For illustrative purposes, the software architecture 200 is described as being implemented by Figure 1 the electronic device 105, such as by the processor and / or memory of the electronic device 105; however, the software architecture 200 can be implemented by any other electronic device, including the electronic device 115 and / or the electronic device 120. However, not all of the depicted components may be used in all specific embodiments, and one or more specific embodiments may include additional or different components compared to those shown in the figure. Variations in the arrangement and type of these components can be made without departing from the spirit or scope of the claims listed herein. Additional components, different components, or fewer components can be provided.

[0044] The software architecture 200 implemented on the electronic device 105 includes a framework. As used herein, a framework may refer to a software environment that provides specific functionality as part of a larger software platform to facilitate the development of software applications, and may provide one or more application programming interfaces (APIs) that developers can utilize to programmatically design a computer-generated reality environment and process operations for such a computer-generated reality environment.

[0045] As shown, a recording framework 230 is provided. The recording framework 230 may provide functionality for recording a computer-generated reality environment provided by an input modality as discussed above. An event detector 220 is provided that receives information corresponding to inputs from various input modalities. A system manager 210 is provided to monitor resources from the electronic device 105 and determine a quality of service metric based on the available resources. The system manager 210 may make a decision to select a particular hardware component corresponding to a respective input modality to activate and / or deactivate according to the quality of service metric, e.g., to release processing resources, conserve power resources, etc. For example, a camera used to track facial expressions may be turned off, or another camera used to track gestures may be turned off.

[0046] In one or more specific implementations, when a particular hardware is deactivated, the electronic device 105 may provide a notification to warn the user that a particular input modality is unavailable. Similarly, the electronic device 105 may provide a notification to warn the user that a particular input modality is available when the particular hardware is activated.

[0047] Figure 3A An example of facial expression tracking to initiate computer-generated reality recording in accordance with a specific implementation of the present subject matter is shown. The following discussion pertains to components of the electronic device 105 that include various cameras or image sensors to implement facial tracking of a user's face.

[0048] In a specific implementation, the electronic device 105 may utilize various sensors to track the facial expressions of a user 301 using the electronic device 105. As shown, different regions within the user's face may be tracked by sensors of the electronic device 105. For example, a camera may track the movement of the right eyebrow 310 and the left eyebrow 312 of the user 301. Another camera may track the movement of the region 302 including the right eye and the region 304 including the left eye. Different cameras may track the movement of a first region 308 (e.g., including the apex and nostrils of the nose) and / or a second region 306 (e.g., including the dorsum and / or bridge of the nose). Yet another camera may track the mouth 316 of the user 301. Additionally, a particular camera may track the mandible 314 including the chin of the user 301.

[0049] Although the various cameras are discussed above, it should be understood that the same camera can track more than one part of a user's face and still be within the scope of the subject technology. For example, the same camera can be used to track a user's jaw 314 and the user's mouth 316 of user 301.

[0050] The information from the various cameras can be analyzed independently or used in combination by the event detector 220. The event detector 220 can use this information to detect a facial expression of the user's face. In response to the detected facial expression, the event detector 220 can send a request to the recording framework 230 to initiate a recording within the computer-generated reality environment, e.g., a recording of a particular region of interest or field of view.

[0051] Different types of facial expressions can correspond to various emotions of the user 301. In an example, the electronic device 105 determines that the detected facial expression corresponds to a particular emotion (e.g., surprise, anger, happiness) and initiates a recording of the computer-generated reality environment in response to the emotion based on the detected facial expression.

[0052] Figure 3B and Figure 3C An example of tracking a gaze direction to initiate a computer-generated reality recording according to a particular implementation of the subject technology is shown. The following discussion relates to components of the electronic device 105 that includes various cameras to enable tracking of the gaze direction of the eyes of a user's face.

[0053] As Figure 3B shown, an image captured by at least one camera of the electronic device 105 is analyzed to determine the relative position of the user's eyes within the field of view. In a particular implementation, the electronic device 105 can distinguish the user's pupils and can use the relative position of the pupils relative to the eye position to determine the gaze direction. For example, in Figure 3C , the electronic device 105 can use the detected position of the user's pupils relative to the user's eyes and determine the region on the display of the electronic device 105 that the user is looking at within the field of view 320. Additionally, in a particular implementation, the electronic device 105 can also detect a movement such as the user closing his or her eyes for a particular period of time, which can be used to initiate a recording within the computer-generated reality environment.

[0054] The event detector 220 can analyze the above information to determine the gaze direction. The event detector 220 can use this information to determine the gaze direction of the user's eyes. In response to the determined gaze direction, the event detector 220 can send a request to the recording framework 230 to initiate recording within the computer-generated reality environment. For example, in response to determining that the user's gaze direction is in a particular direction or towards a particular object or person in the current scene of the computer-generated reality environment, the event detector 220 can send such a request to the recording framework 230 to initiate recording.

[0055] Figure 4A , Figure 4B and Figure 4C illustrate examples of determining a region of interest within a computer-generated reality environment and initiating recording based on the region of interest according to some specific implementations of the present subject matter technology. The following discussion pertains to components of the electronic device 105.

[0056] In a specific implementation, one or more input modalities can be used to identify the region of interest, which is the focus of recording within the computer-generated reality environment. For example, the user can perform a gesture or some other interaction (e.g., pressing a button or switch on the electronic device 105, providing a hotword or voice) to identify the region of interest. It should also be understood that the user can combinatorially utilize one or more input modes to identify the region of interest and / or initiate recording. Additionally, as described above, recording can be initiated when an event occurring within the scene (such as the presence of a person in the scene) is detected.

[0057] As Figure 4A shown, scene 410 depicts a sports event (e.g., ice hockey) taking place within a computer-generated reality environment. In this example, through the use of a specific input modality, the user has selected a region of interest 404 corresponding to the puck in the current scene of the computer-generated reality environment. In scene 410, the puck is moving towards the person 402 corresponding to the first hockey player. The event detector 220 has detected the presence of a particular person 406 (e.g., a star hockey player), and in response, initiates recording of the computer-generated reality environment by sending a request to the recording framework 230.

[0058] In Figure 4BIn [description], scene 420 shows that specific person 406 has moved to a position different from the initial position in scene 410. The recording framework 230 continues to record the computer-generated reality environment and focuses on the region of interest 404 because the puck moves closer to the hockey stick of person 402 in scene 420. Thus, the region of interest 404 may be moving or in motion, and the recording moves or tracks the region of interest. In a specific implementation, the recording framework 230 records the entire scene 420, despite the region of interest 404. During a future playback of the recording, the presentation of the recording may focus on the region of interest 404 corresponding to the puck.

[0059] In Figure 4C [description], scene 430 shows that specific person 406 has moved to a position different from the position in scene 420 and is completely within the view frame of scene 430. The recording framework 230 continues to record the computer-generated reality environment and focuses on the region of interest 404 because the puck moves across the ice rink in scene 430.

[0060] Figure 5A 、 Figure 5B and Figure 5C show examples of providing annotations to various objects or entities within a computer-generated reality environment according to some specific implementations of the present subject technology. The following discussion relates to components of the electronic device 105.

[0061] In Figure 5A [description], scene 502 is rendered to the user and includes various objects or entities within a computer-generated reality environment. In Figure 5BIn [the example], event detector 220 detects the presence of person 504, animal 508, and vehicle 506. In the example, detection of an object occurs when the recorded video stream is passed to event detector 220. Event detector 220 forwards information corresponding to detected person 504, animal 508, and vehicle 506 to recording framework 230. Based on the received information, recording framework 230 generates annotation 512 corresponding to person 504. Recording framework 230 also generates annotation 514 corresponding to vehicle 506. Additionally, recording framework 230 generates annotation 516 corresponding to animal 508. Alternatively, event detector 220 may generate the above annotations, and recording framework 230 may store the annotations as metadata having coordinates (and / or other information) corresponding to the annotation. The above annotations may be stored as metadata associated with the recording of the scene content, such as, in the example, adding metadata to be included as part of the recording of the content (e.g., a modified version of the content record now including the metadata). In the example, the object is identified and recognized as a person, animal, etc., and metadata is stored in association with the identified object. In a particular implementation, such metadata corresponding to the annotation may be stored in the memory of electronic device 105, and / or included in a computer-generated reality record, and / or stored separately in a different electronic device (e.g., a server or infrastructure device). It should also be understood that different sets of annotations may be applied to a given computer-generated reality record, enabling various uses of different annotations in combination with the playback of the record.

[0062] In Figure 5C [the example], electronic device 105 renders an update to scene 502, which now displays annotation 512 corresponding to person 504, annotation 514 corresponding to vehicle 506, and annotation 516 corresponding to animal 508. As shown, the annotations are rendered as part of scene 502, which may include elements corresponding to the physical environment that are mixed with digitally generated content (e.g., the annotations). In some particular implementations, such annotations may be provided in a different format or not displayed in the scene. For example, the annotations may be provided to the user in audio form, serving as a narration of the computer-generated reality environment the user is currently experiencing.

[0063] Figure 6 A flowchart of an example process 600 for initiating recording of content within the field of view of electronic device 105 in accordance with one or more particular implementations is shown. For purposes of explanation, process 600 is described herein primarily with reference to Figure 1 and Figure 2 electronic device 105. However, process 600 is not limited to Figure 1 and Figure 2in the electronic device 105, and one or more boxes (or operations) of process 600 may be performed by one or more other components of other suitable devices. Further for purposes of explanation, the boxes of process 600 are described herein as occurring sequentially or linearly. However, multiple boxes of process 600 may occur in parallel. Additionally, the boxes of process 600 need not be performed in the order shown, and / or one or more boxes of process 600 need not be performed and / or may be replaced by other operations.

[0064] As Figure 6 shown, the electronic device 105 determines an operation mode (610) at least in part based on whether the electronic device 105 is communicatively coupled to an associated base device. Based on the determined operation mode, the electronic device 105 identifies a set of input modalities (612) for initiating recording of content within the field of view of the electronic device 105. The electronic device 105 monitors sensor information (614) generated by at least one sensor included in or communicatively coupled to the electronic device 105. When the monitored sensor information indicates that at least one input modality of the identified set of input modalities has been triggered, the electronic device 105 initiates recording of content within the field of view of the electronic device 105 (616).

[0065] Figure 7 FIG. shows a flowchart of an exemplary process 700 for updating a set of input modalities for initiating recording of content on an electronic device 105 in accordance with one or more specific embodiments. For purposes of explanation, process 700 is described herein primarily with reference to Figure 1 and Figure 2 the electronic device 105 of. However, process 700 is not limited to Figure 1 and Figure 2 the electronic device 105 in, and one or more boxes (or operations) of process 700 may be performed by one or more other components of other suitable devices. Further for purposes of explanation, the boxes of process 700 are described herein as occurring sequentially or linearly. However, multiple boxes of process 700 may occur in parallel. Additionally, the boxes of process 700 need not be performed in the order shown, and / or one or more boxes of process 700 need not be performed and / or may be replaced by other operations.

[0066] As Figure 7 shown, the electronic device 105 detects that the operation mode of the electronic device 105 has changed (710). In response to detecting the change, the electronic device 105 updates a set of input modalities for initiating recording of content based on the changed operation mode (712).

[0067] Figure 8Flowchart 800 shows an exemplary process for determining a quality of service metric associated with an operating mode of an electronic device 105 in accordance with one or more particular implementations.

[0068] As Figure 8 shown, the electronic device 105 determines a quality of service metric associated with an operating mode of the electronic device 105 (810). At least in part based on the quality of service metric, the electronic device 105 selects at least one input modality (812). The electronic device 105 provides the at least one input modality as the set of input modalities for initiating a recording of content within the field of view of the electronic device 105 (814).

[0069] As described above, one aspect of the present technology is the collection and use of data obtained from various sources. The present disclosure anticipates that, in some instances, the collected data may include personal information data that uniquely identifies or can be used to contact or locate a specific person. Such personal information data can include demographic data, location-based data, telephone numbers, email addresses, social network identifiers, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.

[0070] The present disclosure recognizes that the use of such personal information data in the technologies of the present invention can be used to benefit users. The present disclosure also anticipates uses of personal information data that are beneficial to users. For example, health and fitness data can be used to provide insights into a user's overall health condition or can be used as positive feedback for an individual using the technology to pursue health goals.

[0071] The present disclosure contemplates that entities responsible for collecting, analyzing, disclosing, transmitting, storing, or otherwise using such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. Such policies should be readily accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of those legitimate uses. In addition, such collection / sharing should be done after receiving informed consent from the user. Further, such entities should consider taking any necessary steps to safeguard and secure access to such personal information data and to ensure that others with access to the personal information data comply with their privacy policies and procedures. Additionally, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices. Further, policies and practices should be adjusted to account for the specific types of personal information data being collected and / or accessed and to apply applicable laws and standards including specific considerations of the jurisdiction. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws such as the Health Insurance Portability and Accountability Act (HIPAA); while health data in other countries may be subject to other regulations and policies and should be handled accordingly. Thus, different privacy practices should be maintained for different types of personal data in each country.

[0072] Notwithstanding the foregoing, the present disclosure also contemplates embodiments where users selectively block the use or access of personal information data. That is, the present disclosure contemplates that hardware elements and / or software elements may be provided to prevent or block access to such personal information data. For example, the technology may be configured to allow a user to select to participate in a "opt-in" or "opt-out" of the collection of personal information data either during registration for a service or at any time thereafter. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notices related to the access or use of personal information. For example, a user may be notified at the time of downloading an application that their personal information data will be accessed and then reminded again just prior to the personal information data being accessed by the application.

[0073] In addition, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of inadvertent or unauthorized access or use. Once data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. Further, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated, when appropriate, by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of the data stored (e.g., collecting location data at the city level rather than at the address level), controlling how the data is stored (e.g., aggregating data across users), and / or other methods.

[0074] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that the various embodiments may also be implemented without access to such personal information data. That is, the various embodiments of the present inventive technology will not fail to operate properly due to the lack of all or a portion of such personal information data. For example, content may be selected and delivered to a user based on non-personal information data or a small amount of personal information, such as content requested by a device associated with the user, other non-personal information, or publicly available information.

[0075] Figure 9 An electronic system 900 is shown that can be utilized to implement one or more specific implementations of the present subject matter technology. The electronic system 900 can be Figure 1 the electronic device 105, the electronic device 104, the electronic device 110, the electronic device 115, and / or the server 120 as shown and / or can be a part thereof. The electronic system 900 can include various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 900 includes a bus 908, one or more processing units 912, a system memory 904 (and / or cache), a ROM 910, a permanent storage device 902, an input device interface 914, an output device interface 906, and one or more network interfaces 916, or subsets and variant forms thereof.

[0076] The bus 908 generally represents the entire system bus, peripheral device bus, and chipset bus that communicatively connect many internal devices of the electronic system 900. In one or more specific implementations, the bus 908 communicatively connects one or more processing units 912 with the ROM 910, the system memory 904, and the permanent storage device 902. The one or more processing units 912 retrieve instructions to be executed and data to be processed from these various memory units in order to perform the processes disclosed in the present subject matter. In different specific implementations, the one or more processing units 912 can be a single processor or a multi-core processor.

[0077] The ROM 910 stores static data and instructions required by the one or more processing units 912 and other modules of the electronic system 900. On the other hand, the permanent storage device 902 can be a read-write memory device. The permanent storage device 902 can be a non-volatile memory unit that stores instructions and data even when the electronic system 900 is turned off. In one or more specific implementations, a mass storage device (such as, a magnetic disk or an optical disk and its corresponding disk drive) can be used as the permanent storage device 902.

[0078] In one or more embodiments, a removable storage device (such as a floppy disk, flash drive, and their corresponding disk drives) may be used as the permanent storage device 902. Like the permanent storage device 902, the system memory 904 may be a read-write memory device. However, unlike the permanent storage device 902, the system memory 904 may be a volatile read-write memory, such as random access memory. The system memory 904 may store any of the instructions and data that one or more processing units 912 may need during operation. In one or more embodiments, the processes disclosed by this subject matter are stored in the system memory 904, the permanent storage device 902, and / or the ROM 910. One or more processing units 912 retrieve the instructions to be executed and the data to be processed from these various memory units in order to execute the processes of one or more embodiments.

[0079] The bus 908 is also connected to an input device interface 914 and an output device interface 906. The input device interface 914 enables a user to transmit information to and select commands for the electronic system 900. Input devices that may be used with the input device interface 914 may include, for example, an alphanumeric keyboard and a pointing device (also referred to as a "cursor control device"). The output device interface 906 may enable, for example, the display of images generated by the electronic system 900. Output devices that may be used with the output device interface 906 may include, for example, a printer and a display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a flexible display, a flat panel display, a solid state display, a projector, or any other device for outputting information. One or more embodiments may include a device that serves as both an input device and an output device, such as a touch screen. In these embodiments, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including acoustic, voice, or tactile input.

[0080] Finally, as Figure 9 shown, the bus 908 also couples the electronic system 900 to one or more networks and / or to one or more network nodes through one or more network interfaces 916, such as Figure 1 the electronic device 110 shown in. In this way, the electronic system 900 may be part of a computer network (such as a LAN, a wide area network ("WAN"), or an intranet), or may be part of a network of networks (such as the Internet). Any or all components of the electronic system 900 may be used with the subject matter disclosed herein.

[0081] The above functions can be implemented in computer software, firmware or hardware. This technology can be implemented using one or more computer program products. The programmable processor and computer can be included in or packaged as a mobile device. The process and logical flow can be executed by one or more programmable processors and one or more programmable logic circuits. General and special computing devices and storage devices can be interconnected through a communication network.

[0082] Some specific implementations include electronic components that store computer program instructions in a machine-readable or computer-readable medium (also referred to as a computer-readable storage medium, machine-readable medium, or machine-readable storage medium), such as microprocessors, storage devices, and memories. Some examples of such computer-readable media include RAM, ROM, compact disc read-only memory (CD-ROM), recordable compact disc (CD-R), rewritable compact disc (CD-RW), digital versatile disc read-only (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid state disk drives, read-only and recordable discs, ultra density optical discs, any other optical or magnetic medium, and floppy disks. The computer-readable medium can store a computer program that can be executed by at least one processing unit and includes an instruction set for performing various operations. Examples of computer programs or computer code include machine code, such as that produced by a compiler, and files that include higher-level code that can be executed by a computer, electronic component, or microprocessor using an interpreter.

[0083] Although the above discussion mainly relates to microprocessors or multi-core processors that execute software, some specific implementations are performed by one or more integrated circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some specific implementations, such integrated circuits execute instructions stored on the circuit itself.

[0084] As used in this specification and any claims of this patent application, the terms "computer", "server", "processor", and "memory" all refer to electronic or other technical devices. These terms exclude humans or groups of humans. For the purposes of this specification, the term display or displaying means display on an electronic device. As used in this specification and any claims of this patent application, the terms "computer-readable medium" and "computer-readable media" are completely limited to tangible physical objects that store information in a form readable by a computer. These terms do not include any wireless signals, wired download signals, and any other transient signals.

[0085] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user and a keyboard and a pointing device by which the user can provide input to the computer. The display device can be, for example, a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and the pointing device can be, for example, a mouse or a trackball. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input received from the user can be in any form, including acoustic, speech, or tactile input. In addition, the computer can interact with the user by sending documents to and receiving documents from the devices used by the user; for example, by sending a web page to a web browser in response to a request received from the web browser on a user client device.

[0086] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, such as a data server, or includes a middleware component, such as an application server, or includes a front-end component, such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), the Internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network).

[0087] The computing system can include clients and servers. The clients and servers are generally remote from each other and can interact through a communication network. The relationship between the client and the server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, the server transmits data (e.g., an HTML page) to the client device (e.g., to display data to a user interacting with the client device and to receive user input from the user interacting with the client device). Data generated at the client device (e.g., the result of a user interaction) can be received at the server.

[0088] Those skilled in the art will recognize that the various illustrative blocks, modules, elements, components, methods, and algorithms described herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, the various illustrative blocks, modules, elements, components, methods, and algorithms have been described generally in terms of functionality above. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality may be implemented differently for each particular application. The various components and blocks may be arranged differently (e.g., arranged in a different order, or partitioned in a different manner) without departing from the scope of the claimed subject matter.

[0089] It should be understood that the specific order or hierarchical structure of steps in the processes disclosed herein is an illustration of exemplary methods. Based on design preference, it should be understood that the specific order or hierarchical structure of steps in a process may be rearranged. Some of the steps in the process may be performed simultaneously. The appended method claims present elements of the various steps in a sample order and are not meant to be limited to the specific order or hierarchical structure presented.

[0090] The foregoing description has been provided to enable a person skilled in the art to practice the various aspects described herein. The foregoing description provides various examples of the claimed subject matter, and the claimed subject matter is not limited to these examples. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein but are to be accorded the full scope consistent with the language of the claims, where the singular forms of elements are not intended to mean "only one" but rather "one or more" unless specifically stated otherwise. Unless specifically stated otherwise, the term "some" means one or more. Pronouns in the masculine gender (e.g., his) include the feminine and neuter genders (e.g., her and its), and vice versa. Headings and subheadings (if any) are used only for convenience and do not limit the claimed invention described herein.

[0091] As used herein, the term website may include any aspect of a website, including one or more web pages, one or more servers for hosting or storing web-related content, etc. Thus, the term website may be used interchangeably with the terms web page and server. The predicate words "configured to", "operable to", and "programmed to" do not imply any particular tangible or intangible modification to a subject matter but are intended to be used interchangeably. For example, a component or a processor configured to monitor and control operations may also mean that the processor is programmed to monitor and control operations or that the processor is operable to monitor and control operations. Similarly, a processor configured to execute code may be interpreted as a processor programmed to execute code or operable to execute code.

[0092] As used herein, the term "automatically" may include performance by a computer or machine without user intervention; e.g., by instructions responsive to a predicate action of a computer or machine or other initiating mechanism. The word "exemplary" is used herein to mean "serving as an example or illustration". Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs.

[0093] Phrases such as "aspect" do not mean that such aspect is essential to the subject technology or that such aspect applies to all configurations of the subject technology. The disclosure related to an aspect may apply to all configurations, or one or more configurations. An aspect may provide one or more examples. Phrases such as "aspect" may refer to one or more aspects, and vice versa. Phrases such as "embodiment" do not mean that such embodiment is essential to the subject technology or that such embodiment applies to all configurations of the subject technology. The disclosure related to an embodiment may apply to all embodiments, or one or more embodiments. An embodiment may provide one or more examples. Phrases such as "embodiment" may refer to one or more embodiments, and vice versa. Phrases such as "configuration" do not mean that such configuration is essential to the subject technology or that such configuration applies to all configurations of the subject technology. The disclosure related to a configuration may apply to all configurations or one or more configurations. A configuration may provide one or more examples. Phrases such as "configuration" may refer to one or more configurations, and vice versa.

Claims

1. A method, comprising: Determining a quality of service metric associated with an operation of an electronic device; Based on the determined quality of service metric, selecting a first set of user input modalities or a second set of user input modalities for initiating recording of content within a field of view of the electronic device, the second set of user input modalities being different from the first set of user input modalities; Monitoring sensor information generated by at least one sensor included in the electronic device or communicatively coupled to the electronic device; And When the selected first set of user input modalities is selected, when the monitored sensor information indicates that a first user input corresponding to at least one of the selected first set of user input modalities has been received from a user, or when the second set of user input modalities is selected, when the monitored sensor information indicates that a second user input corresponding to at least one of the selected second set of user input modalities has been received from a user, initiating recording of the content within the field of view of the electronic device.

2. The method according to claim 1, further comprising: Detecting that the quality of service metric has changed; And In response to detecting the change, updating a set of user input modalities for initiating recording of content based on the changed quality of service metric.

3. The method according to claim 1, wherein the quality of service metric is determined at least in part based on available computing resources or available power in the electronic device, the available power including an amount of battery power.

4. The method according to claim 1, further comprising: Deactivating a specific input modality corresponding to at least one sensor in the electronic device, wherein the specific input modality includes facial expression, gaze direction, eye position, gesture, hardware input, voice, or recognition of an object or a person in a scene.

5. The method according to claim 1, further comprising: When the quality of service metric is below a threshold, deactivating a specific sensor in the electronic device, wherein the specific sensor includes a camera, an inertial measurement unit, a microphone, or a touch sensor.

6. The method according to claim 1, further comprising: Determining a region of interest in the field of view of the electronic device.

7. The method according to claim 6, wherein the region of interest is determined based on a gesture or an indicator corresponding to the region of interest.

8. The method according to claim 1, further comprising: Generating an annotation corresponding to the recording of the content; And Adding the annotation as metadata to the recording of the content.

9. A system, comprising: A processor; A memory device containing instructions that, when executed by the processor, cause the processor to perform operations including the following: Determining a quality of service metric associated with an operation of an electronic device; Based on the determined quality of service metric, selecting a first set of input modalities or a second set of input modalities for initiating recording of content within a field of view of the electronic device, the second set of input modalities being different from the first set of input modalities; Monitoring sensor information generated by at least one sensor included in the electronic device or communicatively coupled to the electronic device; And When the selected first set of input modalities is selected, when the monitored sensor information indicates that a first user input corresponding to at least one of the selected first set of input modalities has been received from the user, or when the second set of input modalities is selected, when the monitored sensor information indicates that a second user input corresponding to at least one of the selected second set of input modalities has been received from the user, initiate recording of the content within the field of view of the electronic device.

10. The system according to claim 9, wherein the memory device includes additional instructions that, when executed by the processor, further cause the processor to perform additional operations, the additional operations further including: Detecting that the quality of service metric has changed; And In response to detecting the change, updating a set of input modalities for initiating recording of content based on the changed quality of service metric.

11. The system according to claim 9, wherein the quality of service metric is determined at least in part based on available computing resources or available power in the electronic device, the available power including the amount of battery power.

12. The system according to claim 9, wherein the memory device includes additional instructions that, when executed by the processor, further cause the processor to perform additional operations, the additional operations further including: Deactivating a specific input modality corresponding to at least one sensor in the electronic device, where the specific input modality includes facial expression, gaze direction, eye position, gesture, hardware input, voice, or recognition of an object or person in the scene.

13. The system according to claim 9, wherein the memory device includes additional instructions that, when executed by the processor, further cause the processor to perform additional operations, the additional operations further including: When the quality of service metric is below a threshold, deactivating a specific sensor in the electronic device, where the specific sensor includes a camera, an inertial measurement unit, a microphone, or a touch sensor.

14. The system according to claim 9, wherein the memory device includes additional instructions that, when executed by the processor, further cause the processor to perform additional operations, the additional operations further including: Determining a region of interest in the field of view of the electronic device, where the region of interest is determined based on a gesture or indicator corresponding to the region of interest.

15. The system according to claim 9, wherein the memory device includes additional instructions that, when executed by the processor, further cause the processor to perform additional operations, the additional operations further including: Generating an annotation corresponding to the recording of the content; And Adding the annotation as metadata to the recording of the content.

16. A non-transitory machine-readable medium including instructions that, when executed by a computing device, cause the computing device to perform operations including the following: Determining a quality of service metric associated with the operation of an electronic device; Based on the determined quality of service metric, select a first set of input modalities or a second set of input modalities for initiating recording of content within the field of view of the electronic device, the second set of input modalities being different from the first set of input modalities; Monitor sensor information generated by at least one sensor included in or communicatively coupled to the electronic device; And When the selected first set of input modalities is selected, when the monitored sensor information indicates that a first user input corresponding to at least one of the selected first set of input modalities has been received from the user, or when the second set of input modalities is selected, when the monitored sensor information indicates that a second user input corresponding to at least one of the selected second set of input modalities has been received from the user, initiate recording of the content within the field of view of the electronic device.

17. The non-transitory machine-readable medium according to claim 16, wherein the operation further comprises: Detecting that the quality of service metric has changed; And In response to detecting the change, updating a set of input modalities for initiating recording of content based on the changed quality of service metric.

18. The non-transitory machine-readable medium according to claim 16, wherein the quality of service metric is determined at least in part based on available computing resources or available power in the electronic device, the available power including an amount of battery power.

19. The non-transitory machine-readable medium according to claim 16, further comprising: When the quality of service metric is below a threshold, deactivate a specific sensor in the electronic device, wherein the specific sensor includes a camera, an inertial measurement unit, a microphone, or a touch sensor.

20. The non-transitory machine-readable medium according to claim 16, further comprising: Determining a region of interest in the field of view of the electronic device.

Citation Information

Patent Citations

  • Method and apparatus for determining interaction mode

    CN102823145A

  • Display control method and apparatus

    CN105229570A

  • Method and apparatus for reducing power consumption in a mobile electronic device

    CN105684525A

  • Emotional / cognitive state presentation

    CN109154861A