Providing suggestions based on saved content

JP2026148569APending Publication Date: 2026-09-17APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026036844
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2026-02-10
Filing Date
2026-03-09
Publication Date
2026-09-17

Smart Images

  • Figure 2026148569000001_ABST
    Figure 2026148569000001_ABST
Patent Text Reader

Abstract

Provide suggestions based on saved content. [Solution] A technique is disclosed for determining and presenting action suggestions based on content such as a live view of a 3D scene and a saved view of a 3D scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to presenting action suggestions determined based on views of a three-dimensional scene.

Background Art

[0002] Development of computer systems for interacting with and / or providing three-dimensional scenes has expanded significantly in recent years. Examples of three-dimensional scenes (e.g., environments) include physical scenes and augmented reality scenes.

Summary of Invention

[0003] Example methods are disclosed herein. An example of a method comprises: receiving, at a first computer system in communication with one or more image sensors, a user input corresponding to a request to save an object in a three-dimensional (3D) scene; in response to receiving the user input corresponding to the request to save the object in the 3D scene, capturing a view of the 3D scene via the one or more image sensors; and after the view of the 3D scene is captured via the one or more image sensors, displaying, at a second computer system in communication with a display generation component, a first user interface of a first application via the display generation component, wherein displaying the first user interface of the first application comprises displaying a first suggested graphical element, the first suggested graphical element is selectable to cause the second computer system to execute a first action suggestion via a second application different from the first application, and the first action suggestion is determined based on performing image recognition on the view of the 3D scene and based on context information different from the view of the 3D scene.

[0004] Examples of non-temporary computer-readable storage media are disclosed herein. One or more examples of non-temporary computer-readable storage media store one or more programs configured to be executed by one or more processors of one or more computer systems that are in communication with one or more image sensors and display generation components. One or more programs receive user input from a first computer system among one or more computer systems in response to a request to save objects in a three-dimensional (3D) scene, and in response to receiving user input in response to a request to save objects in a 3D scene, the first computer system captures a view of the 3D scene via one or more image sensors, and after the view of the 3D scene has been captured via one or more image sensors, the second computer system among one or more computer systems includes instructions to display a first user interface of a first application via a display generation component, the display of the first user interface of the first application includes displaying a first proposed graphical element, the first proposed graphical element is selectable to cause the second computer system to execute a first action proposal via a second application different from the first application, the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.

[0005] Examples of computer systems are described herein. One or more examples of computer systems are configured to communicate with one or more image sensors and display generation components. One or more computer systems include one or more processors and memory for storing one or more programs configured to run by one or more processors, the one or more programs include instructions for a first computer system of the one or more computer systems to receive user input corresponding to a request to save objects in a three-dimensional (3D) scene, and for the first computer system to capture a view of the 3D scene via one or more image sensors in response to receiving user input corresponding to a request to save objects in a 3D scene, after the view of the 3D scene has been captured via one or more image sensors, for a second computer system of the one or more computer systems to display a first user interface of a first application via a display generation component, the display of the first user interface of the first application includes displaying a first proposed graphical element, the first proposed graphical element is selectable to cause the second computer system to execute a first action proposal via a second application different from the first application, the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.

[0006] An example of one or more computer systems is configured to communicate with one or more image sensors and display generation components. The example of one or more computer systems includes means for a first computer system among the one or more computer systems to receive user input corresponding to a request to save objects in a three-dimensional (3D) scene; means for the first computer system to capture a view of the 3D scene via one or more image sensors in response to receiving user input corresponding to a request to save objects in a 3D scene; and means for a second computer system among the one or more computer systems to display a first user interface of a first application via a display generation component after the view of the 3D scene has been captured via one or more image sensors, wherein displaying the first user interface of the first application includes displaying a first proposed graphical element, the first proposed graphical element is selectable to cause the second computer system to execute a first action proposal via a second application different from the first application, the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.

[0007] By determining and presenting suggestions according to the techniques described herein, a computer system can suggest reasonable actions to the user and efficiently execute those actions. Action suggestions may also be presented for content previously saved by the user, thereby allowing the user to save content of interest and later view reasonable suggestions for that saved content of interest. In this way, the user-device interface becomes more efficient and accurate (for example, by suggesting accurate and reasonable actions, by reducing the number of user inputs required to initiate an action, by reducing the number of user inputs required to cancel and / or undo the results of an incorrect action, and / or by helping the user remember and / or execute a desired action), which further reduces power consumption and improves the battery life of the device by allowing the user to use the device more quickly and efficiently.

[0008] In some examples, the computer system (e.g., the first computer system and / or the second computer system) is a desktop computer with an associated display. In some examples, the computer system is a portable device (e.g., a handheld device such as a notebook computer, tablet computer, or smartphone). In some examples, the computer system is a personal electronic device (e.g., a wearable electronic device such as a wristwatch or head-mounted device). In some examples, the computer system has a touchpad. In some examples, the computer system has one or more cameras. In some examples, the computer system has a display generating component (e.g., a display device such as a head-mounted display, display, projector, touch-sensitive display (also known as a “touchscreen” or “touchscreen display”), or other devices or components that present visual content to the user on or within the display generating component itself, or that are generated from the display generating component and are visible elsewhere). In some examples, the computer system does not have a display generating component and does not present visual content to the user. In some examples, the computer system has a touch-sensitive display (also known as a “touchscreen” or “touchscreen display”). In some examples, the computer system has one or more eye-tracking components. In some examples, the computer system has one or more hand-tracking components. In some examples, the computer system has one or more output devices, the output devices include one or more tactile output generators and / or one or more audio output devices. In some examples, the computer system has one or more processors, memory, and one or more modules, programs, or instruction sets stored in memory for performing various functions described herein.In some examples, the user interacts with the computer system through stylus and / or finger touch and gestures on a touch-sensitive surface, the movement of the user's eyes and hands or body in space captured by cameras and other motion sensors, and / or voice input captured by one or more audio input devices. The executable instructions that perform these functions are optionally contained in temporary computer-readable storage media and / or non-temporary computer-readable storage media, or other computer program products configured to be executed by one or more processors.

[0009] It should be noted that the various examples described herein can be combined with any other examples described herein. The features and advantages described herein are not exhaustive, and many additional features and advantages will become apparent to those skilled in the art, particularly in light of the drawings, specification and claims. Furthermore, it should be noted that the language used herein has been selected solely for readability and explanatory purposes and not to define or limit the subject matter of the invention.

[0010] For a better understanding of the various embodiments described, please refer to the following “Modes for Carrying Out the Invention” in conjunction with the following drawings, where similar reference numbers refer to corresponding parts throughout those drawings. [Brief explanation of the drawing]

[0011] [Figure 1] This block diagram shows the operating environment of a computer system for interacting with a three-dimensional (3D) scene, according to several embodiments.

[0012] [Figure 2] This is a block diagram of user-responsive components of a computer system, with several examples.

[0013] [Figure 3A] Here are some examples of block diagrams of computer system controllers.

[0014] [Figure 3B] FIG. 1 is a diagram illustrating a block diagram of components of a controller according to various examples.

[0015] [Figure 3C] FIG. 1 is a diagram illustrating a block diagram of components of a controller according to various examples.

[0016] [Figure 4] FIG. 1 is a diagram illustrating an architecture of a foundation model according to some examples.

[0017] [Figure 5A] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5B] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5C] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5D] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5E] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5F] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5G] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5H] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5I] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5J] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5K] FIG. 1 is a diagram illustrating a technique for providing a proposal according to various examples. [Figure 5L]FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5M] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5N] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5O] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5P] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5Q] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5R] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples. [Figure 5S] FIG. 1 is a diagram illustrating techniques for providing proposals according to various examples.

[0018] [Figure 6] FIG. 1 is a flow diagram illustrating a method for providing action proposals according to various examples. DETAILED DESCRIPTION OF EMBODIMENTS OF THE INVENTION

[0019] FIGS. 1 to 4 describe examples of computer systems and techniques for interacting with a three-dimensional scene. FIGS. 5A to 5S illustrate techniques for providing proposals. FIG. 6 is a flow diagram of a method for providing action proposals. FIGS. 5A to 5S are used to describe the method of FIG. 6.

[0020] Furthermore, in any method described herein that is conditional on one or more conditions being met by one or more steps, it should be understood that the method described can be repeated in multiple iterations such that all the conditions that the steps of the method are conditional on are met in different iterations of the method. For example, if a method requires that a first step be performed if a condition is met, and a second step be performed if the condition is not met, a person skilled in the art will understand that the steps described in the claim are repeated in an unspecified order until the conditions are met and then not met. Thus, a method described in one or more steps that depends on one or more conditions being met can be rewritten as a method that is repeated until each of the conditions described in the method is met. However, this is not required for claims of a system or computer-readable medium that include instructions for performing a contingency operation based on the satisfaction of the corresponding one or more conditions, and therefore it is possible to determine whether the contingency is met without explicitly repeating the steps of the method until all the conditions that the steps of the method are conditional on are met. Those skilled in the art will understand that, as with methods involving incidental steps, a system or computer-readable storage medium may repeat the steps of the method as many times as necessary to ensure that all incidental steps are performed.

[0021] Figure 1 is a block diagram illustrating the operating environment of a computer system 101 for interacting with a 3D scene, in several examples. In Figure 1, the user interacts with a 3D scene 105 through an operating environment 100 which includes the computer system 101. In some examples, the computer system 101 includes a controller 110 (e.g., a processor in a portable electronic device or remote server), a user-responsive component 120, one or more input devices 125 (e.g., an eye-tracking device 130, a hand-tracking device 140, and / or other input devices 150), one or more output devices 155 (e.g., a speaker 160, a tactile output generator 170, and other output devices 180), one or more sensors 190 (e.g., an image sensor, a light sensor, a depth sensor, a tactile sensor, an orientation sensor, a proximity sensor, a temperature sensor, a location sensor, a motion sensor, a speed sensor, an audio sensor, etc.), and one or more peripheral devices 195 (e.g., a home appliance, a wearable device, etc.). In some examples, one or more of the input device 125, output device 155, sensor 190, and peripheral device 195 are integrated with the user-responsive component 120 (for example, within a head-mounted device or handheld device).

[0022] While Figure 1 shows features suitable for operating environment 100, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more suitable embodiments of the examples disclosed herein.

[0023] Hardware: There are many different types of electronic systems that enable a person to perceive and / or interact with various three-dimensional scenes. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (e.g., contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may include speakers and / or other audio output devices integrated into the head-mounted system to provide audio output. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). Alternatively, a head-mounted system may be configured to operate without displaying content, for example, by providing output to the user via haptic and / or auditory means. The head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment, and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to the person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, holographic medium, optical coupler, optical reflector, or any combination thereof. In one example, the transparent or translucent display may be configured to be selectively opaque.A projection-based system may employ retinal projection technology to project graphical images onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as holograms or onto physical surfaces.

[0024] In some examples, the user-responsive component 120 is configured to provide visual components of a three-dimensional scene. In some examples, the user-responsive component 120 includes a preferred combination of software, firmware, and / or hardware. The user-responsive component 120 is described in more detail below with reference to Figure 2. In some examples, the functionality of the controller 110 is provided by and / or combined with the user-responsive component 120. In some examples, the user-responsive component 120 provides the user with an augmented reality (XR) experience while the user is virtually and / or physically present in the scene 105.

[0025] In some examples, the user-responsive component 120 is worn on a part of the user's body (e.g., the user's head or the user's hand). In some examples, the user-responsive component 120 includes one or more XR displays provided for displaying XR content. In some examples, the user-responsive component 120 surrounds the user's field of view. In some examples, the user-responsive component 120 is a handheld device (such as a smartphone or tablet) configured to present XR content, and the user holds the device, which has a display directed towards the user's field of view and a camera directed towards scene 105. In some examples, the handheld device is optionally placed in an enclosure worn on the user's head. In some examples, the handheld device is optionally placed on a support in front of the user (e.g., a tripod). In some examples, the user-responsive component 120 is an XR chamber, enclosure, or room configured to present XR content in which the user is not wearing or holding the user-responsive component 120. Many user interfaces described with reference to one type of hardware for displaying XR content (e.g., a handheld device or a device on a tripod) may also be implemented on another type of hardware for displaying XR content (e.g., a head-mounted device (HMD) or other wearable computing devices). For example, a user interface showing interaction with XR content triggered based on interaction occurring in the space in front of a handheld or tripod-mounted device may also be implemented similarly to an HMD, where the interaction occurs in the space in front of the HMD and the XR content response is displayed through the HMD. Similarly, a user interface showing interaction with XR content triggered based on the movement of a handheld or tripod-mounted device relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)) may also be implemented similarly to an HMD, where the movement is triggered by the movement of the HMD relative to a physical environment (e.g., Scene 105 or a part of the user's body (e.g., the user's eyes, head, or hands)).

[0026] Figure 2 is a block diagram of user-responsive components 120 in several examples. While certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more appropriate embodiments of the examples disclosed herein. Furthermore, Figure 2 is more intended to illustrate the function of various features that may exist in a particular implementation, in contrast to the structural outlines of the examples described herein. As will be recognized by those skilled in the art, the components shown separately can be combined, and some components can be separated. For example, some functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various examples. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some examples, will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0027] In some examples, the user-responsive component 120 (e.g., HMD) includes one or more processing units 202 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 206, one or more communication interfaces 208 (e.g., USB, FIREWIRE®, THUNDERBOLT®, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, BLUETOOTH®, ZIGBEE®, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, one or more XR displays 212, one or more optional in-facing and / or out-facing image sensors 214, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0028] In some examples, one or more communication buses 204 include circuits that interconnect system components and control communication between system components. In some examples, one or more I / O devices and sensors 206 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a thermometer, one or more biosensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.).

[0029] In some examples, one or more XR displays 212 are configured to provide an XR experience to the user. In some examples, one or more XR displays 212 correspond to holographic, digital light processing (DLP), liquid crystal displays (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistors (OLET), organic light-emitting diodes (OLED), surface conduction electron emission displays (SED), field emission displays (FED), quantum dot light-emitting diodes (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some examples, one or more XR displays 212 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic. For example, a user-responsive component 120 (e.g., HMD) includes a single XR display. In another example, the user-responsive component 120 includes an XR display for each of the user's eyes. In some examples, one or more XR displays 212 can present XR content. In some examples, one or more XR displays 212 are omitted from the user-responsive component 120. For example, user-responsive component 120 does not include any components configured to display content (or any components configured to display XR content), and user-responsive component 120 provides output via audio and / or haptic output types.

[0030] In some examples, one or more image sensors 214 are configured to acquire image data corresponding to at least a portion of the user's face, including the user's eyes (and may also be referred to as eye-tracking cameras). In some examples, one or more image sensors 214 are configured to acquire image data corresponding to at least a portion of the user's hands and optionally, at least a portion of the user's arms (and may also be referred to as hand-tracking cameras). In some examples, one or more image sensors 214 are configured to face forward to acquire image data corresponding to a scene that the user would view if a user-responsive component 120 (e.g., an HMD) were not present (and may also be referred to as a scene camera). One or more optional image sensors 214 may include one or more RGB cameras (e.g., complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors), one or more infrared (IR) cameras, one or more event-based cameras, and / or similar.

[0031] Memory 220 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some examples, memory 220 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes non-temporary computer-readable storage media. In some examples, memory 220, or the non-temporary computer-readable storage media of memory 220, including an optional operating system 230 and XR experience module 240, stores the following programs, modules, and data structures, or subsets thereof:

[0032] The operating system 230 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks. In some examples, the XR experience module 240 is configured to present XR content to the user via one or more XR displays 212 or one or more speakers. For this purpose, in various examples, the XR experience module 240 includes a data acquisition unit 242, an XR presentation unit 244, an XR map generation unit 246, and a data transmission unit 248.

[0033] In some examples, the data acquisition unit 242 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from at least the controller 110 in Figure 1. For this purpose, in various examples, the data acquisition unit 242 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0034] In some examples, the XR presentation unit 244 is configured to present XR content via one or more XR displays 212 or one or more speakers. For this purpose, in various examples, the XR presentation unit 244 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0035] In some examples, the XR map generation unit 246 is configured to generate an XR map (e.g., a 3D map of an augmented reality scene or a map of a physical environment in which computer-generated objects can be placed) based on media content data. For this purpose, in various examples, the XR map generation unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0036] In some examples, the data transmission unit 248 is configured to transmit data (e.g., presentation data, location data, sensor data, etc.) to at least the controller 110, and optionally to one or more of the input device 125, output device 155, sensor 190, and / or peripheral device 195. For this purpose, in various examples, the data transmission unit 248 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0037] Although the data acquisition unit 242, XR presentation unit 244, XR map generation unit 246, and data transmission unit 248 are shown residing on a single device (e.g., user-responsive component 120 in Figure 1), in other examples any combination of the data acquisition unit 242, XR presentation unit 244, XR map generation unit 246, and data transmission unit 248 may reside in separate computing devices.

[0038] Returning to Figure 1, the controller 110 is configured to manage and adjust the user experience with respect to the 3D scene. In some examples, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with respect to Figures 3A-3C.

[0039] In some examples, the controller 110 is a computing device that is local to or remote to the scene 105 (e.g., the physical environment). For example, the controller 110 is a local server located within the scene 105. In another example, the controller 110 is a remote server located outside the scene 105 (e.g., a cloud server, a central server, etc.). In some examples, the controller 110 is communicably coupled to components of the computer system 101 (e.g., output device 155 and / or user-responsive component 120) configured to provide output to the user via one or more wired or wireless communication channels (e.g., BLUETOOTH®, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In some examples, the controller 110 is contained within an enclosure (e.g., a physical housing) of a component of the computer system 101 (e.g., user-responsive component 120) configured to provide output to the user, or shares the same physical enclosure or support structure as a component of the computer system 101 configured to provide output to the user.

[0040] In some examples, the various components and functions of the controller 110, as described below with respect to Figures 3A-3C, 4, 5A-5S, and 6, are distributed across multiple devices. For example, a first set of the components of the controller 110 (and their associated functions) is implemented on a remote server system relative to scene 105, while a second set of the components of the controller 110 (and their associated functions) is local to scene 105. For example, the second set of components is implemented within a portable electronic device (e.g., a wearable device such as an HMD) located within scene 105. It will be understood that the specific manner in which the various components and functions of the controller 110 are distributed across various devices may vary based on the various implementations of the examples described herein.

[0041] Figure 3A is a schematic block diagram of controller 110 in several examples. While certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more appropriate embodiments of the examples disclosed herein. Furthermore, Figure 3A is more intended to illustrate the function of various features that may be present in a particular implementation, in contrast to the structural schematics of the examples described herein. As will be recognized by those skilled in the art, the components shown separately can be combined, and some components can be separated. For example, some functional modules shown separately in Figure 3A can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various examples. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation and, in some examples, will depend in part on a particular combination of hardware, software, and / or firmware selected for a particular implementation.

[0042] In some examples, the controller 110 includes one or more processing units 302 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing cores, and / or similar), one or more input / output (I / O) devices 306, and one or more communication interfaces 308 (e.g., Universal Serial Bus (USB), FireWire®, Thunderbolt®, IEEE 802.3x, IEEE 802.11x, IEEE 802.11x, etc.). This includes 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth®, ZIGBEE®, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0043] In some examples, one or more communication buses 304 include circuits that interconnect system components and control communication between system components. In some examples, one or more I / O devices 306 include at least one of the following: a keyboard, mouse, touchpad, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, and / or the like.

[0044] Memory 320 includes high-speed random access memory such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some examples, memory 320 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some examples, memory 320 or the non-temporary computer-readable storage medium of memory 320 stores the following programs, modules, and data structures, or subsets thereof, including the operating system 330, a save-for-later application 331, an application programming interface (API) 332, an application 333, and a three-dimensional (3D) experience module 340.

[0045] The operating system 330 includes instructions for handling various basic system services and instructions for performing hardware-dependent tasks.

[0046] The saveforwarder application 331 is configured to access content stored in an associated repository (e.g., data storage), organize the content (e.g., by determining category 375 in Figure 3C) to determine action suggestions (e.g., 374 in Figure 3C) based on the content, and / or determine user information (e.g., 376 in Figure 3C) based on the content. An example of the functionality of the saveforwarder application 331 is described in detail below with respect to the generative action suggestion unit 370 in Figure 3C, and illustrated below in Figures 5A to 5S. In some examples, the associated repository is implemented in memory 320, at least partially. In some examples, the repository is implemented in a distributed manner. For example, the repository corresponds to storage associated with a particular user's cloud storage account.

[0047] Various different types of content can be stored in the repository associated with the saveforator application 331. Examples of such content include various types of data (e.g., image data (e.g., captured via the image sensor 214), screenshots, documents, calendar items, music, movies, products, web pages, physical objects, virtual objects, messages, etc.) and / or individual representations of various types of data. Individual representations of specific types of data include, for example, metadata describing the content of the data, e.g., natural language descriptions of image data and / or natural language summaries of documents. Content is stored in the repository in response to a computer system (e.g., 101, 500, and / or 502) receiving a specific type of user input, referred to herein as “saveforator input”. In some examples, saveforator input includes natural language input (e.g., “save for later”), touch input, gesture input, gaze input, and / or inputs that activate hardware elements. In some examples, the saveforator input corresponds to the selection of a saveforator graphical element (e.g., 506 in Figures 5A to 5J), or to a request to save content to the repository associated with the saveforator application 331.

[0048] In some examples, the controller 110 handles image data captured in response to a save fore-later input differently from image data captured in response to another camera capture input (e.g., an input requesting the activation of a physical or virtual shutter button to capture an image). Image data captured in response to a save fore-later input may be referred to as a "save fore-later view," and image data captured in response to another camera capture input may be referred to as a "camera image." Therefore, the user provides different types of input for capturing a save fore-later view and for capturing a camera image. For example, while the computer system is displaying a view of a 3D scene, the computer system simultaneously displays a camera capture graphical element (e.g., 508 in Figures 5A-5J) and a save fore-later view graphical element (e.g., 506 in Figures 5A-5J). In response to receiving an input to select a camera capture graphical element, the computer system captures a camera image of the 3D scene, and in response to receiving an input to select a save fore-later view graphical element, the computer system captures a save fore-later view of the 3D scene. As another example, in response to the activation of a first hardware element, the computer system captures a camera image of the 3D scene, and in response to the activation of another second hardware element, the computer system captures a save for later view of the 3D scene. As yet another example, in response to natural language input indicating the intent to capture a camera (e.g., "take a picture"), the computer system captures a camera image of the 3D scene, and in response to natural language input indicating the intent to save for later (e.g., "save for later"), the computer system captures a save for later view of the 3D scene.In some examples, the SaveForeLator view is processed by default (e.g., without receiving any further user input in addition to the SaveForeLator input) using one or more of the processes described below with respect to the Generative Action Proposal Unit 370 (e.g., to determine the Action Proposal 374, Category 375, and / or User Information 376), while the camera image is not processed by default using one or more of such processes. In some examples, the SaveForeLator view is available by default through the user interface of the SaveForeLator application 331, while the camera image is not available by default for display through the user interface of the SaveForeLator application 331. Similarly, in some examples, the camera image is available by default for display through the user interface of another photo application (e.g., an application that allows a user to view and / or edit captured photos and / or videos), while the SaveForeLator view is not available by default for display through the user interface of that other photo application. For example, a computer system displays a dedicated save for later view user interface (e.g., Figure 5L) to display the save for later view, and a separate dedicated photo application user interface (e.g., Figure 5K) to display the captured images and videos.

[0049] Figure 3A illustrates that the saveforwarder application 331 is a separate software component from the operating system 330; however, in some examples, the functionality of the saveforwarder application 331 is implemented by the operating system 330. Therefore, in some examples, the functionality of the saveforwarder application 331 is provided at the operating system level, and such functionality does not need to be implemented in a separate software module from the operating system 330.

[0050] API 332 provides an interface that allows other applications 333 to access and / or use features provided by and / or information determined by the SaveForeLetter application 331 (e.g., action suggestions 374, categories 375, and / or user information 376). Similarly, API 332 provides an interface that allows the SaveForeLetter application 331 to access and / or use features provided by and / or information determined by the application 333. For example, by making an API call via API 332, the SaveForeLetter application 331 can cause application 333 to initiate an action suggestion determined by the SaveForeLetter application 331 (e.g., sending a message, making a phone call, setting a calendar entry, booking a flight, etc.). As another example, by making an API call via API 332, application 333 can retrieve user information 376 determined by the SaveForeLetter application 331 for use by application 333.

[0051] In some examples, API 332 operates to protect privacy by restricting and / or preventing the exposure of a user's personal information (e.g., stored in a repository associated with the SaveForeLetter application 331) to application 333. For example, one or more protocols of API 332 (defined, e.g., by the syntax and / or parameters of permitted API calls) prohibit API calls from application 333 that request access to certain types of functionality of the SaveForeLetter application 331 and / or certain information associated with the SaveForeLetter application 331. As an example, API 332 prohibits API calls from application 333 that request access to content stored in the repository, but permits API calls from application 333 that request access to information determined from the stored content (e.g., abstracted and / or generalized information). As a concrete example, suppose the stored content includes multiple SaveForeLetter views that collectively represent the user's favorite movies. Based on the SaveForeLator view, the SaveForeLator application 331 infers the user's favorite movies, for example, in accordance with the techniques described below with respect to the generative action suggestion unit 370. API 332 allows the movie application to request the inferred favorite movies via API calls (for example, to later suggest watching and / or purchasing the movies through the movie application), but does not allow the movie application to access the SaveForeLator view on which the movies were inferred.

[0052] In some examples, API 332 implements different protocols for different types of applications 333. For example, API 332 allows first-party applications (e.g., applications pre-installed on the computer system at the time of purchase or provided via operating system update files) to make API calls requesting certain types of information from the saveforator application 331 (and allows the saveforator application 331 to respond to such calls by providing the requested type of information), but prohibits third-party applications (e.g., applications provided via an application store, downloaded over a network, and / or read from a storage device) to make API calls requesting such types of information from the saveforator application 331. In some examples, API 332 includes multiple different APIs, each allowing individual applications 333 to access different functions of the saveforator application 331 and / or different information determined by the saveforator application 331. For example, one of the APIs may be exposed to a first-party application, another to a third-party application, and so on, allowing different types of applications to access different features and / or data associated with the saveforator application 331.

[0053] Although Figure 3A illustrates API 332 as a separate software component from the operating system 330, in some examples API 332 is implemented as part of the operating system 330. In some examples API 332 is partially implemented by firmware, microcode, or other low-level logic that runs partially within the computer system's hardware.

[0054] Application 333 includes one or more applications for performing various functions. Examples include web browser applications, fitness applications, health applications, media applications, navigation applications, calendar applications, digital payment applications, camera applications, weather information applications, photo editing applications, word processing applications, drawing applications, application stores, and online shopping applications. Application 333 may include first-party and third-party applications.

[0055] In some examples, the 3D Experience Module 340 is configured to manage and adjust the user experience provided by the computer system 101 with respect to a 3D scene. For example, the 3D Experience Module 340 is configured to acquire data corresponding to the 3D scene (e.g., data generated by the computer system 101 and / or data from the data acquisition unit 341 described later) and to cause the computer system 101 to perform actions for the user based on the data (e.g., provide suggestions, display content, etc.). For this purpose, in various examples, the 3D Experience Module 340 includes a data acquisition unit 341, a tracking unit 342, an adjustment unit 346, a data transmission unit 348, a digital assistant (DA) unit 350, a live action suggestion unit 360, and a generative action suggestion unit 370.

[0056] In some examples, the data acquisition unit 341 is configured to acquire data (e.g., presentation data, interaction data, sensor data, location data, etc.) from one or more of the user-responsive components 120, input devices 125, output devices 155, sensors 190, and peripheral devices 195. For this purpose, in various embodiments, the data acquisition unit 341 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0057] In some examples, the tracking unit 342 is configured to map scene 105 and track the location of the user (and / or any portable devices held or worn by the user). For this purpose, in various examples, the tracking unit 342 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0058] In some examples, the tracking unit 342 includes an eye-tracking unit 343. The eye-tracking unit 343 includes instructions and / or logic for tracking the position and movement of the user's gaze (or more broadly, the user's eyes, face, or head) using data acquired from the eye-tracking device 130. In some examples, the eye-tracking unit 343 tracks the position and movement of the user's gaze relative to the physical environment, relative to the user (e.g., the user's hands, face, or head), relative to a device worn or held by the user, and / or relative to content displayed by the user-responsive component 120.

[0059] The eye-tracking device 130 is controlled by the eye-tracking unit 343 and includes various hardware and / or software components configured to perform eye-tracking techniques. For example, the eye-tracking device 130 includes at least one eye-tracking camera (e.g., an infrared (IR) or near-infrared (NIR) camera) and an illumination source (e.g., an IR or NIR light source such as an array or ring of LEDs) that emits light (e.g., IR or NIR light) toward the user's eyes. The eye-tracking camera may be directed toward the user's eyes to receive IR or NIR light directly from the reflected light source, or alternatively, it may be directed toward a mirror that reflects IR or NIR light from the eyes toward the eye-tracking camera. The eye-tracking device 130 optionally captures images of the user's eyes (e.g., as a video stream captured at 60-120 frames / second), analyzes the images to generate eye-tracking information, and communicates the eye-tracking information to the eye-tracking unit 343. In some examples, both of the user's eyes are tracked separately by their respective eye-tracking cameras and illumination sources. In some cases, only one of the user's eyes is tracked by a separate eye-tracking camera and lighting source.

[0060] In some examples, the tracking unit 342 includes a hand tracking unit 344. The hand tracking unit 344 includes instructions and / or logic for tracking the position and / or movement of one or more parts of the user's hand using hand tracking data acquired from the hand tracking device 140. The hand tracking unit 344 tracks the position and / or movement relative to the scene 105, relative to the user (e.g., the user's head, face, or eyes), relative to a device worn or held by the user, relative to content displayed by the user-responsive component 120, and / or relative to a coordinate system defined relative to the user's hand. In some examples, the hand tracking unit 344 analyzes the hand tracking data to identify hand gestures (e.g., pointing gestures, pinch gestures, clenching gestures, and / or grasping gestures) and / or identify content corresponding to the hand gestures (e.g., physical or virtual content), such as content selected by the hand gestures. In some examples, the hand gestures are air gestures. Air gestures are gestures detected without the user touching (or independently of) an input element that is part of a device (e.g., a computer system 101, one or more input devices 125, a hand tracking device 140, device 500, and / or device 502), and are based on detected movements of a part of the user's body in the air (e.g., head, one or more arms, one or more hands, one or more fingers, and / or one or more legs), including the user's body movement relative to an absolute reference (e.g., the angle of the user's arm relative to the ground, or the distance of the user's hand relative to the ground), the user's body movement relative to another part of the user's body (e.g., the movement of the user's hand relative to the user's shoulder, the movement of one of the user's hands relative to the user's other hand, and / or the movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movements of a part of the user's body (e.g., a tap gesture including the movement of the hand in a predetermined posture by a predetermined amount and / or speed, or a shake gesture including a predetermined speed or a rotation amount of the part of the user's body).

[0061] The hand tracking device 140 is controlled by the hand tracking unit 344 and includes various hardware and / or software components configured to perform hand tracking and hand gesture recognition techniques. For example, the hand tracking device 140 includes one or more image sensors (e.g., one or more IR cameras, 3D cameras, depth cameras, and / or color cameras) that capture three-dimensional information (e.g., a depth map) representing the hand of a human user. One or more image sensors capture an image of the hand with sufficient resolution to distinguish the fingers and their respective positions. In some examples, one or more image sensors project a speckled pattern onto the environment including the hand and capture an image of the projected pattern. In some examples, one or more image sensors capture a temporal sequence of hand tracking data (e.g., captured three-dimensional information and / or captured images of the projected pattern), and the hand tracking device 140 communicates the temporal sequence of hand tracking data to the hand tracking unit 344 for further analysis, for example, to identify hand gestures, hand poses, and / or hand movements.

[0062] In some examples, the hand tracking device 140 includes one or more hardware input devices configured to be worn and / or held (or otherwise attached) by each of the user's one or more hands. In such examples, the hand tracking unit 344 tracks the position, orientation, and / or movement of the user's hand based on tracking the position, orientation, and / or movement of each hardware input device. The hand tracking unit 344 tracks the position, orientation, and / or movement of each hardware input device optically (e.g., via one or more image sensors) and / or based on data obtained from sensors contained within the hardware input device (e.g., accelerometers, magnetometers, gyroscopes, inertial measurement units, etc.). In some examples, the hardware input device includes one or more physical controls (e.g., buttons, touch-sensitive surfaces, pressure-sensitive surfaces, knobs, joysticks, etc.). In some examples, instead of, or in addition to, performing a particular function in response to the detection of a distinct type of hand gesture, the computer system 101 also performs that particular function in response to user input selecting a distinct physical control of the hardware input device. For example, the computer system 101 interprets a pinch hand gesture input as a selection of the focused element, and / or the selection of a physical button on a hardware device as a selection of the focused element.

[0063] In some examples, the adjustment unit 346 is configured to manage and adjust the experience provided to the user via the user-responsive components 120, one or more output devices 155, and / or one or more peripheral devices 195. For this purpose, in various examples, the adjustment unit 346 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0064] In some examples, the data transmission unit 348 is configured to transmit data (e.g., presentation data, location data, etc.) to a user-responsive component 120, one or more input devices 125, an output device 155, a sensor 190, and / or peripheral devices 195. For this purpose, in various examples, the data transmission unit 348 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0065] The Digital Assistant (DA) unit 350 includes instructions and / or logic for providing DA functionality to the computer system 101. Thus, the DA unit 350 provides DA functionality to the user of the computer system 101 while the user and / or their avatar are present in a three-dimensional scene. For example, the DA performs various tasks related to the three-dimensional scene, either in advance or upon request from the user. In some examples, the DA unit 350 performs at least some of the following: converting speech input to text (e.g., using a speech-to-text (STT) processing unit 352); identifying the user's intent expressed in natural language input received from the user; actively extracting and acquiring information necessary to fully satisfy the user's intent (e.g., by removing ambiguity in terminology in natural language input and / or by acquiring information from a data acquisition unit 341); determining a task flow to satisfy the identified intent; and executing that task flow to satisfy the identified intent.

[0066] In some examples, the DA unit 350 includes a natural language processing (NLP) unit 351 configured to identify user intents. The NLP unit 351 retrieves n best text representation candidates (singular or plural) ("word sequences (singular or plural)" or "token sequences (singular or plural)") generated by the STT processing unit 352 and attempts to associate each of these text representation candidates with one or more user intents recognized by the DA. In some examples, the user intent represents a task, which is executable by the DA and has an associated task flow implemented in the task flow processing unit 353. This associated task flow is a set of programmed actions and steps that the DA takes to execute that task. The scope of the DA's capabilities depends, in some examples, on the number and types of task flows implemented in the task flow processing unit 353, in other words, on the number and types of user intents recognized by the DA.

[0067] In some examples, once the NLP unit 351 identifies a user intent based on a user request, the NLP unit 351 causes the task flow processing unit 353 to perform the necessary actions to satisfy the user request. For example, the task flow processing unit 353 executes a task flow corresponding to the identified user intent in order to perform a task that satisfies the user request. In some examples, performing a task includes causing the computer system 101 to provide output (e.g., graphic output, audio output, and / or haptic output) indicating the performed task.

[0068] The Live Action Suggestion Unit 360 is configured to cause a computer system to present action suggestions (e.g., display output and / or audio output) based on detected objects in the 3D scene. Action suggestions are presented while the user is immersed in the 3D scene, for example, while the user is perceiving a live view of the 3D scene. In some examples, the user perceives a live view of the 3D scene by browsing a displayed live view of the 3D scene through a user interface that displays a live view of the 3D scene via pass-through video, such as the user interface of a camera application. In some examples, the user perceives a live view of the 3D scene by directly browsing the 3D scene without the help of a display, for example, when the computer system is not configured to display content. In some examples, the user perceives a live view of the 3D scene by browsing the 3D scene through a transparent or translucent medium that allows light representing the 3D scene to physically pass through the medium.

[0069] Figure 3B illustrates a block diagram of the live action suggestion unit 360 in several examples. As shown, the live action suggestion unit 360 is configured to determine an action suggestion 362 based on input image data 361. To this end, the live action suggestion unit 360 includes an object detection unit 363, a suggestion gating unit 364, and a suggestion decision unit 365.

[0070] The object detection unit 363 is configured to detect objects (e.g., text or other objects) present in a 3D scene represented by the image data 361. For example, the object detection unit 363 implements surface scanning techniques and / or text detection techniques to detect objects. In some examples, the object detection unit 363 further implements object classification techniques to determine the type and / or subtype of the detected objects (e.g., using computer vision models and / or text classification models). For example, the object detection unit 363 classifies an object detected in the 3D scene as text and further determines the type of text, such as a phone number, name, physical address, email address, or social media identifier. In another example, the object detection unit 363 classifies an object detected in the 3D scene as a person, plant, food, landmark, body of water, mountain, etc.

[0071] The proposal decision unit 365 is configured to determine an action proposal 362 based on the type and / or subtype of the object detected by the object detection unit 363. For example, the proposal decision unit 365 implements a predetermined rule for mapping the object type and / or subtype to an action proposal 362. For example, if the object is text, the corresponding action proposal 362 is to read the text aloud through a text-to-speech process. For example, if the object is text in a foreign language, the corresponding action proposal 362 is to translate the text into the user's native language. For example, if the object is a telephone number, the corresponding action proposal 362 is to call the telephone number and / or add the telephone number to the user's contact list. For example, if the object is a plant, the corresponding action proposal 362 is to obtain further information about the plant, such as obtaining its identification information through a web search.

[0072] The proposal gating unit 364 determines whether one or more proposal criteria are met, and if one or more criteria are met, it is configured to cause the computer system to present the corresponding action proposal 362 (e.g., display output and / or audio output). If one or more proposal criteria are not met, the proposal gating unit 364 causes the computer to refrain from presenting the action proposal 362. In some examples, a proposal criterion is met when the confidence score of the action proposal 362 (e.g., the confidence that the object detection unit 363 classified the corresponding object) exceeds a threshold, and a proposal criterion is not met when the confidence score of the action proposal 362 does not exceed a threshold. Therefore, in some examples, the proposal gating unit 364 enables the presentation of action proposals 362 corresponding to objects identified / classified with relatively high confidence.

[0073] In some examples, the suggestion gating unit 364 determines whether the suggestion gating criteria are met by comparing the representation of the image data 361 with representations of one or more sets of gating words. The gating words describe conditions about the image data 361 that, if met, prevent the presentation of action suggestions 362 for the image data 361. For example, suppose the gating word is “no word”. If the image data 361 satisfies the “no word” condition (meaning the image data 361 does not depict a word), then no action suggestion 362 (or at least an action suggestion 362 determined based on the text depicted by the image data 361) is determined and / or presented. This can favorably prevent the presentation of potentially inappropriate suggestions for image data 361 that depict characters but not words. For example, if image data 361 depicts a keyboard but not words, the computer system will not present action suggestions for reading the characters on the keyboard and / or for adding the characters on the keyboard to the user's contact list. As another example, suppose the gating word is “flat surface”. If the image data 361 satisfies the condition of a "flat surface" (meaning that the image data 361 primarily depicts a flat surface), no action proposal 362 for the image data 361 is determined and / or presented. This advantageously prevents the presentation of actions that may be inappropriate for the image data 361, such as an action proposal to perform a web search on a flat wall.

[0074] In some examples, the representation of image data 361 is an image embedding (e.g., a vector representation) of image data 361 generated by an image encoder component of a computer vision model. In some examples, the proposed gating unit 364 does not further process image data 361 using further components of the computer vision model, such as components configured to classify image 361 and / or determine a natural language description of image 361. Thus, generating the image embedding may have a relatively low computational cost compared to the full processing of image data 361 using a computer vision model. In some examples, the representation of a set of one or more gating words is a text embedding of the gating words, similarly generated by a text encoder. The proposed gating unit 364 compares the representation of image data 361 to the representation of the gating words by comparing the image embedding to this text embedding (e.g., via cosine similarity). If the embedding matches (for example, defined by comparing the cosine similarity score to a threshold) (meaning that image data 361 satisfies the conditions of "no words" or "flat surface"), the proposed criteria are not met, and therefore no action proposal 362 is determined and / or presented for image data 361. If the embedding does not match (meaning that image data 361 does not satisfy the conditions of "no words" or "flat surface"), the proposed criteria are met, and therefore an action proposal 362 is determined and / or presented for image data 361.

[0075] In some examples, the proposed gating unit 364 generates a representation (e.g., an image embedding) of the image data 361 when it determines that a set of image stability criteria is met. For example, the proposed gating unit 364 generates an image embedding if the image data 361 was captured during a period when the corresponding image sensor was relatively stable, such as when the movement of the image sensor (and / or the physical component housing the image sensor) is below a threshold amount. In another example, the proposed gating unit 364 analyzes the image data 361 to determine if it is a stable image (e.g., captured during a period of relatively little user movement) and generates an image embedding only for stable images.

[0076] In some examples, the set of gating words is determined based on the image data 361. For example, if the suggestion gating unit 364 detects text (e.g., characters) based on the image data 361, the gating word is set to "no word". Otherwise, the gating unit 364 sets the gating word to "flat surface". In this way, the set of gating words can change based on the content depicted by the image data 361, thereby allowing inappropriate action suggestions to be filtered out as the view of the 3D scene changes.

[0077] Returning to Figure 3A, the generative action suggestion unit 370 is configured to organize the content stored in the repository (e.g., assign categories) and to determine user information based on the stored content, in order to determine an action suggestion based on the content stored in the repository associated with the saveforwarder application 331.

[0078] Figure 3C illustrates block diagrams of the generative action suggestion unit 370 in various examples. The generative action suggestion unit 370 includes a generative model 371. The generative model 371 is configured to determine action suggestions 374, categories 375 (e.g., one or more categories for one or more of the content items), and / or user information 376 (e.g., user preferences, interests, and / or habits) by processing input content 372 (e.g., one or more content items stored in a repository) together with optional contextual information 373. In some examples, the generative model 371 implements an AI model (e.g., a large-scale language model (LLM), such as a multimodal LLM) to perform such functions. The AI ​​model is based on (e.g., is or is built upon) a foundational model, as described below with respect to Figure 4. In some examples, the generative action suggestion unit 370 is replaced by another unit configured to determine the action suggestion 374, category 375, and / or user information 376 based on the input content 372 and optional context information 373 via a non-AI-based technique. In some examples, the non-AI-based technique is configured to compare an input content item with one or more example content items, each having a predetermined corresponding action suggestion 374, category 375, and / or user information 376.

[0079] Generally, context information 373 provides additional information to specific content items of the input content 372. It will be understood that the generative model 371 may use context information 373 to determine a reasonable / accurate action suggestion 374 for the input content 372, to determine a reasonable / accurate category 375 for the input content 372, and / or to determine reasonable / accurate user information 376 based on the input content 372, for example, more accurate / relevant than determining such information using only the input content 372.

[0080] In some examples, contextual information 373 includes personal information of a user of a computer system (e.g., 500 and / or 502). Examples of personal information include contact data (e.g., contact information of the user and / or other users), email data, message data, calendar data, phone data (e.g., call logs and voicemails), location data, reminder data, photos, videos, health information, workout information, financial information, web search history, navigation history, media data (e.g., music and audiobooks), information related to the user's home (e.g., the status of the user's appliances and home security system, and / or home security system access information), information about the user's daily routine, digital assistant shortcuts, documents (e.g., notes, journal entries, and / or lists), etc. In some examples, personal information includes other content stored in a repository associated with the SaveForeLetter application 331. For example, suppose contextual information 373 includes a screenshot of a recipe stored in the repository. The screenshot shows that the recipe requires celery and onions. Based on such contextual information 373 and input content 372, which is a saveforward view depicting celery and onions, the generative model 371 determines an action suggestion 374 for adding celery and onions to the user's shopping list and ordering celery and onions through a grocery delivery service. As another example, suppose the contextual information 373 includes multiple saveforward views, each depicting promotional information for a particular movie. Based on such contextual information 373 and input content 372, which is a screenshot advertising a particular movie, the generative model 371 determines user information 376 indicating that the user is interested in a particular movie.

[0081] In some examples, context information 373 includes the location of the computer system when the saveforator input for content 372 was received, such as the location of the computer system that received the saveforator input. For example, suppose content 372 is a saveforator view of a restaurant, and context information 373 includes the previous location of the computer system when the saveforator input was received. Based on content 372, the previous location of the computer system, and additional context information 373 indicating that the computer system's current location is near the previous location, the generative model 371 determines an action suggestion 374 for ordering food at the restaurant. In contrast, if context information 373 were not used, the generative model 371 would not determine an action suggestion 374 for ordering food at the restaurant (because such an action suggestion is only valid when the user is near the restaurant), and instead might determine an action suggestion 374 for searching for more information about the restaurant.

[0082] In some examples, context information 373 includes scene data (e.g., audio data and / or image data) representing the 3D scene. In some examples, the computer system captures such scene data representing the 3D scene before, during, and / or after receiving input for the computer system to capture a save for later view of the 3D scene. Such scene data may provide additional information that is valid for determining action proposals 374 for the save for later view, valid for determining categories 374 for the save for later view, and / or valid for determining user information 376 based on the save for later view. For example, suppose content 372 is a save for later view depicting a particular product. Before the computer system receives the corresponding save for later input, the computer system captures scene data depicting that the product is sold in a gift shop and that Christmas music is playing in the 3D scene. Based on content 372 and scene data, the generative model 371 determines an action proposal 374 for purchasing the product and determines that the product's category 375 is Christmas gift. In some examples, the generative action suggestion unit 370 limits the scene data used as contextual information 373 for a particular save for later view to scene data captured within a first predetermined duration (e.g., 10 seconds, 5 seconds, or 1 second) before the corresponding save for later input is received, and / or scene data captured within a second predetermined duration (e.g., 10 seconds, 5 seconds, or 1 second) after the corresponding save for later input is received. For example, the generative action suggestion unit 370 uses scene data captured within 5 seconds before and / or after receiving the save for later input to determine an action suggestion 374 for the corresponding save for later view.

[0083] In some examples, contextual information 373 includes the attention of the computer system user, such as the user's attention when a saveforator input for a particular saveforator view is received. User attention is defined by the user's gaze direction and / or the user's posture (e.g., body posture and / or head posture). Thus, in some examples, user attention is determined based on data detected via sensors 190, 214, 130, and / or 140. User attention can indicate one or more elements of content 372 that are in focus for the user, such as the element the user is paying attention to. For example, suppose content 372 is a saveforator view depicting two different restaurants, and user attention indicates that the user is looking at the first of the two restaurants when the computer system receives the corresponding saveforator input. Based on the content 372 and the user's attention, the generative model 371 determines an action proposal 374 for ordering food at the first restaurant, but does not determine an action proposal for ordering food at the second restaurant (or determines a lower confidence score for the action proposal for ordering food at the second restaurant).

[0084] In some examples, the live action proposal unit 360 determines action proposal 362 more quickly than the generative action proposal unit 370 determines action proposal 374. For example, the live action proposal unit 360 determines action proposal 362 within approximately 0.1 to 0.2 seconds of receiving image data 361, while the generative action proposal unit 370 determines action proposal 374 within approximately 2 seconds of receiving input content 372 (e.g., image data and / or its representation). Such a time difference may be because the live action proposal unit 360 implements a relatively simple / computationally inefficient process (e.g., as described above) to determine action proposal 362, while the generative action proposal unit 370 implements a relatively complex / computationally inefficient process (e.g., multimodal LLM) to determine action proposal 374.

[0085] Due to this time lag, in some examples, action suggestions 362 for a 3D scene may be presented live within the 3D scene (e.g., while the user perceives a live view of the 3D scene), while action suggestions 374 for a 3D scene may not be presented live within the 3D scene. Instead, in some examples, action suggestions 374 are determined and / or presented when the user moves to the user interface of the saveforator application 331. More specifically, in some examples, content 372 (e.g., a saveforator view of a 3D scene) is not analyzed using the process described with respect to the generative action suggestion unit 370 until after the content has been captured (e.g., in response to a saveforator input) and / or after the content (and / or its representation) has been saved to the repository associated with the saveforator application 331. Operating the computer system in such a manner may provide an improved user experience by not embarrassing the user with potentially inappropriate action suggestions. For example, while a user is immersed in a 3D scene containing an object, an action suggestion 362 for the object can be presented quickly before the object moves out of view (e.g., due to the user's movement). Thus, the user may perceive that the action suggestion 362 for an object within their field of view was presented virtually instantaneously (e.g., within 0.1-0.2 seconds of detecting the object). In contrast, an action suggestion 374 for an object may not be presented until later, because the object may move out of view before the time required to decide on the action suggestion 374 (e.g., 2 seconds) has elapsed, and therefore the action suggestion 374 may be less relevant and / or of less interest to the user.

[0086] In some examples, the generative model 371 includes a single generative model configured to determine each of the action suggestion 374, the category 375, and the user information 376. In some examples, the single generative model is prompted to generate each of different types of outputs via different input prompts, such as "determine an action suggestion based on [X] and [Y]", "determine the category of [X] based on [Y]", and "predict information about the user based on [X] and [Y]", where [X] represents content 372 (e.g., one or more content items stored in a repository associated with the saveforwarder application 331) and [Y] represents contextual information 373. In some examples, the generative model 371 includes multiple different generative models, each configured to perform one of the following: determine the action suggestion 374, determine the category 375, and determine the user information 376 (e.g., fine-tuned and / or trained).

[0087] In some examples, the functionality of the generative action suggestion unit 370 described above is implemented by the save-forward application 331. For example, although Figure 3A illustrates that the generative action suggestion unit 370 and the save-forward application 331 are different modules, the functionality of both the generative action suggestion unit 370 and the save-forward application 331 can also be provided by a single module, such as the save-forward application 331.

[0088] In some examples, the 3D EXPERIENCE module 340 accesses one or more artificial intelligence (AI) models configured to perform various functions described herein. The AI ​​models are implemented at least in part on the controller 110 (e.g., locally on a single device or in a distributed manner), and / or the controller 110 communicates with one or more external services that provide access to the AI ​​models. In some examples, one or more components and functions of the DA unit 350, the live action suggestion unit 360, and / or the generative action suggestion unit 370 are implemented using AI models. For example, the DA unit 350 implements one or more AI models to perform speech recognition, intent determination (e.g., natural language processing and / or image processing), object recognition, and / or response generation, and the generative action suggestion unit 370 implements one or more AI models to determine action suggestions 374, categories 375, and / or user information 376.

[0089] In some cases, an AI model is based on (e.g., is or is built upon) one or more foundational models. Generally, a foundational model is a deep learning neural network that is trained on a large training dataset and can be adapted to perform specific functions. Thus, a foundational model can aggregate information learned from large (and optionally, multimodal) datasets and be adapted (e.g., fine-tuned) to perform a variety of downstream tasks that the foundational model may not have been originally designed to perform. Examples of such tasks include language translation, speech recognition, user intent determination (e.g., natural language processing), sentiment analysis, computer vision tasks (e.g., object recognition and scene understanding), question answering, image generation, audio generation, and the generation of computer executable instructions. A foundational model can accept a single type of input (e.g., text data) or multimodal input such as two or more of text data, image data, video data, audio data, sensor data, etc. In some cases, a foundational model is prompted to perform a particular task by providing a natural language description of that task. Examples of foundational models include Open AI, Inc.'s GPT-n series models (e.g., GPT-1, GPT-2, GPT-3®, and GPT-4®), DALL-E, and CLIP, Microsoft Corporation's Florence and Florence-2, Google LLC's BERT, and Meta Platforms, Inc.'s LLaMA, LLaMA-2, and LLaMA-3.

[0090] Figure 4 illustrates Architecture 400 for a foundational model with several examples. Architecture 400 is merely illustrative, and various modifications to it are possible. Thus, the components of Architecture 400 (and the functions associated with them) can be combined, the order of the components (and the functions associated with them) can be changed, components of Architecture 400 can be removed, and other components can be added to Architecture 400. Furthermore, although Architecture 400 is based on transformers, those skilled in the art will understand that Architecture 400 can additionally or alternatively implement other types of machine learning models, such as models based on convolutional neural networks (CNNs) and models based on recurrent neural networks (RNNs).

[0091] Architecture 400 is configured to process input data 402 and generate output data 480 corresponding to a desired task. Input data 402 includes one or more types of data, such as text data, image data, video data, audio data, sensor data (e.g., motion sensors, biosensors, temperature sensors, etc.), computer executable instructions, and structured data (e.g., in the form of XML files, JSON files, or other file types). In some examples, input data 402 includes data from data acquisition unit 341. Output data 480 includes one or more types of data, depending on the task being performed. For example, output data 480 includes one or more of text data, image data, audio data, and computer executable instructions. The input and output data types described above are merely examples, and it will be understood that Architecture 400 can be configured to accept various types of data as input and generate various types of data as output. Such data types can be diverse based on the specific functions configured for the underlying model to perform.

[0092] Architecture 400 includes an embedded module 404, an encoder 408, an embedded module 428, a decoder 424, and an output module 450, the functions of which are described below.

[0093] The embedding module 404 is configured to receive input data 402 and parse the input data 402 into one or more token sequences. The embedding module 404 is further configured to determine the embedding (e.g., vector representation) of each token representing each token in the embedding space, such that, for example, similar tokens are close together in the embedding space and dissimilar tokens are far apart. In some examples, the embedding module 404 includes a position encoder configured to encode position information within the embedding. The individual position information of a given embedding indicates its relative position within the sequence. The embedding module 404 is configured to output embedding data 406 of the input data by aggregating the token embeddings of the input data 402.

[0094] The encoder 408 is configured to map the embedded data 406 to an encoder representation 410. The encoder representation 410 represents contextual information for each token, showing learned information about how each token relates to (e.g., attends) each other. The encoder 408 includes an attention layer 412, a feedforward layer 416, normalization layers 414 and 418, and residual connections 420 and 422. In some examples, the attention layer 412 applies a self-attention mechanism to the embedded data 406 to compute an attention representation (e.g., in matrix form) of the relationships of each token in the sequence to each other. In some examples, the attention layer 412 is multi-headed to compute multiple different attention representations of the relationships of each token to each other, each different representation showing a different learned property of the token sequence. The attention layer 412 is configured to aggregate the attention representations to output attention data 460 showing the interrelationships between tokens from the input data 402. In some examples, the attention layer 412 further masks the attention data 460 to suppress data representing relationships between selected tokens. The encoder 408 then passes the attention data 460 (optionally masked) through the normalization layer 414, the feedforward layer 416, and the normalization layer 418 to generate the encoder representation 410. The residual connections 420 and 422 can help stabilize and shorten the training and / or inference process by allowing the output of the embedding module 404 (i.e., the embedding data 406) to be passed directly to the normalization layer 414, and the output of the normalization layer 414 to be passed directly to the normalization layer 418, respectively.

[0095] Figure 4 illustrates that architecture 400 includes a single encoder 408, but in other examples, architecture 400 includes multiple stacked encoders configured to output an encoder representation 410. Each of the stacked encoders can generate different attention data, which may enable architecture 400 to learn different types of interrelationships between tokens and generate output data 410 based on a more complete set of learned relationships.

[0096] Decoder 424 is configured to receive encoder representation 410 and previous output embedding 430 as input and generate output data 480. Embedding module 428 is configured to generate previous output embedding 430. Embedding module 428 is similar to embedding module 404. Specifically, embedding module 428 tokenizes previous output data 426 (e.g., output data 480 generated by a previous iteration), determines the embedding for each token, and optionally encodes positional information into each embedding to generate previous output embedding 430.

[0097] Decoder 424 includes attention layers 432 and 436, normalization layers 434, 438, and 442, a feedforward layer 440, and residual connections 462, 464, and 466. Attention layer 432 is configured to output attention data 470 indicating the interrelationships between tokens from the previous output data 426. Attention layer 432 is similar to attention layer 412. For example, attention layer 432 applies a multi-head self-attention mechanism to the previous output embedding 430 and optionally masks attention data 470 to suppress data representing relationships between selected tokens (e.g., relationships between one token and a future token) so that architecture 400 does not consider future tokens as context when generating output data 480. Decoder 424 then passes attention data 470 (optionally masked) to normalization layer 434 to generate normalized attention data 470-1.

[0098] The attention layer 436 receives the encoder representation 410 and normalized attention data 470-1 as input to generate encoder-decoder attention data 475. The encoder-decoder attention data 475 correlates the input data 402 to the previous output data 426 by representing the relationship between the output of the encoder 408 and the previous output of the decoder 424. The attention layer 436 allows the decoder 424 to increase the weights of the parts of the encoder representation 410 that have been learned as relatively reasonable for generating the output data 480. In some examples, the attention layer 436 applies a multi-head attention mechanism to the encoder representation 410 and normalized attention data 470-1 to generate the encoder-decoder attention data 475. In some examples, the attention layer 436 further masks the encoder-decoder attention data 475 to suppress the interrelationships between selected tokens.

[0099] Decoder 424 then passes encoder-decoder attention data 475 (optionally masked) through normalization layer 438, feedforward layer 440, and normalization layer 442 to generate further processed encoder-decoder attention data 475-1. Normalization layer 442 then provides the further processed encoder-decoder attention data 475-1 to output module 450. Similar to residual connections 420 and 422, residual connections 462, 464, and 466 can stabilize and shorten the training and / or inference process by allowing the output of the corresponding component to be passed directly as input to the corresponding component.

[0100] Figure 4 illustrates that architecture 400 includes a single decoder 424, but in other examples, architecture 400 includes multiple stacked decoders, each configured to learn / generate different types of encoder-decoder attention data 475. This allows architecture 400 to learn multiple different types of interrelationships between tokens from input data 402 and tokens from output data 480, thereby allowing architecture 400 to generate output data 480 based on a more complete set of learned relationships.

[0101] The output module 450 is configured to generate output data 480 from further processed encoder-decoder attention data 475-1. For example, the output module 450 includes one or more linear layers that apply a learned linear transformation to the further processed encoder-decoder attention data 475-1, and a softmax layer that generates a probability distribution over possible classes (e.g., words or symbols) of output tokens based on the linear transformation data. The output module 450 then selects (e.g., predicts) elements of the output data 480 based on the probability distribution. The architecture 400 then passes the output data 480 as the previous input data 426 to the embedding module 428 to start another iteration of the training and / or inference process for the architecture 400.

[0102] It will be understood that various different AI models can be built based on the components of Architecture 400. For example, some large-scale language models (LLMs) (e.g., GPT-2 and GPT-3®) are decoder-only (e.g., include one or more instances of decoder 424 and do not include encoder 408), some LLMs (e.g., BERT) are encoder-only (e.g., include one or more instances of encoder 408 and do not include decoder 424), and other foundational models (e.g., Florence-2) are encoder-decoder type (e.g., include one or more instances of encoder 408 and one or more instances of decoder 424). Furthermore, it will be understood that foundational models built based on the components of Architecture 400 can be fine-tuned based on reinforcement learning techniques and training data specific to those tasks in order to optimize specific tasks such as extracting meaningful information from image and / or video data, generating code, generating music, or providing meaningful suggestions to a particular user.

[0103] Figures 5A to 5S illustrate techniques for providing suggestions through various examples. Figures 5A to 5D and 5F to 5J illustrate the user view of each 3D scene, while Figures 5E and 5K to 5S illustrate content displayed by device 500 or device 502 (e.g., user interfaces of web pages and applications). In some examples, device 500 provides at least a portion of the 3D scenes in Figures 5A to 5D and 5F to 5J. For example, the 3D scene is an XR scene containing at least some virtual elements generated by device 500. In other examples, the 3D scene is a physical scene.

[0104] The examples in Figures 5K to 5S illustrate that device 502 (a different device from device 500) displays separate content, but in other examples, device 500 displays that separate content instead. Therefore, device 500 and device 502 can be the same device.

[0105] Device 500 implements at least some of the components of the computer system 101 and / or the user-responsive component 120. Device 502 implements at least some of the components of the computer system 101. In the examples of Figures 5A to 5J, device 500 is a tablet device owned by the user, and device 502 is a smartphone device owned by the same user. In the examples of Figures 5A to 5D and Figures 5F to 5J, the user holds device 500 and views each 3D scene via pass-through video displayed by device 500.

[0106] In other examples, at least one of device 500 or 502 is a different type of device. For example, device 500 is instead a wearable device (e.g., earphones, headphones, smartwatch, or HMD (e.g., XR headset or glasses)), and device 502 is instead a tablet device owned by the same user. If device 500 is an HMD, the 3D scenes in Figures 5A-5D and 5F-5J are viewed via the HMD. For example, the 3D scenes in Figures 5A-5D and 5F-5J may be a physical scene viewed via pass-through video, a physical scene viewed directly via optical see-through through the transparent components of the HMD, or a virtual scene viewed via one or more displays of the HMD. In some examples, device 500 does not include a display, and the 3D scenes in Figures 5A-5D and 5F-5J are physical scenes viewed directly by the user.

[0107] In Figures 5A-5D and 5F-5J, the user and device 500 exist within their respective 3D scenes. For example, the 3D scene is either a physical or augmented reality scene, and the user and device 500 physically exist within the 3D scene. In another example, the user's avatar exists within the scene. For example, if the scene is a virtual reality scene, the user's avatar exists within the virtual reality scene.

[0108] In Figure 5A, device 500 displays a view of a 3D scene including a flat surface 504. The view of the 3D scene (e.g., a live view of the 3D scene) is displayed on a camera user interface 505 (e.g., a user interface for a camera application on device 500), which includes a save for later graphical element 506 and a camera capture graphical element (e.g., a shutter button) 508. When selected, the save for later graphical element 506 causes device 500 to capture a save for later view of the current 3D scene. When selected, the camera capture graphical element 508 causes device 500 to capture a camera image of the current 3D scene.

[0109] In Figure 5A, based on the captured image data, device 500 detects a flat surface 504. Device 500 compares the representation of the image data with the representation of the gating word "flat surface," for example, according to the technique described above with respect to the proposed gating unit 364, and determines that the representations match (for example, the image data primarily depicts a flat surface). Since the representations match, device 500 does not propose any action for the flat surface 504.

[0110] In Figure 5B, device 500 displays a view of a 3D scene including the keyboard 510. Based on the captured image data, device 500 detects the keyboard 510 and sets the gating word to "no word". Device 500 compares the representation of the image data with the representation of the gating word and determines that the representations match (e.g., the image data does not contain a word), for example, according to the technique described above with respect to the gating unit 364. Since the representations match, device 500 does not offer any action suggestions for the keyboard 510, such as action suggestions for speaking the characters on the keyboard 510 via text-to-articulation conversion and for summarizing the text on the keyboard 510.

[0111] In Figure 5C, device 500 displays a view of a 3D scene including a business card 512. Business card 512-1 includes a name 512-1, a telephone number 512-2, and an email address 512-3. Based on the captured image data, device 500 detects the business card 512 and its associated text. Device 500 compares the representation of the image data to the representation of the gating word “no word” and determines that the representations do not match (e.g., the image data contains a word). Since the representations do not match, device 500 displays action suggestions 514-1 (add name 512-1 to the user's contact list), 514-2 (call telephone number 512-2), and 514-3 (send an email to email address 512-3) (e.g., action suggestion 362 determined according to the technique described above with respect to the live action suggestion unit 360) within the camera user interface 505, which includes a live view of the 3D scene in Figure 5C.

[0112] In Figure 5C, device 500 receives a touch input 516 to select a SaveForeLator graphical element 506. In response to receiving the touch input 516, device 500 captures a SaveForeLator view 518 (Figures 5L and 5Q) of the 3D scene (including the business card 512) and saves the SaveForeLator view 518 to the repository associated with the SaveForeLator application 331.

[0113] In Figure 5D, device 500 displays a view of a 3D scene including tree 521. Device 500 further displays action suggestions 520 for retrieving additional information about tree 521, for example, via image search provided by a web browser application. Device 500 determines the action suggestion 520 according to the techniques described above with respect to the live action suggestion unit 360. In Figure 5D, device 500 receives a touch input 524 to select a camera capture graphical element 508. In response to receiving the touch input 524, the device captures a camera image 522 (Figure 5K) including tree 521.

[0114] In Figure 5E, device 500 displays a web browser user interface 525. The web browser user interface 525 includes a pasta recipe 527 that requires celery and onions. The web browser user interface 525 further includes a SaveForeLetter graphical element 506. In Figure 5D, device 500 receives a touch input 526 to select the SaveForeLetter graphical element 506, and in response, device 500 saves the recipe 527 to the repository associated with the SaveForeLetter application 331.

[0115] In Figure 5F, device 500 displays a view of a 3D scene including an onion 528 and celery 530. In Figure 5F, device 500 further receives a touch input 532 to select a SaveForeLator graphical element 506. In response to receiving the touch input 532, device 500 captures a SaveForeLator view 534 (Figures 5L and 5M) of the 3D scene including the onion 528 and celery 530, and saves the SaveForeLator view 534 to the repository associated with the SaveForeLator application 331.

[0116] In Figure 5G, device 500 displays a view of a 3D scene including restaurants 536 and 538. Device 500 further receives a touch input 540 to select a SaveForeLator graphical element 506. Device 500 further detects user attention by detecting the user's gaze directed towards the location of restaurant 536 when the touch input 540 is received, for example, indicated by gaze location 542. In response to receiving the touch input 540, device 500 captures a SaveForeLator view 544 (Figures 5L and 5N) of the 3D scene including restaurants 536 and 538, and stores the SaveForeLator view 544 in a repository associated with the SaveForeLator application 331.

[0117] In Figure 5H, device 500 displays a view of a 3D scene that includes a gift shop 546. The 3D scene also includes Christmas music 548 playing in the background. Device 500 detects scene data (e.g., image and audio data) that depicts the gift shop 546 and indicates that Christmas music 548 is playing in the background.

[0118] In Figure 5I, immediately after device 500 displays a view of the 3D scene in Figure 5H, device 500 displays a view of the 3D scene including the gift 550 sold in the gift shop 546. For example, immediately after viewing the scene in Figure 5H, the user enters the gift shop 546 and is now browsing the gift 550 for potential purchase. In Figure 5I, device 500 receives a touch input 552 to select a SaveForeLator graphical element 506. In response to receiving the touch input 552, device 500 captures a SaveForeLator view 554 (Figures 5L and 5O) including the gift 550 and saves the SaveForeLator view 554 to the repository associated with the SaveForeLator application 331.

[0119] In Figure 5J, device 500 displays a view of a 3D scene including a promotional poster 556 for a movie named "Movie X". Device 500 receives a touch input 558 to select a SaveForeLator graphical element 506, and in response captures a SaveForeLator view 560 (Figure 5L) including the promotional poster 556, and saves the SaveForeLator view 560 to the repository associated with the SaveForeLator application 331.

[0120] For simplicity, Figures 5F to 5J illustrate that device 500 does not present live action suggestions determined based on each live view of each 3D scene. However, in some examples, in one or more of Figures 5F to 5J, device 500 presents one or more live action suggestions (e.g., 362) determined for one or more objects in the 3D scenes of Figures 5F to 5J. For example, in Figure 5J, while device 500 is displaying a live view of the corresponding 3D scene, device 500 displays an action suggestion to summarize the text contained in the advertising poster 556.

[0121] In the examples of Figures 5A to 5J, it is explained that the saveforator view and / or camera image is captured in response to the device 500 receiving touch input, but in other examples, the device 500 receives other types of input (e.g., voice input, gaze input, gesture input, air gesture input, and / or input via peripheral devices) to capture the saveforator view or camera image. Furthermore, as described above with respect to Figures 3A to 3C, saving the saveforator view in a repository associated with the saveforator application 331 includes saving the corresponding image data, saving a representation of the image data without saving the image data (e.g., metadata describing the content depicted by the image data), or saving both the image data and the representation of the image data.

[0122] Referring now to Figures 5K to 5S, after the interaction described in Figures 5A to 5J, for example, after the Save Forward View / camera image described has been captured, the content is displayed by device 502 in Figures 5K to 5S.

[0123] In Figure 5K, device 502 displays a user interface 562 for device 502's photo application (e.g., an application that allows the user to view and / or edit captured images and / or videos). User interface 562 includes the user's library of captured images and videos. For example, user interface 562 includes camera images 522 (e.g., captured in response to touch input 524 in Figure 5D) and other camera images 564 (e.g., other recently captured camera images). Notably, none of the save for later views 518, 534, 544, 554, or 560 are displayed in user interface 562 (and / or are not available for display).

[0124] In Figure 5L, device 502 displays the user interface 566 of the SaveForeLator application 331. The user interface 566 includes SaveForeLator views 518, 534, 544, 554, and 560, and recipes 527 (e.g., content captured in response to SaveForeLator inputs 516, 526, 532, 540, 552, and 558). Notably, the camera image 522 is not displayed in the user interface 566 (and is not available for display). Thus, Figures 5K to 5L illustrate how different user interfaces for different applications (e.g., a photo application and the SaveForeLator application 331) provide access to camera images and SaveForeLator content (e.g., SaveForeLator views), respectively. In other examples, a single application (e.g., a photo application or a save for later view application 331) provides access to both camera images (e.g., 522) and save for later views (e.g., 518, 534, 544, 554, and 560), for example, through a first user interface of the single application that provides access to camera images, and through another second user interface of the single application that provides access to save for later views. In some examples, a single user interface of a single application provides access to both camera images and save for later views. For example, the save for later view / content is displayed in the user's photo library simultaneously with the camera images.

[0125] In Figure 5L, the user interface 566 includes an instruction 568 indicating that the gift 550 has been assigned (e.g., classified) to the category “Christmas Gift”. Device 502 determines the category (e.g., as one of the categories 375) according to the techniques described above with respect to the generative action suggestion unit 370. Specifically, based on view 554 of the 3D scene in Figure 5I and scene data of the 3D scene in Figure 5H (e.g., indicating that immediately before view 554 was captured, the 3D scene included the playback of Christmas music 548 and a gift shop 546), Device 502 determines the category “Christmas Gift” for the gift 550.

[0126] In Figure 5L, device 502 receives a touch input 570 to select a SaveForeLator view 534. In response to the touch input 570, in Figure 5M, device 502 displays a user interface 572 (e.g., of the SaveForeLator application 331) including an enlarged version of the SaveForeLator view 534, such as a SaveForeLator view previously captured when, for example, device 500 and the user were present in the scene of Figure 5E. In Figures 5L-5S, the user and device 502 are not present in the 3D scenes of Figures 5A-5D and Figures 5F-5J. For example, the user is instead at home and using device 502 to review previously captured SaveForeLator content (e.g., a SaveForeLator view) for action suggestions.

[0127] In Figure 5M, the user interface 572 includes an action suggestion 574-1 for adding onions 528 and celery 530 to the user's shopping list, and an action suggestion 574-2 for initiating an order for onions and celery via a grocery delivery application. Device 502 determines action suggestions 574-1 and 574-2 (for example, as two of action suggestions 374) in accordance with the techniques described above with respect to the generative action suggestion unit 370. Specifically, device 502 determines action suggestions 574-1 and 574-2 based on the saveforlater view 534 and the recipe 527. For example, based on a saved recipe 527 that requires celery and onions, device 502 infers that the actions “Add to shopping list” and “Order from grocery store” are appropriate for the saveforlater view 534 which includes onions 528 and celery 530.

[0128] In Figure 5N, device 502 displays a user interface 576 (e.g., of the SaveForeLator application 331) including an enlarged view of the SaveForeLator view 544. Device 502 displays the user interface 576 in response to receiving user input to select the SaveForeLator view 544 in Figure 5L. The user interface 576 includes an action suggestion 578 for ordering food at restaurant 536. Device 502 determines the action suggestion 578 (e.g., as one of the action suggestions 374) according to the techniques described above with respect to the generative action suggestion unit 370. Specifically, device 502 determines the action suggestion 578 based on the SaveForeLator view 544, user attention data (Figure 5G) indicating that the user's gaze was directed towards restaurant 536 when the touch input 540 was received (e.g., as indicated by gaze location 542), the location of device 500 when the touch input 540 was received, and the current location of device 502. For example, based on gaze data indicating that the user is more interested in restaurant 536 (than restaurant 538), and location data indicating that the current location of device 502 is near the location of restaurant 536, device 500 infers that ordering food from restaurant 536 (instead of restaurant 538) is reasonable for saveforlater view 544.

[0129] In Figure 5O, device 502 displays a user interface 580 (e.g., of the SaveForeLator application 331) which includes an enlarged view of the SaveForeLator view 554. Device 502 displays the user interface 580 in response to receiving user input to select the SaveForeLator view 554 in Figure 5L. The user interface 580 includes an action suggestion 582 for purchasing a gift 550, for example, from an online retailer. Device 502 determines the action suggestion 582 (e.g., as one of the action suggestions 374) according to the techniques described above with respect to the generative action suggestion unit 370. Specifically, device 502 determines the action suggestion 582 based on the SaveForeLator view 554 and scene data of the 3D scene in Figure 5H (e.g., scene data depicting the gift shop 546).

[0130] In Figure 5O, device 502 receives a touch input 584 to select an action proposal 582. In Figure 5P, device 502 responds to the touch input 584 by initiating the action proposal 582. Specifically, in Figure 5P, device 502 displays a user interface 586 for an online retailer application, allowing the user to purchase a gift 550.

[0131] In Figure 5Q, device 502 displays a user interface 587 (e.g., of the SaveForator application 331) which includes an enlarged view of the SaveForator view 518. Device 502 displays the user interface 587 in response to receiving user input to select the SaveForator view 518 in Figure 5L. The user interface 587 includes an action suggestion 588 for sending Peter B a message to remind him to water the plants. Device 502 determines the action suggestion 588 (e.g., as one of the action suggestions 374) in accordance with the techniques described above with respect to the generative action suggestion unit 370. Specifically, device 502 determines the action suggestion 588 based on the SaveForator view 518 and reminder data indicating that the user has a reminder to “ask Peter B to water the plants.” Device 502, for example, infers an action to send a message to Peter B to remind him to water the plants, based on the reminder data and a save-for-later view 518 depicting Peter B's business card 512.

[0132] Figures 5C and 5Q illustrate that while the user is provided with a live view of the 3D scene, a first type of action suggestion (e.g., 514-1, 514-2, 514-3, and / or 362) is presented for an object in the 3D scene (e.g., 512), whereas a second type of action suggestion (e.g., 588 and / or 374) for the same object is not presented while the user is provided with a live view of the 3D scene. Instead, the second type of action suggestion is determined and / or presented later, for example, after device 500 receives a saveforator input (e.g., 516) about that object, after the corresponding view (e.g., 518) is saved in the repository, and while the user is browsing the user interface (e.g., 587) of the saveforator application 331. As explained above with respect to the generative action suggestion unit 370, the second type of action suggestion may take longer to decide compared to the first type of action suggestion, and therefore, the second type of action suggestion may not be presented while the user is provided with a live view of the 3D scene. For example, action suggestions 514-1, 514-2, and 514-3 in Figure 5C are relatively generic actions for business card 512, which are decided through a relatively inexpensive computation process (for example, as explained with respect to the live action suggestion unit 360). In contrast, action suggestion 588 in Figure 5Q is a personalized action for business card 512, which is decided through a relatively computationally expensive process (for example, as explained with respect to the generative action suggestion unit 370).

[0133] In Figure 5R, device 502 displays a user interface 590 (e.g., a home screen user interface) that is different from the user interface of the Save Fore Later application 331. On the user interface 590, device 502 displays an action suggestion 592 to remind Peter B to water the plants (e.g., as a notification graphical element). Device 502 determines the action suggestion 592 in the same way that it determined the action suggestion 588 (e.g., based on the Save Fore Later view 518 and context information 372).

[0134] Figures 5Q to 5R illustrate how several action suggestions (e.g., 574-1, 574-2, 578, 582, and / or 588) are presented in the user interface of the saveforwarder application 331 (e.g., 572, 576, 580, and / or 587), while other action suggestions (e.g., 592) are presented in other user interfaces (e.g., 590), such as the home screen user interface, the lock screen user interface, and / or the user interface of another application. In some examples, device 502 presents an action proposal within its different user interface (e.g., as a notification on the home screen, lock screen, etc.) if an action proposal meets one or more criteria, such as when the action proposal has a confidence score exceeding a threshold (e.g., determined by generative model 371), when the same action proposal has been determined more than a threshold number of times by generative action proposal unit 370, when the action proposal is initiated by a given application, when the action proposal contacts a given contact (e.g., a person), and / or when the action proposal has an urgency value exceeding a threshold (e.g., determined by generative model 371).

[0135] In Figures 5R to 5S, the generative action suggestion unit 370 has already determined user information (e.g., 376) indicating that the user is interested in movie X, based on the save for later view 560 (Figure 5J). In Figure 5R, device 502 receives touch input 596 to select the icon 594 of the movie streaming application. In Figure 5S, device 502 displays the movie streaming user interface 598 in response to receiving touch input 596. The movie streaming application user interface 598 includes a suggestion 599 for the user to watch a movie named movie X. To display the suggestion 599, the movie streaming application has already requested user information from the save for later application 331 (e.g., via an API call via API 332).

[0136] Other applications may request user information (e.g., 376) in a similar manner. For example, an online shopping application may request information from a saveforwarder application 331 indicating products and / or product categories that the user may be interested in (e.g., to suggest products and / or product categories for purchase). Another example is a music application requesting information from a saveforwarder application 331 indicating the user's media preferences (e.g., to create a personalized playlist for the user). In some examples, an application requests user information via API calls through API 332. In some examples, an application requests user information periodically (e.g., once a day, once an hour, etc.), at installation, at installation of application updates, at application startup (e.g., every time the application is started), and / or when it receives user input instructing the application to retrieve such information. As described above with respect to API 332, the SaveForeLetter application 331 may provide user information (e.g., 376) to the requesting application in a privacy-preserving manner, for example, by providing abstracted and / or generalized user information but not SaveForeLetter views and / or content (e.g., 554, 518, 527, 570, 544, and 560).

[0137] Additional explanations regarding Figures 5A to 5S are provided below with reference to Method 600, which is described in relation to Figure 6.

[0138] Figure 6 is a flowchart illustrating method 600 for providing action suggestions in various examples. In some examples, method 600 is executed in a first computer system (e.g., a first instance of computer system 101, device 500, and / or user-responsive component 120) in communication with one or more image sensors, and in a second computer system (e.g., a second instance of computer system 101 and / or device 502) in communication with display generation components. In some examples, method 600 is governed by instructions stored in one or more non-temporary (or temporary) computer-readable storage media and executed by one or more processors of the first computer system and / or the second computer system, such as 202 and / or 302. In some examples, the operation of method 600 is distributed across multiple computer systems, such as a first computer system, a second computer system, and a separate server system. Some operations of method 600 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0139] Method 600 includes receiving user input (e.g., saveforator input) (e.g., 516, 532, 540, 552, and / or 558) in a first computer system in response to a request to save objects (e.g., 512, 528, 530, 536, 550, and / or 556) in a three-dimensional (3D) scene (602).

[0140] Method 600 includes capturing a view of a 3D scene (e.g., 554, 518, 534, 544, and / or 560) via one or more image sensors in a first computer system in response to receiving user input corresponding to a request to save objects in a 3D scene (604).

[0141] Method 600 includes, in a second computer system, displaying a first user interface of a first application (e.g., 572, 576, 580, and / or 587) via a display generation component after a view of a 3D scene has been captured via one or more image sensors (606), wherein displaying the first user interface of the first application includes displaying first proposed graphical elements (e.g., 574-1, 574-2, 578, 582, and / or 588) (608), and the first proposed graphical The physical element is selectable to cause a second computer system to execute a first action proposal (e.g., 374) via a second application (e.g., 333) different from the first application, the first action proposal being determined based on performing image recognition on a view of a 3D scene (e.g., 554, 518, 534, 544, and / or 560) and on contextual information (e.g., 373) different from the view of the 3D scene (e.g., by a generative action proposal unit 370).

[0142] In some examples, Method 600 includes, in a second computer system, receiving user input (e.g., 584) corresponding to the selection of a first proposed graphical element (e.g., 582) while displaying a first user interface (e.g., 580) of a first application via a display generation component, and initiating a first action proposal via the second application (e.g., as illustrated in Figure 5P) in response to receiving the user input corresponding to the selection of the first proposed graphical element.

[0143] In some examples, the first computer system is the second computer system.

[0144] In some cases, the first computer system is different from the second computer system.

[0145] In some examples, user inputs corresponding to a request to save an object in a 3D scene (e.g., 516, 532, 540, 552, and / or 558) correspond to a first type of user input (e.g., saveforator input), and method 600 includes receiving a user input corresponding to a request to capture an image of a 3D scene (e.g., 524) in a first computer system, which corresponds to a second type of user input (e.g., camera capture input) different from the first type of user input, and capturing individual images of the 3D scene (e.g., 522) via one or more image sensors in response to receiving the user input corresponding to a request to capture an image of a 3D scene.

[0146] In some examples, method 600 includes simultaneously displaying a save graphical element (e.g., 506) and an image capture graphical element (e.g., 508) in a first computer system, where user input corresponding to a request to save an object in a 3D scene corresponds to the selection of the save graphical element, and user input corresponding to a request to capture an image of the 3D scene (e.g., 524) corresponds to the selection of the image capture graphical element.

[0147] In some examples, Method 600 includes displaying a second user interface (e.g., 566) of a first application via a display generation component in a second computer system, which includes displaying a view of a 3D scene (e.g., 554, 518, 534, 544, and / or 560), wherein individual images of the 3D scene (e.g., 522) are not available for display in the second user interface of the first application; and displaying a first user interface (e.g., 562) of a third application different from the first application via a display generation component, which includes displaying individual images of a 3D scene (e.g., 522), wherein views of the 3D scene (e.g., 554, 518, 534, 544, and / or 560) are not available for display in the first user interface of the third application. In some examples, none of the images captured in response to receiving a second type of user input are available for display within the second user interface of the first application. In some examples, none of the images captured in response to receiving a second type of user input are available for display through any user interface of the first application. In some examples, none of the views captured in response to receiving a first type of user input are available for display in the first user interface of the third application. In some examples, none of the views captured in response to receiving a first type of user input are available for display through any user interface of the third application.

[0148] In some examples, method 600 includes displaying a third user interface (e.g., 566) of a first application (e.g., saveforwarder application 331) in a second computer system via a display generation component, which includes displaying a view of a 3D scene, wherein individual images of the 3D scene are not available for display in the third user interface of the first application; and displaying a fourth user interface of the first application via a display generation component, which is different from the third user interface of the first application, which includes displaying individual images of a 3D scene, wherein the view of the 3D scene is not available for display in the fourth user interface of the first application. In some examples, any image captured in response to receiving a second type of user input is not available for display in the third user interface of the first application. In some examples, none of the views captured in response to receiving the first type of user input are available for display within the fourth user interface of the first application.

[0149] In some examples, contextual information (e.g., 373) includes information personal to the user of the first computer system and / or the second computer system.

[0150] In some examples, user-personal information includes a second object (e.g., 554, 518, 527, 534, 544, 560, 512, 528, 530, 536, 538, 550, and / or 556) associated with a first application (e.g., stored in a repository associated with the first application), and the second object is distinct from the object in the 3D scene.

[0151] In some examples, contextual information (e.g., 373) includes the location of a first computer system when user input (e.g., 516, 532, 540, 552, and / or 558) was received in response to a request to save an object in a 3D scene.

[0152] In some examples, method 600 further includes capturing data representing a 3D scene (e.g., the 3D scene in Figure 5H) in a first computer system, wherein the data representing the 3D scene is different from a view of the 3D scene (e.g., 554), and contextual information (e.g., 373) includes the data representing the 3D scene.

[0153] In some examples, method 600 further includes detecting the attention of a user of the first computer system (as described with respect to Figure 5G, for example), and the contextual information (e.g., 373) includes the attention of a user of the first computer system.

[0154] In some examples, method 600 further includes, in a second computer system, displaying an instruction (e.g., 568) that an object (e.g., 550) is assigned to a first category (e.g., category 375) via a display generation component, where the first category is determined based on contextual information (e.g., by a generative action suggestion unit 370).

[0155] In some examples, displaying a first user interface of a first application (e.g., 572) further includes displaying a second proposed graphical element (e.g., 574-1 and / or 574-2), the second proposed graphical element being selectable to cause a second computer system to execute a second action proposal via a fourth application (e.g., 333), the second action proposal being determined based on performing image recognition on a view of a 3D scene (e.g., 534) and on contextual information (e.g., 373).

[0156] In some cases, the fourth application is different from the second application.

[0157] In some examples, Method 600 further includes detecting individual objects (e.g., 512) in a 3D scene via one or more image sensors before receiving user input (e.g., 516, 532, 540, 552, and / or 558) in a first computer system corresponding to a request to save objects in a 3D scene; presenting third action proposals (e.g., 362, 514-1, 514-2, and / or 514-3) determined based on the individual object (e.g., by the Live Action Proposal Unit 360) in response to the detection of individual objects in the 3D scene via one or more image sensors, in accordance with the determination that the proposed criteria are met; and ceasing to present the third action proposals in accordance with the determination that the proposed criteria are not met.

[0158] In some examples, method 600 further includes displaying a camera user interface (e.g., 505) in a first computer system, and presenting a third action suggestion determined based on individual objects, or displaying a third action suggestion in the camera user interface.

[0159] In some examples, a third action proposal (e.g., 362, 514-1, 514-2, and / or 514-3) is determined based on analyzing the view 3D scene using a first type of process (e.g., one or more of the processes described above with respect to the live action proposal unit 360).

[0160] In some examples, the first action proposal (e.g., 374, 574-1, 574-2, 578, 582, and / or 588) is determined based on analyzing the view of the 3D scene using a second type of process different from the first type of process (e.g., one or more processes described above with respect to the generative action proposal unit 370).

[0161] In some examples, a view of a 3D scene (e.g., 518) is not analyzed using a second type of process until after the view of the 3D scene has been captured and a first representation of the view of the 3D scene (and / or the view of the 3D scene) has been saved in association with a first application (e.g., saved in a repository associated with the first application) (as illustrated, for example, with respect to Figures 5C and 5Q).

[0162] In some examples, a representation of one or more sets of words, which describes a condition about a view of a 3D scene and, when the condition is met, prevents the presentation of action suggestions for the view of the 3D scene, is compared (e.g., by the suggestion gating unit 364) to a second representation of the view of the 3D scene (e.g., the 3D scenes in Figures 5A, 5B, and / or 5C), where the suggestion criterion is met if the representation of the set of words does not match the second representation of the view of the 3D scene, and the suggestion criterion is not met if the representation of the set of words matches the second representation of the view of the 3D scene.

[0163] In some examples, a set of one or more words is determined based on the view of the 3D scene (for example, by the suggestion gating unit 364).

[0164] In some examples, a second representation of the 3D scene view is generated according to a determination that the 3D scene view satisfies a set of stability criteria (e.g., by proposed gating unit 364).

[0165] In some examples, method 600 includes, in a second computer system, receiving user input (e.g., 526) in response to a request to save a content object (e.g., 527) in association with a first application (e.g., saveforwarder application 331) while the content object (e.g., 527) is being displayed via a display generation component, and, in response to receiving user input in response to a request to save a content object in association with a first application, causing the content object to be saved in association with a first application (e.g., saved in a repository associated with the first application).

[0166] In some examples, Method 600 includes, in a second computer system, after a view of a 3D scene (e.g., 518) has been captured via one or more image sensors (and optionally stored in a repository associated with a first application), displaying a separate proposal graphical element (e.g., 592) corresponding to a fourth action proposal in a proposal user interface (e.g., 590) different from the first user interface of the first application via a display generating component, according to the determination that the fourth action proposal, which is determined based on performing image recognition on the view of the 3D scene and / or based on contextual information (e.g., 373), satisfies a set of proposal presentation criteria.

[0167] In some examples, the second computer system includes an application programming interface (API) (e.g., 332) that enables a requesting application (e.g., 333) to communicate with a first application (e.g., 331), and method 600 includes, in the second computer system, the requesting application sending an API call to the first application via the API; in response to receiving the API call from the requesting application, the first application providing user information (e.g., 376) determined based on one or more objects (e.g., 560 and / or 556) stored in association with the first application (e.g., stored in a repository associated with the first application); and displaying the user interface of the requesting application, including, after the requesting application has received the user information, receiving user input (e.g., 596) corresponding to a request to display the user interface of the requesting application (e.g., 598); and in response to receiving the user input corresponding to a request to display the user interface of the requesting application, displaying a suggestion (e.g., 599) determined based on the user information.

[0168] In some examples, individual representations of a 3D scene view (and / or a 3D scene view) are stored in a repository associated with a first application (e.g., 331), and the first user interface of the first application is displayed after the individual representations of the 3D scene view (and / or a 3D scene view) have been stored in the repository associated with the first application.

[0169] The above is written with reference to specific embodiments for illustrative purposes. However, the above exemplary discussion is not intended to be exhaustive or to limit the invention to the exact form disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the invention and its practical applications, thereby enabling other persons skilled in the art to best use the invention and the various described embodiments with various modifications suitable for specific applications that may be conceived.

[0170] As described above, one aspect of the technology involves collecting and using data available from various sources to present action suggestions to the user. This disclosure suggests that in some cases, such collected data may include personal information data that uniquely identifies a particular person, or personal information data that can be used to contact a particular person or locate them. Such personal information data may include demographic data, location-based data, telephone numbers, email addresses, Twitter® IDs, home addresses, data or records relating to a user's health or fitness level (e.g., vital signs measurements, medication information, exercise information), date of birth, or any other identifying or personal information.

[0171] This disclosure acknowledges that such use of personal data in the technology may be for the benefit of the user. For example, personal data can be used to provide accurate and / or reasonable action suggestions. Furthermore, other uses of personal data that may benefit the user are also conceivable in this disclosure. For example, health and fitness data can be used to provide insights into the user's overall wellness, or as positive feedback to individuals using the technology to pursue wellness goals.

[0172] This disclosure implies that entities involved in the collection, analysis, disclosure, transmission, storage, or other use of such personal data should adhere to a robust privacy policy and / or privacy practices. Specifically, such entities should implement and consistently use a privacy policy and practices that are generally recognized as meeting or exceeding industry or government requirements for the strict confidentiality of personal data. Such policies should be readily accessible to users and should be updated as data collection and / or use changes. Personal data from users should be collected for the lawful and legitimate use of the entity and should not be shared or sold for any other purpose. Furthermore, such collection / sharing should be carried out only after informing and obtaining the user's consent. In addition, such entities should consider taking all necessary steps to protect and secure access to such personal data and to ensure that others with access to the personal data faithfully adhere to those privacy policies and procedures. Furthermore, such entities may undergo third-party evaluations to demonstrate their compliance with widely accepted privacy policies and practices. Furthermore, policies and practices should be adapted to the specific types of personal data collected and / or accessed, and to applicable laws and standards, including jurisdiction-specific considerations. For example, in the United States, the collection or access to certain health data may be subject to federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA). Health data in other countries, on the other hand, may be subject to other regulations and policies and should be addressed accordingly. Therefore, different privacy practices should be maintained in each country with respect to different types of personal data.

[0173] Notwithstanding the foregoing, the Disclosure also envisions embodiments that allow a user to selectively prevent the use of or access to personal data. That is, the Disclosure envisions that hardware and / or software elements may be provided to prevent or prevent access to such personal data. For example, when presenting an action offer to a user, the Technology may be configured to allow the user to choose to “opt in” or “opt out” of participating in the collection of personal data during or at any time thereafter. In another example, the user may choose not to provide the personal data on which the action offer was determined. In yet another example, the user may choose to limit the length of time such data is retained. In addition to providing “opt-in” and “opt-out” options, the Disclosure envisions providing notices regarding access to or use of personal data. For example, the user may be notified when downloading an app that will access the user’s personal data, and then reminded again immediately before the app accesses the personal data.

[0174] Furthermore, the intent of this disclosure is that personal data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Risks can be minimized by limiting data collection and deleting data when it is no longer needed. In addition, where applicable in certain health-related applications, data anonymization can be used to protect user privacy. Anonymization can be facilitated by removing certain identifiers (e.g., date of birth), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data across users), and / or by other means, where necessary.

[0175] Therefore, while this disclosure broadly covers the use of personal data to implement one or more different disclosed embodiments, it is also conceivable that these different embodiments could be implemented without requiring access to such personal data. In other words, the different embodiments of the technology would not be rendered inoperable by the absence of all or part of such personal data. For example, action suggestions could be generated based on non-personal data or a minimal amount of personal information, such as content requested by a device associated with the user, other non-personal information available for services, or publicly available information.

[0176] [Section 1] It is a method, In a first computer system that is in communication with one or more image sensors, Receiving user input in response to a request to save an object in a 3D scene, and In response to receiving the user input corresponding to the request to save the objects in the 3D scene, capture a view of the 3D scene via one or more image sensors, and In a second computer system that is in communication with a display generation component, The process includes, after the view of the 3D scene is captured via one or more image sensors, displaying a first user interface of a first application via the display generation component, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. The first proposed action is determined based on performing image recognition on the view of the 3D scene and on contextual information different from that of the view of the 3D scene. method. [Section 2] In the second computer system which is in communication with the aforementioned display generation component, While displaying the first user interface of the first application via the display generation component, the system receives user input corresponding to the selection of the first proposed graphical element, In response to receiving the user input corresponding to the selection of the first proposed graphical element, the first action proposal is initiated via the second application, The method described in item 1, further comprising: [Section 3] The method according to any one of items 1 to 2, wherein the first computer system is the second computer system. [Section 4] The method according to any one of items 1 to 2, wherein the first computer system is different from the second computer system. [Section 5] The user input corresponding to the request to save the object in the 3D scene corresponds to a first type of user input, and the method In the first computer system that is in communication with one or more image sensors, Receiving user input corresponding to a request to capture an image of the 3D scene, which corresponds to a second type of user input different from the first type of user input, The method according to any one of claims 1 to 4, further comprising capturing individual images of the 3D scene via one or more image sensors in response to receiving the user input corresponding to the request to capture an image of the 3D scene. [Section 6] In the first computer system that is in communication with one or more image sensors, The further includes simultaneously displaying a save graphical element and an image capture graphical element, wherein the user input corresponding to the request to save the object in the 3D scene corresponds to the selection of the save graphical element, and the user input corresponding to the request to capture an image of the 3D scene corresponds to the selection of the image capture graphical element. The method described in item 5. [Section 7] In the second computer system which is in communication with the aforementioned display generation component, Displaying the second user interface of the first application via the display generation component, which includes displaying the view of the 3D scene, wherein the individual images of the 3D scene are not available for display in the second user interface of the first application, Displaying a first user interface of a third application different from the first application via the display generation component, which includes displaying the individual images of the 3D scene, wherein the view of the 3D scene is not available for display in the first user interface of the third application, and displaying the first user interface of the third application. The method described in any one of items 5 to 6, further including the method described in any one of items 5 to 6. [Section 8] In the second computer system which is in communication with the aforementioned display generation component, Displaying a third user interface of the first application via the display generation component, which includes displaying the view of the 3D scene, wherein the individual images of the 3D scene are not available for display in the third user interface of the first application, Displaying a fourth user interface of the first application via the display generation component, which is different from the third user interface of the first application, and includes displaying the individual images of the 3D scene, wherein the view of the 3D scene is not available for display in the fourth user interface of the first application, The method described in any one of paragraphs 5 to 7, further including the method described in any one of paragraphs 5 to 7. [Section 9] The method according to any one of paragraphs 1 to 8, wherein the context information includes information personal to the users of the first computer system and / or the second computer system. [Section 10] The method according to paragraph 9, wherein the personal information for the user includes a second object stored in association with the first application, and the second object is different from the object in the 3D scene. [Section 11] The method according to any one of claims 1 to 10, wherein the context information includes the location of the first computer system when the user input corresponding to the request to save the object in the 3D scene was received. [Section 12] In the first computer system that is in communication with one or more image sensors, The further includes capturing data representing the 3D scene, wherein the data representing the 3D scene differs from the view of the 3D scene, and the context information includes the data representing the 3D scene. The method described in any one of items 9 to 11. [Section 13] In the first computer system that is in communication with one or more image sensors, Further including detecting the attention of the user of the first computer system, the context information includes the attention of the user of the first computer system. The method described in any one of items 1 to 12. [Section 14] In the second computer system which is in communication with the aforementioned display generation component, The display generation component further includes displaying an indication that the object is assigned to a first category, the first category being determined based on the context information. The method described in any one of items 1 to 13. [Section 15] Displaying the first user interface of the first application further includes displaying the second proposed graphical element. The second proposed graphical element is selectable to cause the second computer system to execute the second action proposal via the fourth application. The second action proposal is determined based on performing image recognition on the view of the 3D scene and on the context information. The method described in any one of items 1 to 14. [Section 16] The method according to paragraph 15, wherein the fourth application is different from the second application. [Section 17] In the first computer system that is in communication with one or more image sensors, Before receiving the user input corresponding to the request to save the object in the 3D scene, To detect individual objects in the 3D scene via one or more image sensors, and, In response to detecting the individual objects in the 3D scene via one or more image sensors, In accordance with the determination that the proposal criteria are met, present a third action proposal determined based on the individual object, and In accordance with the determination that the aforementioned proposal criteria are not met, the submission of the third action proposal will be withdrawn. The method described in any one of items 1 to 16, further including the method described in any one of items 1 to 16. [Section 18] In the first computer system that is in communication with one or more image sensors, Further including displaying a camera user interface and presenting the third action suggestion determined based on the individual objects, including displaying the third action suggestion in the camera user interface, The method described in item 17. [Section 19] The method according to any one of paragraphs 17 to 18, wherein the third proposed action is determined based on analyzing the view 3D scene using the first type of process. [Section 20] The method according to paragraph 19, wherein the first action proposal is determined based on analyzing the view of the 3D scene using a second type of process different from the first type of process. [Section 21] The view of the 3D scene is not analyzed using a second type of process until after the view of the 3D scene has been captured and the first representation of the view of the 3D scene has been saved in association with the first application. The method described in item 20. [Section 22] A set of one or more words is used to describe a condition for the view of the 3D scene, and when the condition is met, the presentation of an action suggestion for the view of the 3D scene is prevented. This set of words is then compared to a second representation of the view of the 3D scene. The proposed criterion is satisfied when the representation of the set of one or more words does not match the second representation of the view of the 3D scene. The proposed criterion is not satisfied if the representation of the set of one or more words matches the second representation of the view of the 3D scene. The method described in any one of items 17 to 21. [Section 23] The method according to paragraph 22, wherein the set of one or more words is determined based on the view of the 3D scene. [Section 24] The method according to any one of items 22 to 23, wherein the second representation of the view of the 3D scene is generated in accordance with the determination that the view of the 3D scene satisfies a set of stability criteria. [Section 25] In the second computer system which is in communication with the aforementioned display generation component, While displaying a content object via the aforementioned display generation component, the system receives user input corresponding to a request to associate and save the content object with the first application, In response to receiving the user input corresponding to the request to save the content object in association with the first application, the representation of the content object is saved in association with the first application, The method described in any one of items 1 to 24, further including the method described in any one of items 1 to 24. [Section 26] In the second computer system which is in communication with the aforementioned display generation component, After the view of the 3D scene is captured via one or more image sensors, a fourth action proposal is made, based on performing image recognition on the view of the 3D scene and / or based on the contextual information, in accordance with the determination that the fourth action proposal satisfies a set of proposal presentation criteria, the display generation component is used to display a separate proposal graphical element corresponding to the fourth action proposal in a proposal user interface different from the first user interface of the first application. The method described in any one of items 1 to 25, further including the method described in any one of items 1 to 25. [Section 27] The second computer system includes an application programming interface (API) that enables a requesting application to communicate with the first application, and the method is In the second computer system which is in communication with the aforementioned display generation component, The requesting application sends an API call to the first application via the API. In response to receiving the API call from the requesting application, the first application provides, via the API, user information determined based on one or more objects stored in association with the first application, and After the requesting application receives the user information, Receiving user input corresponding to a request to display the user interface of the requesting application, and The method according to any one of claims 1 to 26, further comprising displaying the user interface of the requesting application, including displaying a suggestion determined based on the user information, in response to receiving the user input corresponding to the request to display the user interface of the requesting application. [Section 28] The method according to any one of items 1 to 27, wherein individual representations of the views of the 3D scene are stored in a repository associated with the first application, and the first user interface of the first application is displayed after the individual representations of the views of the 3D scene are stored in the repository associated with the first application. [Section 29] One or more non-temporary computer-readable storage media for storing one or more programs configured to be executed by one or more processors of one or more computer systems that are in communication with one or more image sensors and display generation components, wherein the one or more programs include instructions for performing the method described in any one of items 1 to 28. [Section 30] One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more processors, A computer system comprising one or more memories storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method described in any one of the items 1 to 28. [Section 31] One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more computer systems comprising means for performing the method described in any one of paragraphs 1 to 28. [Section 32] One or more non-temporary computer-readable storage media storing one or more programs configured to be executed by one or more processors of one or more computer systems that are in communication with one or more image sensors and display generation components, wherein the one or more programs are A first computer system among the one or more computer systems receives user input corresponding to a request to save an object in a three-dimensional (3D) scene. In response to receiving the user input corresponding to the request to save the objects in the 3D scene, the first computer system captures a view of the 3D scene via one or more image sensors. After the view of the 3D scene is captured via one or more image sensors, a second computer system among the one or more computer systems provides instructions for displaying the first user interface of the first application via the display generation component, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more non-temporary computer-readable storage media, wherein the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene. [Section 33] One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more processors, The system comprises one or more memories that store one or more programs configured to be executed by the one or more processors, A first computer system among the one or more computer systems receives user input corresponding to a request to save an object in a three-dimensional (3D) scene. In response to receiving the user input corresponding to the request to save the objects in the 3D scene, the first computer system captures a view of the 3D scene via one or more image sensors. After the view of the 3D scene is captured via one or more image sensors, a second computer system among the one or more computer systems provides instructions for displaying the first user interface of the first application via the display generation component, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more computer systems, wherein the first proposed action is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from that of the view of the 3D scene. [Section 34] One or more computer systems configured to communicate with one or more image sensors and display generation components, A means for receiving user input corresponding to a request to save an object in a three-dimensional (3D) scene by a first computer system among the one or more computer systems, In response to receiving the user input corresponding to the request to save the objects in the 3D scene, the first computer system provides means for capturing a view of the 3D scene via one or more image sensors, The system comprises means for displaying the first user interface of the first application via the display generation component by a second computer system among the one or more computer systems, after the view of the 3D scene has been captured by the first computer system via one or more image sensors, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more computer systems, wherein the first proposed action is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from that of the view of the 3D scene.

Claims

1. It is a method, In a first computer system that is in communication with one or more image sensors, Receiving user input in response to a request to save an object in a three-dimensional (3D) scene, and In response to receiving the user input corresponding to the request to save the objects in the 3D scene, capture a view of the 3D scene via one or more image sensors, and In a second computer system that is in communication with a display generation component, The process includes, after the view of the 3D scene is captured via one or more image sensors, displaying a first user interface of a first application via the display generation component, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. The first action proposal is determined based on performing image recognition on the view of the 3D scene and on contextual information different from that of the view of the 3D scene. method.

2. In the second computer system which is in communication with the display generation component, While displaying the first user interface of the first application via the display generation component, user input corresponding to the selection of the first proposed graphical element is received, In response to receiving the user input corresponding to the selection of the first proposed graphical element, the first action proposal is initiated via the second application. The method according to claim 1, further comprising:

3. The method according to any one of claims 1 to 2, wherein the first computer system is the second computer system.

4. The method according to any one of claims 1 to 2, wherein the first computer system is different from the second computer system.

5. The user input corresponding to the request to save the object in the 3D scene corresponds to a first type of user input, and the method In the first computer system which is in communication with one or more image sensors, Receiving user input corresponding to a request to capture an image of the 3D scene, which corresponds to a second type of user input different from the first type of user input, The method according to any one of claims 1 to 4, further comprising capturing individual images of the 3D scene via one or more image sensors in response to receiving the user input corresponding to the request to capture an image of the 3D scene.

6. In the first computer system which is in communication with one or more image sensors, The further includes simultaneously displaying a save graphical element and an image capture graphical element, wherein the user input corresponding to the request to save the object in the 3D scene corresponds to the selection of the save graphical element, and the user input corresponding to the request to capture an image of the 3D scene corresponds to the selection of the image capture graphical element. The method according to claim 5.

7. In the second computer system which is in communication with the display generation component, Displaying the second user interface of the first application via the display generation component, which includes displaying the view of the 3D scene, wherein the individual images of the 3D scene are not available for display in the second user interface of the first application, Displaying a first user interface of a third application different from the first application via the display generation component, including displaying the individual images of the 3D scene, wherein the view of the 3D scene is not available for display in the first user interface of the third application, and displaying the first user interface of the third application. The method according to any one of claims 5 to 6, further comprising:

8. In the second computer system which is in communication with the display generation component, Displaying the third user interface of the first application via the display generation component, which includes displaying the view of the 3D scene, wherein the individual images of the 3D scene are not available for display in the third user interface of the first application, Displaying a fourth user interface of the first application via the display generation component, which is different from the third user interface of the first application, and includes displaying the individual images of the 3D scene, wherein the view of the 3D scene is not available for display in the fourth user interface of the first application, The method according to any one of claims 5 to 7, further comprising:

9. The method according to any one of claims 1 to 8, wherein the context information includes information personal to the user of the first computer system and / or the second computer system.

10. The method according to claim 9, wherein the personal information for the user includes a second object stored in association with the first application, and the second object is different from the object in the 3D scene.

11. The method according to any one of claims 1 to 10, wherein the context information includes the location of the first computer system when the user input corresponding to the request to save the object in the 3D scene was received.

12. In the first computer system which is in communication with one or more image sensors, The further includes capturing data representing the 3D scene, wherein the data representing the 3D scene differs from the view of the 3D scene, and the context information includes the data representing the 3D scene. The method according to any one of claims 9 to 11.

13. In the first computer system which is in communication with one or more image sensors, The context information further includes detecting the attention of the user of the first computer system, wherein the context information includes the attention of the user of the first computer system. The method according to any one of claims 1 to 12.

14. In the second computer system which is in communication with the display generation component, The display generation component further includes displaying an indication that the object is assigned to a first category, the first category being determined based on the context information. The method according to any one of claims 1 to 13.

15. Displaying the first user interface of the first application further includes displaying the second proposed graphical element, The second proposed graphical element is selectable to cause the second computer system to execute the second action proposal via the fourth application. The second action proposal is determined based on performing image recognition on the view of the 3D scene and on the context information. The method according to any one of claims 1 to 14.

16. The method according to claim 15, wherein the fourth application is different from the second application.

17. In the first computer system which is in communication with one or more image sensors, Before receiving the user input corresponding to the request to save the object in the 3D scene, To detect individual objects within the 3D scene via one or more image sensors, and In response to detecting the individual objects in the 3D scene via one or more image sensors, In accordance with the determination that the proposal criteria are met, a third action proposal will be presented, which will be determined based on the individual objects, and In accordance with the determination that the aforementioned proposal criteria are not met, the submission of the third action proposal will be withdrawn. The method according to any one of claims 1 to 16, further comprising:

18. In the first computer system which is in communication with one or more image sensors, Further including displaying a camera user interface and presenting the third action proposal determined based on the individual objects, including displaying the third action proposal in the camera user interface, The method according to claim 17.

19. The method according to any one of claims 17 to 18, wherein the third action proposal is determined based on analyzing the view 3D scene using a first type of process.

20. The method according to claim 19, wherein the first action proposal is determined based on analyzing the view of the 3D scene using a second type of process different from the first type of process.

21. The view of the 3D scene is not analyzed using a second type of process until after the view of the 3D scene has been captured and the first representation of the view of the 3D scene has been saved in association with the first application. The method according to claim 20.

22. A set of one or more words, which describes a condition for the view of the 3D scene and prevents the presentation of an action suggestion for the view of the 3D scene when the condition is met, is compared with a second representation of the view of the 3D scene. The proposed criterion is satisfied when the representation of the set of one or more words does not match the second representation of the view of the 3D scene. The proposed criterion is not satisfied if the representation of the set of one or more words matches the second representation of the view of the 3D scene. The method according to any one of claims 17 to 21.

23. The method according to claim 22, wherein the set of one or more words is determined based on the view of the 3D scene.

24. The method according to any one of claims 22 to 23, wherein the second representation of the view of the 3D scene is generated in accordance with a determination that the view of the 3D scene satisfies a set of stability criteria.

25. In the second computer system which is in communication with the display generation component, While displaying a content object via the aforementioned display generation component, the system receives user input corresponding to a request to save the content object in association with the first application, In response to receiving the user input corresponding to the request to save the content object in association with the first application, the representation of the content object is saved in association with the first application. The method according to any one of claims 1 to 24, further comprising:

26. In the second computer system which is in communication with the display generation component, After the view of the 3D scene is captured via one or more image sensors, a fourth action proposal is made, and based on the determination that the fourth action proposal, determined on the view of the 3D scene and / or based on the context information, satisfies a set of proposal presentation criteria, the display generation component is used to display individual proposal graphical elements corresponding to the fourth action proposal in a proposal user interface different from the first user interface of the first application. The method according to any one of claims 1 to 25, further comprising:

27. The second computer system includes an application programming interface (API) that enables a requesting application to communicate with the first application, and the method is In the second computer system which is in communication with the display generation component, The requesting application sends an API call to the first application via the API. In response to receiving the API call from the requesting application, the first application provides, via the API, user information determined based on one or more objects stored in association with the first application, and After the requesting application receives the user information, Receiving user input corresponding to a request to display the user interface of the requesting application, and The method according to any one of claims 1 to 26, further comprising displaying the user interface of the requesting application, including displaying a suggestion determined based on the user information, in response to receiving the user input corresponding to the request to display the user interface of the requesting application.

28. The method according to any one of claims 1 to 27, wherein individual representations of the views of the 3D scene are stored in a repository associated with the first application, and the first user interface of the first application is displayed after the individual representations of the views of the 3D scene are stored in the repository associated with the first application.

29. One or more non-temporary computer-readable storage media for storing one or more programs configured to be executed by one or more processors of one or more computer systems that are in communication with one or more image sensors and display generation components, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 28.

30. One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more processors, A computer system comprising one or more memories storing one or more programs configured to be executed by one or more processors, wherein the one or more programs include instructions for performing the method according to any one of claims 1 to 28.

31. One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more computer systems comprising means for performing the method described in any one of claims 1 to 28.

32. One or more non-temporary computer-readable storage media that store one or more programs configured to be executed by one or more processors of one or more computer systems that are in communication with one or more image sensors and display generation components, wherein the one or more programs are A first computer system among the one or more computer systems receives user input corresponding to a request to save an object in a three-dimensional (3D) scene. In response to receiving the user input corresponding to the request to save the object in the 3D scene, the first computer system captures a view of the 3D scene via one or more image sensors. After the view of the 3D scene is captured via one or more image sensors, a second computer system among the one or more computer systems provides instructions for displaying the first user interface of the first application via the display generation component. Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more non-temporary computer-readable storage media, wherein the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.

33. One or more computer systems configured to communicate with one or more image sensors and display generation components, One or more processors, The system comprises one or more memories that store one or more programs configured to be executed by the one or more processors, A first computer system among the one or more computer systems receives user input corresponding to a request to save an object in a three-dimensional (3D) scene. In response to receiving the user input corresponding to the request to save the object in the 3D scene, the first computer system captures a view of the 3D scene via one or more image sensors. After the view of the 3D scene is captured via one or more image sensors, a second computer system among the one or more computer systems provides instructions for displaying the first user interface of the first application via the display generation component. Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more computer systems in which the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.

34. One or more computer systems configured to communicate with one or more image sensors and display generation components, A means for receiving user input corresponding to a request to save an object in a three-dimensional (3D) scene by a first computer system among the one or more computer systems, In response to receiving the user input corresponding to the request to save the objects in the 3D scene, the first computer system provides means for capturing a view of the 3D scene via one or more image sensors, The system comprises means for displaying the first user interface of the first application via the display generation component by a second computer system among the one or more computer systems, after the view of the 3D scene has been captured by the first computer system via one or more image sensors, Displaying the first user interface of the first application includes displaying the first proposed graphical elements. The first proposed graphical element is selectable to cause the second computer system to execute the first action proposal via a second application different from the first application. One or more computer systems in which the first action proposal is determined based on performing image recognition on the view of the 3D scene and based on contextual information different from the view of the 3D scene.