Techniques for capturing media
Patent Information
- Application Number
- CN202580016625.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2025-03-03
- Publication Date
- 2026-09-22
Smart Images

Figure CN122804407A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 562,635, filed March 7, 2024, entitled “TECHNIQUES FOR CAPTURING MEDIA”, which is incorporated herein by reference in its entirety for all purposes. Background Technology
[0002] Computer systems typically capture media, including still photos and videos. The environment used for capture is usually designed based on how the user holds the computer. Summary of the Invention
[0003] Existing technologies for capturing media are often cumbersome and inefficient. For example, some existing technologies use complex and time-consuming user interfaces that may include multiple buttons or keystrokes. Some existing technologies require more time than necessary, wasting user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0004] Therefore, this technology provides electronic devices with faster and more efficient methods and interfaces for capturing media. Such methods and interfaces can optionally complement or replace other methods of media capture. These methods and interfaces reduce the cognitive burden on users and result in more efficient human-machine interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time between battery charging cycles.
[0005] In some embodiments, a method is described that is executed at a computer system communicating with a media capture component and a microphone. In some embodiments, the method includes: detecting a first input via the microphone corresponding to a request to capture media; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of one or more conditions satisfying, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of one or more conditions satisfying, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0006] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a microphone. In some embodiments, the one or more programs include instructions for performing: detecting a first input corresponding to a request to capture media via the microphone; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of one or more conditions satisfying, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of one or more conditions satisfying, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0007] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a microphone. In some embodiments, the one or more programs include instructions for performing: detecting a first input via the microphone corresponding to a request to capture media; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of one or more conditions satisfying, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of one or more conditions satisfying, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0008] In some embodiments, a computer system configured to communicate with a media capture component and a microphone is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing: detecting a first input via the microphone corresponding to a request to capture media; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of one or more conditions satisfying, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of one or more conditions satisfying, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0009] In some embodiments, a computer system configured to communicate with a media capture component and a microphone is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting a first input corresponding to a request to capture media via the microphone; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of conditions satisfying one or more conditions, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of conditions satisfying one or more conditions, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0010] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to be executed by one or more processors of a computer system communicating with a media capture component and a microphone. In some embodiments, the one or more programs include instructions for performing: detecting a first input corresponding to a request to capture media via the microphone; and after detecting the first input corresponding to the request to capture media: capturing media via the media capture component in response to a first set of one or more conditions satisfying, based on determining that the first input corresponds to a first instruction; and capturing media via the media capture component in response to a second set of one or more conditions satisfying, based on determining that the first input corresponds to a second instruction different from the first instruction, wherein the second set of one or more conditions differs from the first set of one or more conditions.
[0011] In some embodiments, a method is described that is performed at a computer system communicating with a media capture component and a motion component. In some embodiments, the method includes: when capturing video via the media capture component: based on determining that a first set of one or more capture conditions is satisfied, moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode, wherein moving the portion of the computer system in the first motion mode causes the framing of the video to change as the portion of the computer system moves; and based on determining that a second set of one or more capture conditions is satisfied, wherein the second set of one or more capture conditions differs from the first set of one or more capture conditions, moving the portion of the computer system via the motion component in a second motion mode different from the first motion mode, wherein moving the portion of the computer system in the second motion mode causes the framing of the video to change as the portion of the computer system moves.
[0012] In some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a motion component. In some embodiments, the one or more programs include instructions for performing the following operations when capturing video via the media capture component: moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode according to a first set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the first motion mode causes the framing of the video to change as the portion of the computer system moves; and moving the portion of the computer system via the motion component in a second motion mode, different from the first motion mode, according to a second set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the second motion mode causes the framing of the video to change as the portion of the computer system moves.
[0013] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a motion component. In some embodiments, the one or more programs include instructions for performing the following operations when capturing video via the media capture component: moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode according to a first set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the first motion mode causes the framing of the video to change as the portion of the computer system moves; and moving the portion of the computer system via the motion component in a second motion mode, different from the first motion mode, according to a second set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the second motion mode causes the framing of the video to change as the portion of the computer system moves.
[0014] In some embodiments, a computer system configured to communicate with a media capture component and a motion component is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing the following operations when capturing video via the media capture component: moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode according to a first set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the first motion mode changes the framing of the video as the portion of the computer system moves; and moving the portion of the computer system via the motion component in a second motion mode different from the first motion mode according to a second set of one or more capture conditions determined to be satisfied, wherein moving the portion of the computer system in the second motion mode changes the framing of the video as the portion of the computer system moves.
[0015] In some embodiments, a computer system configured to communicate with a media capture component and a motion component is described. In some embodiments, the computer system includes components for performing each of the following steps: when capturing video via the media capture component: based on determining that a first set of one or more capture conditions is satisfied, moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode, wherein moving the portion of the computer system in the first motion mode causes the framing of the video to change as the portion of the computer system moves; and based on determining that a second set of one or more capture conditions is satisfied, wherein the second set of one or more capture conditions differs from the first set of one or more capture conditions, moving the portion of the computer system via the motion component in a second motion mode different from the first motion mode, wherein moving the portion of the computer system in the second motion mode causes the framing of the video to change as the portion of the computer system moves.
[0016] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with a media capture component and a motion component. In some embodiments, the one or more programs include instructions for performing the following operations when capturing video via the media capture component: moving a portion of the computer system, including the media capture component, via the motion component in a first motion mode according to a first set of conditions determined to be satisfied, wherein moving the portion of the computer system in the first motion mode causes the framing of the video to change as the portion of the computer system moves; and moving the portion of the computer system via the motion component in a second motion mode, different from the first motion mode, according to a second set of conditions determined to be satisfied, wherein moving the portion of the computer system in the second motion mode causes the framing of the video to change as the portion of the computer system moves.
[0017] In some embodiments, a method is described that is performed at a computer system communicating with a media capture component, an input component, and an output component. In some embodiments, the method includes: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to the one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that the one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that the one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0018] In some embodiments, a non-transitory computer-readable storage medium is described, which stores one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component, an input component, and an output component. In some embodiments, the one or more programs include instructions for performing: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to the one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0019] In some embodiments, a transient computer-readable storage medium is described that stores one or more programs executed by one or more processors of a computer system configured to communicate with a media capture component, an input component, and an output component. In some embodiments, the one or more programs include instructions for performing: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to the one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0020] In some embodiments, a computer system configured to communicate with a media capture component, an input component, and an output component is described. In some embodiments, the computer system includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to the one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that the one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that the one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0021] In some embodiments, a computer system configured to communicate with a media capture component, an input component, and an output component is described. In some embodiments, the computer system includes components for performing each of the following steps: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0022] In some embodiments, a computer program product is described. In some embodiments, the computer program product includes one or more programs configured to execute by one or more processors of a computer system communicating with a media capture component, an input component, and an output component. In some embodiments, the one or more programs include instructions for performing: detecting via the input component a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; preparing to capture media via the media capture component in response to detecting the first set of one or more inputs corresponding to the one or more instructions; and when preparing to capture media via the media capture component: providing a first composition guide via the output component based on determining that one or more instructions include first content, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; and providing a second composition guide via the output component based on determining that one or more instructions include second content different from the first content, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
[0023] Executable instructions for performing these functions may optionally be included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Attached Figure Description
[0024] To better understand the various described embodiments, reference should be made to the following detailed description in conjunction with the accompanying drawings, wherein the same reference numerals in all the drawings indicate corresponding parts.
[0025] Figure 1 This is a block diagram illustrating a computer system according to some implementation schemes.
[0026] Figures 2A to 2C These are illustrations of exemplary components and user interfaces of an electronic device according to some implementation schemes.
[0027] Figure 3 This is a block diagram illustrating exemplary components of a device according to some implementation schemes.
[0028] Figure 4 This is a functional diagram of an exemplary actuator device according to some implementation schemes.
[0029] Figure 5 This is a functional diagram of an exemplary intelligent agent system based on some implementation schemes.
[0030] Figures 6A to 6M An exemplary user interface for capturing media is illustrated according to some implementation schemes.
[0031] Figure 7 This is a flowchart illustrating a method for selectively capturing media according to some implementation schemes.
[0032] Figure 8 This is a flowchart illustrating a method for repositioning a camera according to some implementation schemes.
[0033] Figure 9 This is a flowchart illustrating a method for providing design guidance based on some implementation schemes. Detailed Implementation
[0034] The following description illustrates exemplary methods, components, parameters, etc. While specific examples are set forth below, it should be understood that such implementations should not be construed as limiting the scope of this disclosure to the explicit description of the examples set forth herein, but rather as providing illustrative examples.
[0035] One or more steps of the method described herein may depend on satisfying one or more conditions. In some embodiments, the method is performed through multiple iterative processes. In some embodiments, the conditional steps may be satisfied in different iterations of the same process and still remain within the scope of the method described herein. For example, for a given method comprising two steps depending on different conditions, those skilled in the art will understand that the given method should be considered performed even if the process is repeated multiple times until the conditional step is satisfied. In some embodiments, multiple iterations of the process are not required to practice the claims as set forth herein. For example, the claims of an electronic device, system, or computer-readable medium may be performed without iteratively repeating the process. In some embodiments, the claims of an electronic device, system, or computer-readable medium include instructions for performing one or more steps depending on satisfying one or more conditions. Because such instructions are stored in one or more processors and / or one or more memory locations, the claims of an electronic device, system, or computer-readable medium may include logic for determining whether one or more conditions have been satisfied without requiring the steps of the process to be repeated.
[0036] Although numerical descriptors such as "first" and / or "second" are used below to describe elements, these elements do not correspond to sequential or different representations and should not be limited to the stated numerical terms. In some embodiments, these terms are used only as prefixes to distinguish references to one element from references to another. For example, "first" device and "second" device can be two separate references to the same device. Conversely, for example, "first" device and "second" device can be references to two different devices (e.g., not the same device and / or not the same type of device). For example, a first computer system and a second computer system do not correspond to first and second in time and are merely used to distinguish the two computer systems. Therefore, without departing from the scope of the various described embodiments, a first computer system may be referred to as a second computer system, and a second computer system may be referred to as a first computer system.
[0037] In the description of various elements and examples, the use of certain terms is to provide a productive description of the following topics and should not be construed as restrictive. As used in describing the various examples herein, the singular forms “a,” “an,” and “the” should not be construed as excluding or precluding the plural forms unless the context clearly indicates otherwise. Similarly, “and / or” is used to cover any and all possible combinations of one or more associated listed items. For example, “x and / or y” should be construed as including “x” or “y” as well as “x and y” as possible permutations. Furthermore, the use of the terms “comprising” and / or “including” in this specification specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0038] When describing choices and / or logical possibilities, the term "if" may optionally be interpreted, depending on the context, as meaning "when," "in response to determination," "in response to detection," or "according to determination." Similarly, depending on the context, the phrases "if it is determined..." or "if [the stated condition or event] is detected" may optionally be interpreted as meaning "in response to determination," "in response to detection," "in response to detection," or "according to determination."
[0039] The processes described below enhance device operability and make user-device and / or user-device interfaces more efficient through various technologies (e.g., by assisting users in providing appropriate input and reducing user errors when operating / interacting with the device), including by providing users with improved feedback (e.g., visual, tactile, acoustic, and / or haptic feedback), reducing the amount of input required to perform operations, providing additional control options without cluttering the user interface with additional displayed controls, performing operations when a set of conditions is met without requiring additional input (e.g., user input), and / or additional technologies (such as improving the security and / or privacy of the computer system and reducing the aging of one or more parts of the user interface on the display). These technologies also reduce power consumption and extend device battery life by enabling users to use the device faster and more efficiently.
[0040] Figure 1 A block diagram depicts a computer system 100 (e.g., an electronic device and / or electronic system) comprising a set of electronic components communicating (e.g., connected) with each other (e.g., wired or wireless). It should be understood that computer system 100 is merely one example of a computer system that can be used to perform the functions described below, and one or more other computer systems can be used to perform the functions described below. Furthermore, although... Figure 1 The computer architecture of computer system 100 is described, but other computer architectures of computer systems (e.g., including more components, similar components, and / or fewer components) may be used to perform the functionality described herein.
[0041] In some implementations, computer system 100 may correspond to (e.g., is and / or includes) a system-on-a-chip, a server system, a personal computer system, a smartphone, a smartwatch, a wearable device, a tablet computer, a laptop computer, a fitness tracker, a head-mounted display (HMD) device, a desktop computer, public equipment (e.g., smart speakers, connected thermostats and / or additional home-based computer systems), accessories (e.g., switches, lights, speakers, air conditioners, heaters, window covers, fans, locks, media playback devices, televisions, etc.), controllers, hubs and / or sensors.
[0042] In some embodiments, the sensor includes one or more hardware components capable of detecting (e.g., sensing, generating, and / or processing) information about the physical environment near the sensor. For example, the sensor may be configured to detect information around the sensor, detect information in one or more directions extending outward from the sensor, and / or detect information based on contact between the sensor and elements of the physical environment. In some embodiments, the hardware components of the sensor include sensing components (e.g., temperature and / or image sensors), transmitting components (e.g., radio and / or laser transmitters), and / or receiving components (e.g., laser and / or radio receivers). In some implementations, the sensors include angle sensors, breakage sensors, flow sensors, force sensors, gas sensors, humidity or moisture sensors, glass breakage sensors, chemical sensors, contact sensors, non-contact sensors, image sensors (e.g., RGB cameras and / or infrared sensors), particle sensors, photoelectric sensors (e.g., ambient light and / or sunlight), positioning sensors (e.g., GPS), precipitation sensors, pressure sensors, proximity sensors, radiation sensors, inertial measurement units, leak sensors, liquid level sensors, metal sensors, microphones, motion sensors, distance or depth sensors (e.g., RADAR, LiDAR), speed sensors, temperature sensors, time-of-flight sensors, torque sensors, ultrasonic sensors, vacancy sensors, presence sensors, voltage and / or current sensors, conductivity sensors, resistivity sensors, capacitance sensors, and / or water sensors. Although in Figure 1 Only a single computer system is depicted, but the functionality described below can be implemented using two or more computer systems operating together. Additionally, in some embodiments, computer system 100 includes one or more sensors as described above, and captures information about the physical environment by combining data from one sensor with data from one or more additional sensors (e.g., which are part of the computer and / or one or more additional computer systems).
[0043] like Figure 1As illustrated, computer system 100 comprises processor subsystem 110, memory 120, and I / O interface 130. Memory 120 corresponds to system memory that communicates with processor subsystem 110. Electronic components constituting computer system 100 are electrically connected via interconnect 150, which allows communication between components of computer system 100. For example, interconnect 150 may be a system bus, one or more memory locations, and / or additional electrical channels for connecting multiple components of computer system 100. Additionally, I / O interface 130 is connected to I / O device 140 via wired and / or wireless connections. In some embodiments, computer system 100 includes a component comprising I / O interface 130 and I / O device 140, such that the functionality of each component is included within that component. Furthermore, it should be understood that computer system 100 may include one or more I / O interfaces that communicate with one or more I / O devices. In some embodiments, computer system 100 comprises multiple processor subsystems 100s, each processor subsystem being electrically connected via interconnect 150.
[0044] In some embodiments, processor subsystem 110 includes one or more processors or separate processing units capable of executing instructions (e.g., programs, systems, and / or interrupts) to perform the functionality described herein. For example, operating system-level and / or application-level instructions executed by processor subsystem 110. In some embodiments, processor subsystem 110 includes one or more components (e.g., implemented as hardware, software, and / or combinations thereof) capable of supporting, interpreting, and / or executing machine learning instructions and / or operations. For example, computer system 100 may perform operations locally based on a machine learning model. Alternatively or additionally, computer system 100 may communicate with (e.g., perform computations thereto and / or execute corresponding instructions) a remote interactive knowledge base (e.g., processing resources implementing machine learning models, artificial intelligence models, and / or large language models) to perform operations that may otherwise be outside the set of capabilities of computer system 100. For example, computer system 100 may determine a set of inputs (e.g., instructions, data, and / or parameters) to an interactive knowledge base for performing desired machine learning operations.
[0045] The memory 120, which communicates with the processor subsystem 110, can be implemented using a variety of different physical, non-transitory memory media. In some embodiments, the computer system 100 includes multiple memory components and / or various types of memory components, each of which is directly and / or connected to the processor subsystem 110 via interconnect 150. For example, the memory 120 can be implemented using removable flash drives, storage arrays, storage area networks (e.g., SANs), flash memory, hard disk storage devices, optical drive storage devices, floppy disk storage devices, removable disk storage devices, random access memory (e.g., SDRAM, DDR SDRAM, RAM-SRAM, EDO RAM, and / or RAMBUS RAM) and / or read-only memory (e.g., PROM and / or EEPROM). Additionally, in some embodiments, the processor subsystem 110 and / or interconnect 150 are connected to a memory controller, which is electrically connected to the memory 120.
[0046] In some embodiments, the instructions may be executed by processor subsystem 110. In this example, memory 120 may include a computer-readable medium (e.g., a non-transitory or transient computer-readable medium) that can be used to store (e.g., configured to store, assigned to store, and / or store) instructions executable by processor subsystem 110. In some embodiments, each instruction stored by memory 120 and executed by processor subsystem 110 corresponds to an operation for performing the functionality described herein. For example, memory 120 may store program instructions to implement methods 700, 800, and 900 described below (…). Figure 7 , Figure 8 and Figure 9 Related functionality.
[0047] As mentioned above, I / O interface 130 may be one or more types of interfaces that enable computer system 100 to communicate with other devices. In some embodiments, I / O interface 130 includes a bridge chip (e.g., a southbridge) connecting the front-side bus to one or more back-side buses. In some embodiments, I / O interface 130 enables communication with one or more I / O devices (exemplified as I / O device 140) via one or more corresponding buses or other interfaces. For example, I / O devices may include one or more: physical user interface devices (e.g., physical keyboard, mouse, and / or joystick), storage devices (e.g., as described above with respect to memory 120), network interface devices (e.g., to a local area network or wide area network), sensor devices (e.g., as described above with respect to sensors), and / or auditory and / or visual output devices (e.g., screens, speakers, lamps, and / or projectors). In some embodiments, the visual output device is referred to as a display component. For example, a display component may be configured to provide visual output, such as displaying images on a physical visual medium via an LED display or image projection. As used herein, “display” content includes content that is displayed by sending data (e.g., image data and / or video data) to an integrated or external display component via a wired or wireless connection to visually generate content (e.g., video data rendered and / or decoded by a display controller).
[0048] In some embodiments, computer system 100 includes a component that integrates I / O device 140 with other components (e.g., a component including I / O interface 130 and I / O device 140). In some embodiments, I / O device 140 is separate from other components of computer system 100 (e.g., it is a discrete component). In some embodiments, I / O device 140 includes a network interface device that allows computer system 100 to connect to a network or other computer system (e.g., communicate with it) via wired or wireless means. In some embodiments, the network interface device may include Wi-Fi, Bluetooth, NFC, USB, Thunderbolt, Ethernet, etc. For example, computer system 100 may utilize NFC connectivity to facilitate banking, credit, financial, token (e.g., fungible or non-fungible tokens) and / or cryptocurrency transactions between computer system 100 and another nearby computer system.
[0049] In some embodiments, I / O device 140 includes components for detecting users (e.g., users, people, animals, another computer system different from the computer system, and / or objects) and / or input from the detected users (e.g., tap input and / or non-tap input (e.g., verbal input, acoustic requests, acoustic commands, acoustic statements, swipe input, press and drag input, gaze input, air gestures, and / or mouse clicks)). In some embodiments, I / O device 140 enables computer system 100 to identify users associated with and / or who do not have accounts within the environment. For example, computer system 100 may detect known users (e.g., users corresponding to accounts) and access information about the users using the known users' accounts. In some embodiments, as part of computer system 100's user detection, computer system 100 detects that a user's account is associated with a group of users (e.g., included in and / or identified relative to that group of users). For example, computer system 100 may access information associated with accounts in a family defined as a group of accounts in response to detecting such a member. In some implementations, the user's account may be linked to additional accounts and / or additional computer systems. For example, computer system 100 may detect such additional computer systems and / or detect such computer systems used to detect users. In some implementations, computer system 100 detects unknown users and enables guest accounts for unknown users to use computer system 100.
[0050] In some embodiments, I / O device 140 includes one or more cameras. In some embodiments, the camera includes an image sensor (e.g., one or more optical sensors and / or one or more depth camera sensors) that provides computer system 100 with the ability to detect user and / or user gestures (e.g., hand gestures and / or air gestures) as input. In some embodiments, air gestures are gestures detected without the user touching an input element that is part of the device (or independently of an input element that is part of the device) and based on detected movement of a part of the user's body through the air (including movement of the user's body relative to an absolute reference (e.g., the angle of the user's arm relative to the ground or the distance of the user's hand relative to the ground), movement relative to another part of the user's body (e.g., movement of the user's hand relative to the user's shoulder, movement of one of the user's hands relative to the user's other hand, and / or movement of the user's fingers relative to another finger or part of the user's hand), and / or absolute movement of a part of the user's body (e.g., a tapping gesture that includes the hand moving a predetermined amount and / or speed in a predetermined pose, or a shaking gesture that includes a predetermined speed or amount of rotation of a part of the user's body)). In some implementations, one or more cameras enable computer system 100 to send image and / or video information to an application. For example, image data captured by a camera can enable computer system 100 to complete a video call by sending video data to an application used to perform the video call.
[0051] In some embodiments, I / O device 140 includes one or more microphones. For example, the microphone may be used by 100 to obtain data and / or information from a user without contact input. In some embodiments, the microphone enables computer system 100 to detect verbal and / or voice input from a user. In some embodiments, computer system 100 utilizes voice input to enable personal assistant functionality. For example, a user makes a request to computer system 100 to perform an action and / or obtain information from the user. In some embodiments, computer system 100 utilizes voice input (e.g., in conjunction with one or more other input and / or output technologies) to request and / or detect information from a user without requiring physical contact between the user and computer system 100.
[0052] In some embodiments, I / O device 140 includes physical input media for a user to interact directly with computer system 100. In some embodiments, the physical input media includes one or more physical buttons (e.g., tactilely pressable buttons and / or touch-sensitive non-pressable components) on and / or connected to computer system 100, mouse and keyboard input methods (e.g., connected to computer system 100 together with and / or separately from one or more I / O interfaces), and / or touch-sensitive display components.
[0053] In some embodiments, I / O device 140 includes one or more components for outputting information (e.g., display components, audio generation components, speakers, haptic output devices, displays, projectors, and / or touch-sensitive displays). In some embodiments, computer system 100 uses I / O device 140 to transmit information and / or the state of computer system 100. In some embodiments, I / O device 140 includes haptic output components. For example, the haptic output component may be a haptic generation component that enables computer system 100 to transmit information to a user who is in contact with computer system 100 (e.g., holding, touching, and / or near it). In some embodiments, I / O device 140 includes one or more components for outputting visual output (e.g., video, images, animations, 3D rendering, augmented reality overlay, motion graphics, data visualization, digital art, etc.). For example, displaying content from one or more applications and / or system applications, and / or displaying widgets corresponding to one or more applications (e.g., controls displaying real-time information and / or data).
[0054] In some implementations, I / O device 140 includes one or more components for outputting audio (e.g., smart speaker, home theater system, soundbar, headphones, earphones, earbuds, speaker, TV speaker, augmented reality headphone speaker, audio jack, optical audio output, Bluetooth audio output, HDMI audio output, audio sensor, etc.). In some implementations, computer system 100 is capable of outputting audio through one or more speakers. For example, computer system 100 outputs audio-based content and / or information to a user. In some implementations, one or more speakers enable spatial audio (e.g., audio output corresponding to the environment (e.g., computer system 100 detects materials and / or objects in the environment and / or computer system 100 changes audio modes, intensities, and / or waveforms to compensate for changing environmental characteristics)).
[0055] Figure 2 to Figure 5Exemplary components and user interfaces of an electronic device 200 according to some embodiments are illustrated. The electronic device 200 (sometimes referred to herein as device 200) may include one or more features of a computer system 100. (Referring to Figures 2 to...) Figure 5 In the described example, device 200 is a laptop computer. In some embodiments, device 200 is not limited to a laptop computer, and those skilled in the art will recognize that device 200 can be one or more other devices (e.g., one or more of the components and / or functions described herein with respect to device 200). For example, device 200 can be a public device (such as a smart display, smart speaker, and / or television) and / or a personal device (such as a smartphone, smartwatch, tablet, desktop computer, fitness tracker, and / or head-mounted display). In some embodiments, the public device is configured to provide functionality to multiple users (e.g., simultaneously and / or at different times). In such embodiments, the public device can be managed and / or set by a single user. In some embodiments, the personal device is configured to provide functionality to a single user (e.g., once, such as when a single user logs into the personal device).
[0056] Figures 2A to 2C An example is shown of a device 200 located in three different physical locations. For example... Figure 2A As illustrated, device 200 is a laptop computer (also referred to herein as a "laptop"), which includes a base portion 200-2 (e.g., as shown in the image). Figure 2A The device 200 is horizontally placed on a surface such as a table and connected to a base portion 200-2 at a connection 200-3 (e.g., one or more connection points, motor arms, hinges, and / or joints). This connection allows the display portion 200-1 to pivot and / or change orientation relative to the base portion 200-2. For example, the device 200 may pivot at the connection 200-3 to rotate the display portion 200-1 and / or the device 200 to one or more positions corresponding to the “closed” internal state (e.g., as described below regarding...). Figure 2C(Further description). In some embodiments, the positioning corresponding to the "off" internal state is the positioning of the device 200 in a predetermined pose. For example, the predetermined pose may include a display portion 200-1 positioned parallel to the base portion 200-2 or forming a predetermined angle (e.g., 60 degrees) with respect to the base portion 200-2. In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., an area facing downwards, not visible, and / or obscuring the displayed content). In some embodiments, in the "off" internal state, the area of the device 200 in which content is displayed is not positioned in a manner corresponding to (e.g., indicating, associating with, and / or configured to accompany) the "off" internal state (e.g., instead positioned in a manner corresponding to the "on" internal state). For example, when not in a "closed" internal state, device 200 can be positioned within a range of different open positions (e.g., where display portion 200-1 is not parallel to base portion 200-2, and where the area where the content displayed by device 200 is visible and / or unobstructed). It should be recognized that display portion 200-1 being parallel to base portion 200-2 is an example of positioning corresponding to a "closed" internal state of device 200 (e.g., closed positioning). In some embodiments, another configuration may set another orientation of display portion 200-1 relative to base portion 200-2 as a closed positioning of device 200, such as... Figure 2C exemplified.
[0057] Figure 2A The left side illustrates display screen 200-4 (representing the area where device 200 displays content), and the right side illustrates device 200 in the corresponding pose. For example... Figure 2A As illustrated, device 200 is in a first position (e.g., display portion 200-1 is perpendicular to base portion 200-2, forming a 90-degree angle). Figure 2A In this context, display screen 200-4 represents the content currently being displayed (e.g., via a display component) when device 200 is first activated. Figure 2A In this embodiment, display screen 200-4 illustrates the device 200 in an "on" internal state (e.g., operable, powered, awake, higher power and / or more resource-intensive than the "off" state, and / or activated). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces (e.g., user interface objects, windows, application user interfaces, system user interfaces, controls, and / or other visual content). In some embodiments, the device 200 displays (e.g., via display screen 200-4) one or more user interfaces while in an "on" internal state. For example, in Figure 2A In this configuration, device 200 is in an "on" internal state, and display screen 200-4 shows a desktop user interface 200-5, including an application window. In some embodiments, the user interface includes (and / or) one or more user interface objects (e.g., windows, icons, and / or other graphical objects). For example, the user interface (e.g., 200-5) may include one or more graphical objects that are different from and / or the same as the application window.
[0058] Figure 2B Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2B As illustrated, device 200 is in a second position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming an angle of 120 degrees (e.g., more than). Figure 2A (at a larger angle). Figure 2B In the diagram, display screen 200-4 represents the content being displayed when device 200 is in the second position. Display screen 200-4 illustrates the internal state of device 200 being "on" (e.g., with...). Figure 2A (The top diagram shows the same internal state). Figure 2B In the process, device 200 displays (e.g., via display screen 200-4) a desktop user interface 200-5 (e.g., with...). Figure 2A (The same as shown in the image). In some implementations, device 200 displays a different user interface (e.g., different from desktop user interface 200-5). For example, although... Figure 2B Example of device 200 in a state of being with Figure 2A Different positioning displays and Figure 2A The same desktop user interface 200-5 exists, but device 200 may display different user interfaces. In some embodiments, device 200 displays a user interface corresponding to (e.g., based on, due to, caused by, involved in, and / or configured to accompany) a physical state (e.g., positioning, location, and / or orientation), including content specific to a particular angle or specific to the current context.
[0059] Figure 2C Display screen 200-4 is illustrated on the left, and device 200 in the corresponding pose is illustrated on the right. Figure 2C As illustrated, device 200 is in a third position (e.g., display portion 200-1 is at an angle relative to base portion 200-2 (e.g., via connection 200-3), forming a 60-degree angle (e.g., compared to...). Figure 2A and Figure 2B (smaller angles)). Figure 2C In the diagram, display screen 200-4 shows the content being displayed when device 200 is in the third position. Figure 2C In the diagram, displays 200-4 illustrate an internal state in which device 200 is "off" (e.g., not operating, not powered, not woken up, not activated, powered off, asleep, hibernating, inactive, and / or disabled). In some embodiments, device 200 does not display (e.g., via displays 200-4) one or more user interfaces (e.g., no visual content is displayed) when it is in the "off" internal state. In some embodiments, device 200 displays (e.g., via displays 200-4) one or more user interfaces (e.g., the same as and / or different from one or more user interfaces displayed when it is in the "on" internal state) (e.g., a user interface specific to the "off" state and / or a way of displaying a user interface not specific to the "off" internal state). Figure 2C In this case, display screen 200-4 is blank because nothing is displayed on the monitor of device 200 (e.g., display screen 200-4 is off and / or does not display the user interface) (e.g., desktop user interface 200-5 is not displayed on display screen 200-4).
[0060] In some embodiments, device 200 includes one or more components (referred herein also as “moving components”) that enable device 200 to perform (e.g., cause and / or control) movement (and / or be moved). For example, performing movement may include a portion of mobile device 200 (e.g., less or all components of the device moving), all of mobile device 200 (e.g., the entire device (including all its components) moving, such as by changing position), and / or moving one or more other devices and / or components (e.g., communicating with device 200 and / or the moving components of device 200). For example, device 200 may move automatically (e.g., pivot), cause and / or control movement of display portion 200-1 relative to base portion 200-2, such as moving to... Figures 2A to 2CAny of the illustrated locations. In some embodiments, device 200 performs movement based on its internal state. Performing movement based on internal state enables device 200 to perform new (e.g., otherwise unavailable) interactions. For example, such new interactions of device 200 can be configured using special features, functions, patterns, and / or procedures that leverage device 200's ability to perform movement. Examples of such interactions include using movement to (e.g., to a user) convey the device's internal state (e.g., on, off, sleep, and / or hibernate) to assist user input (e.g., shorten the distance to the user) and / or enhance the device's interactive behavior (e.g., moving in a specific manner during interaction with the user, conveying information such as importance and / or direction of attention). In some embodiments, the performed movement corresponds to (e.g., caused by, responded to, and / or determined and / or performed based on) one or more of the following: detected input, detected context (e.g., environmental context and / or user context), and / or the device 200's internal state (e.g., internal state and / or a set of multiple internal states). For example, device 200 can move the display portion, causing device 200 to move from a position where... Figure 2A The illustrated first positioning moves to the position where Figure 2B The illustrated second positioning. In this example, device 200 can detect that the user has repositioned relative to device 200 (e.g., the user stands up), and in response, device 200 can perform a movement to the second positioning such that the display is at an optimized viewing angle based on the height and / or angle of the user's eye relative to the display of device 200. As another example, device 200 can perform a movement such that device 200 moves from a position... Figure 2A The illustrated first positioning moves to the position where Figure 2C The illustrated third location. In this example, device 200 may perform a movement to a third location in response to detecting an internal state with reduced activity (e.g., an "off" internal state as described above). In this way, movement of device 200 to one or more locations can indicate the internal state of device 200.
[0061] Figures 2A to 2C An example is illustrated of a device 200 having a display portion capable of moving with one degree of freedom via a connection 200-3 (e.g., a hinge) connecting the display portion 200-1 to a base portion 200-2. In some embodiments, the device 200 includes one or more components having one or more degrees of freedom. For example, a moving component of the device 200 (e.g., an output component that causes and / or allows movement) (e.g., Figure 5Device 200-26C may include multiple degrees of freedom (e.g., six degrees of freedom including three translational components and three rotational components). For example, device 200 may be implemented to move the display portion by telescopic forward or backward movement (e.g., display portion 200-1 moves forward in space relative to the base portion while the base portion 200-2 remains stationary (e.g., to shorten and / or lengthen the user's viewing distance)). As yet another example, device 200 may be implemented to move the display portion to rotate about an axis perpendicular to the hinge, such that the display portion can rotate to position the display to follow the user as the user walks around device 200. Although Figures 2A to 2C The example shown illustrates a hinge, but other moving components may be included in device 200, such as actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), movable bases, rotatable components, and / or rotatable bases. In some embodiments, one or more moving components may enable device 200 to move in different ways, such as rotation (e.g., 0 to 360 degrees), lateral movement (e.g., to the right, left, down, up, and / or any combination thereof), and / or tilting (e.g., 0 to 360 degrees).
[0062] Figure 3 An exemplary block diagram of device 200 is illustrated. In some embodiments, device 200 includes... Figure 1 A, Figure 1 B. Figure 3 and Figure 5 B describes some or all of the components. For example... Figure 3 As illustrated, device 200 has a bus 200-13 that operatively couples I / O segments 200-12 (also referred to as I / O sub-segments and / or I / O interfaces) to processor 200-11 and memory 200-10. For example... Figure 3 As illustrated, I / O section 200-12 is connected to output device 200-16 (also referred to herein as "output component"). In some embodiments, output device 200-16 includes one or more visual output devices (e.g., display components such as monitors, displays, projectors, and / or touch-sensitive displays), one or more tactile output devices (e.g., devices that cause vibration and / or other tactile outputs), one or more audio output devices (e.g., speakers), and / or one or more moving components (e.g., actuators, motors, mechanical linkages, devices that cause and / or allow movement, and / or one or more moving components as described above). Figure 3As illustrated, output device 200-16 includes two exemplary moving components (e.g., a movement controller 200-17 and an actuator 200-18). Actuator 200-18 can be any component that performs (e.g., partial and / or overall) physical movement of a device (e.g., device 200 and / or devices coupled to and / or in contact with that device). Movement controller 200-17 can be any component (e.g., a control device) that controls actuator 200-18 (e.g., provides control signals to it). For example, movement controller 200-17 can provide control signals that actuate actuator 200-18 (e.g., cause physical movement). In some embodiments, movement controller 200-17 includes one or more logic components (e.g., a processor), one or more feedback components (e.g., sensors), and / or one or more control components (e.g., for applying control signals, such as relays, switches, and / or control lines). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in the same device and / or component (e.g., a dedicated onboard motion controller 200-17 attached to the actuator 200-18). In some embodiments, the motion controller 200-17 and the actuator 200-18 are embodied in different devices and / or components (e.g., one or more processors 200-11 may serve as the motion controller 200-17 for the actuator 200-18). In some embodiments, the motion controller 200-17 and / or the actuator 200-18 are embodied in a device (or one or more devices) other than device 200 (e.g., device 200 is coupled to (e.g., temporarily and / or removably) another device and may instruct the motion controller 200-17 and / or the actuator 200-18 to control the other device). Actuator 200-18 can be used to induce one or more types of mechanical movement (e.g., linear and / or rotary movement) in one or more ways (e.g., using electric, magnetic, hydraulic and / or pneumatic power). Examples of actuator 200-18 may include electromechanical actuators, linear actuators and / or rotary actuators.
[0063] like Figure 3As illustrated, I / O section 200-12 is connected to input device 200-14. In some embodiments, input device 200-14 includes one or more visual input devices (e.g., cameras and / or light sensors), one or more physical input devices (e.g., buttons, sliders, switches, touch-sensitive surfaces, and / or rotatable input mechanisms), one or more audio input devices (e.g., microphones), and / or other input devices (e.g., accelerometers, pressure sensors (e.g., contact strength sensors), distance sensors, temperature sensors, GPS sensors, accelerometers, orientation sensors (e.g., compasses), gyroscopes, motion sensors, and / or biometric sensors). Furthermore, I / O section 200-12 may be connected to communication unit 200-15 for receiving application and operating system data using Wi-Fi, Bluetooth, Near Field Communication (NFC), cellular, and / or other wireless (and / or wired) communication technologies.
[0064] The memory 200-10 of device 200 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 200-11, cause the computer processors to perform, for example, the techniques described below, including methods 700, 800, and 900. Figures 7 to 9 A computer-readable storage medium can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transient computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media can include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical discs based on CD, DVD, and Blu-ray technologies, and persistent solid-state storage such as flash memory and solid-state drives. Electronic device 200 is not limited to Figure 3 The components and configurations may include, but may include, other and / or additional components in a variety of possible configurations, all of which are intended to fall within the scope of this disclosure.
[0065] Figure 4 A functional diagram of actuator 200-18B according to some embodiments is illustrated. As described above, actuator 200-18B can be any component that performs physical movement. In some embodiments, actuator 200-18B is operated using inputs including control signal 200-18A and / or energy source 200-18B. For example, actuator 200-18B can be a rotary actuator that converts electrical energy into rotational movement. This rotational movement can cause the aforementioned... Figures 2A to 2CThe movement of the display portion of the described device 200 (e.g., the counterclockwise rotation of the actuator moves the device 200 to a position with a large angle) Figure 2B The illustrated second positioning), and the clockwise (e.g., counterclockwise) rotational movement of the actuator moves the device 200 to a positioning with a smaller angle (e.g., Figure 2C (The illustrated third positioning). Control signal 200-18A may indicate one or more start and / or stop commands, movement and / or actuation direction, movement and / or actuation speed, movement and / or actuation time, target positioning (e.g., pose and / or position) of movement and / or actuation, and / or one or more other characteristics of movement and / or actuation. In some embodiments, the control signal and the energy source are the same signal and / or input. In some embodiments, one or more additional components (e.g., mechanical and / or electrical) (e.g., removably or permanently) are coupled to actuator 200-18B to influence movement and / or actuation (e.g., mechanical linkages such as lead screws, gears, and / or other components for changing (e.g., switching) the characteristics of movement and / or actuation). In some embodiments, actuator 200-18B includes one or more feedback components (e.g., a positioning sensor, an encoder, an overcurrent sensor, and / or a force sensor) that form part of a feedback loop for modifying and / or stopping movement and / or actuation (e.g., slowing down actuation upon reaching a target position and / or stopping actuation if physical resistance to actuation is detected via a sensor). In some embodiments, one or more feedback components are included (e.g., partially and / or entirely) in a motion controller (e.g., motion controller 200-13) operatively coupled to the actuator.
[0066] Now turn attention to the functionality (e.g., features and / or capabilities) of one or more devices (e.g., computer system 100 and / or electronic device 200). One such functionality is the implementation of an “intelligent agent,” which may alternatively be referred to as a software intelligent agent, intelligent intelligent agent, interactive intelligent agent, virtual assistant, intelligent virtual assistant, interactive virtual assistant, personal assistant, intelligent personal assistant, interactive personal assistant, intelligent interactive personal assistant, and / or artificial intelligence (AI) assistant. In some implementations, an intelligent agent refers to one or more functions implemented in hardware and / or software (e.g., local and / or remote) on an intelligent agent system (e.g., a single device and / or multiple devices). In some implementations, the intelligent agent performs operations to perceive the environment, acquire knowledge, retrieve knowledge, learn skills, interact with a user, and / or perform tasks. The agent may, for example, perform these (and / or other) operations in response to user input and / or automatically (e.g., at an appropriate time determined based on the perceived context). An incomplete list of exemplary operations that an intelligent agent can be used with and / or employed therewith includes: tracking a user’s eyes, face, and / or body (e.g., to move with the user and / or identify the user’s intentions and / or activities); detecting, identifying, and / or classifying users in the environment; detecting and / or responding to input (e.g., verbal input, air gestures, and / or physical input, such as touch input and / or force input to physical hardware components (e.g., buttons, knobs, and / or sliders); detecting context (e.g., user context, operational context, and / or environmental context); moving (e.g., changing pose, orientation, orientation, and / or location); performing one or more operations in response to input, context, and / or stimuli (e.g., objects or events that elicit one or more responsive operations on the device (e.g., outside and / or inside the device)); providing intelligent interaction capabilities (e.g., in part due to one or more machine learning (“ML”) models, such as large language models (“LLM”)) to respond to and / or perform operations; and / or (e.g., automatically and / or intelligently) performing tasks (e.g., a set of operations for achieving a specific goal). In some implementations, the agent performs actions in response to contactless input (e.g., air gestures and / or natural language commands). The foregoing list is intended to exemplify actions that can be performed by an agent, but is not intended to be an exhaustive list. Other actions fall within the expected scope of the agent's capabilities. Furthermore, for the purposes of this disclosure, the agent need not include all the functionalities mentioned herein, but may include fewer or more functionalities (e.g., the agent may be implemented on an agent system that does not have mobile functionality but otherwise includes an intelligent personal assistant capable of interacting with a user).
[0067] In some embodiments, a user is one or more of a user, person, object, and / or animal in an environment (e.g., a device) that is perceived (e.g., detected by the device, one or more other devices, and / or one or more of its components). In some embodiments, a user is an entity that is perceived (e.g., detected by the device, one or more other devices, and / or one or more of its components). In some embodiments, an entity is something distinguishable from surrounding entities (e.g., components of the environment and / or other users) and / or something that is considered to be a discrete logical construct via one or more components (e.g., a sensing component and / or other components). In some embodiments, a user is physical and / or virtual. For example, a physical user may represent a user standing in front of the device and perceived by the device. As another example, a virtual user may represent a digital image in a virtual scene perceived by the device (e.g., a digital image detected in a media stream received by the device and / or captured by the device's camera). Although presented above as examples of “user,” the terms and / or concepts referred to as “person,” “object,” and / or “animal” are used interchangeably with “user” throughout this disclosure unless otherwise expressly stated. For example, the use of the term “user” can also be understood to refer to “user” unless otherwise expressly stated.
[0068] As an example, and to revisit Figures 2A to 2C An agent, at least partially implemented on device 200, can perform operations that cause the display portion 200-1 of device 200 to move relative to the base portion 200-2. For example, agent detection (e.g., perceiving and determining its occurrence) includes the context of a user standing up (e.g., based on face detection and tracking); and in response, the agent causes device 200 to open and / or device 200 to open the display portion 200-1 to a greater angle. As another example, the agent can detect verbal input corresponding to (e.g., interpreted as and / or implying including) a request to move the display (e.g., “Please move my display” or “Please enter sleep mode”); and in response, the agent causes device 200 to move and / or device 200 to move the display portion 200-1.
[0069] Figure 5 A functional diagram of an exemplary intelligent agent system 200-20A is shown. Figure 5 As shown, the agent system 200-20A has a dashed box boundary that surrounds the input component 200-22, the agent component 200-24, and the output component 200-26. In some embodiments, the agent system 200-20A includes more than Figure 5Fewer, more, and / or different components are illustrated. In some embodiments, the agent system 200-20 is implemented on a single device (e.g., computer system 100 and / or device 200). In some embodiments, the agent system 200-20 is implemented on multiple devices. In some embodiments, in Figure 5 One or more components of the agent system 200-20 illustrated and / or described with respect to this figure are external to but operatively coupled to the agent system (e.g., accessories, external devices, external sensors, external actuators, external display components, external speakers, and / or external databases). In some embodiments, one or more components of the agent system 200-20 are local to one or more other components of the agent system 200-20. In some embodiments, one or more components of the agent system 200-20 are remote from one or more other components of the agent system 200-20.
[0070] In some implementations, input components 200-22 include components for performing sensing and / or communication functions of the agent system 200-20. For example... Figure 5 As illustrated, input components 200-22 include one or more sensors 200-22A. The one or more sensors 200-22A may include any components for detecting data corresponding to the physical environment. Examples of the one or more sensors 200-22A may include: cameras, light sensors, microphones, accelerometers, positioning sensors, pressure sensors, temperature sensors, olfactory sensors, and / or contact sensors. This list is not intended to be exhaustive, and the one or more sensors 200-22A may include other sensors not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to detect data corresponding to the physical environment. Figure 5 As illustrated, input component 200-22 includes one or more communication components 200-22B. The one or more communication components 200-22B may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the intelligent agent system 200-20. Communication components 200-22B may be between different devices and / or between components within the same device. Communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, input component 200-22 includes more than Figure 5 The components illustrated herein may be fewer, more, and / or different. In some implementations, input components 200-22 are implemented in hardware and / or software.
[0071] In some implementations, agent components 200-24 include components that manage and / or perform the functions of the agents in agent system 200-20. For example... Figure 5 As illustrated, agent components 200-24 include the following functional components: task flow, coordination and / or orchestration component 200-24A, management component 200-24B, perception component 200-24C, evaluation component 200-24D, interaction component 200-24E, policy and decision-making component 200-24F, knowledge component 200-24G, learning component 200-24H, model component 200-24I, and API component 200-24J. Each of these components is briefly described below. It is important to note that this list of agent components 200-24 is not intended to be exhaustive, and agent components 200-24 may include other functional components not explicitly identified herein, which may be used (e.g., to process, store, and / or transform) any function of the agent, such as those described herein. In some embodiments, agent components 200-24 include more than Figure 5 Fewer, more, and / or different components are illustrated. In some implementations, the agent components 200-24 are implemented in hardware and / or software.
[0072] In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between various components. For example, operations may include managing a data processing task flow to move from perception components 200-24C (e.g., those detecting speech input) to model components 200-24I (e.g., for processing the detected speech input using a large language model to determine the content and / or intent of the speech input). In some implementations, task flow, coordination, and / or orchestration components 200-24A perform operations that enable the agent to manage coordination between one or more external components (e.g., resources). For example, Figure 5 Examples of external components (such as external databases 200-30) are illustrated. In some embodiments, management component 200-24B includes functionality performed by the operating system of the device implementing the intelligent agent system 200-20. In some embodiments, management component 200-24B includes functionality performed by one or more applications of the device implementing the intelligent agent system 200-20.
[0073] In some embodiments, management components 200-24B perform operations that enable the agent system to handle management tasks, such as managing system and / or component updates, managing user accounts, and managing system settings and / or component settings. In some embodiments, management components 200-24B include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, management components 200-24B include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0074] In some embodiments, the sensing components 200-24C perform operations that enable the agent to perceive environmental input. For example, operations may include detecting existing context and / or environmental conditions, detecting the presence of a user (e.g., a user, person, object, and / or animal in the environment), detecting input including verbal input, detecting input including air gestures, detecting facial expressions, detecting user characteristics (e.g., visible and / or invisible), and / or detecting verbal and / or physical cues. In some embodiments, the sensing components 200-24C include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the sensing components 200-24C include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0075] In some embodiments, the evaluation component 200-24D performs operations that enable the agent to process evaluation data (e.g., to determine context, such as user context, environmental context, and / or operational context). For example, the operations may include evaluating data collected from the perception component 200-24C, the knowledge component 200-24G, the external database 200-30, and / or the remote processing resource 200-32. In some embodiments, the evaluation component 200-24D includes functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the evaluation component 200-24D includes functionality performed by one or more applications of the device implementing the agent system 200-20.
[0076] This document refers to an environmental context (also referred to herein as "context of the environment" and / or "context corresponding to the environment"). In some embodiments, an environmental context is a context based on one or more characteristics of the environment (e.g., user, location, time, weather, and / or lighting). For example, an environmental context may include rain outside, daytime, and / or the device currently being in a park. In some embodiments, the device (e.g., using an agent) uses one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device) to determine the environmental context (e.g., currently true, happening, and / or applicable).
[0077] This document refers to user context (also referred to herein as "user context" and / or "context corresponding to a user") (and / or user context). In some embodiments, user context is a context based on one or more characteristics of a user (and / or a user). For example, user context may include a user's appearance and / or clothing, personality, actions, behaviors, movement, location, and / or pose. In some embodiments, a device (e.g., using an agent) determines user context (e.g., currently true, happening, and / or applicable) using one or more of detected inputs (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, a device determines user context based on historical context and / or learned user characteristics, wherein one or more user characteristics are learned and / or stored by the device over a period of time.
[0078] This document refers to an operational context (also referred to herein as "the context of operation" and / or "operational context"). In some embodiments, an operational context is a context based on one or more characteristics of the device's operation (e.g., the device and / or one or more other devices that determine and / or access the operational context). For example, an operational context may include the internal state of the device (and / or one or more components of the device), the device's internal dialogue (e.g., the device's understanding of the context), the operations performed by the device, and applications and / or processes executed on the device (e.g., running and / or opening). In some embodiments, the device (e.g., using an agent) uses one or more of the following to determine the operational context (e.g., currently true, happening, and / or applicable): detected input (e.g., via one or more input components) and / or received data (e.g., from one or more other devices and / or components communicating with the device). In some embodiments, the device (e.g., using an agent) uses one or more internal states (e.g., accessed, retrieved, and / or queried by the device's processes) to determine the operational context (e.g., currently true, happening, and / or applicable).
[0079] In some embodiments, the interaction components 200-24E perform operations that enable the agent to manage and / or perform interactions with a user. For example, operations may include determining an appropriate interaction model for a specific context and / or in response to specific inputs. In some embodiments, the interaction components 200-24E include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the interaction components 200-24E include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0080] In some embodiments, the policy and decision components 200-24F perform operations that enable the agent to take actions based on available data. For example, operations may include determining which operations to perform and / or which functional components to utilize in response to detected context. In some embodiments, the policy and decision components 200-24F include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the policy and decision components 200-24F include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0081] In some embodiments, knowledge components 200-24G perform operations that enable the agent to access and use the stored knowledge. For example, operations may include indexing, storing, and / or retrieving data from a data repository, database, and / or other resource. In some embodiments, knowledge components 200-24G include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, knowledge components 200-24G include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0082] In some embodiments, the learning component 200-24H performs operations that enable the agent to learn through experience. For example, operations may include observing and / or tracking data, including preferences, routines, user characteristics, and / or environmental characteristics, in a way that allows the data to inform future actions of the agent and / or its components (e.g., when performing tasks and / or interacting with a user). In some embodiments, the learning component 200-24H includes functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, the learning component 200-24H includes functionality performed by one or more applications of the device implementing the agent system 200-20.
[0083] In some embodiments, model components 200-24I perform operations that enable the agent to apply an ML model (e.g., a large language model (LLM)) to process data. For example, operations may include storing the ML model, executing the ML model, training and / or retraining the ML model, and / or otherwise managing aspects of implementing the ML model. In some embodiments, model components 200-24I include functionality performed by the operating system of the device implementing the agent system 200-20. In some embodiments, model components 200-24I include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0084] In some implementations, the agent system 200-20 responds to natural language input. For example, the agent system 200-20 responds to natural language input in the form of statements, questions, commands, and / or requests. In some implementations, the agent system 200-20 outputs text and / or speech output provided in natural language or mimicking a natural language style. For example, the agent system 200-20 may use a speech response indicating the current outside temperature at the user's location (e.g., "It's 18 degrees outside") to handle the natural language question "How hot is it outside?". In some implementations, the agent system 200-20 responds to natural language input by providing information (e.g., weather, travel, and / or calendar information) and / or performing tasks (e.g., opening a document, searching a database, and / or opening an application).
[0085] In some implementations, the intelligent agent system 200-20 includes and / or relies on one or more data models to process inputs (e.g., natural language input, gesture input, visual input, and / or other data input) and / or provide outputs (e.g., information output via natural language output, visual output, audio output, and / or text output). Such data models may include user data (e.g., data based on a specific interaction and / or from the user with whom the interaction took place) and / or global data (e.g., general data based on the interaction and / or data from many users) and / or be trained using user data and / or global data. For example, user data (e.g., preferences, prior use of language and / or phrases, calendar entries, contact lists, and / or activity data) can be used to better infer user intent and / or provide responses more likely to resolve user requests. In some implementations, the data models used by the intelligent agent system 200-20 include one or more machine learning components (e.g., hardware and / or software) (e.g., one or more neural networks), are used by one or more machine learning components, and / or are implemented using one or more machine learning components. Such machine learning components can be used to process spoken input to determine words and / or phrases therein, one or more contexts corresponding to the words, user intent corresponding to the words, one or more confidence scores, and / or a set of one or more actions to be taken in response to the spoken input. Similar operations can be performed to process other types of input, such as visual input, data input, and / or text input. Such data models may include machine learning and / or data processing models, including but not limited to natural language processing models, language models, speech recognition models, object recognition models, visual processing models, ontology, task flow models, and / or intent recognition models (e.g., for determining user intent).
[0086] In some implementations, application programming interface (API) components 200-24J perform operations that enable the agent to interface with services, devices, and / or components. For example, operations may include relaying data (e.g., requests, responses, and / or other messages) between data interfaces (e.g., between software programs, between system processes and application processes, between system processes, between application processes, between communication protocols, between clients and servers, between file systems, and / or between components on different sides of a trust boundary). In some implementations, the data interfaces served by API components 200-24J are local (e.g., for a device, such as two application processes exchanging data) and / or remote (e.g., from a device, such as interfaced with a web service via a remote server). In some implementations, API components 200-24J include functionality performed by the operating system of the device implementing the agent system 200-20. In some implementations, API components 200-24J include functionality performed by one or more applications of the device implementing the agent system 200-20.
[0087] In some implementations, output components 200-26 include components for performing the output functions of the agent system 200-20. A brief description follows. Figure 5 The exemplary output components are illustrated herein. In some embodiments, output components 200-26 include... Figure 5 The components illustrated may include fewer components, more components, and / or different components. In some implementations, the input components are implemented in hardware and / or software.
[0088] like Figure 5 As illustrated, output components 200-26 include one or more visual output components 200-26A. One or more visual output components 200-26A may include any component used for outputting (e.g., generating, creating, and / or displaying) and / or causing visual output (e.g., visually perceptible output, such as a graphical user interface, playback of visual media content, and / or lighting). Examples of one or more visual output components 200-26A may include: display components, projectors, head-mounted displays (HMDs), light-emitting diodes (“LEDs”), and / or components that create visually perceptible effects (e.g., movement). This list is not intended to be exhaustive, and one or more visual output components 200-26A may include other visual output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output visual output.
[0089] like Figure 5As illustrated, output components 200-26 include one or more audio output components 200-26B. One or more audio output components 200-26B may include any component for outputting (e.g., generating and / or creating) and / or causing audio output (e.g., audibly perceptible output, such as sound, music, speech, and / or audio media content). Examples of one or more audio output components 200-26B may include: speakers, audio amplifiers, tone generators, and / or components that produce audibly perceptible effects (e.g., movement, such as vibration). This list is not intended to be exhaustive, and one or more audio output components 200-26B may include other audio output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) the output audio output.
[0090] like Figure 5 As illustrated, output components 200-26 include one or more motion output components 200-26C (also referred to herein as "motion components"). One or more motion output components 200-26C may include any component for outputting (e.g., generating and / or creating) and / or causing motion output (e.g., output including physical movement of a device and / or another device / component). Examples of one or more motion output components 200-26C may include: motion controllers, actuators, mechanical linkages, electromechanical devices, and / or components that generate physical movement. This list is not intended to be exhaustive, and one or more motion output components 200-26C may include other motion output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output motion output. Figure 5 As illustrated, output components 200-26 include one or more haptic output components 200-26D. One or more haptic output components 200-26D may include any component for outputting (e.g., generating, creating, and / or displaying) and / or causing haptic output (e.g., output using haptically perceptible means, such as vibration, pressure, texture, and / or shape). Examples of one or more haptic output components 200-26D may include: speakers, components that generate vibrations, components that generate texture changes, components that generate pressure changes, and / or components that create perceptible haptic effects. This list is not intended to be exhaustive, and one or more haptic output components 200-26D may include other haptic output components not explicitly identified herein that detect, generate, and / or otherwise provide data that can be used (e.g., processed, stored, and / or transformed) to output haptic output.
[0091] like Figure 5As illustrated, output components 200-26 include one or more communication components 200-26E. The one or more communication components 200-26E may include any component (e.g., antenna, modem, network interface component, encoder, decoder, and / or communication protocol stack) for transmitting and / or receiving communications internal and / or external to the agent system 200-20. In some embodiments, communication may be between different devices and / or between components within the same device. In some embodiments, communication may include control signals and / or data (e.g., messages, instructions, files, application data, and / or media streams). In some embodiments, one or more communication components 200-26E include one or more features of one or more communication components 200-22B (e.g., as described above). In some embodiments, one or more communication components 200-26E are identical to one or more communication components 200-22B (e.g., handling communication inputs and outputs and therefore considered as one or more of input and output components and / or both).
[0092] Throughout this disclosure, reference may be made to moving output (e.g., referred to in various forms such as: movement, device movement, moving output, device motion, motion output, and / or motion output). In some embodiments, output movement (e.g., an output that causes movement) refers to movement of an electronic device (e.g., a portion or component thereof relative to another portion and / or the entire electronic device). For example, refer again to Figure 2B The movable output can refer to the device 200 actuating the movable component 200-3 to move the display part 200-1 to... Figure 2B The illustrated location (e.g., from) Figure 2A (Positioning within). In some embodiments, the motion output is not (e.g., excluding and / or not only including) tactile output (e.g., tactile motion output). In some embodiments, the motion output is not (e.g., excluding and / or not only including) vibration output. In some embodiments, the motion output is not (e.g., excluding and / or not only including) oscillatory motion (e.g., movement of an actuator that causes vibration solely by repeatedly moving a component along a path within the device). In some embodiments, the motion output includes (e.g., requiring and / or causing) a change in the position and / or pose of at least a portion (and / or all) of a component or electronic device. In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device from a first position and / or a first pose to a second position and / or a second pose. For example, relative to Figures 2A to 2C ,exist Figure 2A , Figure 2B and Figure 2CIn each of these embodiments, display portion 200-1 is shown in a different position (e.g., in space) and pose (e.g., relative to base portion 200-2). In some embodiments, the motion output includes an output that moves at least a portion (and / or all) of a component or electronic device to a third position and / or a third pose (e.g., from a first position and / or a first pose and / or from a second position and / or a second pose). In some embodiments, the third position and / or the third pose is the same as the first position and / or the first pose and / or the second position and / or the second pose. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement is made to return to... Figure 2A The illustrated first positioning. For example, the motion output may include... Figure 2A Equipment 200 from Figure 2A Starting from the first position shown in the example, move to Figure 2B The second position is illustrated, and the movement continues until it stops at... Figure 2C The illustrated third position.
[0093] Throughout this disclosure, electronic devices can be exemplified (and / or described) as being in different positions and / or poses at different times. For example, Figure 2A Example of device 200 in the first position, Figure 2B An example is shown of device 200 in the second position, and Figure 2A A device 200 in a third position is illustrated. In some embodiments, the electronic device moves itself between such positions and / or poses (e.g., using a movement output). For example, device 200 moves from a first position to a second position under its own power (e.g., using a power supply and one or more actuators to induce movement). Specifically, any examples of electronic devices illustrated and / or described herein in different positions and / or poses (e.g., at different times) should be understood to cover scenarios where the device moves itself between such positions and / or poses (e.g., unless otherwise explicitly stated).
[0094] Throughout this disclosure, reference may be made to “performing output,” “causing output,” and / or “output” (e.g., via one or more output generating devices and / or via one or more output generating components) (and / or similar phrases). In some embodiments, the output (e.g., or variations thereof) includes (and / or) output movement (e.g., moving the output as described above).
[0095] Throughout this disclosure, references may be made to “display,” “cause display,” and / or “output visual content” (e.g., via one or more display components) (and / or similar phrases). In some embodiments, display (e.g., or variations thereof) includes displaying visual content in conjunction with output movement (e.g., moving output as described above).
[0096] Throughout this disclosure, reference may be made to "output audio," "output that causes audio," and / or "provide audio output" (e.g., via one or more audio generation components and / or via one or more audio output devices) (and / or similar phrases). In some embodiments, outputting audio (e.g., or variations thereof) includes outputting audio content in conjunction with output movement (e.g., movement output as described above).
[0097] Throughout this disclosure, reference may be made to the movement (and / or similar phrases) of a digital avatar (e.g., or other representations of a displayed user, agent, and / or role) (e.g., via one or more display components). In some embodiments, moving a digital avatar (e.g., or a variant thereof) includes movement in conjunction with output movement (e.g., movement output as described above) to display visual content. For example, displaying an avatar nodding in agreement may include an electronic device moving in a manner similar to avatar movement (e.g., simulating a nod). In some embodiments, moving a digital avatar (e.g., or a variant thereof) includes movement with output movement (e.g., movement output as described above) without displaying visual content. For example, a device may perform a simulated nod without moving the displayed avatar's movement output (e.g., the avatar does not move relative to the display). Figure 5As illustrated, agent system 200-20 may optionally interface with external components such as external database 200-30, remote processing component 200-32, and / or remote management component 200-34. In some embodiments, external database 200-30 represents one or more functions that provide data storage resources accessible to agent system 200-20. In some embodiments, access to data in external database 200-30 is provided directly to agent system 200-20 (e.g., the agent system manages the database) and / or indirectly to agent system 200-20 (e.g., the database is managed by a different system, but the data stored therein can be provided and / or stored for use by agent system 200-20). In some embodiments, external database 200-30 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., is a database of web services accessible to different agent systems), and / or a combination of dedicated and non-dedicated database resources. In some embodiments, remote processing component 200-32 represents one or more components that serve as data processing resources accessible to agent system 200-20. In some embodiments, access to remote processing component 200-32 is provided directly to agent system 200-20 (e.g., the agent system manages the processing resources) and / or indirectly to agent system 200-20 (e.g., processing resources managed by a different system, but which can provide data processing for the benefit of agent system 200-20). In some embodiments, remote processing component 200-32 is dedicated to agent system 200-20 (e.g., for its use only), not dedicated to agent system 200-20 (e.g., a processing resource of a web service accessible to a different agent system), and / or a combination of both dedicated and non-dedicated processing resources. Examples of data processing include processing image data (e.g., for feature extraction and / or object detection), processing audio data (e.g., for processing natural language speech input via a large language model), and / or training machine learning algorithms and / or models. In some implementations, remote management component 200-34 represents functions including management functions and / or functions related to management functions. For example, such management functions may include providing component updates (e.g., software and / or firmware updates) to agent system 200-30, managing accounts (e.g., associated permissions, access controls, and / or preferences), synchronizing between different agent systems and / or their components (e.g., enabling agents accessible via multiple devices of a user to provide a consistent user experience across such devices), managing cooperation with other services and / or agent systems, error reporting, managing backup resources to maintain agent system reliability and / or agent availability, and / or other functions required for agent system 200-20 to perform operations (such as those described herein).
[0098] The above text is about Figure 5 The various components of the described intelligent agent system 200-20 represent functional blocks that represent functionality. This functionality can be implemented on the same and / or different hardware (e.g., physical components) and / or by the same and / or different software. For example, a functional block can be implemented using one or more physical components, devices (e.g., computer system 100 and / or device 200), and / or software programs. In other words, each functional block does not necessarily represent a single, discrete physical component, device, and / or software program, but can be implemented using one or more of these. Furthermore, the intelligent agent system 200-20 may include multiple implementations of the functionality represented by the respective functional blocks. For example, the agent system 200-20 may include multiple different model components representing ML models used in different contexts, multiple different API components representing different APIs for different services, and / or multiple different visual output components for outputting different types of visual output.
[0099] Now let’s turn our attention to a discussion of the concepts that may arise regarding the operation of intelligent agents.
[0100] As discussed throughout, an intelligent agent may be able to interact with a user. In some implementations, this capability includes the ability to process explicit requests, commands, and / or statements. In some implementations, explicit requests, commands, and / or statements include and / or are interpreted as instructions relating to completing a task (e.g., displaying X, completing task Y, and / or performing operation Z). In some implementations, the intelligent agent includes the ability to process implicit requests, commands, and / or statements. In some implementations, implicit requests, commands, and / or statements do not include explicit requests, commands, and / or statements. For example, “I like to go to Europe” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays an itinerary in response to the statement. As another example, “This picture is for my grandmother” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 displays a suggestion to modify the picture. As another example, “I am tired” can be interpreted as an implicit request, command, and / or statement, and in response to detection, device 200 initiates a meditation session with a sleep meditation application. As another example, "I miss my grandfather" can be interpreted as an implicit request, command, and / or statement, which, upon detection, device 200 can initiate a real-time communication session with the grandfather (e.g., a phone call, video call, and / or text messaging session). In some implementations, implicit requests are more likely to be processed based on one or more current contexts, operational contexts, and / or user contexts, while explicit requests are less likely to be processed based on one or more current contexts, operational contexts, and / or user contexts. For example, the phrase "Call my grandfather" can be an explicit request, and in response to detecting such a request, device 200 will initiate a real-time communication session with the grandfather regardless of one or more current contexts, operational contexts, and / or user contexts. However, the phrase "I miss my grandfather" can be an implicit request, and in response to detecting such a request, device 200 can display a list of gifts to buy for the grandfather if the user has recently been discussing buying gifts, or can call the grandfather in a different context that does not include the user's recent discussions about buying gifts. In some implementations, a request can include one or more explicit requests and one or more implicit requests. In some implementations, implicit requests are responded to independently of explicit requests; in other implementations, responses to implicit requests depend on explicit requests.
[0101] This document may refer to responses of an intelligent agent output by a device. In some embodiments, the response includes an audio component (e.g., audio output, acoustic output, sound and / or speech) (also referred to herein as a “speech response,” “audio response,” and / or “acoustic response”) and / or a visual component (e.g., a display and / or movement of a representation and / or avatar). In some embodiments, the response includes a motion component (e.g., movement of the device). In some embodiments, the response includes a tactile component (e.g., touch and / or vibration).
[0102] This document may refer to internal dialogue, internal context, and / or operational context, which may refer to the dynamic context or dynamic decision-making process of a device, the internal state of device 200, and / or internal data of the device based in part on its decisions. In some embodiments, internal dialogue includes a set of one or more rules, features, detections, and / or observations used by a computer system to generate responses to one or more commands, questions, and / or statements. In some embodiments, the set of one or more rules, features, detections, and / or observations is learned and / or generated via deep learning and / or one or more machine learning algorithms and / or using one or more machine learning and / or system agents. In some embodiments, internal dialogue is generated in real time. In some embodiments, internal dialogue is stored locally and / or via cloud storage. In some embodiments, internal dialogue can be modified, updated, and / or deleted. In some embodiments, internal dialogue is generated based on other internal dialogues.
[0103] This document may refer to (e.g., the personality and / or behavior of an agent, user, and / or role) of a person or entity (or a representation of personality / behavior). In some embodiments, personality and / or behavior refers to one or more characteristics that a device detects, understands, conforms to, applies, and / or tracks. In some embodiments, personality or behavior is used as the basis for performing operations. For example, an agent may detect a user's personality and respond in a personality-based manner (e.g., outputting different responses in response to different user personalities). As another example, an agent may output a response having characteristics corresponding to one or more characteristics corresponding to personality and / or behavior (e.g., outputting responses in different ways depending on the agent's personality). In some embodiments, such characteristics represent and / or simulate a user's personality, such as how the user acts and / or speaks. In some embodiments, such characteristics approximate a user's personality.
[0104] In some embodiments, the intelligent agent is a system intelligent agent. In some embodiments, the system intelligent agent is an intelligent agent corresponding to an operating system originating from the device (e.g., the device implementing the intelligent agent) and / or a process controlled by the device's operating system. In some embodiments, the intelligent agent is an application intelligent agent. In some embodiments, the application intelligent agent is an intelligent agent corresponding to an application originating from the device (e.g., the device implementing the intelligent agent) (e.g., installed on and / or executed by the device) and / or a process controlled by the device's application.
[0105] This document may refer to representations of agents (e.g., and / or users (e.g., people, objects, and / or animals) and / or user interface objects (e.g., animated characters)) (e.g., digital avatars and / or digital avatar representations). In some embodiments, the representation of an agent refers to a set of output characteristics (e.g., visual and / or audio) of the agent (and / or user and / or user interface object). For example, the representation of an agent may include (and / or correspond to) a set of one or more visual characteristics (e.g., facial features of an animated face) and / or one or more audio characteristics (e.g., language and speech characteristics of audio output). In some embodiments, the representation (e.g., of the agent) is used to represent the agent's output. For example, a device implementing an interactive agent outputs audio in the agent's voice and displays an animated face of the agent moving in a manner that simulates the agent speaking the audio output. In this way, the user can feel that they are having a normal conversation with the agent. In some embodiments, the representation of an agent includes (or does not include) personality and / or behavioral characteristics (e.g., as described above). For example, the representation of an agent may include (and / or correspond to) a set of visual characteristics (e.g., facial features of an animated face) and a set of personality characteristics. In some implementations, the representation of the intelligent agent includes a set of user characteristics corresponding to the user's visual representation (e.g., representations of the user's appearance, voice, and / or personality used as digital images that appear to move and / or speak). In some implementations, the representation is a facial representation (e.g., a user interface object outputting features that simulate a human face and / or facial expressions (e.g., for conveying information to a viewer)).
[0106] In some implementations, a role (e.g., the role of an agent and / or digital avatar) refers to a specific set of characteristics represented. For example, an avatar may embody the characteristics (e.g., use, application, interaction with, and / or output) of fictional and / or non-fictional characters (e.g., from movies, shows, books, TV series, and / or popular culture).
[0107] In some implementations, (e.g., agent and / or digital avatar) speech refers to one or more characteristics corresponding to sound outputs that are similar to (e.g., representing, imitating, and / or reproducing) spoken language (e.g., attributable to and / or simulated as output by an agent and / or digital avatar). For example, device 200 may output sentences that sound different depending on the speech used. In some implementations, a particular character and / or digital avatar may be configured to use a particular speech (e.g., have a corresponding speech). In some implementations, the particular speech may mimic a user's speech.
[0108] In some implementations, the appearance (e.g., of an agent and / or digital avatar) refers to a set of one or more characteristics corresponding to the visual output representing the digital avatar (and / or agent). For example, device 200 may output an avatar having a set of facial features that form an appearance similar to a specific character from a movie.
[0109] In some implementations, the expression of a digital avatar refers to one or more characteristics corresponding to a specific visual appearance of a user, digital avatar, and / or intelligent agent. For example, device 200 may output an avatar having a set of facial features arranged in a specific manner to give the appearance of a facial expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a frown is an expression of sadness, a smile is an expression of happiness, and / or wide eyes are an expression of surprise). As another example, device 200 may output a digital avatar having a set of body features (e.g., arms and / or legs) arranged in a specific manner to give the appearance of a body expression (e.g., which can be used as a form of nonverbal communication to the user) (e.g., a gesture is an expression of agreement, covering the eyes is an expression of fear, and shrugging is an expression of lack of knowledge). In some implementations, expressions include movement of the digital avatar (e.g., a nod is an expression of agreement and / or disagreement). In some implementations, device 200 may be movable via a motion component to indicate expressions accompanying or not accompanying movement of the digital avatar. In some implementations, the agent performs one or more operations that depend on the user's facial expressions (e.g., detecting whether a person is sad and responding with a kind statement or question). In some implementations, facial expressions (e.g., whether they are used, how they are used, and / or how they are output) depend on personality. For example, a first-person sex may use more specific facial expressions than a second-person sex. As another example, a first-person sex's expressions (e.g., frowning, smiling, and / or widening their eyes) may look different from a second-person sex's expressions (and / or similar and / or equivalent expressions) (e.g., a first-person sex smiles with their teeth showing, but a second-person sex smiles without showing their teeth).
[0110] In some implementations, an agent (e.g., a digital image of the agent and / or an agent system implementing the agent (e.g., hardware and / or software)) mimics characteristics of another user, agent, and / or role (e.g., in terms of personality, behavior, facial expressions, and / or voice). In some implementations, mimicry includes mirroring the user (e.g., copying the use of phrases and / or movements detected from a user interacting with the agent). In some implementations, simulating user characteristics includes attempting to reproduce the user's characteristics (e.g., in exactly the same way and / or in a way that is similar to the characteristics but not an exact reproduction of the characteristics). For example, an agent mimicking voice and / or facial expressions does not require the agent to have exactly the same voice and / or facial expressions as the user being mimicked (e.g., simply to resemble the user's voice and / or facial expressions).
[0111] In some implementations, components and / or devices use (e.g., performing actions, making decisions, and / or determining context based on them) learned characteristics (e.g., characteristics of the context, user, and / or environment learned by the device over time (e.g., via detection, prior experience, and / or feedback (e.g., from one or more users))). For example, characteristics learned over time may include user routines. In such an example, if a particular user requests a summary of any new messages for that user from the agent at the same time every day, the agent may learn to automate actions based on the characteristics of the learned routines (e.g., what data is needed, when data is needed, and / or for which user). In some implementations, the learned characteristics enable the agent (and / or device) to improve its understanding (and / or response to) of the context, user, and / or environment, and / or its understanding of context, user, and / or environment that is otherwise not (and / or will not) understood (e.g., not responded to or responded to incorrectly). In some implementations, the learned characteristics are formed using reinforcement learning (e.g., by and / or for the agent). In some implementations, the learned features correspond to one or more confidence levels, determinism, and / or rewards (e.g., shaped by one or more reward functions). In some implementations, the learned features (and / or how they are used to influence the output of the agent and / or device) can change over time (e.g., confidence levels, determinism, and / or rewards change over time). For example, the output of a device before learning a set of learned features may differ from the output of a device after learning a set of learned features. In some implementations, components and / or devices use the learned knowledge. For example, similar to what is described above regarding learned features, learned knowledge may refer to information used to update (e.g., enhance, add to, and / or expand) the device's knowledge base (e.g., for use by agents implemented thereon). In some implementations, multiple sets of learned features for a user may be stored and / or used. In some implementations, different sets of learned features for different users may be stored and / or used.
[0112] This document may refer to interactions with an agent (and / or device). In some implementations, an interaction refers to a set of one or more inputs and / or outputs from a device implementing the agent and one or more users. For example, an interaction may be a user input (e.g., “Please turn on the light”) and a corresponding output (e.g., turning on the light and / or the device’s response “OK”). In some implementations, an interaction may include multiple inputs / outputs performed by one or more parties to the interaction (e.g., a device and / or a user). For example, an interaction may include a first user input (e.g., “Please turn on the light”) and a corresponding first output (e.g., “Which lights?”), and also include a second user input (e.g., “Kitchen light”) and a second output from the device (e.g., “OK”). In some implementations, which inputs and / or outputs are considered together as an interaction is based on logical and / or contextual grouping (e.g., interactions within the previous thirty (30) seconds and / or interactions related to turning on the light). As those skilled in the art will understand, interactions may be considered in an implementation-dependent manner (e.g., determining when an interaction is complete may involve determining whether the user is still present (e.g., is still talking) and / or whether the user is still talking about the light or has moved on to a different topic). In some implementations, the interaction is the current interaction (e.g., ongoing, currently occurring, and / or active). In some implementations, the interaction is a previous interaction. The examples above describe a device that engages in dialogue with a user. In some implementations, the dialogue is between two or more users (e.g., users in an environment). For example, the device may detect dialogue between users (e.g., users directing speech and responses to each other, rather than to the device).
[0113] In some implementations, the agent (and / or device) determines and / or performs an action based on an intent corresponding to the user. For example, the device detects user input and outputs a response that depends on the intent of the user input. For example, the device detects user input including a pointing gesture detected along with a verbal command to “turn on the light,” and in response, the device turns on the light determined to correspond to the intent of the input (e.g., the light the pointing gesture is pointing to). In some implementations, intent is determined using one or more of the following (e.g., determined by the device that detects the input and / or by one or more other devices): one or more inputs, knowledge (e.g., knowledge about the user learned based on observed behavior, personality, and history of interactions), learned characteristics, and / or context. In some implementations, intent is determined based on one or more types of input (e.g., verbal input, visual input via a camera, and / or contextual input).
[0114] Now turn our attention to the implementation of user interfaces (“UIs”) and associated processes on electronic devices (such as computer system 100 and / or electronic device 200).
[0115] Figures 6A to 6M Exemplary user interfaces for capturing content based on detected input, according to some implementation schemes, are illustrated. The user interfaces in these figures are used to illustrate the processes described below, including... Figures 7 to 9 The process in.
[0116] Figures 6A to 6M The left side illustrates a computer system 600 as a tablet computer displaying various user interface objects. It should be recognized that the computer system 600 can be other types of computer systems, such as smartphones, smartwatches, laptops, utilities, smart speakers, accessories, personal gaming systems, desktop computers, fitness trackers, and / or head-mounted display (HMD) devices. In some embodiments, the computer system 600 includes one or more input devices and / or sensors (e.g., cameras, LiDAR sensors, motion sensors, infrared sensors, touch-sensitive surfaces, physical input mechanisms (such as buttons or sliders), and / or microphones) and / or communicates with such input devices and / or sensors. Such sensors can be used to detect the presence of a user in the environment, attention, statements from the user, inputs corresponding to the user, requests from the user, and / or instructions from the user. It should be recognized that while some embodiments described herein refer to input as voice input, other types of input can be used with the techniques described herein, such as touch input detected via touch-sensitive surfaces and / or air gestures detected via cameras (e.g., cameras communicating with the computer system 600 (e.g., wireless and / or wired communication)). In some embodiments, computer system 600 includes one or more output devices (e.g., a display screen, projector, touch-sensitive display, speaker, and / or movable components) and / or communicates with said one or more output devices. Such output devices are used to present information and / or cause different visual changes to computer system 600. In some embodiments, computer system 600 includes one or more movable components (e.g., actuators, movable bases, rotatable components, and / or rotatable bases) and / or communicates with said one or more movable components. Such movable components are used to change the positioning (e.g., location and / or orientation) of the entire computer system 600 or a portion of the computer system 600 (e.g., including one or more sensors, input components, and / or output components). In some embodiments, computer system 600 includes one or more components and / or features described above with respect to computer system 100 and / or device 200. In some embodiments, computer system 600 includes as described above with respect to... Figure 5The description refers to one or more intelligent agents and / or the functionality of intelligent agents. In some embodiments, computer system 600 is, includes, implements, or communicates with one or more intelligent agent systems, as described above. Figure 5 As described, one or more operations are performed (and / or cause to be performed) by an agent.
[0117] Figures 6A to 6M The right side includes Figure 618. Figure 618 is a visual aid representing the physical space and / or environment including computer system 600 and a first user. Figure 618 includes a computer system representation 620 representing computer system 600 and a user representation 622 representing the first user. The positioning of computer system representation 620 and user representation 622 within Figure 618 represents the real-world positioning of computer system 600 relative to the first user. Figure 618 includes a detection field representation 654 representing the detection field and / or field of view (sometimes referred to as the detection field of camera 602) of one or more camera sensors (e.g., camera 602) of computer system 600. The detection field corresponds to the detection field of one or more forward sensors of computer system 600. In some embodiments, one or more other sensors of computer system 600 have a similar sensor configuration to those of computer system 600. Figure 6A The detection field representation 652 illustrated herein is different from the detection field represented by 654 (e.g., overlapping but smaller or larger and / or non-overlapping). In this embodiment, there is one user in the environment. In some embodiments, there is more than one user in the environment. Figure 618 also includes a detection field range representation. Detection field range representation 652 indicates the range of the detection field of camera 602 (e.g., one or more sensors of computer system 600) (e.g., sometimes referred to as the global detection field) (e.g., via a different set (such as a wide-angle setting) and / or a different set of one or more sensors compared to those used with respect to detection field representation 654) and / or the limits of the range of motion of detection field representation 654 (e.g., via a movable sensor and / or one or more moving components described above).
[0118] Figures 6A to 6MAn example is illustrated of a process in which a user requests a computer system 600 to capture media in the form of photographs and / or videos. In some embodiments, the request includes instructions on how the computer system 600 should capture still photographs and / or videos. In such embodiments, the computer system 600 guides the user differently based on (1) the specificity of the request and / or (2) prior interactions with the user. In some embodiments, the request also includes instructions on how the computer system should mimic the style of a media artist and / or other predefined styles when capturing media. When capturing media based on style, the computer system 600 uses different settings (such as camera and / or movement) to capture media according to the requested style. In some embodiments, the computer system 600 moves one or more camera sensors to change the framing or alignment of the user, features, and / or objects within the detection field of one or more camera sensors and / or the computer system 600 to recreate the requested style. In some embodiments, the computer system 600 outputs composition instructions to the user to help recreate the requested style. For example, in some implementations, in order to recreate the requested style, computer system 600 instructs the user to move the object to the right, to the left, and / or to move it within the detection field of one or more camera sensors.
[0119] like Figure 6A As illustrated, computer system 600 includes a camera 602 and a display 604, both of which are forward-facing (e.g., facing the first user). Figure 6A As illustrated, the display 604 occupies most of the front-facing side of the computer system 600. For example... Figure 6A As illustrated, camera 602 is located above display 604. In some embodiments, camera 602 is located at different positions relative to display 604. For example, in some embodiments, camera 602 is located near one of the side edges of display 604. Figure 6A As illustrated, camera 602 is located external to display 604. In some embodiments, camera 602 is located internally or behind display 604. In some embodiments, camera 602 includes one or more camera sensors. In some embodiments, camera 602 includes multiple camera sensors with different designs and capabilities (e.g., wide-angle, fisheye, macro, telephoto, and / or standard lenses).
[0120] like Figure 6AAs illustrated, computer system 600 displays camera user interface 606 via display 604. Camera user interface 606 includes a live view 608 that includes a representation of objects and / or individuals within the detection field of one or more camera sensors (e.g., camera 602) connected to and / or communicating with computer system 600 and / or the computer system. Displaying live view 608 within camera user interface 606 allows a user of computer system 600 to view a preview of content to be included in resulting media items captured by computer system 600. In some embodiments, live view 608 includes representations of the detection fields of multiple different camera sensors. In some embodiments, live view 608 occupies a large portion of camera user interface 606 and optionally extends to the edges of camera user interface 606. In some embodiments, live view 608 occupies the entire camera user interface 606.
[0121] like Figure 6A As illustrated, the camera user interface 606 includes a viewfinder 610, camera controls 612, and a content indicator 614. For example... Figure 6A As illustrated, computer system 600 displays camera controls 612 and a content indicator 614 below viewfinder 610. Viewfinder 610 indicates the content that will be captured by computer system 600 during the media capture process performed by computer system 600 (e.g., content within viewfinder 610 will be visible in the resulting media captured by computer system 600, and content outside viewfinder 610 will not be visible in the resulting media captured by computer system 600). In some embodiments, computer system 600 displays a live view 608 extending beyond viewfinder 610 to indicate content within the detection field of one or more camera sensors but not visible in the resulting media captured by computer system 600. In some embodiments, computer system 600 moves viewfinder 610 within camera user interface 606 to indicate different viewfinder scales when (e.g., by the user and / or by computer system 600) selects different camera modes. In some implementations, in response to the detection of touch input at a location corresponding to camera control 612, computer system 600 initiates a media capture operation to capture content located within viewfinder 610.
[0122] like Figure 6A As illustrated, computer system 600 displays a representation (e.g., first content 616) of recently captured media content (e.g., an image of a flower) within content indicator 614. In some embodiments, in response to detecting touch input at a location corresponding to content indicator 614, computer system 600 displays a magnified version of the recently captured media content.
[0123] like Figure 6AAs illustrated in Figure 618, user representation 622 is within detection field representation 654, thereby indicating that the first user is within the detection field of one or more camera sensors (e.g., the detection field of camera 602). Therefore, in Figure 6A At this location, the first user is within the detection field of one or more camera sensors. For example... Figure 6A As illustrated, because the first user is within the detection field of one or more camera sensors, the computer system 600 displays a user image 624 within a live view 608. The user image 624 is a representation of the first user detected by one or more camera sensors (e.g., camera 602). Figure 6A At this point, computer system 600 detects a first verbal input 605a from a first user, which corresponds to a request from the first user for computer system 600 to capture a first image. It should be understood that the request to capture the first image can be other types of input besides verbal input, such as a tap input detected via display 604 and / or an air gesture (e.g., air pointing, air swiping, pinch-off gesture and / or pinch-on gesture) detected via camera 602.
[0124] like Figure 6B As illustrated, in response to detecting a first verbal input 605a, computer system 600 outputs first audio content 626 corresponding to the query from computer system 600 to the first user regarding the style in which the first image is captured (e.g., how the first user wishes to be framed, at what zoom level the computer system 600 should capture the still image, and / or whether the computer system 600 should track the first user). The content included in the first audio content 626 is based on the content detected in the first verbal input 605a. Given the simplicity of the first verbal input 605a, computer system 600 outputs the first audio content 626, which includes content for extracting additional information from the first user regarding the type of photograph the first user wishes the computer system 600 to capture. Examples of additional information that computer system 600 can query from the first user include the type of photograph, the style of the photograph, the orientation of the photograph, and / or the zoom level of the photograph. As explained in more detail below, when the first user provides more information to computer system 600 when prompting computer system 600 to capture a photograph, computer system 600 does not attempt to extract additional information or attempts to extract less information from the first user.
[0125] exist Figure 6BAt a point after the first audio content 626 is output, the computer system 600 detects a compositional instruction 605b from the first user, which corresponds to an instruction from the first user to the computer system 600 regarding the style in which the first user wishes to capture the first image. The compositional instruction 605b instructs the computer system 600 to capture the first user in a style corresponding to an individual named Kyle. Kyle is an individual and / or artist whose subject is not tracked before and / or during media capture. In response to the detection of the compositional instruction 605b, the computer system 600 configures itself not to track the first user when the first user moves within the environment. In some embodiments, Kyle is a known artist with a recognizable artistic style who has not previously interacted with the computer system 600. In some embodiments, Kyle is a second user known to the computer system 600 whose image capture habits are known to the computer system 600.
[0126] In some implementations, composition guidelines are based on an art style specifically requested by the user (e.g., a particular artist's style, art movement, art period, art technique, and / or art genre). For example, in some implementations, the first user specifies that they want the captured media to resemble a suspense novel cover. In some implementations, composition guidelines are based on specific conditions indicated by the user. For example, in some implementations, the first user specifies that they want to capture a black-and-white portrait of themselves taken from above with a soft-focus filter. As another example, in some implementations, the first user specifies that they want an image of their hand to showcase their new championship ring, thereby instructing the computer system 600 to ensure that the words on the ring are legible.
[0127] In some implementations, the composition guide includes instructions for the computer system 600 to begin and / or stop the media capture process once a user is detected in a specific pose. For example, in some implementations, the composition guide includes instructions such as "start recording video once I sit down," causing the computer system 600 to wait until it detects that the first user has sat down before capturing media content. As another example, in some implementations, the composition guide includes instructions such as "stop taking pictures when I stand up," causing the computer system 600 to stop capturing media content when it detects that the first user has stood up. In some implementations, the composition guide includes instructions for the computer system 600 to begin the media capture process once it detects that the first user has performed a gesture. In some implementations, the composition guide includes instructions for the computer system 600 to begin the media capture process once it detects that the first user is looking in the direction of the computer system 600.
[0128] In some implementations, the computer system 600 detects a framing guide that includes instructions indicating time boundaries for capturing media content. In some implementations, the indicated time boundaries are based on camera-detectable input, such as pose, gesture, gaze, and / or facial expressions (e.g., a first user raising three fingers, indicating that the computer system 600 should initiate a media capture operation after three seconds). In some implementations, the time boundaries include information about when the computer system 600 should begin capturing media content. For example, in some implementations, the framing guide includes instructions such as “take a photo in forty-five seconds,” causing the computer system 600 to wait forty-five seconds before initiating media capture. In some implementations, the framing guide includes information about when the computer system 600 should stop capturing media content. For example, in some implementations, the framing guide includes instructions such as “stop capturing video in thirty seconds,” causing the computer system 600 to stop capturing media thirty seconds after initiating media capture. In some implementations, the indicated time boundaries include information about the intervals at which the computer system 600 should capture different media. For example, in some implementations, the composition guidelines include information such as "take a picture every six seconds," causing the computer system 600 to capture one image every six seconds. As another example, in some implementations, the composition guidelines include information such as "capture two seconds of video each time I hit the ball," causing the computer system 600 to capture two seconds of video each time the computer system 600 detects the first user hitting the ball. In some implementations, the indicated time boundaries include a combination of the aforementioned media content capture start, stop, and / or interval information. For example, in some implementations, the composition guidelines include information such as "start taking pictures when I pick up the flower, one picture per second, for two minutes."
[0129] In some implementations, the method of capturing media is automatically selected by the computer system 600. In some implementations, the method of capturing media is automatically selected by the computer system 600 based on the detected scene. For example, if the computer system 600 detects that the first user is wearing a graduation cap and gown, the computer system 600 automatically selects to capture the media using portrait settings that highlight the graduate, such as blurring the background and increasing color saturation so that the gown does not appear faded. As another example, if the computer system 600 detects that the first user is posing too dramatically in strong light, the computer system 600 automatically selects to capture the media in black and white using a grainy filter to mimic the look of an old film. In some implementations, composition guidelines are automatically selected based on media capture settings (such as whether to capture the media in black and white or in color). In some implementations, composition guidelines are automatically selected based on system default media capture rules, such as selecting to capture the media with a certain amount of color, light, and / or contrast balance.
[0130] In some embodiments, as described in more detail below, computer system 600 moves a portion of computer system 600 via one or more moving components to track a first user. In such embodiments, moving a portion of computer system 600 and thereby moving camera 602 causes (e.g., represented by detection field representation 654) the detection field of camera 602 to move. In some embodiments, computer system 600 uses multiple types of movement (e.g., simultaneous or sequential movement) to move a portion of computer system 600, including directional movement (e.g., left, right, up, down, forward and / or backward), rotational movement (e.g., yaw, roll and / or pitch), and / or positioning movement (e.g., extending, shortening, folding, tilting and / or tilting).
[0131] exist Figure 6C As indicated by User 622 being positioned on the right side within Figure 618, the first user moves to the right within the detection field of camera 602 in one or more camera sensors. Figure 6C At this point, it is determined that the first user has moved to the right within the detection field of one or more camera sensors. Figure 6C Because computer system 600 is configured not to track the first user (e.g., in response to detecting composition guide 605b), computer system 600 does not move a portion of itself in response to the first user moving to the right (e.g., a first movement pattern). That is, because composition guide 605b requests computer system 600 to capture images like Kyle, and because Kyle does not track the movement of its subject when capturing media, computer system 600 does not track the first user when configured to capture media based on composition guide 605b. In some embodiments, as described in more detail below, computer system 600 automatically moves a portion of itself in response to detecting movement of the first user.
[0132] like Figure 6C As illustrated, because the first user is within the right-hand portion of the detection field of camera 602, and computer system 600 does not move any part of computer system 600, computer system 600 displays user image 624 within the right-hand portion of live view 608.
[0133] exist Figure 6D As indicated by User 622 being positioned on the left side within Figure 618, the first user moves to the left within the detection field of one or more camera sensors. Figure 6D At this point, it is determined that the first user has moved to the left within the detection field of computer system 600. Figure 6DBecause computer system 600 has configured itself not to track the first user (e.g., in response to detecting composition guide 605b), computer system 600 does not move a portion of computer system 600 to move the detection field of camera 602 in one or more camera sensors based on the movement of the first user. Figure 6D As illustrated, because the first user is within the left portion of the detection field 654 of the computer system 600, and the computer system 600 does not move any part of the computer system 600, the computer system 600 displays the user image 624 within the left portion of the live view 608.
[0134] exist Figure 6D At a time after the first user moves to the left within the detection field of camera 602, the first user stops moving and maintains a pose. Examples of the first user maintaining a pose include the first user holding facial features, limbs, extremities, and / or torso in a certain position for a period of time. A predetermined time period for the first user to maintain a pose is determined after the time after detecting the first user's leftward movement within the detection field of camera 602. In some embodiments, computer system 600 detects the first user's maintained pose based on the rate of change of the first user's position detected by camera 602 over time. In some embodiments, computer system 600 automatically (e.g., without intermediate user input) captures a photograph of the first user based on determining that the first user has stopped moving and / or maintained a pose for a predetermined time period.
[0135] exist Figure 6D Based on determining that the first user maintains a pose for a predetermined period of time, the computer system 600 outputs second audio content 628 corresponding to a query from the computer system 600 to the first user asking if the first user is ready to capture a first image (e.g., are you ready?). In some embodiments, the computer system 600 outputs the second audio content 628 in response to detecting a predetermined and / or known pose of the first user (e.g., a pose that the computer system 600 identifies as a pose the first user frequently assumes and / or a pose that the computer system 600 has been trained to recognize and seek). In some embodiments, the computer system 600 outputs the second audio content 628 in response to detecting a gaze and / or gesture of the first user in a specific direction. Figure 6DAt a time after the output of the second audio content 628, the computer system 600 detects a third verbal input 605d from the first user corresponding to a positive response (e.g., yes) to the second audio content 628. It should be understood that a positive response can be other types of input besides verbal input, such as a tap input detected via the display 604 and / or an air gesture (e.g., air pointing, air swiping, pinch-off gesture, and / or pinch-on gesture) detected via the camera 602.
[0136] exist Figure 6E At that point, in response to the detection of the third verbal input 605d, the computer system 600 outputs the third audio content 630 corresponding to the countdown. Figure 6E In response to the computer system 600 reaching the end of a countdown within the third audio content 630, the computer system 600 captures a first image including the content within the viewfinder 610. After capturing the first image, the computer system 600 displays a representation of the first image (e.g., second content 632) in the content indicator 614 and stops displaying a representation of the previously captured media content (e.g., first content 616). In some embodiments, in response to detecting that a first user has maintained a pose for a predetermined time threshold, the computer system 600 outputs the third audio content 630 corresponding to the countdown without outputting the second audio content 628 or without detecting the third verbal input 605d. In some embodiments, in response to detecting the third verbal input 605d after outputting the first audio content 626, the computer system 600 captures media content without outputting the third audio content 630. In some embodiments, the computer system 600 displays a countdown.
[0137] In some embodiments, while computer system 600 is outputting third audio content 630 corresponding to a countdown, computer system 600 detects a verbal command and / or air gesture from a first user corresponding to a request to capture media. For example, if the first user does not want to wait until the countdown ends before capturing the media, such as when the first user finds it difficult to maintain a pose, the first user can verbally request computer system 600 to capture the media. In some embodiments, in response to detecting a verbal command and / or air gesture from the first user corresponding to a request to capture media while outputting third audio content 630, computer system 600 stops outputting third audio content 630 and captures the media. In some embodiments, while computer system 600 is outputting third audio content 630 corresponding to a countdown, computer system 600 detects a verbal command and / or air gesture corresponding to a request to cancel media capture. For example, the first user can change their preferred composition of the media as seen in live view 608 and perform an air gesture corresponding to a request to cancel media capture. In some embodiments, in response to detecting a verbal command and / or air gesture corresponding to a request to cancel media capture operation when outputting third audio content 630, computer system 600 stops outputting third audio content 630 and does not capture media. In some embodiments, in response to detecting a verbal command and / or air gesture corresponding to a request to cancel media capture operation when outputting third audio content 630, computer system 600 outputs (e.g., visual output and / or audible output) an indication that computer system 600 will not capture media.
[0138] In some implementations, the composition guidelines include conditions for the computer system 600 to perform the media capture process. In some implementations, if the conditions included in the composition guidelines are met, the computer system 600 automatically (e.g., without intermediate user input) captures the media content. In some implementations, if the conditions included in the composition guidelines are not met, the computer system 600 does not capture the media. In some implementations, while the computer system 600 is outputting third audio content 630 corresponding to a countdown, the computer system 600 detects a change in the environment that causes the conditions included in the composition guidelines for the desired media to no longer be met. Examples of environmental changes that may cause the conditions included in the composition guidelines to no longer be met include changes in lighting (e.g., the sun hides behind clouds, blinds are opened, a flashlight is turned on and / or lights are turned off), changes in one or more of the characteristics and / or positioning of the first user (e.g., the first user stops smiling, looks away from computer system 600, changes positioning and / or moves out of the detection field of camera 602), and / or physical changes in the environment (e.g., a family pet runs across the detection field of one or more camera sensors, a sanitation vehicle appears in the background, and / or an object falls over). In some embodiments, in response to detecting an environmental change that causes the conditions included in the composition guidelines to no longer be met when outputting the third audio content 630, computer system 600 stops outputting the third audio content 630 and does not capture the media.
[0139] exist Figure 6F As indicated by the positioning of user representation 622 together with the right portion of detection field representation 654, the first user moves to the right within the detection field of one or more camera sensors. Figure 6F At this point, it is determined that the first user has moved to the right within the detection field of one or more camera sensors. Figure 6F Because computer system 600 has configured itself not to track the first user (e.g., in response to detecting drawing guide 605b), computer system 600 does not move a portion of computer system 600 in response to the first user moving to the right. Figure 6F As illustrated, because the first user moves to the right within the detection field of one or more camera sensors, and the computer system 600 does not move any part of the computer system, the computer system 600 displays the user image 624 on the right side of the live view 608. Figure 6F At this point, the computer system 600 detects that the first user maintains the second pose for a predetermined amount of time.
[0140] exist Figure 6FIn response to detecting that a first user holds a second pose for a predetermined amount of time after the first image is captured, computer system 600 captures a second image. That is, after computer system 600 captures the first image, it captures the second image in response to detecting that the first user holds a second pose without detecting another verbal command and / or air gesture. In some embodiments, computer system 600 detects that the first user is in a series of poses and, in response, captures media content corresponding to each detected pose. In some embodiments, computer system 600 detects that the first user holds a second pose within a predetermined threshold amount of time after capturing the first image and, in response, captures additional media content. In some embodiments, computer system 600 does not detect that the first user holds a second pose within a predetermined threshold amount of time after capturing the first image and, in response, does not capture additional media content. For example, if computer system 600 detects that the first user holds a second pose for too long after capturing the first image, then computer system 600 does not capture the second image.
[0141] In some implementations, after capturing media content, computer system 600 detects the pose of a first user (which is a pose in a known set of poses), and in response, computer system 600 captures media content. For example, if computer system 600 detects a pose from a set of poses that computer system 600 has been configured to detect, then computer system 600 captures media. In some implementations, after capturing media content, computer system 600 detects the pose of a first user (which is not a pose in a known set of poses), and in response, computer system 600 does not capture additional media content.
[0142] like Figure 6F As shown, in response to capturing a second image, computer system 600 displays a representation of the second image (e.g., third content 634) in content indicator 614 and stops displaying a representation of the first image in content indicator 614. Figure 6FAt this point, computer system 600 detects a composition guide 605f from a first user that corresponds to the first user's command to capture a photograph in a specific style (e.g., "capture a photograph like Jane"). Jane is the individual who tracks their subject before and / or during media capture. Furthermore, Jane is the individual who captures the subject once it is centered within the detection field of one or more camera sensors. In response to the detection of composition guide 605f, computer system 600 configures itself to track the first user as they move within the environment and to capture an image of the subject once it is centered within the detection field of one or more camera sensors. In some embodiments, composition guide 605f is similar to composition guide 605b.
[0143] exist Figure 6G In this process, because computer system 600 is configured to track the first user, computer system 600 moves a portion of itself in a second movement mode until the first user is centered within the detection field of one or more camera sensors (e.g., a response different from the response of computer system 600 to composition guide 605b). The second movement mode differs from the first movement mode (e.g., as referenced above). Figure 6C (As described). In some embodiments, the second movement mode is the same as the first movement mode. In some embodiments, the computer system 600 automatically moves a portion of the computer system 600 in response to detecting movement of the first user.
[0144] In some embodiments, the movement pattern includes two or more types of movement (e.g., directional movement, rotational movement, and / or positioning movement) (e.g., lateral movement (lateral movement, forward movement, backward movement, and / or vertical movement) and / or rotational movement (e.g., clockwise rotation and / or counterclockwise rotation)). For example, in some embodiments, the movement pattern includes rotating a portion of the computer system 600 to the left while extending a portion of the computer system 600 forward. As another example, in some embodiments, the movement pattern includes movement in two different lateral directions (e.g., right and left and / or up and down). In some embodiments, different movement patterns have different movement speeds. In some embodiments, different movement patterns have the same movement speed. In some embodiments, a movement pattern has more than one movement speed. For example, a movement pattern may begin with a slow movement and end with a fast movement. In some embodiments, different movement patterns have different user and / or object following parameters (e.g., head, torso, hand, and / or within a three-part region). For example, the computer system 600 follows the user's hand at a distance based on the current movement pattern of a portion of the computer system 600. In some implementations, different movement patterns have the same user and / or object following parameters. In some implementations, a movement pattern has more than one user and / or object following parameter. For example, a movement pattern may have a following parameter for the user's torso and a following parameter for the user's head. In some implementations, the computer system 600 moves a portion of itself using multiple types of movement (e.g., simultaneous or sequential movement), including directional movement (e.g., left, right, up, down, forward, and / or backward), rotational movement (e.g., yaw, roll, and / or pitch), and / or positioning movement (e.g., extension, retraction, folding, tilting, and / or leaning).
[0145] In some embodiments, computer system 600 moves a portion of itself to improve the composition of the media content. For example, computer system 600 moves a portion of itself so that the user is more easily detected by camera 602. In some embodiments, improving the composition of the media content includes changing the composition to meet compositional guidelines requested by a first user. For example, in some embodiments, in response to detecting a first user request to capture an image in an artist's style, which typically has a subject captured at a specific downward angle, computer system 600 moves a portion of itself such that camera 602 is positioned above the user. In some embodiments, improving the composition of the media content includes changing the perspective of the composition to meet compositional guidelines automatically selected by computer system 600. For example, in some embodiments, in response to detecting a first user request to capture an image of a user and a second user on a beach, computer system 600 automatically selects an art style based on the rule of thirds and automatically moves a portion of itself to align the two users along the left third of the detection field of one or more camera sensors and to align the horizon along the bottom third of the detection field of one or more camera sensors.
[0146] In some embodiments, in response to a portion of the mobile computer system 600 (e.g., to improve the composition of media content), the computer system 600 outputs audio content explaining to a first user why and / or how the computer system 600 moves a portion of itself. For example, in some embodiments, in response to a portion of the mobile computer system 600 capturing image content of a home and / or while a portion of the mobile computer system 600 is capturing image content of a home, the computer system 600 outputs audio content such as “I am moving to get everyone in the frame.” In some embodiments, the computer system 600 moves a portion of itself to follow a subject based on specific user instructions. For example, the computer system 600 moves a portion of itself in response to detecting that a first user says, “Follow me in the room and take a picture whenever I pose with a book or pen.” In some embodiments, moving a portion of computer system 600 to follow a user (e.g., to keep the user within and / or centered within the detection field of camera 602) includes following the first user entirely or a portion of the first user, such as the first user's head, eyes, shoulders, torso, hands, and / or feet. In some embodiments, moving a portion of computer system 600 to follow the first user causes computer system 600 to move in different movement patterns depending on the portion of the user being followed. For example, in some embodiments, as part of following the first user's head, computer system 600 moves a portion of computer system 600 in a uniform manner at the level of the first user's head, while as part of following the first user's hands, computer system 600 moves a portion of computer system 600 in a more sweeping manner before changing height according to the height of the first user's hands. In some embodiments, computer system 600 moves a portion of computer system 600 while capturing a panoramic image.
[0147] exist Figure 6G The system determines that the first user is located within the center of the detection field of one or more camera sensors. Based on this determination, the computer system 600 stops moving a portion of itself. Figure 6G As illustrated, computer system 600 displays a user image 624 centered within viewfinder 610. In some embodiments, in response to detecting that a first user is centered within the detection field of camera 602, computer system 600 repeats... Figures 6D to 6EThe steps described herein are: (1) outputting audio content corresponding to asking the first user whether they are ready to capture image content, (2) detecting verbal input corresponding to an affirmative response, (3) outputting audio content corresponding to a countdown, and / or (4) capturing an image once the computer system 600 reaches the end of the countdown.
[0148] exist Figure 6G At the location where the first user is detected to be centered within the detection field of camera 602, the time after which the first user maintains the pose for a predetermined amount of time is determined. Figure 6G Based on determining that the first user maintains the pose for a predetermined amount of time, the computer system 600 automatically (e.g., without intermediate user input) captures a third image without outputting audio content corresponding to a query to the first user asking if they are ready to capture image content. Figure 6G As illustrated, in response to capturing a third image, computer system 600 displays a representation of the third image (e.g., fourth content 636) within content indicator 614 and stops displaying a representation of the previously captured media content (e.g., third content 634). In some embodiments, when capturing media content, computer system 600 automatically moves a portion of itself to meet composition guidelines based on specific conditions indicated by the user. For example, in some embodiments, if a detected composition guideline instructs the computer system to track a first user based on the time of day, then computer system 600 tracks the first user at a distance based on the time of day. In some embodiments, before and / or while capturing media content, computer system 600 configures camera settings of camera 602 such that computer system 600 can meet the detected composition guidelines. For example, in some embodiments, computer system 600 configures camera 602 to capture photographs only in black and white mode in response to a detected composition guideline instructing computer system 600 to capture a still photograph.
[0149] In some embodiments, computer system 600 automatically captures a third image in response to detecting that a first user is in a predetermined and / or known pose. In some embodiments, computer system 600 automatically captures a third image in response to detecting that the first user maintains the pose for a predetermined amount of time and that the first user is centered within the detection field of camera 602 (e.g., detecting that media content satisfies a composition guide within composition guide 605f). In some embodiments, computer system 600 automatically captures image content in response to detecting that the first user is in a predetermined and / or known pose and that the first user is centered within the detection field of camera 602. In some embodiments, computer system 600 captures a third image after determining that the first user is centered within the detection field of one or more camera sensors in response to detecting the first user's gaze and / or gesture in a particular direction.
[0150] In some implementations, if the media capture operation corresponds to a still photograph or a live photograph (e.g., a photographic operation in which computer system 600 captures media data before and / or after capturing a photograph), then computer system 600 does not move a portion of computer system 600 while capturing image content. In some implementations, if computer system 600 moves a portion of computer system 600 (e.g., to improve the composition of media content, when following a user, and / or in response to detecting user input), then in the case where the media capture operation corresponds to a still photograph or a live photograph, computer system 600 stops moving a portion of computer system 600 while capturing image content, as described above. Figures 6F to 6G As described herein. For example, in some embodiments, when a user changes position and pose, computer system 600 moves a portion of itself to follow the user's head. When the user changes to a new position and / or pose, computer system 600 stops moving a portion of itself while capturing image content and resumes moving the portion of itself after the media capture operation is complete. In some embodiments, computer system 600 moves a portion of itself while capturing media. For example, in some embodiments, if the media is expected to be media in which a first user is running, computer system 600 moves a portion of itself while capturing media to keep the first user within the detection field of one or more camera sensors.
[0151] exist Figure 6H At, Figure 6G At the time following the capture of the third image, Figure 618 includes an object representation 638 to the right and next to the user representation 622, which indicates that the object is located next to the first user within the environment. Figure 6HAs indicated by diagram 618, user representation 622 and object representation 638 are located within detection field representation 654. Therefore, in Figure 6H At this location, the first user and object are situated within the detection field of one or more camera sensors. For example... Figure 6H As illustrated, because the first user and the object are within the detection field of camera 602, computer system 600 displays user image 624 and object image 640 in live view 608. Object image 640 is a representation of the object detected by one or more camera sensors (e.g., camera 602). Figure 6H At this point, computer system 600 detects a composition guide 605h from the first user that corresponds to the first user's command to computer system 600 to capture video content (e.g., to capture video).
[0152] exist Figure 6I At this point, in response to the detection of the mapping instruction 605h, the computer system 600 configures itself to track the first user. Figure 6I At that point, in response to the detection of composition guidance 605h, computer system 600 outputs composition instruction 642. More specifically, in Figure 6I At this point, computer system 600 outputs a composition instruction 642 stating "move the triangle to the left." Due to the simplicity and / or universality of the composition guide 605h, computer system 600 outputs composition instruction 642. Composition guide 605h does not contain any specific instructions and / or guidance from computer system 600 regarding how video should be shot. Therefore, in Figure 6I The composition instruction 642 determines the best way to capture video.
[0153] In some implementations, the composition instructions output by the computer system 600 are based on one or more objects detected by the computer system 600 within the detection field and / or environment (e.g., objects, furniture, lighting, users, and / or background types) of one or more camera sensors. For example, in Figure 6I In response to the detection of composition guidance 605h, computer system 600 outputs composition instructions 642, which include instructions corresponding to a first user and an object detected by computer system 600 within the detection field of camera 602. In some embodiments, the composition instructions are based on the position of the user and / or object within the environment. For example, in some embodiments, computer system 600 outputs audio content corresponding to composition instructions (such as "stand in front of the triangle").
[0154] In some embodiments, the framing instructions output by the computer system 600 include cues for one or more users to move within the environment. For example, in some embodiments, in response to detecting framing guidance, the computer system 600 outputs framing instructions that include instructions for a first user to walk to the left within the environment. In some embodiments, the framing instructions output by the computer system 600 are based on the distance between the edge of the detection field of one or more camera sensors and the location of the subject (e.g., the user's face, the user's body, and / or the user's limbs). For example, based on determining a predetermined distance (e.g., 0.1 inches to 24 inches) from the edge of the detection field of one or more camera sensors to the user's head, the computer system 600 outputs audio content corresponding to instructions for the user to center themselves within the detection field of one or more camera sensors.
[0155] In some embodiments, the composition instructions output by the computer system 600 are based on a horizontal line detected within the detection field of the camera 602. For example, in some embodiments, the computer system 600 outputs composition instructions that include instructions for a first user to straighten their shoulders so that the first user's shoulders are horizontal within the detection field of one or more camera sensors.
[0156] In some embodiments, the composition instructions output by the computer system 600 are based on the positioning of the first user's face. For example, in some embodiments, the computer system 600 outputs composition instructions that include instructions for the first user to tilt their chin upwards. In some embodiments, the composition instructions output by the computer system 600 are based on the positioning of the first user's face relative to a fixed reference point. For example, in some embodiments, the computer system 600 outputs composition instructions that include instructions for the first user to move their face until it is aligned with the edge of a triangle. As another example, in some embodiments, the computer system 600 outputs composition instructions that include instructions for the first user to orient their face within the right third of the detection field of one or more camera sensors for a more dynamic layout. In some embodiments, the composition instructions output by the computer system 600 are based on the positioning of the first user's face relative to their body. For example, in some embodiments, the computer system 600 outputs composition instructions that include instructions for the first user to move their face forward until it is above the first user's knees to create foreground and background in the media that draw attention to the user's face.
[0157] In some implementations, the composition instructions output by computer system 600 are more detailed when the instructions and / or composition guidelines are more general than when they are more specific. For example, in some implementations, computer system 600 outputs more detailed composition instructions in response to detecting a composition guideline such as "take a picture of me and my dog in front of a tree" than when computer system 600 detects a composition guideline such as "use a soft-focus filter and take a picture of me and my dog in front of a tree with my gaze shifted away and looking to the left." In some implementations, the composition instructions output by computer system 600 are less detailed when the instructions and / or composition guidelines are more general than when they are more specific. In some implementations, the composition instructions output by computer system 600 are different for a first user and for a second user, even if computer system 600 detects the same composition guideline from both users. For example, if both the first user and the second user request the computer system 600 to capture videos of them dancing in a style that mimics a musical from the 1940s, the computer system 600 will output a composition instruction such as "squat" to the first user, and the computer system 600 will output a composition instruction such as "bend at the knees" to the second user.
[0158] In some embodiments, the composition instructions output by the computer system 600 include one or more prompts for changing the lighting in the environment, such as changing the amount of light (e.g., increasing or decreasing the amount of light), changing the type of lighting (e.g., direct versus indirect lighting), changing the color of the lighting, and / or changing the color temperature of the lighting (e.g., giving the lighting a warmer or cooler hue)). In some embodiments, when the computer system 600 communicates with a lighting system in the environment (e.g., an intelligent lighting system), the computer system 600 changes the lighting in the environment in response to detecting composition instructions from a first user. In some embodiments, the computer system 600 requests permission before changing the lighting in the environment. In some embodiments, the computer system 600 is granted permission from the first user to change the lighting in the environment before the computer system 600 detects composition instructions from the user. In some embodiments, the computer system 600 changes the lighting in the environment automatically (e.g., without detecting user input).
[0159] exist Figure 6J As indicated by the positioning of user representation 622 and object representation 638 in Figure 618, the first user moves the object and themselves to the left within the environment. Figure 6J At this point, determine that the first user will move the object and themselves to the left. Figure 6JBased on the determination that the first user moves the object and themselves to the left, the computer system 600 moves a portion of the computer system 600 in a third movement mode (e.g., rotates a portion of the computer system 600 to the left) such that, when initiating the video capture operation, the first user is centered within the detection field of one or more camera sensors. Figure 6J At this point, a set of one or more criteria is determined (e.g., the first user is gazing at computer system 600, the first user is in a specific pose, and / or the first user is performing a gesture). Figure 6J At a given location, based on determining that a set of one or more criteria is met, computer system 600 initiates video capture. In some implementations, computer system 600 automatically selects an art style for capturing video based on the detection of one or more conditions. For example, when low ambient brightness is detected, computer system 600 captures video using an art style suitable for dark conditions.
[0160] exist Figure 6J and Figure 6K During this process, when the first user moves to the right within the environment, the computer system 600 tracks the first user. When the first user moves to the right within the environment, the computer system 600 captures video of the first user. Figure 6K At this point, computer system 600 completes (for example, terminates) the video capture process.
[0161] exist Figure 6K As indicated by the positioning of user representation 622 in Figure 618 and object representation 638 in Figure 618, the first user has moved to the right, thus leaving the object behind. Figure 6K As illustrated, as part of completing the video capture process, computer system 600 displays a representation of the newly captured video content (e.g., fifth content 644) and stops displaying representations of previously captured media content (e.g., fourth content 636). Figure 6K At this point, computer system 600 detects a composition guide 605k from the first user, corresponding to the style John is instructed to use to capture video content. John is an individual who captures video of a subject by magnifying its appearance. Furthermore, John typically captures video of a subject while the individual is moving within the environment. Composition guide 605k is more specific than composition guide 605h. That is, composition guide 605k instructs computer system 600 to shoot video in a specific style (e.g., similar to John), while composition guide 605h generally instructs computer system 600 to shoot video.
[0162] exist Figure 6LIn response to detecting composition guide 605k, computer system 600 outputs composition instruction 646 corresponding to the instruction given by computer system 600 to the first user (e.g., "Walk slowly along the wall to your left"). Due to the difference in level of detail between composition guide 605h and composition guide 605k... Figure 6L Composition command 646 and Figure 6H The composition instruction 642 is different at that location. Figure 6L Because composition guide 605k is more specific than composition guide 605h, computer system 600 does indeed provide the first user with instructions on how to frame the environment (e.g., instructions for moving objects within the environment to prepare for video).
[0163] exist Figure 6L At the location, such as user image 624 within live view 608 compared to... Figure 6K As indicated by the larger description, in response to the detection of composition guide 605k (e.g., composition guide within composition guide 605k), computer system 600 zooms in on the first user via camera 602. In some embodiments, computer system 600 outputs composition instructions based on elements detected in the physical environment. For example, in some embodiments, when computer system 600 detects that the physical environment is dark, computer system 600 outputs composition instructions commanding the first user to increase the brightness of the physical environment.
[0164] exist Figure 6L and Figure 6M Between, the first user moves to the left within the environment. Figure 6L and Figure 6MBetween these points, it is determined that the first user is moving to the left. Based on the determination that the first user is moving to the left within the environment, computer system 600 initiates a video capture operation and tracks the user as the first user moves. In some embodiments, computer system 600 initiates the video capture operation based on the determination that the first user is following a composition instruction 646. In some embodiments, when capturing video content, computer system 600 automatically (e.g., without intermediate user input) moves a portion of computer system 600 to meet the detected composition instructions. For example, in some embodiments, when capturing video content, computer system 600 moves a portion of computer system 600 in a manner that creates the impression that the first user is being followed by an animal, based on detected composition instructions from the first user. In some embodiments, when capturing video content, computer system 600 automatically moves a portion of computer system 600 based on conditions detected in the environment. For example, when capturing video media, in response to computer system 600 detecting that parents are helping their baby walk, computer system 600 moves a portion of computer system 600 close to the ground to capture video of the baby walking, moves slowly with the baby, and then moves back to obtain video of the parents and baby walking together. In some implementations, when capturing media content, computer system 600 moves a portion of itself based on one or more settings of camera 602. For example, when the active settings of camera 602 require a user to be in the left third of the detection field of one or more camera sensors before capturing media, computer system 600 automatically moves a portion of itself until the first user is aligned in the left or right third of the detection field of one or more camera sensors. As another example, when the active settings of camera 602 result in capturing media in black and white only, computer system 600 moves a portion of itself differently than when the active settings of camera 602 result in capturing media in color.
[0165] exist Figure 6M At this point, computer system 600 has completed capturing the video of the first user. (As follows) Figure 6M As illustrated, after computer system 600 has completed capturing video content, computer system 600 displays a representation of the newly captured video content (e.g., sixth content 648) and stops displaying representations of previously captured media content (e.g., fifth content 644). Figure 6M Even though the computer system 600 has completed capturing video content, it maintains the display of the first user's view at the increased zoom level. In some embodiments, as part of completing video content capture, the computer system 600 reduces the first user's zoom level.
[0166] Figure 7This is a flowchart illustrating a method (e.g., process 700) for selectively capturing media according to some embodiments. Some operations in process 700 may be optionally combined, the order of some operations may be optionally changed, and some operations may be optionally omitted.
[0167] As described below, Process 700 provides an intuitive way to selectively capture media. Process 700 reduces the cognitive load on the user, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to interact with such devices faster and more efficiently, saving power and increasing the time between battery charging.
[0168] In some embodiments, process 700 is performed at a computer system (e.g., 600) that communicates with a media capture component (e.g., 602) (e.g., sensors, environmental sensors, the capture component, cameras (e.g., periscope cameras, telephoto cameras, wide-angle cameras, and / or ultra-wide-angle cameras), depth sensors, microphones, heart monitors, and / or temperature sensors) and a microphone. In some embodiments, the computer system is a telephone, watch, tablet, fitness tracker, wearable device, display, portable computer system, accessory, speaker, light, head-mounted display (HMD), and / or personal computing device. In some embodiments, the media capture component includes and / or a microphone. In some embodiments, the microphone is different from the media capture component.
[0169] The computer system detects (702) (and / or receives) a first input (e.g., a verbal media capture instruction) corresponding to a request to capture media (e.g., images, video, and / or audio) via a microphone (e.g., 605a, 605f, 605h, and / or 605k). In some embodiments, the first input is detected when the computer system is in a media capture mode (e.g., a mode in which the computer system is configured to capture images, video, and / or audio). In some embodiments, when in media capture mode, the computer system displays a user interface corresponding to the media capture mode via a display component in communication with the computer system. In some embodiments, when in media capture mode, the computer system displays a live preview of the output (e.g., images, video, and / or audio) of the media capture component via a display component in communication with the computer system.
[0170] After detecting a first input (e.g., 605a, 605f, 605h and / or 605k) corresponding to a request to capture media (704) (and / or in combination with and / or in response to detecting the first input (e.g., 605a, 605f, 605h and / or 605k) corresponding to a request to capture media (e.g., 605a, 605f, 605h and / or 605k)) (and / or when in media capture mode) and based on determining that the first input (e.g., 605a, 605f, 605h and / or 605k) corresponds to a first instruction (and / or a first set of one or more instructions), the computer system captures (706) media (e.g., 616, 632, 634, 636, 644 and / or 648) (e.g., images, videos and / or audio) via a media capture component (e.g., 602) in response to a first set of one or more conditions being satisfied (e.g., not capturing media when a first set of one or more conditions is not satisfied and / or capturing media only when a first set of one or more conditions is satisfied). In some embodiments, in response to detecting a first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a first instruction, the computer system captures media when, and / or in response to, a first set of one or more conditions being met. In some embodiments, in response to detecting a first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a first instruction, the computer system is configured to capture media when, and / or in response to, a first set of one or more conditions being met. In some embodiments, after detecting a first input corresponding to a request to capture media (and / or in combination with or in response to detecting a first input corresponding to a request to capture media) (and / or when in media capture mode) and based on determining that the first input corresponds to a second instruction different from the first instruction, the computer system does not capture media (e.g., images, video, and / or audio) via a media capture component when, and / or in response to, a first set of one or more conditions being met. In some implementations, in response to detecting a first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a second instruction, the computer system does not capture media when a second set of one or more conditions is met, or in response to a second set of one or more conditions being met. In some implementations, in response to detecting a first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a second instruction, the computer system is not configured to capture media when a second set of one or more conditions is met, or in response to a second set of one or more conditions being met, and / or is configured not to capture media when a second set of one or more conditions is met, or in response to a second set of one or more conditions being met.In some implementations, after detecting a first input corresponding to a request to capture media (and / or in combination with or in response to detecting a first input corresponding to a request to capture media) (and / or when in media capture mode), media is not captured if a first set of one or more conditions is not met.
[0171] After detecting a first input corresponding to a request to capture media (704) and determining that the first input (e.g., 605a, 605f, 605h, and / or 605k) corresponds to a second instruction different from a first instruction (and / or a second set of one or more instructions different from a first set of one or more instructions), the computer system captures (708) media (e.g., 616, 632, 634, 636, 644, and / or 648) (e.g., images, videos, and / or audio) via a media capture component (e.g., 602) in response to the second set of one or more conditions being satisfied (e.g., not capturing media when the second set of one or more conditions is not satisfied and / or capturing media only when the second set of one or more conditions is satisfied), wherein the second set of one or more conditions is different from the first set of one or more conditions. In some embodiments, in response to detecting a first input corresponding to a request to capture media and / or determining that the first input corresponds to a second instruction, the computer system captures media when the second set of one or more conditions is satisfied, and / or in response to the second set of one or more conditions being satisfied. In some embodiments, in response to detecting a first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a second instruction, the computer system is configured to capture media when, and / or in response to, a second set of conditions being met. In some embodiments, after detecting the first input corresponding to a request to capture media (and / or in conjunction with or in response to detecting the first input corresponding to a request to capture media) (and / or when in media capture mode) and based on determining that the first input corresponds to a first instruction, the computer system does not capture media (e.g., images, video, and / or audio) via the media capture component when, and / or in response to, a second set of conditions being met. In some embodiments, in response to detecting the first input corresponding to a request to capture media and / or based on determining that the first input corresponds to a first instruction, the computer system does not capture media when, and / or in response to, a second set of conditions being met. In some implementations, in response to detecting a first input corresponding to a request to capture media and / or in response to determining that the first input corresponds to a first instruction, the computer system is not configured to capture media when a second set of one or more conditions is met, and / or in response to a second set of one or more conditions being met, and / or is configured not to capture media when a second set of one or more conditions is met, and / or in response to a second set of one or more conditions being met.In some implementations, after detecting a first input corresponding to a request to capture media (and / or in conjunction with or in response to detecting the first input corresponding to a request to capture media) (and / or when in media capture mode), media is not captured and / or the computer system does not capture media if a second set of one or more conditions is not met. In some implementations, the computer system displays and / or saves a representation of the media based on determining that the first input corresponds to a first instruction (and / or a first set of one or more instructions). In some implementations, the computer system displays and / or saves a representation of the media based on determining that the first input corresponds to a second instruction (and / or a first set of one or more instructions). Selectively capturing media when a predetermined set of conditions is met (e.g., the first instruction where the first input corresponds to the second instruction) automatically allows the computer system to perform the media capture process based on specific guidelines expressed by the user, such that the media capture process is customized according to the user's expectations, thereby performing operation when the set of conditions has been met without requiring further user input.
[0172] In some embodiments, the computer system (e.g., 600) communicates (e.g., wireless and / or wired communication) with a first moving component (e.g., actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), a movable base, a rotatable component, a motor, a lift, a level, and / or a rotatable base). In some embodiments, after detecting a first input corresponding to a request to capture media (e.g., 605a, 605f, 605h, and / or 605k) (e.g., and when the computer system is in the media capture mode of the media capture component) (e.g., in response to detecting the first input corresponding to a request to capture media), the computer system moves a portion of the computer system via the first moving component (e.g., as described above regarding...). Figure 6G(As described) (e.g., a portion of a computer system including a media capture component) (e.g., the computer system translating and / or rotating). In some embodiments, a portion of the computer system moves about a single axis of the moving component. In some embodiments, a portion of the computer system moves about two or more axes of the moving component. In some embodiments, a portion of the computer system moves in two different ways (e.g., the computer system rotates, tilts, and / or translates). In some embodiments, a portion of the computer system moves in two different directions (e.g., left and up, tilting up and moving right, or back and right). In some embodiments, a portion of the computer system moves in a single direction. In some embodiments, a portion of the computer system stops moving when capturing and / or about to capture media. In some embodiments, a portion of the computer system continues to move while capturing media. In some embodiments, a portion of the computer system (e.g., a movable arm, hinge, and / or base) moves, while in some embodiments, another portion of the computer system does not move.
[0173] In some implementations, after detecting a first input (e.g., 605a, 605f, 605h and / or 605k) corresponding to a request to capture media, the computer system moves the positioning of the media capture component (e.g., 602) via a first moving component, such that the field of view (e.g., 654) of the media capture component (e.g., 602) moves from the first positioning to a second positioning different from the first positioning (e.g., as described above). Figure 6G (As described herein). In some embodiments, the mobile computer system causes the user's framing to change as the computer system moves (e.g., when the computer system starts moving, the user is in the left portion of the media capture component's field of view, and when the computer system stops moving, the user is in the right portion of the media capture component's field of view). In some embodiments, the computer system moves with the user. In some embodiments, the computer system moves the positioning of the media capture component to keep the user within the media capture component's field of view. Moving the positioning of the media capture component via a first moving component after detecting a first input corresponding to a request to capture media, such that the field of view of the first moving component moves from a first positioning to a second positioning, allows the computer system to better position the media capture component to capture media, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when a set of conditions has been met without requiring further user input.
[0174] In some embodiments, after detecting a first input (e.g., 605a, 605f, 605h, and / or 605k) and based on determining that the first input includes one or more instructions (e.g., audible instructions and / or gesture-based instructions) for the computer system (e.g., 600) to move in a first manner (e.g., the first input includes a voice command guiding the computer system to move in a first manner) (e.g., the first manner includes translation in one or more directions and / or rotation about one or more axes of a moving component), a portion of the computer system (e.g., 600) (e.g., a portion of the computer system including a media capture component) moves in the first manner. In some embodiments, after detecting a first input and based on determining that the first input includes one or more instructions for the computer system to move in a second manner different from the first manner (e.g., the first input includes a voice command guiding the computer system to move in a second manner) (e.g., the second manner includes translation in one or more directions and / or rotation about one or more axes of a moving component), a portion of the computer system moves in a second manner different from the first manner (e.g., not the first manner) (e.g., as described above regarding...). Figure 6C and Figure 6G (As described). In some embodiments, the computer system moves in a first mode and a second mode based on instructions that determine the first input includes corresponding to a first mode and a second mode. In some embodiments, the computer system moves sequentially in the first mode and the second mode. In some embodiments, the computer system moves concurrently in the first mode and the second mode. In some embodiments, the first mode and the second mode are the same. In some embodiments, moving in the first mode and / or the second mode causes a portion of the computer system to translate and / or rotate. In some embodiments, moving in the first mode and / or the second mode causes a portion of the computer system to move in a corresponding movement pattern (e.g., determined by the computer system and / or by the user of the computer system). In some embodiments, when a portion of the computer system moves in the first mode, the portion of the computer system moves at a different speed and / or direction compared to when a portion of the computer system moves in the second mode. Moving in a corresponding mode based on which instructions are included in the first input automatically allows the computer system to move in various modes based on user instructions, thereby performing operations when a set of conditions has been met without further user input.
[0175] In some embodiments, detecting a first input (e.g., 605a, 605f, 605h, and / or 605k) corresponding to a request to capture media includes capturing one or more verbal instructions (e.g., instructions included in 605a, 605f, 605h, and / or 605k) via a microphone. In some embodiments, the first and / or second instructions include one or more verbal instructions (e.g., spoken instructions and / or audible instructions). In some embodiments, the computer system determines that the verbal instruction is provided by a primary user (e.g., a user registered with the computer system and / or a target user). In some embodiments, the computer system determines that the verbal instruction is provided by a non-primary user (e.g., a user not registered with the computer system and / or a user who is not a target user). Capturing media via the media capture component after detecting one or more verbal instructions allows the computer system to perform media capture operations without displaying a corresponding user interface, thereby providing additional control options without cluttering the user interface with additional displayed controls and providing improved feedback (e.g., upon detection of the first input).
[0176] In some implementations, detecting a first input corresponding to a request to capture media (e.g., 605a, 605f, 605h, and / or 605k) includes capturing one or more gesture-based instructions (e.g., as described above) via one or more input devices (e.g., media capture components and / or another type of device or sensor, and / or the gesture being performed). Figure 6A (as described herein) (e.g., one or more air gestures involving the movement of a part of the user's body in the air). In some embodiments, the first and / or second instructions include gesture-based instructions (e.g., the user points in a direction, the user makes a hand gesture in a direction, the user faces their body in a direction, and / or the user walks in a direction). In some embodiments, the instructions are a combination of verbal instructions and gesture-based instructions. In some embodiments, the computer system determines that the gesture-based instructions are provided by a primary user (e.g., a user registered with the computer system and / or a target user). In some embodiments, the computer system determines that the gesture-based instructions are provided by a non-primary user (e.g., a user not registered with the computer system and / or a user who is not a target user). Capturing media via a media capture component after detecting one or more gesture-based instructions allows the computer system to perform media capture operations without displaying a corresponding user interface, thereby providing additional control options without cluttering the user interface with additional displayed controls and providing improved feedback.
[0177] In some implementations, in response to a first set of conditions satisfying one or more conditions, a portion of a computer system (e.g., 600) is moved via a media capture component (e.g., 616, 632, 634, 636, 644 and / or 648) (e.g., images, videos and / or audio) including a portion of the media capture component (e.g., 602) moving a portion of the computer system (e.g., 600) (e.g., a portion of the computer system including a portion of the media capture component) by capturing media (e.g., 616, 632, 634, 636, 644 and / or 648) (e.g., images, videos and / or audio) via a media capture component (e.g., 602) by a first set of compositional rules determined according to a first set of conditions satisfying one or more conditions and / or not according to a second set of conditions satisfying one or more conditions (e.g., whether a portion of the user is within a specific area of the field of view (e.g., the upper third, middle third and / or lower third)) (e.g., whether one or more objects are captured at a specific zoom level, with a specific filter, with a specific color and / or with a specific amount of light)). In some implementations, capturing media (e.g., images, videos, and / or audio) via a media capture component in response to a second set of conditions satisfying one or more includes moving a portion of the computer system (e.g., as described above) in a fourth manner different from a third-party approach (e.g., based on a second set of composition rules determined according to a second set of conditions satisfying one or more and / or not according to a first set of conditions satisfying one or more). Figure 6L (As described herein). In some embodiments, moving in a third manner causes a portion of the computer system to translate and / or rotate. In some embodiments, moving in a third manner causes a portion of the computer system to move in a corresponding movement pattern. In some embodiments, when a portion of the computer system moves in a third manner, it moves at a different speed and / or direction compared to when a portion of the computer system moves in a first and / or second manner.
[0178] In some embodiments, the computer system (e.g., 600) communicates with a second moving component (e.g., an actuator (e.g., a pneumatic actuator, a hydraulic actuator, and / or an electric actuator), a movable base, a rotatable component, a motor, a lift, a level, and / or a rotatable base). In some embodiments, when media (e.g., 616, 632, 634, 636, 644, and / or 648) is captured via a media capture component (e.g., 602) and based on the determination that the media being captured is a first type of media (e.g., video media or still photographs), the computer system moves (e.g., translates and / or rotates) a portion of the computer system (e.g., a portion of the computer system including the media capture component) via the second moving component during (and / or while) the media (e.g., the first type of media) capture (e.g., during the capture of the media (e.g., the first type of media)). Figure 6G and Figure 6M(As described above). In some implementations, when media is captured via a media capture component and it is determined that the media being captured is a second type of media (e.g., video media or still photographs) different from the first type of media, the computer system abandons moving a portion of the computer system via a second moving component during the capture of the media (the second type of media) (e.g., as described above). Figure 6C (As described herein). In some embodiments, upon determining that the computer system has stopped capturing a first type of media, a portion of the computer system stops moving. In some embodiments, upon determining that a portion of the computer system has stopped capturing a first type of media, a portion of the computer system continues to move. In some embodiments, a portion of the computer system moves such that the media capture component can track the user while the media capture component is capturing media. In some embodiments, a portion of the computer system moves while capturing both a first type of media and a second type of media. Selectively moving a portion of the computer system based on the type of media being captured automatically allows the computer system to indicate whether it is capturing video or still images, thereby performing operations when a set of conditions has been met without requiring further user input and providing improved feedback.
[0179] In some implementations, the first type of media (e.g., 616, 632, 634, 636, 644, and / or 648) is video or panoramic photographs (e.g., photographs showing an environmental field of view larger than the field of view of the media capture component). In some implementations, the second type of media (e.g., 616, 632, 634, 636, 644, and / or 648) is still photographs (e.g., or a collection of one or more animated images and / or photographs (e.g., photographs including a representation of the field of view of the media capture component immediately before a request to capture a photograph is detected and a representation of the field of view of the media capture component immediately after a request to capture a photograph is detected)).
[0180] In some implementations, the first input (e.g., 605a, 605f, 605h and / or 605k) includes instructions (e.g., indications, voice commands, directives and / or orders) for the computer system (e.g., 600) to capture delayed media (e.g., 616, 632, 634, 636, 644 and / or 648) until the input detected by the camera is detected (e.g., as described above). Figure 6B (as described in the description).
[0181] In some implementations, the input detected by the camera includes (e.g., a collection of one or more cameras external to a computer system via a media capture component and / or a view detected as described above) Figure 6B(As described herein). In some embodiments, the computer system outputs an indication that a gaze has been detected (e.g., a haptic indication, a graphical indication, and / or an audio indication). In some embodiments, the gaze is directed toward the computer system. In some embodiments, the gaze is directed away from the computer system. In some embodiments, the gaze lasts for a predetermined period of time (e.g., 1 to 15 seconds). Delaying media capture until the computer system detects a gaze allows the computer system to perform media capture operations without requiring the user to move out of the field of view of the media capture component to select user interface objects displayed by the computer system, thereby providing additional control options without cluttering the user interface with additional displayed controls.
[0182] In some implementations, the input detected by the camera includes detected gestures (e.g., a smiling gesture, a thumbs-up gesture, an index finger pointing gesture, and / or a gesture with both hands raised above the head) (e.g., as described above). Figure 6B (As described herein). In some embodiments, the gesture is directed (e.g., aiming, focusing, pointing) at the computer system. In some embodiments, the gesture is directed (e.g., aiming, focusing, pointing) at an object (e.g., an individual, an inanimate object, and / or an animal). In some embodiments, the computer system outputs an indication that the gesture has been detected (e.g., a haptic indication, a graphical indication, and / or an audio indication). Delaying media capture until the computer system detects the gesture allows the computer system to perform media capture operations without requiring the user to move out of the field of view of the media capture component to select a user interface object displayed by the computer system, thus providing additional control options without cluttering the user interface with additional displayed controls.
[0183] In some implementations, the input detected by the camera includes detected pose (e.g., as described above). Figure 6B (as described herein) (e.g., seated posture, standing posture, posture involving two or more individuals (e.g., human pyramid and / or arm interlocking) and / or kneeling posture). In some embodiments, the computer system outputs an indication of detected posture (e.g., haptic indication, graphical indication, and / or audio indication). In some embodiments, the computer system outputs a graphical representation of the posture. Delaying media capture until the computer system detects the posture allows the computer system to perform media capture operations without requiring the user to move out of the field of view of the media capture component to select user interface objects displayed by the computer system, thereby providing additional control options without cluttering the user interface with additional displayed controls.
[0184] In some implementations, the first input (e.g., 605a, 605f, 605h, and / or 605k) includes a set of one or more time-based instructions indicating one or more media capture parameters (e.g., when to start media capture, when to end media capture, and / or how long to capture media) (e.g., time boundaries and / or time guidelines) (e.g., as described above). Figure 6B (As described herein). In some embodiments, the computer system outputs an indication that a time-based instruction has been detected (e.g., a visual indication, a tactile indication, and / or an audible indication) (e.g., a series of beeps indicating the number of seconds until the media capture operation is initiated, a voice indication of the length of the media capture operation (e.g., "The video will be captured for 5 seconds" and / or "The video will last for 30 seconds")), and optionally indicates the associated time duration associated with the time-based instruction. Capturing media after detecting a set of one or more time-based instructions allows the computer system to capture the media according to a time guide included in the first input, such that when capturing the media, the user included in the media is correctly framed and aligned within the field of view of the media capture component, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing an operation when the set of conditions has been met without requiring further user input.
[0185] In some implementations, a set of one or more time-based instructions (e.g., audible instructions and / or gesture-based instructions) includes one or more indications (e.g., as described above) of when the capture of media (e.g., 616, 632, 634, 636, 644 and / or 648) will be initiated. Figure 6B (as described herein) (e.g., "initiate media capture after 10 seconds" and / or "initiate media capture at 3:30 PM"). In some embodiments, initiating media capture is based on timing (e.g., time of day, countdown, timing since the user last interacted with the computer system, and / or timing since the user last viewed the computer system). In some embodiments, initiating media capture is based on the amount of time the user has been in a specific pose. In some embodiments, initiating media capture is based on the amount of time that has elapsed since the user performed the gesture. Capturing media after detecting one or more indications of when to initiate media capture allows the computer system to initiate media capture at an appropriate time, such that the content of the media is correctly framed and aligned within the field of view of the media capture component during the media capture process, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when a set of conditions has been met without requiring further user input.
[0186] In some implementations, a set of one or more time-based instructions includes one or more indications (e.g., audible instructions and / or gesture-based instructions) of when the capture of media (e.g., 616, 632, 634, 636, 644 and / or 648) will stop (e.g., audible instructions and / or gesture-based instructions) (e.g., “stop capturing media in 10 seconds” and / or “stop capturing media at 3:30 p.m.”) (e.g., as described above in...). Figure 6B (As described herein). In some implementations, stopping media capture is based on timing (e.g., time of day, countdown, timing since the user last interacted with the computer, and / or timing since the user last viewed the computer system). In some implementations, stopping media capture is based on the amount of time the user has been in a specific pose. In some implementations, stopping media capture is based on the amount of time that has elapsed since the user performed the gesture. Capturing media after detecting one or more indications that media capture will stop allows the computer system to stop capturing media at the appropriate time, such that the media capture component captures only the desired content, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing actions when the set of conditions has been met without requiring further user input.
[0187] In some implementations, the set of one or more time-based instructions includes one or more indications (e.g., audible instructions and / or gesture-based instructions) of the time interval (e.g., 1 second to 60 seconds) between the captures of individual (and / or different) media items (e.g., 616, 632, 634, 636, 644 and / or 648) (e.g., as described above). Figure 6B (As described herein). In some embodiments, individual media items are media items of different types. In some embodiments, individual media items of the same type are media items of the same type. In some embodiments, a second media item is automatically captured (e.g., without intermediate user input) after the first media item, based on the determination that a time interval has expired. In some embodiments, the computer system outputs (e.g., audibly outputs and / or displays) a countdown timer for the time interval after capturing the initial media item. Capturing media after detecting one or more indications of the time interval between captures of individual media items allows the computer system to pause for a sufficient amount of time between captures of different media items, allowing the content of the media item to be realigned within the field of view of the media capture component, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when a set of conditions has been met without further user input.
[0188] In some implementations, the first input (e.g., 605a, 605f, 605h and / or 605k) includes compositional guidance (e.g., instructions on how and / or why the media should be captured) relating to the captured media (e.g., 616, 632, 634, 636, 644 and / or 648) (e.g., instructions on how and / or why the media should be captured). This guidance may include instructions on how users in the media should be spatially oriented, how users should be positioned relative to each other, and / or how objects should be spatially oriented within the field of view of the media capture component (e.g., rule of thirds, golden ratio, golden triangle, spatial rules, odd number rules, and / or black-and-white users). Figure 6F (As described herein). In some embodiments, the capture is based on determining that the composition guidelines are followed, satisfying a first set of one or more conditions and / or a second set of one or more conditions. In some embodiments, the media is captured only when the composition guidelines are met. In some embodiments, the media is captured even when the composition guidelines are not followed. Capturing the media after detecting an indication to the composition guidelines allows the computer system to capture the media based on the user's spatial arrangement within the field of view of the media capture, thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when the set of conditions has been met without requiring further user input.
[0189] In some implementations, a first set of one or more conditions (and / or a second set of one or more conditions) does not include conditions corresponding to input (e.g., tap input, swipe input, voice command, rotation of a rotatable input mechanism, and / or air hand gesture) corresponding to input from a first user (e.g., input performed by the user) detected by a computer system, media capture component, external media capture component, and / or external computer system (e.g., input performed by the user). Figure 6E (As described in the description). Capturing media without detecting input corresponding to the first user allows the computer system to perform the media capture process without requiring the first user to physically interact with the user interface object displayed by the computer system (e.g., select the user interface object), thereby providing additional control options without cluttering the user interface with additional displayed controls.
[0190] In some implementations, the first set of one or more conditions includes conditions that are met when (e.g., by a computer system and / or by a media capture component) it is determined that a person in the field of view (e.g., 654) of the media capture component (e.g., 602) has stopped moving (e.g., in some way (e.g., stopped walking, stopped jumping, stopped exercising, and / or stopped running in place) and / or stopped moving completely) for a threshold amount of time (e.g., 1 second to 60 seconds). Figure 6D(As described herein). In some implementations, the progress of the threshold time amount's expiration begins based on the determination that the person has not moved for a predefined period of time. In some implementations, the progress of the threshold time amount's expiration is paused based on the determination that the person changes from not moving to moving. In some implementations, the progress of the threshold time amount's expiration restarts based on the determination that the person has begun to move. Capturing media based on the determination that the person has stopped moving for a threshold time amount allows the computer system to perform the media capture process without requiring the person to physically interact with a user interface object displayed by the computer system (e.g., select that user interface object), thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when a set of conditions has been met without requiring further user input.
[0191] In some implementations, a first set of one or more conditions includes when (e.g., by a computer system and / or by a media capture component) it is determined that a person is positioned in a corresponding pose within the field of view (e.g., 654) of the media capture component (e.g., 602) (e.g., as described above). Figure 6D The conditions described herein (e.g., when a person is within the field of view of the media capture component) (e.g., the person is sitting, the person is kneeling, the distance between the person and one or more corresponding users is below or above a threshold). In some implementations, the conditions are met based on determining that the person has been positioned in the corresponding pose for a predetermined time period (e.g., 1 second to 30 seconds). Capturing media based on determining that the person is positioned in the corresponding pose allows the computer system to perform the media capture process without requiring the person to physically interact with a user interface object displayed by the computer system (e.g., select the user interface object), thereby providing additional control options without cluttering the user interface with additional displayed controls, and performing operations when the set of conditions has been met without requiring further user input.
[0192] In some implementations, before capturing media (e.g., 616, 632, 634, 636, 644 and / or 648) via a media capture component (e.g., 602) in response to a first set of conditions being met and before capturing media (e.g., 616, 632, 634, 636, 644 and / or 648) based on determining that the first set of conditions will continue to be met, the computer system displays a countdown (e.g., 630) of a time period that must elapse before the media is captured (e.g., as described above). Figure 6E(As described herein). In some embodiments, the first set of one or more conditions includes conditions that have been met based on the expiration of a defined time period (e.g., 1 second to 30 seconds). In some embodiments, the computer system displays a series of graphical elements representing the progress of the time period. A countdown displaying the time period that must elapse to meet the specified set of conditions (e.g., to continue meeting the first set of one or more conditions) automatically allows the computer system to indicate to the user the amount of time before the computer system will initiate media capture, thereby performing an operation when the set of conditions has been met without requiring further user input.
[0193] In some implementations, when a countdown to a time period that must elapse before the media (e.g., 616, 632, 634, 636, 644, and / or 648) is captured is displayed, the computer system detects a first set of conditions that are not being met. In some implementations, in response to the detection of a first set of conditions that are not being met, the computer system interrupts the display of the countdown to the time period that must elapse before the media is captured (e.g., stops displaying, pauses displaying, stops showing the elapsed time and / or countdown, and / or stops animation). In some implementations, a first set of conditions that are not met is defined as a criterion that is not met during a predefined time period. In some implementations, the progress of a predefined time period is interrupted based on the determination that the criterion is not met during a predefined time period. Interrupting the display of the countdown to the time period that must elapse before the media is captured in response to the detection of a first set of conditions that are not being met allows the computer system to provide the user with an indication that a first set of conditions is no longer met, thereby providing improved visual feedback and offering additional control options without cluttering the user interface with additional displayed controls.
[0194] In some implementations, when a countdown to the time period that must elapse before the media (e.g., 616, 632, 634, 636, 644, and / or 648) is captured is displayed, the computer system detects a request to capture the media before the end of that time period, prior to its capture (e.g., an explicit request and / or a direct request (e.g., "Capture the media anyway," "Capture the media now," "Ignore the countdown and take a picture immediately")). In some implementations, in response to detecting a request to capture the media before the end of the time period, prior to its capture, the computer system interrupts the display of the countdown to the time period that must elapse before the media is captured (e.g., stops displaying, pauses displaying, stops showing the elapsed time and / or countdown, and / or stops animation) (e.g., as described above). Figure 6E(As described herein). In some embodiments, a first set of conditions, one or more, are not met based on the detection of a request to capture media. In some embodiments, the progress of a predefined time period is interrupted (e.g., the countdown is paused or stopped) in response to the detection of a request to capture media. In some embodiments, the progress of the predefined time period is restarted after the detection of a request to capture media. Interrupting the display of a countdown to the time period that must elapse before the media is captured in response to the detection of a request to capture media before the end of the time period allows the computer system to instruct the computer system to continue capturing media regardless of the countdown state, thereby providing improved visual feedback and providing additional control options without cluttering the user interface with additional displayed controls.
[0195] In some implementations, the media is a first media item (e.g., 616, 632, 634, 636, 644, and / or 648). In some implementations, after capturing the first media item (e.g., immediately after capturing the first media item or within a predefined time period (e.g., 1 second to 600 seconds) after capturing the first media item), the computer system detects a change in the human pose in the field of view (e.g., 654) of the media capture component (e.g., a first user changing from sitting to standing, or vice versa) via the media capture component (e.g., 602). In some implementations, in response to detecting a change in the human pose in the field of view of the media capture component, the computer system captures a second media item (e.g., 616, 632, 634, 636, 644, and / or 648) via the media capture component (e.g., as described above). Figure 6F (as described herein) (e.g., different from the first media item) (e.g., video and / or still photograph). In some embodiments, the computer system outputs an indication that the second media item will be captured (e.g., haptic indication, graphical indication, and / or audio indication). In some embodiments, changes in human pose are detected automatically (e.g., without intermediate user input). In some embodiments, in response to the detection of a change in human pose, the media capture component does not capture the second media item based on determining that the change in pose does not satisfy a set of one or more criteria. Capturing the second media item via the media capture component in response to the detection of a change in human pose allows the computer system to perform the media capture process without requiring the human to physically interact with (e.g., select) the corresponding user interface displayed by the computer system, thereby providing additional control options without cluttering the user interface with additional displayed controls.
[0196] In some implementations, the media is a third media item (e.g., 616, 632, 634, 636, 644, and / or 648). In some implementations, after capturing the third media item (e.g., immediately following the capture of the third media item or within a predefined time period of capturing the third media item (e.g., 1 second to 600 seconds)), the computer system detects a change in the pose of the corresponding person in the field of view (e.g., 654) of the media capture component (e.g., 602) (e.g., the corresponding person changes from sitting to standing, or vice versa). In some embodiments, in response to detecting a change in the pose of a person and determining that the person is positioned in one or more target poses (e.g., a predefined pose (e.g., a pose predefined by the computer system or by the person), a pose corresponding to a first input, a pose that positions the person substantially within the field of view of the media capture component, or a pose that matches the pose of the user within the field of view of the media capture component), the computer system captures a fourth media item (e.g., 616, 632, 634, 636, 644, and / or 648) via the media capture component (e.g., different from or separate from the first media item). In some embodiments, in response to detecting a change in the pose of a person and determining that the person is not positioned in one or more target poses, the computer system abandons capturing the fourth media item. In some embodiments, the computer system outputs an indication that the fourth media item will be captured (e.g., a haptic indication, a graphical indication, and / or an audio indication). In some embodiments, the fourth media item and the third media item are the same type of media item, or the fourth media item and the third media item are different types of media items. In some implementations, the computer system stops capturing the fourth media item when it is determined that the person is no longer positioned in the target pose. Selectively capturing the fourth media item based on whether the person is in the target pose automatically allows the computer system to perform the media capture process based on the detected pose of the person (e.g., or abandon the media capture process), thereby performing the operation when the set of conditions is met without further user input and providing additional control options without cluttering the user interface with additional displayed controls.
[0197] In some implementations, the media is a fifth media item (e.g., 616, 632, 634, 636, 644, and / or 648). In some implementations, after capturing the fifth media item (e.g., immediately after capturing the third media item or within a predefined time period (e.g., 1 second to 600 seconds) of capturing the third media item), the computer system detects a change in the pose of the corresponding person in the field of view (e.g., 654) of the media capture component (e.g., 602) (e.g., the corresponding person changes from sitting to standing, or vice versa). In some implementations, in response to detecting a change in the pose of the corresponding person and based on determining that the change in the pose of the corresponding person is detected within a threshold amount of time (e.g., 1 second to 60 seconds) of capturing the fifth media item, the computer system captures a sixth media item (e.g., 616, 632, 634, 636, 644, and / or 648) (e.g., still photographs and / or videos) (e.g., different from and / or separate from the fifth media item) via the media capture component. In some implementations, in response to detecting a change in the pose of the person in question and based on determining that no change in the pose of the person in question has been detected within a threshold time period for capturing the fifth media item, the computer system abandons capturing a sixth media item (and in some implementations, any media item) via the media capture component. In some implementations, the computer system outputs an indication (e.g., a haptic indication, a graphical indication, and / or an audio indication) that the fifth media item will be captured. In some implementations, the fifth and sixth media items are media items of the same type. In some implementations, the fifth and sixth media items are media items of different types. In some implementations, the computer system stops capturing the sixth media item based on determining that the person in question is no longer positioned in the target pose. Selectively capturing the sixth media item based on whether a change in the pose of the person in question has been detected within a threshold time period for capturing the fifth media item automatically allows the computer system to intelligently perform additional media capture procedures based on changes in the pose of the fifth media item (e.g., or intelligently abandon additional media capture procedures), thereby performing operations when a set of conditions has been met without requiring further user input and providing additional control options without cluttering the user interface with additional displayed controls.
[0198] It should be noted that the above text regarding process 700 (for example, Figure 7 The details of the process described herein also apply in a similar manner to the other methods described herein. For example, process 800 may optionally include one or more characteristics of the various methods described above with reference to process 700. For example, the movement pattern of process 800 may be used to locate the camera after the first input of process 700 is detected. For the sake of brevity, these details will not be repeated herein.
[0199] Figure 8This is a flowchart illustrating a method for repositioning a camera (e.g., process 800) according to some embodiments. Some operations in process 800 may be optionally combined, the order of some operations may be optionally changed, and some operations may be optionally omitted.
[0200] As described below, Process 800 provides an intuitive way to reposition the camera. Process 800 reduces the cognitive load on the user, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to interact with such devices faster and more efficiently, saving power and increasing the time between battery charging.
[0201] In some embodiments, process 800 is performed at a computer system (e.g., 600) that communicates with a media capture component (e.g., 602) (e.g., a periscope camera, telephoto camera, wide-angle camera, and / or ultra-wide-angle camera) and a moving component (e.g., actuators (e.g., pneumatic actuators, hydraulic actuators, and / or electric actuators), a movable base, a rotatable component, a motor, a lift, a level, and / or a rotatable base) (e.g., different from and separate from the media capture component). In some embodiments, the computer system is a telephone, watch, tablet, fitness tracker, wearable device, display, portable computer system, accessory, speaker, light, head-mounted display (HMD), and / or personal computing device. In some embodiments, the computer system includes the media capture component and / or the moving component.
[0202] When video (and / or media) is captured via a media capture component (e.g., 602) (and / or a microphone) (802) and according to a first set of conditions determined to be met, the computer system moves (804) (e.g., physically moves) a portion of the computer system (e.g., 600) including the media capture component (e.g., 602) in a first movement mode via a movement component (e.g., via, using, and / or utilizing) a first movement mode, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves (e.g., as described above in...). Figure 6G , Figure 6J , Figure 6M(As described herein). In some embodiments, prior to capturing video via the media capture component, the computer system detects input corresponding to the request to capture media via input components (e.g., camera, depth sensor, microphone, hardware input mechanism, rotatable input mechanism, heart monitor, temperature sensor, and / or touch-sensitive surface) that communicate with the computer system. In some embodiments, in response to detecting input corresponding to the request to capture media, the computer system initiates video capture via the media capture component.
[0203] When video is captured via a media capture component (802) and according to a determination of a second set of one or more capture conditions, wherein the second set of one or more capture conditions differs from a first set of one or more capture conditions, the computer system moves (806) via a movement component in a second movement mode different from the first movement mode (e.g., physically moves) a portion of the computer system (e.g., 600) (e.g., physical parts, camera, display components, display, center and / or specific parts of the display and / or hardware buttons) in a second movement mode different from the first movement mode, wherein moving the portion of the computer system in the second movement mode causes the framing of the video (e.g., 644 and / or 648) to change as the portion of the computer system moves (e.g., as described above in...). Figure 6G , Figure 6J , Figure 6M (as described herein) (e.g., without moving a portion of the computer system in a first movement mode). In some embodiments, when capturing video via a camera, the computer system moves (e.g., physically moves) a portion of the computer system via a moving component, wherein the movement of the first portion of the computer system is performed in a first movement mode when a first set of one or more capture conditions is met, and in combination with and / or in response to the satisfaction of the first set of one or more capture conditions, and wherein the movement of the first portion of the computer system is performed in a second movement mode when a second set of one or more capture conditions is met, and in combination with and / or in response to the satisfaction of the second set of one or more capture conditions. Selectively moving a portion of the computer system in the appropriate mode when a specified set of conditions is met (e.g., the first set of one or more capture conditions is met or the second set of one or more capture conditions is met) automatically allows the computer system to intelligently reposition itself, thereby improving the appearance of the media capture process and the resulting media items, and performing operations when the set of conditions has been met without requiring further user input.
[0204] In some implementations, the computer system (e.g., 600) is configured to capture video using a first artistic style (e.g., a capture style used by a particular video, photographic, and / or graphic editor to capture video) (e.g., a style associated with one or more visual characteristics such as hue, color, shadow, the user's position within the field of view (e.g., the user's head or another body part being in the bottom third, middle third, or top third of the field of view), zoom amount, the amount of focus applied to different objects in the field of view, focusing more on the background and / or objects in the background, focusing on the foreground and / or objects in the foreground, the artist's selective choice of how to depict their subject based on characteristics such as form, color, and / or composition), a set of visual and / or audio characteristics of the media associated with the artist, the artist's selective choice of how to create the media (e.g., using the appropriate medium, using the appropriate lighting setup, the spatial orientation of the computer system), a first set satisfying one or more capture conditions (and a second set not satisfying one or more capture conditions) (e.g., as described above in...). Figure 6L (As described above). In some embodiments, based on determining that the media capture component and / or computer system is configured to capture video using (e.g., based on, according to) a first set of visual characteristics (e.g., hue, color, shadow, user positioning within the field of view of the media capture component, zoom amount, focus on foreground, focus on background), a first set satisfying one or more capture conditions (e.g., and a second set not satisfying one or more capture conditions). In some embodiments, based on determining that the computer system is configured to capture video using a second art style different from the first art style, a second set satisfying one or more capture conditions (e.g., as described above). Figure 6L(as described herein) (e.g., and not satisfying one or more capture conditions). In some embodiments, the media capture component and / or computer system are configured to capture video using (e.g., based on, according to) a second set of visual characteristics that differs from the first set of visual characteristics (e.g., hue, color, shadow, user positioning within the field of view of the media capture component, zoom amount, focus on foreground, focus on background), satisfying one or more capture conditions and not satisfying one or more camera conditions of the first set. In some embodiments, the first set and / or the second set of visual characteristics are based on user preferences. In some embodiments, the first set and / or the second set of visual characteristics are based on the preferences of an individual other than the user. In some embodiments, the computer system displays an indication of the appropriate set of visual characteristics to be used for capturing video, while the media capture component is configured to capture video using the appropriate set of visual characteristics. Selectively moving in a corresponding manner when a specified set of conditions is met (e.g., the computer system is configured to capture video using a first art style or a second art style) automatically allows the computer system to instruct the user on what style the computer system will capture video in, thereby performing the operation when the set of conditions has been met without requiring further user input.
[0205] In some implementations, the computer system detects input (e.g., tap input, swipe input, rotation of a rotatable input mechanism, voice command, gaze, and / or air gesture) corresponding to a corresponding art style (e.g., 605b and / or 605f) (e.g., before and / or when video (and / or media) is captured via the media capture component). In some implementations, after detecting input corresponding to a corresponding art style and in response to the occurrence of a triggering condition (e.g., user input or automatic trigger) corresponding to the captured video (e.g., 644 and / or 648), and based on determining that the corresponding art style corresponds to a first art style, the computer system uses the first art style (e.g., as described above). Figure 6J , Figure 6K , Figure 6L and Figure 6M The video is captured using the second art style (as described above) instead of the second art style and / or in cases where the computer system is not configured to capture video using the second art style. In some implementations, after detecting input corresponding to the corresponding art style and in response to the occurrence of a trigger condition corresponding to the captured video, and based on determining that the corresponding art style corresponds to the second art style, the computer system uses the second art style (e.g., as described above). Figure 6J , Figure 6K , Figure 6L and Figure 6MThe video is captured using the style described herein (instead of the first art style and / or in cases where the computer system is not configured to capture video using the first art style). In some embodiments, the media capture component and / or the computer system is configured to capture video using a second set of visual characteristics in response to the computer system detecting input corresponding to a user. Capturing video using the corresponding art style automatically allows the computer system to capture video using a user-preferred / requested style after detecting input corresponding to the corresponding art style and in response to the occurrence of a first triggering condition corresponding to the captured video, thereby performing the operation without further user input when the set of conditions has been met.
[0206] In some embodiments, the computer system detects the occurrence of one or more corresponding conditions (e.g., before and / or when video (and / or media) is captured via the media capture component) without detecting input from the user (e.g., 605a, 605f, 605h, and / or 605k). In some embodiments, after detecting the occurrence of one or more corresponding conditions without detecting input from the user and in response to the occurrence of a trigger condition (e.g., user input or automatic trigger) corresponding to the captured video (e.g., 644 and / or 648), and based on determining that one or more corresponding conditions correspond to a first art style, the computer system uses the first art style (instead of a second art style and / or captures video without configuring the computer system to use a second art style) to capture video. In some implementations, after detecting the occurrence of one or more corresponding conditions without detecting input from the user, and in response to the occurrence of a trigger condition corresponding to the captured video, and based on determining that one or more corresponding conditions correspond to a second art style, the computer system uses the second art style (instead of the first art style and / or captures video using the first art style when the computer system is not configured to do so) to capture video (e.g., as described above). Figure 6J (As described in the description). After detecting the occurrence of one or more corresponding conditions without detecting user input, using the appropriate art style to capture video allows the computer system to use an art style that optimally complements one or more corresponding conditions to capture video, thereby performing operations when the set of conditions has been met without requiring further user input.
[0207] In some implementations, a first set of one or more media capture settings of a computer system (e.g., 600) that determines whether the computer system is configured to capture a specific type of media (e.g., still photos, videos, a series of moving images, panoramic photos, and / or portraits) is active, a first set that satisfies one or more capture conditions (e.g., and a second set that does not satisfy one or more capture conditions) (e.g., as described above) Figure 6L and Figure 6M (As described above). In some implementations, based on determining that a second set of one or more media capture settings of the computer system, which differs from a first set of one or more media capture settings, is active (e.g., and the first set of one or more media capture settings of the computer system is inactive), a second set satisfies one or more capture conditions (e.g., and the first set does not satisfy one or more capture conditions) (e.g., as described above). Figure 6L and Figure 6M (As described herein). In some embodiments, the computer system displays an indication of which corresponding media capture setting of the computer system is active. In some embodiments, the corresponding media capture setting corresponds to the type of media captured via the media capture component (e.g., still photos or videos). In some embodiments, the corresponding media setting corresponds to the configuration of the media capture component (e.g., the media capture component is configured to capture portraits, panoramas, time-lapse photos, and / or slow-motion videos). Moving a portion of the computer system in a corresponding manner (e.g., a first manner or a second manner) based on which set of one or more media capture settings is active allows the computer system to instruct which guidelines (e.g., settings) to follow when capturing video, thereby performing operations when the set of conditions has been met without further user input and providing improved feedback.
[0208] In some embodiments, the computer's first media capture setting corresponds to capturing media (e.g., video and / or still photographs) with one or more colors (e.g., capturing media in shades (e.g., black and white)). In some embodiments, the computer system (e.g., 600)'s second media capture setting corresponds to capturing media with one or more colors different from the first set of one or more colors (e.g., more colors than black and white) (e.g., as described above). Figure 6L and Figure 6M(As described herein). In some embodiments, the computer system has a corresponding media capture setting in which a first portion of the media is captured without color, and a second portion of the media is captured with color. In some embodiments, the first set of colors does not include one or more colors from the second set of colors. In some embodiments, the second set of colors includes one or more colors from the first set of colors. In some embodiments, color filters and / or black-and-white filters are applied to images captured with the first set of colors, and color filters and / or black-and-white filters are not applied to images captured with the second set of colors.
[0209] In some implementations, a portion of the media capture component (e.g., 602) that moves the computer system (e.g., 600) in a first movement mode includes performing a first type of movement (e.g., translational movement and / or rotational movement and / or motion) via the movement component, and performing a second type of movement different from the first type of movement (e.g., as described above) via the movement component. Figure 6G (As described herein). In some embodiments, moving the computer system in a second movement mode as part of a media capture component includes performing a third type of movement via the movement component, and performing a fourth type of movement different from the third type of movement via the movement component. In some embodiments, the first type of movement, the second type of movement, the third type of movement, and the fourth type of movement are different types of movement. In some embodiments, the first type of movement and the third type of movement are the same type of movement. In some embodiments, the first type of movement and the fourth type of movement are the same type of movement. In some embodiments, the computer system performs two or more types of movement concurrently. In some embodiments, the computer system performs two or more types of movement serially. Performing the first type of movement and the second type of movement as part of moving the computer system in a first movement mode allows the computer system to better frame the content of the video when one or more conditions change, thereby performing operations when the set of conditions has been met without further user input, providing improved feedback (e.g., video is being captured), and providing additional control options without cluttering the user interface with additional displayed controls.
[0210] In some implementations, the first type of movement of a portion of the computer system (e.g., 600) including a media capture component (e.g., 602) is in a first lateral (e.g., sideways, leftward, and / or lateral) direction (e.g., as described above in...). Figure 6GThe movement is described in the description (e.g., along the lateral axis of the moving component). In some embodiments, a second type of movement of the computer system including the media capture component is movement along the lateral axis of the moving component in a second lateral direction (e.g., lateral, rightward, and / or sideways) opposite to the first lateral direction. In some embodiments, two or more types of movement (e.g., first type of movement and second type of movement) include movement in the same lateral direction at different magnitudes. In some embodiments, two or more types of movement include movement in the same lateral direction at different speeds. In some embodiments, two or more types of movement include movement in the same lateral direction while rotating in different directions. In some embodiments, two or more types of movement include movement in the same lateral direction while rotating at different speeds. Moving the computer system including the media capture component in the first lateral direction as part of moving the computer system in the first movement mode allows the computer system to maintain framing of the user while the user moves in the lateral direction during video capture, thereby performing operations when a set of conditions has been met without further user input, providing improved feedback (e.g., video is being captured), and providing additional control options without cluttering the user interface due to additional displayed controls.
[0211] In some implementations, the first type of movement of a computer system (e.g., 600) including a portion of a media capture component (e.g., 602) includes movement in a first vertical direction (e.g., as described above). Figure 6GThe movement is described in the description (e.g., upward and / or downward) (e.g., along the vertical axis of the moving component). In some embodiments, the second type of movement includes movement along the vertical axis of the moving component in a second vertical direction opposite to the first vertical direction (e.g., upward and / or downward). In some embodiments, two or more types of movement (e.g., the first type of movement and the second type of movement) include movement in the same vertical direction at different magnitudes. In some embodiments, two or more types of movement include movement in the same vertical direction at different speeds. In some embodiments, two or more types of movement include movement in the same vertical direction while rotating in different directions. In some embodiments, two or more types of movement include movement in the same vertical direction while rotating at different speeds. As part of moving the computer system in the first vertical direction in the first movement mode, the inclusion of a media capture component allows the computer system to maintain framing of the user while the user moves in the vertical direction during video capture, thereby performing operations when a set of conditions has been met without further user input, providing improved feedback (e.g., video is being captured), and providing additional control options without cluttering the user interface due to additional displayed controls.
[0212] In some implementations, the first type of movement of a computer system (e.g., 600) including a portion of a media capture component (e.g., 602) includes movement in a first longitudinal (e.g., forward and / or backward) direction (e.g., towards and / or away from the user) (e.g., along the longitudinal axis of the moving component) (e.g., as described above). Figure 6G(As described herein). In some embodiments, the second type of movement includes movement along the longitudinal axis of the moving component in a second longitudinal direction (e.g., forward and / or backward) opposite to the first longitudinal direction (e.g., movement toward and / or away from the user). In some embodiments, two or more types of movement include movement in the same longitudinal direction at different magnitudes. In some embodiments, two or more types of movement include movement in the same longitudinal direction at different speeds. In some embodiments, two or more types of movement include movement in the same longitudinal direction while rotating in different directions. In some embodiments, two or more types of movement include movement in the same longitudinal direction while rotating at different speeds. As part of moving the computer system in the first longitudinal direction in the first movement mode, including a media capture component, the computer system can maintain and / or change the distance between a portion of the computer system and the content of the video, thereby performing operations when a set of conditions has been met without further user input, providing improved feedback (e.g., video is being captured), and providing additional control options without cluttering the user interface due to additional displayed controls.
[0213] In some embodiments, one or more of the first type of movement of the computer system (e.g., 600) including a portion of a media capture component (e.g., 602) and the second type of movement of the computer system including a portion of a media capture component include rotational movement (e.g., as described above). Figure 6G (As described herein). In some embodiments, the first type of movement includes rotation (e.g., yaw, roll, and / or pitch rotation) about a first axis of the moving component. In some embodiments, the second type of movement includes rotation (e.g., yaw, roll, and / or pitch rotation) about a second axis of the moving component. In some embodiments, the first axis and the second axis are different. In some embodiments, the first axis and the second axis are the same. In some embodiments, the computer system rotates about the first axis and the second axis at different speeds. In some embodiments, the computer system rotates about the first axis and the second axis at the same speed. In some embodiments, the computer system translates about a first corresponding axis of the moving component while rotating about a second corresponding axis of the moving component.
[0214] In some embodiments, when a portion of the computer system including the media capture component moves in a first movement mode, a portion of the computer system (e.g., 600) including the media capture component (e.g., 602) moves at a first rate (e.g., measured in feet per second, meters per second, and / or inches per second) (e.g., velocity and / or acceleration). In some embodiments, when a portion of the computer system including the media capture component moves in a second movement mode, the portion of the computer system including the media capture component moves at a second rate (e.g., measured in feet per second, meters per second, and / or inches per second) different from the first rate (e.g., greater than or less than the first speed). In some embodiments, when a portion of the computer system moves in a corresponding movement mode, the portion of the computer system accelerates and / or decelerates. In some embodiments, when a portion of the computer system moves in a corresponding movement mode, the portion of the computer system neither accelerates nor decelerates. When a set of conditions is met (e.g., a first set of one or more conditions or a second set of one or more conditions), a portion of the computer system, including a media capture component, moves at a corresponding rate, automatically allowing the computer system to indicate which capture conditions are met, thereby performing an operation when the set of conditions has been met without requiring further user input.
[0215] In some implementations, when a portion of a computer system (e.g., 600) including a media capture component (e.g., 602) moves in a first movement mode, the framing of video (e.g., 644 and / or 648) is altered (e.g., tracking the user and / or following the user) based on (e.g., according to and / or using) a first set of one or more tracking parameters (e.g., how closely the user is tracked, how long the user is tracked, how long the user is outside the frame, criteria for stopping tracking the user, the magnification level of the media capture component, and / or the user's positioning within the framing of the media capture component). In some implementations, when a portion of the computer system, including a media capture component, moves in a second movement mode, the framing of the video is altered based on (e.g., according to and / or using) a second set of one or more tracking parameters that differs from a first set of tracking parameters (e.g., how closely the user is tracked, how long the user is tracked, how long the user is outside the frame, criteria for stopping tracking the user, the magnification level of the media capture component, and / or the positioning of the user and / or the user's body parts within the frame of one or more cameras, such as the user's head remaining in the middle third, top third, or bottom third of the field of view of the media capture component, and / or the middle third, top third, or bottom third of the field of view corresponding to the top third, middle third, and / or bottom third of the media) (e.g., tracking the user and / or following the user). Figure 6G(As described herein). In some implementations, the framing of the video changes differently when the framing changes based on a first set of one or more tracking parameters compared to when the framing changes based on a second set of one or more tracking parameters. Differentiating the framing of the video based on whether a portion of the computer system's media capture component moves relative to a first or second movement mode allows the computer system to instruct itself to move relative to either the first or second movement mode via the framing of the video, thereby providing improved visual feedback.
[0216] It should be noted that the above text regarding process 800 (for example, Figure 8 The details of the process described herein also apply in a similar manner to the other methods described herein. For example, process 900 may optionally include one or more characteristics of the various methods described above with reference to process 800. For example, schematic guidance for process 900 may be provided before or after moving the computer system in the moving mode of process 800. For the sake of brevity, these details will not be repeated herein.
[0217] Figure 9 This is a flowchart illustrating a method (e.g., process 900) for providing diagramming guidance according to some implementation schemes. Some operations in process 900 may be optionally combined, the order of some operations may be optionally changed, and some operations may be optionally omitted.
[0218] As described below, Process 900 provides an intuitive way to offer design guidance. Process 900 reduces the cognitive load on users, thereby creating a more efficient human-computer interface. For battery-powered computing devices, it enables users to interact with such devices faster and more efficiently, saving power and increasing the time between battery charging.
[0219] In some implementations, process 900 is performed at a computer system (e.g., 602) that communicates with media capture components (e.g., 602) (e.g., sensors, environmental sensors, capture components, input components (e.g., cameras, depth sensors, microphones, hardware input mechanisms, rotatable input mechanisms, heart monitors, temperature sensors, and / or touch-sensitive surfaces), cameras (e.g., periscope cameras, telephoto cameras, wide-angle cameras, and / or ultra-wide-angle cameras), depth sensors, microphones, heart monitors, and / or temperature sensors), input components (e.g., 602) (e.g., cameras, depth sensors, microphones, hardware input mechanisms, rotatable input mechanisms, heart monitors, temperature sensors, and / or touch-sensitive surfaces), and output components (e.g., display components (e.g., displays, projectors, and / or touch-sensitive displays), audio components (e.g., smart speakers, home theater systems, soundbars, headphones, earphones, earbuds, speakers, TV speakers, augmented reality headphone speakers, audio jacks, optical audio outputs, Bluetooth audio outputs, and / or HDMI audio outputs), speakers, and / or haptic output devices). In some embodiments, the computer system is a telephone, watch, tablet, fitness tracker, wearable device, display, portable computer system, accessory, speaker, light, head-mounted display (HMD), and / or personal computing device. In some embodiments, the media capture component includes and / or input and / or output components (e.g., input and output components). In some embodiments, the output component is different from the media capture component and / or input component. In some embodiments, the input component is different from the media capture component and / or output component.
[0220] The computer system detects (902) a first set of one or more inputs (e.g., 605a, 605f, 605h, and / or 605k) corresponding to one or more instructions (e.g., instructions included in 605a, 605f, 605h, and / or 605k) comprising one or more spoken words (e.g., verbal input (e.g., spoken words, voice, audible requests, audible commands, and / or audible statements) and / or non-verbal input (e.g., swipe input, press and drag input, gaze input, air gestures, and / or mouse clicks)) via an input component (e.g., 602). In some embodiments, the first set of one or more inputs is detected when in a media capture mode (e.g., a mode in which the computer system is configured to capture images, video, and / or audio). In some embodiments, the first set of one or more inputs is detected when the computer system displays a user interface via a display component in communication with the computer system (e.g., corresponding to the media capture mode). In some implementations, a first set of one or more inputs is detected when the computer system displays a live preview of the output (e.g., images, videos, and / or audio) of the media capture component via a display component that communicates with the computer system.
[0221] In response to the detection of a first set of one or more inputs (e.g., 605a, 605f, 605h and / or 605k) corresponding to one or more instructions, the computer system prepares (904) to capture media (e.g., images, video and / or audio) via a media capture component (e.g., 602). In some embodiments, the preparation to capture media occurs before capturing one or more images and / or media, or simultaneously with displaying a representation of the field of view of one or more cameras and / or one or more media capture components.
[0222] When preparing to capture media via the media capture component (906) (and / or in conjunction with preparing to capture media via the media capture component) and in accordance with determining one or more instructions including first content (e.g., one or more first sets of spoken words), the computer system provides (908) (and / or outputs) first compositional guidance (e.g., 642 and / or 646) via the output component (e.g., visual, tactile and / or audio output by the computer system, which includes instructions for changing the spatial orientation of one or more subjects within the field of view of the media capture component (e.g., these instructions command one or more subjects to move closer, move further, move to the left, move down and / or move up within the field of view of the media capture component, and / or these instructions command one or more subjects to change the spatial relationship between one or more subjects)), wherein the first compositional guidance includes one or more recommendations for changing the spatial arrangement of one or more objects (e.g., 640) in the field of view (e.g., 654) of the media capture component (e.g., 602).
[0223] When preparing to capture media via the media capture component (906) and determining that one or more instructions include second content (e.g., one or more spoken words) different from the first content (e.g., different from a first set of one or more spoken words), the computer system provides (910) (and / or outputs) a second composition guide (e.g., 642 and / or 646) different from the first composition guide (e.g., 642 and / or 646) via the output component (in combination with the first composition guide provided via the output component, in the case of the first composition guide provided via the output component, and / or in the case of the first composition guide not being provided by the output component), wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects (e.g., 640) in the field of view (e.g., 654) of the media capture component (e.g., 602). Selectively outputting composition guides when a specified set of conditions is met (e.g., one or more instructions include the first content or the second content) allows the computer system to automatically adapt the composition guides to the proficiency of the user of the computer system, such that less proficient users are given more guidance than more proficient users, thereby performing operations without further user input when the set of conditions has been met.
[0224] In some implementations, when preparing to capture media via a media capture component (e.g., 602) (and / or in some implementations, before, after, or in combination with it) and based on determining that the corresponding subject (e.g., 622) is located at a first position within the field of view (e.g., 654) of the media capture component (e.g., centered within the field of view of the media capture component, to the left or right of the center of the field of view of the media capture component, and / or above or below the center of the field of view of the media capture component), the computer system provides third compositional guidance (e.g., 642 and / or 646) via an output component (e.g., which includes one or more recommendations for moving one or more objects within and / or outside the field of view of the media capture component). In some embodiments, when preparing to capture media via the media capture component and based on determining that the subject is located at a second position different from the first position within the field of view of the media capture component, the computer system provides a fourth composition guide (e.g., 642 and / or 646) via an output component, different from the third composition guide (e.g., including one or more recommendations for moving one or more objects within and / or outside the field of view of the media capture component). In some embodiments, the third composition guide is different from or the same as the first and / or second composition guides. In some embodiments, the fourth composition guide is different from or the same as the first and / or second composition guides. Selectively providing composition guidance based on the user's positioning allows the computer system to automatically provide instructions on the user's positioning relative to the computer system, thereby performing operations when a set of conditions has been met without further user input.
[0225] In some embodiments, when preparing to capture media (e.g., 605a, 605f, 605h, and / or 605k) via a media capture component (e.g., 602) (and / or in some embodiments, before, after, or in conjunction with it), and based on determining a first set of detected elements (e.g., in the field of view of the media capture component and / or in the environment (e.g., the environment of the computer system and / or the user's environment)) (e.g., living and / or inanimate elements), the computer system provides a fifth mapping guide (e.g., 642 and / or 646) via an output component (e.g., which includes one or more recommendations for changing the spatial arrangement of the first set of elements). In some embodiments, when preparing to capture media via the media capture component and based on determining a second set of elements different from the first set of elements (e.g., in the field of view of the media capture component and / or in the environment (e.g., the environment of the computer system and / or the user's environment)), the computer system provides a sixth mapping guide (e.g., 642 and / or 646) different from the fifth mapping guide (e.g., which includes one or more recommendations for changing the spatial arrangement of the second set of elements). In some implementations, the fifth composition guide may be different from or the same as the first and / or second composition guides. In some implementations, the sixth composition guide may be different from or the same as the first and / or second composition guides. Selectively providing composition guidance based on which elements are detected automatically allows the computer system to provide indications of which elements are within the field of view of the media capture component, thereby performing operations when a set of conditions has been met without requiring further user input.
[0226] In some implementations, a first set of inputs (e.g., 605a, 605f, 605h and / or 605k) is determined to have a first level of detail, and a first composition guide (e.g., 642 and / or 646) (e.g., including first content based on determining one or more instructions) (or a second composition guide including second content based on determining one or more instructions) has a first level of information (e.g., guidance, instruction, and / or directive) (e.g., as described above). Figure 6I (As described above). In some embodiments, based on determining that a first set of one or more inputs has a second level of detail greater than a first level of detail, a first composition guide (e.g., based on determining that one or more instructions include first content) (or based on determining that one or more instructions include second content, a second composition guide) has a second information content greater than a first information content (e.g., guidance, instruction, and / or instruction) (e.g., as described above). Figure 6I(As described herein) (e.g., a first amount of information commands the corresponding subject to move closer to the computer system, while a second amount of information commands the corresponding subject to move 8 feet closer to the computer system, and / or the first amount of information commands the corresponding subject to move to the left, while the second amount of information commands the corresponding subject to move to the left until the corresponding subject is positioned in front of an object (e.g., a chair, sofa, and / or television). Selectively providing a first compositional guide with a specific amount of information automatically allows the computer system to customize the content of its guide based on input, thereby performing an operation when a set of conditions has already been met without requiring further user input.
[0227] In some embodiments, when preparing to capture media (e.g., 605a, 605f, 605h and / or 605k) via a media capture component (e.g., 602) (and / or in some embodiments, before, after, or in combination with it) (and in some embodiments, including first content based on determining one or more instructions, or including second content based on determining one or more instructions) and based on determining that a first set of one or more inputs (e.g., 605a, 605f, 605h and / or 605k) corresponds to a first body (e.g., 622) (e.g., the first set of one or more inputs is performed by the first body and / or the first set of one or more inputs refers to the first body) (e.g., an individual registered with the computer system or an individual not registered with the computer system) (e.g., a subject, user, person, animal, and / or object), the computer system provides an eighth composition guide (e.g., 642 and / or 646) via an output component (e.g., which includes one or more individual-specific recommendations for changing the spatial arrangement of the first body and / or other individuals positioned within the field of view of the media capture component). In some implementations, when preparing to capture media via the media capture component and based on determining that a first set of one or more inputs corresponds to a second body different from the first body (e.g., 622) (e.g., the first set of one or more inputs is performed by the second body and / or the first set of one or more inputs refers to the second body) (e.g., an individual registered with the computer system or an individual not registered with the computer system) (e.g., a subject, user, person, animal, and / or object), the computer system provides a ninth composition guide (e.g., 642 and / or 646) via an output component, different from the eighth composition guide (e.g., including one or more individual-specific recommendations for changing the spatial arrangement of the second body and / or other individuals positioned within the field of view of the media capture component). In some implementations, the composition guide is based on which individual provided the first set of one or more inputs and the content of the first set of one or more inputs. Selectively providing a type of composition guide when a set of conditions is met automatically allows the computer system to customize the content of its guide on a per-user basis, thereby performing operations when the set of conditions has been met without requiring further user input.
[0228] In some implementations, the first composition guide (e.g., 642 and / or 646) (and / or the second composition guide) includes one or more recommendations (e.g., audible and / or visual instructions) regarding one or more characteristics of lighting in the environment (e.g., the user's environment and / or the computer system's environment) within the field of view (e.g., 602) of the media capture component (e.g., 654) that should be altered (e.g., the amount of lighting, the brightness of lighting, the color of lighting, the hue of lighting, and / or the hue of lighting). Figure 6I(As described herein). In some embodiments, the second mapping guidance includes instructions for changing one or more characteristics of light in the environment. In some embodiments, the computer system automatically (e.g., without intermediate user input) changes one or more characteristics of lighting in response to detecting a first set of one or more inputs. Providing one or more recommendations on how to change one or more characteristics of lighting in the environment within the field of view of the media capture component when preparing to capture media allows the computer system to provide indications about the state of the computer system (e.g., the computer system is preparing to capture media) and the state of the environment within the field of view of the media capture component (e.g., the current characteristics of the lighting are not optimal for capturing media), thereby providing improved feedback and providing additional control options without cluttering the user interface with additional displayed controls.
[0229] In some implementations, the first composition guide (e.g., 642 and / or 646) (and / or the second composition guide) includes (e.g., audible and / or visual instructions) one or more recommendations (e.g., as described above) that should change the amount of light in the environment (e.g., ambient light, surrounding ambient light, one or more lights in the environment not physically coupled to the computer system) in the field of view (e.g., 654) of the media capture component (e.g., 602). Figure 6I (as described in the description) (e.g., increasing or decreasing the amount of light). Providing one or more recommendations on how to change the amount of light in the environment within the field of view of the media capture component when preparing to capture media via the media capture component allows the computer system to provide indications about the state of the computer system (e.g., the computer system is preparing to capture media) and the state of the environment within the field of view of the media capture component (e.g., there is too little or too much light in the environment), thereby providing improved feedback and offering additional control options without cluttering the user interface with additional displayed controls.
[0230] In some implementations, the first compositional guidance (e.g., 642 and / or 646) (and / or the second compositional guidance) includes (e.g., audible and / or visual instructions) one or more recommendations (e.g., as described above) regarding the type of light in the environment (e.g., direct light, indirect light, artificial light, and / or natural light) in the field of view (e.g., 654) of the media capture component (e.g., 602). Figure 6I(As described herein). In some embodiments, instructions for changing one or more characteristics of lighting in the environment include instructions for changing two or more types of light in the environment. Providing one or more recommendations on the type of light in the environment in the field of view of the media capture component when preparing to capture media via the media capture component allows the computer system to provide indications about the state of the computer system (e.g., the computer system is preparing to capture media) and the state of the environment in the field of view of the media capture component (e.g., the type of lighting in the environment is not optimal for capturing media), thereby providing improved feedback and providing additional control options without cluttering the user interface with additional displayed controls.
[0231] In some implementations, the first composition guide (e.g., 642 and / or 646) (and / or the second composition guide) includes one or more recommendations for one or more colors of light in the environment that should be altered in the field of view (e.g., 654) of the media capture component (e.g., 602) (e.g., as described above). Figure 6I (As described in the description). When preparing to capture media via the media capture component, providing one or more recommendations for changing one or more colors of light in the environment within the field of view of the media capture component allows the computer system to provide indications about the state of the computer system (e.g., the computer system is preparing to capture media) and the state of the environment within the field of view of the media capture component (e.g., the color of light in the environment is not optimal for capturing media), thereby providing improved feedback and providing additional control options without cluttering the user interface with additional displayed controls.
[0232] In some embodiments, the computer system (e.g., 600) communicates with a collection of one or more external lights (e.g., lighting fixtures, point light sources, spotlights, and / or one or more light sources). In some embodiments, when preparing to capture media (e.g., 605a, 605f, 605h, and / or 605k) via a media capture component (e.g., 602) (and / or, in some embodiments, before, after, or in conjunction with it) and without detecting input corresponding to the subject (e.g., user input and / or input performed by the subject) (e.g., automatically), the computer system transmits instructions to the collection of one or more external lights, wherein transmitting instructions to the collection of one or more external lights causes one or more characteristics of the lighting in the environment within the field of view (e.g., 654) of the media capture component (e.g., as described above). Figure 6I(As described herein). In some embodiments, in response to the computer system receiving confirmation from the subject that it has approved changing one or more characteristics of the lighting in the environment, an instruction is transmitted to a set of one or more external lights. In some embodiments, the instruction is transmitted to a set of one or more external lights even if the computer system has not received confirmation from the subject that it has approved changing one or more characteristics of the lighting in the environment. When preparing to capture media, transmitting instructions to a set of one or more external lights without detecting input corresponding to the subject automatically allows the computer system to control one or more characteristics of the lighting in the environment within the field of view of the media capture component to improve media capture, thereby performing operations without further user input when a set of conditions has been met (e.g., the computer system is preparing to capture media).
[0233] In some implementations, the first composition guidance (e.g., 642 and / or 646) includes one or more recommendations (e.g., audible cues and / or visual cues) for that one or more objects should be moved (e.g., 640) (e.g., living objects and / or inanimate objects) (e.g., someone should move the object). Figure 6I (As described in the description). When a computer system is ready to capture media, providing one or more recommendations on which one or more objects should be moved allows the computer system to indicate the state of the computer system (e.g., the computer system is preparing to capture media) and cause the movement of objects in the environment to improve the capture of media and the resulting media items, thereby providing improved feedback, and performing operations without further user input when a set of conditions (e.g., the computer system is preparing to capture media) has been met.
[0234] In some implementations, the first compositional guidance (e.g., 642 and / or 646) includes one or more recommendations (e.g., audible and / or visual cues) for the subject (e.g., 622) to move from a first position (e.g., to a specific position) (e.g., towards the computer system and / or media capture component, away from the computer system and / or media capture component, to the left of the computer system and / or media capture component, and / or to the right of the computer system and / or media capture component) to a second position (e.g., in the environment). Figure 6L(As described herein). In some embodiments, prompts for moving the subject include instructions to rotate (e.g., bend over) and / or perform translational movements (e.g., move to the right and / or move to the left). Providing one or more recommendations for the subject to move when the computer system is ready to capture media allows the computer system to indicate the state of the computer system (e.g., the computer system is preparing to capture media) and recommend movements of the subject in the environment to improve media capture and the resulting media item, thereby providing improved feedback and performing actions without further user input when a set of conditions has been met (e.g., the computer system is preparing to capture media).
[0235] In some embodiments, when preparing to capture media (e.g., 605a, 605f, 605h and / or 605k) via a media capture component (e.g., 602) (and / or in some embodiments, before, after, or in conjunction with it), and based on determining that a portion of the subject (e.g., 622) (e.g., the subject's face, head, torso, arm, leg, and / or limb) is positioned within the field of view (e.g., 654) of the media capture component satisfies a first set of one or more positioning criteria relative to the field of view of the media capture component (e.g., a portion of the subject is to the left of the center of the field of view of the media capture component, a portion of the subject is to the right of the center of the field of view of the media capture component ... The subject is positioned below the center of the field of view of the media capture component, with a portion of the subject above the center of the field of view, and / or a portion of the subject at the center of the field of view of the media capture component (e.g., the subject continues to be in a corresponding portion of the field of view and / or a portion of the capture media corresponding to the field of view of one or more cameras, such as the middle third, top third, and / or bottom third (or second, fourth, fifth, sixth, etc.) of the field of view of one or more cameras). The computer system provides a tenth composition guide (e.g., 642 and / or 646) via an output component (e.g., the tenth composition guide includes one or more recommendations for changing the spatial positioning of one or more users in the field of view of one or more cameras so that the user's head is centered within one or more thirds of the field of view of one or more cameras) (e.g., as described above). Figure 6I(As described above). In some embodiments, when preparing to capture media via a media capture component and based on determining that a portion of the subject is positioned within the field of view of the media capture component, a second set of one or more positioning criteria differs from a first set of positioning criteria relative to the field of view of the media capture component (e.g., a portion of the user is to the left of the center of the field of view of the media capture component, a portion of the user is to the right of the center of the field of view of the media capture component, a portion of the user is below the center of the field of view of the media capture component, a portion of the user is above the center of the field of view of the media capture component, and / or a portion of the user is at the center of the field of view of the media capture component), the computer system provides an eleventh composition guide (e.g., 642 and / or 646) via an output component, different from the tenth composition guide (e.g., as described above). Figure 6I (As described in the description). Selectively providing compositional guidance based on the positioning of a portion of the subject automatically allows the computer system to provide customized guidance for adjusting the positioning of that portion of the subject relative to the field of view of the media capture component, thereby improving the media capture process and enabling operation to be performed when a set of conditions is already met without further user input.
[0236] In some implementations, a tenth composition guide (e.g., 642 and / or 646) (and / or an eleventh composition guide) is provided by referring to the positioning of a portion of the subject relative to a fixed reference point. Figure 6I (as described in the description) (e.g., the middle third, bottom third, or top third of the field of view (or a second, fourth, fifth, sixth, etc.) and / or a portion of the captured media corresponding to a portion of the field of view) (e.g., centering the face horizontally and / or vertically). Providing compositional guidance by referencing the positioning of a portion of the subject relative to a fixed reference point automatically allows the computer system to provide indications of the spatial relationship between a portion of the subject and the fixed reference point, thereby performing operations without further user input when a set of conditions has been met.
[0237] In some implementations, a tenth compositional guide (e.g., 642 and / or 646) is provided relative to (e.g., according to, depending on, based on) the spatial relationship between a portion of the subject and the body of the subject (e.g., upper torso, lower torso, arms, and / or legs). (e.g., as described above) Figure 6I(as described herein) (e.g., mirroring and / or framing from head to body and from subject to scene). In some embodiments, compositional guidance is provided relative to the lower boundary of the target area of the media capture component's field of view when the body of the subject is above a portion of the subject. In some embodiments, compositional guidance is provided relative to the upper boundary of the target area of the media capture component's field of view when the body of the subject is below a portion of the subject. Providing compositional guidance by referencing the positioning of a portion of the subject relative to the body of the subject automatically allows the computer system to provide indications of the spatial relationship between the portion of the subject and the body of the subject, thereby performing operations without further user input when a set of conditions has been met.
[0238] In some embodiments, when preparing to capture media (e.g., 605a, 605f, 605h and / or 605k) via a media capture component (e.g., 602) (and / or in some embodiments, before, after, or in conjunction with it), and based on a first spatial orientation (e.g., relative to the spatial orientation of the media capture component, relative to the user, and / or relative to one or more objects positioned within the field of view of the media capture component), the computer system provides a twelfth composition guide (e.g., 642 and / or 646) via an output component (e.g., as described above). Figure 6I (as described above) (e.g., including one or more recommended guidelines for repositioning the horizontal plane and / or the media capture component so that the horizontal plane is level with the media capture component) (e.g., camera movement so that lines (shoulder, horizon) in the field of view are horizontal). In some embodiments, when preparing to capture media via the media capture component and based on determining that the horizontal plane in the field of view of the media capture component has a second spatial orientation different from the first spatial orientation (e.g., relative to the spatial orientation of the media capture component, relative to the user, and / or relative to one or more objects positioned within the field of view of the media capture component), the computer system provides a thirteenth composition guideline via an output component that is different from the twelfth composition guideline (e.g., 642 and / or 646) (e.g., as described above). Figure 6I(As described herein) (e.g., including one or more recommended guidelines for repositioning the horizontal plane and / or the media capture component such that the horizontal plane is not angular within the field of view of the media capture component). In some embodiments, the computer system detects the horizontal plane in the field of view of the media capture component when preparing to capture media via the media capture component. Selectively providing compositional guidance based on the orientation of the horizontal plane in the field of view of the media capture component automatically allows the computer system to provide guidance such that the plane in the field of view of the media capture component is correctly aligned with the field of view of the media capture component, thereby performing operations without further user input when a set of conditions has been met.
[0239] In some implementations, when preparing to capture media (e.g., 605a, 605f, 605h and / or 605k) via a media capture component (e.g., 602) (and / or in some implementations, before, after, or in conjunction with it) and based on the determination that there is a first distance amount (e.g., 0.1 feet to 10 feet) between the boundary of the field of view (e.g., 654) of the media capture component and a portion of the corresponding subject (e.g., 622) (e.g., the tip of a finger, the end of an arm, the shoulder, the sole of the user's foot, the top of the user's head, the edge of a desk, and / or the edge of a bookcase), the computer system provides a fourteenth composition guide (e.g., 642 and / or 646) via an output component (e.g., including one or more recommended guides for repositioning the user such that the distance between the user's limbs and the boundary of the field of view is increased or decreased such that the horizontal plane is not angled within the field of view of the media capture component). In some implementations, when preparing to capture media via the media capture component and based on determining a second distance between the boundary of the media capture component's field of view and a portion of the corresponding subject, the computer system provides a fifteenth composition guide (e.g., 642 and / or 646) via an output component, different from the fourteenth composition guide (e.g., including one or more recommended guides for repositioning the user such that the distance between the user's limbs and the boundary of the field of view increases or decreases such that the horizontal plane is not angled within the media capture component's field of view). Selectively providing composition guidance when predefined conditions are met au...
Claims
1. A method, the method comprising: At the computer system that communicates with the media capture components and microphone: The first input corresponding to the request to capture media is detected via the microphone; as well as After detecting the first input corresponding to the request to capture media: Based on determining that the first input corresponds to a first instruction, media is captured via the media capture component in response to a first set of conditions satisfying one or more conditions; as well as Based on determining that the first input corresponds to a second instruction different from the first instruction, media is captured via the media capture component in response to a second set of conditions satisfying one or more conditions, wherein the second set of conditions differs from the first set of conditions.
2. The method of claim 1, wherein the computer system communicates with the first mobile component, the method further comprising: After detecting the first input corresponding to the request to capture media, a portion of the computer system is moved via the first moving component.
3. The method according to claim 2, further comprising: After detecting the first input corresponding to the request to capture media, the positioning of the media capture component is moved via the first moving component, such that the field of view of the media capture component moves from a first positioning to a second positioning different from the first positioning.
4. The method according to any one of claims 2 to 3, wherein: After the first input is detected: Based on the determination that the first input includes one or more instructions for the computer system to move in a first manner, a portion of the computer system moves in the first manner; and Based on the determination that the first input includes one or more instructions for the computer system to move in a second manner different from the first manner, the portion of the computer system moves in the second manner different from the first manner.
5. The method of claims 1 to 4, wherein detecting the first input corresponding to the request to capture the media includes capturing one or more verbal instructions via the microphone.
6. The method of any one of claims 1 to 5, wherein detecting the first input corresponding to the request to capture the medium comprises capturing one or more gesture-based instructions via one or more input devices.
7. The method of any one of claims 1 to 6, wherein capturing media via the media capture component in response to the first set of one or more conditions includes moving a portion of the computer system in a third manner, and wherein capturing media via the media capture component in response to the second set of one or more conditions includes moving the portion of the computer system in a fourth manner different from the third manner.
8. The method according to any one of claims 1 to 7, wherein the computer system communicates with the second mobile component, the method further comprising: When the media is captured via the media capture component: Based on the determination that the media being captured is of a first type, a portion of the computer system is moved via the second moving component during the capture of the media; as well as If it is determined that the media being captured is a second type of media different from the first type of media, the movement of the portion of the computer system via the second moving component is abandoned during the capture of the media.
9. The method of claim 8, wherein the first type of media is a video or panoramic photograph, and wherein the second type of media is a still photograph.
10. The method according to any one of claims 1 to 9, wherein the first input includes an instruction for the computer system to delay the capture of the media until an input detected by the camera is detected.
11. The method of claim 10, wherein the input detected by the camera includes the detection of a gaze.
12. The method according to any one of claims 10 to 11, wherein the input detected by the camera includes the detection of a gesture.
13. The method according to any one of claims 10 to 12, wherein the input detected by the camera includes detected pose.
14. The method of any one of claims 1 to 13, wherein the first input comprises a set of one or more time-based instructions indicating one or more media capture parameters.
15. The method of claim 14, wherein the set of one or more time-based instructions includes one or more indications of when media capture will be initiated.
16. The method of any one of claims 14 to 15, wherein the set of one or more time-based instructions includes one or more indications of when the capture of the media will cease.
17. The method of any one of claims 14 to 16, wherein the set of one or more time-based instructions includes one or more indications of time intervals between the capture of individual media items.
18. The method of any one of claims 1 to 17, wherein the first input includes instructions for compositional guidance related to capturing the media.
19. The method according to any one of claims 1 to 18, wherein the first set of one or more conditions does not include conditions corresponding to the detection of input corresponding to a first user.
20. The method of claim 19, wherein the first set of one or more conditions includes a condition satisfied when it is determined that a person in the field of view of the media capture component has stopped moving for a threshold amount of time.
21. The method of any one of claims 19 to 20, wherein the first set of one or more conditions includes conditions satisfied when it is determined that a person in the field of view of the media capture component is positioned in a corresponding pose.
22. The method according to any one of claims 19 to 21, further comprising: Before capturing the media via the media capture component in response to the first set of conditions being met and before capturing the media based on the determination that the first set of conditions being met continues to be met, a countdown of the time period that must elapse before the media is captured is displayed.
23. The method of claim 22, further comprising: When the countdown for the time period that must elapse before the media is captured is displayed, the first set of conditions that have not been met is detected. as well as In response to the detection that the first set of conditions has not been met, the display of the countdown for the time period that must have elapsed before the media is captured is interrupted.
24. The method according to any one of claims 22 to 23, further comprising: When the countdown of the time period that must elapse before the media is captured is displayed, a request to capture the media before the end of the time period and before the media is captured is detected. as well as In response to the detection of the request to capture the media before the end of the time period and before the media is captured, the display of the countdown for the time period that must elapse before the media is captured is interrupted.
25. The method according to any one of claims 1 to 24, wherein the medium is a first media item, the method further comprising: After capturing the first media item, the change in the pose of the person in the field of view of the media capture component is detected via the media capture component; as well as In response to detecting a change in the pose of the person in the field of view of the media capture component, a second media item is captured via the media capture component.
26. The method according to any one of claims 1 to 25, wherein the medium is a third media item, the method further comprising: After capturing the third media item, the change in the pose of the corresponding person in the field of view of the media capture component is detected; as well as In response to detecting the change in the pose of the corresponding person: Based on determining that the corresponding person is positioned in one or more target poses, a fourth media item is captured via the media capture component; and If it is determined that the corresponding person is not located in the one or more target poses, the capture of the fourth media item is abandoned.
27. The method according to any one of claims 1 to 26, wherein the medium is a fifth media item, the method further comprising: After capturing the fifth media item, the change in the pose of the corresponding person in the field of view of the media capture component is detected; as well as In response to detecting the change in the pose of the corresponding person: A sixth media item is captured via the media capture component based on the determination that a change in the pose of the corresponding person is detected within a threshold time period for capturing the fifth media item; as well as If no change in the pose of the person is detected within the threshold time period for capturing the fifth media item, the capture of the sixth media item is abandoned via the media capture component.
28. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a microphone, the one or more programs including instructions for performing the method according to any one of claims 1 to 27.
29. A computer system configured to communicate with a media capture component and a microphone, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 1 to 27.
30. A computer system configured to communicate with a media capture component and a microphone, the computer system comprising: Components for performing the method according to any one of claims 1 to 27.
31. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a microphone, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 27.
32. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a microphone, the one or more programs including instructions for performing the following operations: The first input corresponding to the request to capture media is detected via the microphone; and After detecting the first input corresponding to the request to capture media: Based on determining that the first input corresponds to a first instruction, media is captured via the media capture component in response to a first set of conditions satisfying one or more conditions; and Based on determining that the first input corresponds to a second instruction different from the first instruction, media is captured via the media capture component in response to a second set of conditions satisfying one or more conditions, wherein the second set of conditions differs from the first set of conditions.
33. A computer system configured to communicate with a media capture component and a microphone, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: The first input corresponding to the request to capture media is detected via the microphone; and After detecting the first input corresponding to the request to capture media: Based on determining that the first input corresponds to a first instruction, media is captured via the media capture component in response to a first set of conditions satisfying one or more conditions; as well as Based on determining that the first input corresponds to a second instruction different from the first instruction, media is captured via the media capture component in response to a second set of conditions satisfying one or more conditions, wherein the second set of conditions differs from the first set of conditions.
34. A computer system configured to communicate with a media capture component and a microphone, the computer system comprising: A component for detecting a first input corresponding to a request to capture media via the microphone; and The component is used after detecting the first input corresponding to the request to capture the media: Based on determining that the first input corresponds to a first instruction, media is captured via the media capture component in response to a first set of conditions satisfying one or more conditions; as well as Based on determining that the first input corresponds to a second instruction different from the first instruction, media is captured via the media capture component in response to a second set of conditions satisfying one or more conditions, wherein the second set of conditions differs from the first set of conditions.
35. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a media capture component and a microphone, the one or more programs comprising instructions for performing the following operations: The first input corresponding to the request to capture media is detected via the microphone; and After detecting the first input corresponding to the request to capture media: Based on determining that the first input corresponds to a first instruction, media is captured via the media capture component in response to a first set of conditions satisfying one or more conditions; and Based on determining that the first input corresponds to a second instruction different from the first instruction, media is captured via the media capture component in response to a second set of conditions satisfying one or more conditions, wherein the second set of conditions differs from the first set of conditions.
36. A method, the method comprising: At the computer system that communicates with the media capture component and the motion component: When video is captured via the media capture component: Based on a first set of conditions that satisfy one or more capture conditions, a portion of the computer system, including the media capture component, is moved in a first movement mode via the moving component, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves. as well as Based on a second set of conditions that satisfy one or more capture conditions, wherein the second set of capture conditions is different from the first set of capture conditions, the portion of the computer system is moved via the moving component in a second moving mode different from the first moving mode, wherein moving the portion of the computer system in the second moving mode causes the framing of the video to change as the portion of the computer system moves.
37. The method of claim 36, wherein: The first set of conditions, which determines that the computer system is configured to capture video using a first art style, satisfies one or more capture conditions; as well as The second set of conditions is determined to be based on the computer system being configured to capture video using a second art style different from the first art style, satisfying one or more capture conditions.
38. The method of claim 37, further comprising: Detect input that corresponds to the given art style; as well as After detecting the input corresponding to the corresponding art style and in response to the occurrence of a trigger condition corresponding to the captured video: Based on the determination that the corresponding art style corresponds to the first art style, the first art style is used to capture video; as well as Based on the determination that the corresponding art style corresponds to the second art style, the second art style is used to capture video.
39. The method of claim 37, further comprising: Detect the occurrence of one or more corresponding conditions without detecting input from the user; as well as After detecting the occurrence of one or more corresponding conditions in the absence of input from the user, and in response to the occurrence of the trigger condition corresponding to the captured video: Based on determining that one or more corresponding conditions correspond to the first art style, the first art style is used to capture video; as well as Based on the determination that one or more corresponding conditions correspond to the second art style, the second art style is used to capture video.
40. The method according to any one of claims 36 to 39, wherein: Based on determining that a first set of one or more media capture settings of the computer system is active, the first set satisfies one or more capture conditions; as well as The second set of one or more media capture settings, which is different from the first set of one or more media capture settings of the computer system, is active and satisfies one or more capture conditions.
41. The method of claim 40, wherein the first media capture setting of the computer corresponds to the capture of media with a first set of one or more colors, and wherein the second media capture setting of the computer system corresponds to the capture of media with a second set of one or more colors that is different from the first set of one or more colors.
42. The method of any one of claims 36 to 41, wherein moving the portion of the computer system including the media capture component in the first movement mode includes performing a first type of movement via the movement component and performing a second type of movement different from the first type of movement via the movement component.
43. The method of claim 42, wherein the first type of movement of the computer system including the portion of the media capture component is movement in a first lateral direction.
44. The method of any one of claims 42 to 43, wherein the first type of movement of the computer system including the portion of the media capture component comprises movement in a first vertical direction.
45. The method of any one of claims 42 to 44, wherein the first type of movement of the computer system including the portion of the media capture component comprises movement in a first longitudinal direction.
46. The method according to any one of claims 42 to 45, wherein one or more of the first type of movement of the computer system including a portion of the media capture component and the second type of movement of the computer system including a portion of the media capture component include rotational movement.
47. The method according to any one of claims 36 to 46, wherein: When the portion of the computer system including the media capture component moves in the first movement mode, the portion of the computer system including the media capture component moves at a first rate; and When the portion of the computer system including the media capture component moves in the second movement mode, the portion of the computer system including the media capture component moves at a second rate different from the first rate.
48. The method according to any one of claims 36 to 47, wherein: When the portion of the computer system including the media capture component moves in the first movement mode, the framing of the video changes based on a first set of one or more tracking parameters; and When the portion of the computer system including the media capture component moves in the second movement mode, the framing of the video changes based on a second set of one or more tracking parameters that differs from the first set of one or more tracking parameters.
49. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a motion component, the one or more programs including instructions for performing the method according to any one of claims 36 to 48.
50. A computer system configured to communicate with a media capture component and a motion component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 36 to 48.
51. A computer system configured to communicate with a media capture component and a motion component, the computer system comprising: Components for performing the method according to any one of claims 36 to 48.
52. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a motion component, the one or more programs comprising instructions for performing the method according to any one of claims 36 to 48.
53. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component and a motion component, the one or more programs including instructions for performing the following operations: When video is captured via the media capture component: Based on a first set of conditions satisfying one or more capture criteria, a portion of the computer system, including the media capture component, is moved in a first movement mode via the moving component, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves; and Based on a second set of conditions that satisfy one or more capture conditions, wherein the second set of capture conditions is different from the first set of capture conditions, the portion of the computer system is moved via the moving component in a second moving mode different from the first moving mode, wherein moving the portion of the computer system in the second moving mode causes the framing of the video to change as the portion of the computer system moves.
54. A computer system configured to communicate with a media capture component and a motion component, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: When video is captured via the media capture component: Based on a first set of conditions that satisfy one or more capture conditions, a portion of the computer system, including the media capture component, is moved in a first movement mode via the moving component, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves. as well as Based on a second set of conditions that satisfy one or more capture conditions, wherein the second set of capture conditions is different from the first set of capture conditions, the portion of the computer system is moved via the moving component in a second moving mode different from the first moving mode, wherein moving the portion of the computer system in the second moving mode causes the framing of the video to change as the portion of the computer system moves.
55. A computer system configured to communicate with a media capture component and a motion component, the computer system comprising: For use with the following components when capturing video via the media capture component: Based on a first set of conditions that satisfy one or more capture conditions, a portion of the computer system, including the media capture component, is moved in a first movement mode via the moving component, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves. as well as Based on a second set of conditions that satisfy one or more capture conditions, wherein the second set of capture conditions is different from the first set of capture conditions, the portion of the computer system is moved via the moving component in a second moving mode different from the first moving mode, wherein moving the portion of the computer system in the second moving mode causes the framing of the video to change as the portion of the computer system moves.
56. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a media capture component and a motion component, the one or more programs comprising instructions for performing the following operations: When video is captured via the media capture component: Based on a first set of conditions satisfying one or more capture criteria, a portion of the computer system, including the media capture component, is moved in a first movement mode via the moving component, wherein moving the portion of the computer system in the first movement mode causes the framing of the video to change as the portion of the computer system moves; and Based on a second set of conditions that satisfy one or more capture conditions, wherein the second set of capture conditions is different from the first set of capture conditions, the portion of the computer system is moved via the moving component in a second moving mode different from the first moving mode, wherein moving the portion of the computer system in the second moving mode causes the framing of the video to change as the portion of the computer system moves.
57. A method, the method comprising: At the computer system that communicates with the media capture component, input component, and output component: The input component detects a first set of one or more inputs corresponding to one or more instructions, including one or more spoken words; In response to detecting the first set of one or more inputs corresponding to the one or more instructions, preparation is made to capture media via the media capture component; as well as When preparing to capture media via the media capture component: Based on determining that the one or more instructions include first content, a first composition guide is provided via the output component, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; as well as Based on the determination that the one or more instructions include second content different from the first content, a second composition guide different from the first composition guide is provided via the output component, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
58. The method according to claim 57, further comprising: When preparing to capture media via the media capture component: Based on determining that the corresponding subject is located at a first position within the field of view of the media capture component, a third composition guide is provided via the output component; as well as Based on determining that the corresponding subject is located at a second position different from the first position within the field of view of the media capture component, a fourth composition guide different from the third composition guide is provided via the output component.
59. The method according to any one of claims 57 to 58, further comprising: When preparing to capture media via the media capture component: Based on the first set of detected elements, a fifth graphing guide is provided via the output component; as well as Based on a second set of elements that are determined to be different from the first set of elements, a sixth composition guide, different from the fifth composition guide, is provided via the output component.
60. The method according to any one of claims 57 to 59, wherein: Based on the determination that the first set of one or more inputs has a first level of detail, the first mapping guidance has a first amount of information; and Based on the determination that the first set of one or more inputs has a second level of detail greater than the first level of detail, the first composition guide has a second amount of information greater than the first amount of information.
61. The method according to any one of claims 57 to 60, the method further comprising: When preparing to capture media via the media capture component: Based on the determination that the first set of one or more inputs corresponds to the first body, an eighth mapping guide is provided via the output component; and Based on the determination that the first set of one or more inputs corresponds to a second individual different from the first individual, a ninth compositional guide, different from the eighth compositional guide, is provided via the output component.
62. The method of any one of claims 57 to 61, wherein the first mapping guidance includes one or more recommendations that should change one or more characteristics of the lighting in the environment within the field of view of the media capture component.
63. The method of claim 62, wherein the first mapping guidance includes one or more recommendations that should change the amount of light in the environment in the field of view of the media capture component.
64. The method of any one of claims 62 to 63, wherein the first mapping guidance includes one or more recommendations that should change the type of light in the environment in the field of view of the media capture component.
65. The method of any one of claims 62 to 64, wherein the first mapping guidance includes one or more recommendations that should change one or more colors of light in the environment within the field of view of the media capture component.
66. The method according to any one of claims 62 to 65, wherein the computer system communicates with a collection of one or more external lights, the method further comprising: When preparing to capture media via the media capture component and in the absence of detected input corresponding to the subject, an instruction is transmitted to the set of one or more external lights, wherein transmitting the instruction to the set of one or more external lights causes one or more characteristics of the lighting in the environment within the field of view of the media capture component to change.
67. The method according to any one of claims 57 to 66, wherein the first mapping guidance includes one or more recommendations that one or more objects should be moved.
68. The method according to claims 57 to 67, wherein the first composition guidance includes one or more recommendations that the subject should move from a first position to a second position.
69. The method according to any one of claims 57 to 68, the method further comprising: When preparing to capture media via the media capture component: A tenth composition guide is provided via the output component based on a first set of positioning criteria, which determines that the positioning of a portion of the subject within the field of view of the media capture component satisfies one or more positioning criteria relative to the field of view of the media capture component; as well as Based on the determination that the positioning of the portion of the corresponding subject within the field of view of the media capture component satisfies a second set of one or more positioning criteria that differ from a first set of one or more positioning criteria relative to the field of view of the media capture component, an eleventh composition guide, different from the tenth composition guide, is provided via the output component.
70. The method of claim 69, wherein the tenth compositional guidance is provided with reference to the positioning of the portion of the corresponding subject relative to a fixed reference point.
71. The method according to any one of claims 69 to 70, wherein the tenth compositional guide is provided relative to the spatial relationship between the portion of the respective subject and the body of the respective subject.
72. The method according to any one of claims 57 to 71, further comprising: When preparing to capture media via the media capture component: Based on determining that the horizontal plane in the field of view of the media capture component has a first spatial orientation, a twelfth composition guide is provided via the output component; as well as Based on the determination that the horizontal plane in the field of view of the media capture component has a second spatial orientation different from the first spatial orientation, a thirteenth composition guide different from the twelfth composition guide is provided via the output component.
73. The method according to any one of claims 57 to 72, further comprising: When preparing to capture media via the media capture component: Based on determining a first distance between the boundary of the field of view of the media capture component and a portion of the corresponding subject, a fourteenth composition guide is provided via the output component; as well as Based on determining a second distance between the boundary of the field of view of the media capture component and the corresponding portion of the subject, a fifteenth composition guide, different from the fourteenth composition guide, is provided via the output component.
74. The method according to any one of claims 57 to 73, wherein the computer system communicates with the mobile component, the method further comprising: When preparing to capture media via the media capture component, a portion of the computer system is moved via the moving component.
75. The method according to claim 74, further comprising: The output component outputs an instruction on why the part of the computer system is moved.
76. The method according to any one of claims 57 to 75, the method further comprising: After preparing to capture media via the media capture component and based on a set of one or more conditions determined to be met, the media is captured via the media capture component without detecting intermediate input.
77. The method according to any one of claims 57 to 76, the method further comprising: After preparing to capture media via the media capture component, detect the corresponding input from the user; as well as In response to detecting the corresponding input from the user, a second media item is captured via the media capture component.
78. The method according to any one of claims 57 to 77, wherein the computer system communicates with the display component, the method further comprising: When preparing to capture media via the media capture component, a representation of the field of view of the media capture component is displayed via the display component.
79. A non-transitory computer-readable medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component, an input component, and an output component, the one or more programs including instructions for performing the method according to any one of claims 57 to 78.
80. A computer system configured to communicate with a media capture component, an input component, and an output component, the computer system comprising: One or more processors; and A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the method according to any one of claims 57 to 78.
81. A computer system configured to communicate with a media capture component, an input component, and an output component, the computer system comprising: Components for performing the method according to any one of claims 57 to 78.
82. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component, an input component, and an output component, the one or more programs comprising instructions for performing the method according to any one of claims 57 to 78.
83. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with a media capture component, an input component, and an output component, the one or more programs including instructions for performing the following operations: The input component detects a first set of one or more inputs corresponding to one or more instructions, including one or more spoken words; In response to detecting the first set of one or more inputs corresponding to the one or more instructions, preparation is made to capture media via the media capture component; as well as When preparing to capture media via the media capture component: Based on determining that the one or more instructions include first content, a first composition guide is provided via the output component, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; as well as Based on the determination that the one or more instructions include second content different from the first content, a second composition guide different from the first composition guide is provided via the output component, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
84. A computer system configured to communicate with a media capture component, an input component, and an output component, the computer system comprising: One or more processors; and The memory stores one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing the following operations: The input component detects a first set of one or more inputs corresponding to one or more instructions, including one or more spoken words; In response to detecting the first set of one or more inputs corresponding to the one or more instructions, preparation is made to capture media via the media capture component; as well as When preparing to capture media via the media capture component: Based on determining that the one or more instructions include first content, a first composition guide is provided via the output component, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; as well as Based on the determination that the one or more instructions include second content different from the first content, a second composition guide different from the first composition guide is provided via the output component, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
85. A computer system configured to communicate with a media capture component, an input component, and an output component, the computer system comprising: A component for detecting, via the input component, a first set of one or more inputs corresponding to one or more instructions including one or more spoken words; The component is used for: preparing to capture media via the media capture component in response to detecting a first set of one or more inputs corresponding to the one or more instructions; and For use with the following components when preparing to capture media via the media capture component: Based on determining that the one or more instructions include first content, a first composition guide is provided via the output component, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; as well as Based on the determination that the one or more instructions include second content different from the first content, a second composition guide different from the first composition guide is provided via the output component, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.
86. A computer program product comprising one or more programs configured to execute by one or more processors of a computer system in communication with a media capture component, an input component, and an output component, the one or more programs comprising instructions for performing the following operations: The input component detects a first set of one or more inputs corresponding to one or more instructions, including one or more spoken words; In response to detecting the first set of one or more inputs corresponding to the one or more instructions, preparation is made to capture media via the media capture component; as well as When preparing to capture media via the media capture component: Based on determining that the one or more instructions include first content, a first composition guide is provided via the output component, wherein the first composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component; as well as Based on the determination that the one or more instructions include second content different from the first content, a second composition guide different from the first composition guide is provided via the output component, wherein the second composition guide includes one or more recommendations for changing the spatial arrangement of one or more objects in the field of view of the media capture component.