Method and user interface for auditory features
By achieving faster and more effective methods and interfaces for providing auditory features in electronic devices, the problems of complex operation and inefficiency in the prior art are solved, and the user experience and equipment efficiency are improved.
Patent Information
- Application Number
- CN202510086263.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-18
- Filing Date
- 2022-05-19
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, when using electronic devices to provide auditory features, there are usually problems of complex operation and inefficiency, resulting in wasting user time and equipment energy.
Improve the efficiency of providing auditory features by implementing faster and more efficient methods and interfaces in a computer system, including parallel playback of audio media items according to a parallel audio standard set when playing audio media items, or performing predetermined actions through voice input in a user interface.
Reduces the cognitive burden of users and improves the efficiency of the human-machine interface, especially in battery-powered devices, extends battery life and saves device energy.
Smart Images

Figure CN119987712A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 202280043272.2, application date May 19, 2022, and name “Method and User Interface for Auditory Features”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to U.S. Patent Application No. 17 / 504,174, filed on October 18, 2021, entitled “METHODS AND USER INTERFACES FOR AUDITORY FEATURES,” U.S. Provisional Application No. 63 / 197,452, filed on June 6, 2021, entitled “METHODS AND USER INTERFACES FOR AUDITORY FEATURES,” and U.S. Provisional Application Serial No. 63 / 190,765, filed on May 19, 2021, entitled “METHODS AND USER INTERFACES FOR AUDITORY FEATURES.” All of these applications are incorporated herein by reference in their entirety. Technical Field
[0004] The present disclosure relates generally to computer user interfaces and, more particularly, to techniques for providing auditory features. Background Art
[0005] Personal electronic devices allow users to implement various functions of electronic devices. In some cases, such functions provide auditory features to users and / or allow users to interact with electronic devices using auditory input. Summary of the invention
[0006] However, some techniques for providing auditory features using electronic devices are often cumbersome and inefficient. For example, some prior art techniques use complex and time-consuming user interfaces that may include multiple keystrokes or keystrokes. As another example, the operability provided by some prior art techniques using auditory input may be limited. Therefore, prior art techniques require more time than necessary, which results in wasted user time and device energy. This latter consideration is particularly important in battery-powered devices.
[0007] Therefore, the present technology provides a faster and more efficient method and interface for providing auditory features for electronic devices. Such methods and interfaces optionally supplement or replace other methods for providing auditory features. Such methods and interfaces reduce the cognitive burden on users and produce more efficient human-computer interfaces. For battery-powered computing devices, such methods and interfaces save power and increase the time interval between battery charges, for example, by reducing the number and / or time of inputs required to operate such devices.
[0008] An exemplary method is described herein. An exemplary method includes: at a computer system in communication with one or more input devices: while playing an audio media item of a first type, receiving a request to play an audio media item of a second type via the one or more input devices; upon determining that a parallel audio standard set is satisfied, playing in parallel: the audio media item of the first type; and the audio media item of the second type; and upon determining that the parallel audio standard set is not satisfied: stopping playing the audio media item of the first type; and playing the audio media item of the second type.
[0009] An exemplary method includes, at a computer system that communicates with a display generating component and one or more input devices: while displaying a user interface including a group of user interface objects via the display generating component, receiving a first voice input associated with a first predetermined action via the one or more input devices; and in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing a first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0010] An exemplary method includes, at a computer system that communicates with a display generating component and one or more input devices: executing a sound registration process, the sound registration process comprising: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input satisfies a sound input standard; based on determining that the group of one or more sound inputs satisfies the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not satisfy the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0011] An exemplary non-transitory computer-readable storage medium configured to be executed by one or more processors of a computer system is described herein. An exemplary non-transitory computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with one or more input devices and includes instructions for the following operations: while playing an audio media item of a first type, receiving a request to play an audio media item of a second type via one or more input devices; based on determining that a parallel audio standard set is met, playing in parallel: an audio media item of the first type; and an audio media item of the second type; and based on determining that the parallel audio standard set is not met: stopping playing the audio media item of the first type; and playing the audio media item of the second type.
[0012] An exemplary non-transitory computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with a display generation component and one or more input devices and includes instructions for the following operations: when a user interface including a group of user interface objects is displayed via the display generation component, receiving a first voice input associated with a first predetermined action via one or more input devices; and in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing a first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0013] An exemplary non-transitory computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with a display generating component and one or more input devices and includes instructions for the following operations: executing a sound registration process, the sound registration process including: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input satisfies a sound input standard; based on determining that the group of one or more sound inputs satisfies the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not satisfy the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0014] An exemplary transient computer-readable storage medium configured to be executed by one or more processors of a computer system is described herein. An exemplary non-transitory computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with one or more input devices and includes instructions for the following operations: while playing an audio media item of a first type, receiving a request to play an audio media item of a second type via one or more input devices; based on determining that a parallel audio standard set is met, playing in parallel: an audio media item of the first type; and an audio media item of the second type; and based on determining that the parallel audio standard set is not met: stopping playing the audio media item of the first type; and playing the audio media item of the second type.
[0015] An exemplary transient computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with a display generation component and one or more input devices and includes instructions for the following operations: when a user interface including a group of user interface objects is displayed via the display generation component, receiving a first voice input associated with a first predetermined action via one or more input devices; and in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing a first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0016] An exemplary transient computer-readable storage medium configured to be executed by one or more processors of a computer system communicates with a display generating component and one or more input devices and includes instructions for the following operations: performing a sound registration process, the sound registration process including: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input satisfies a sound input standard; based on determining that the group of one or more sound inputs satisfies the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not satisfy the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0017] An exemplary computer system is described herein. An exemplary computer system is configured to communicate with one or more input devices and includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: while playing a first type of audio media item, receiving a request to play a second type of audio media item via one or more input devices; based on determining that a parallel audio standard set is met, playing in parallel: the first type of audio media item; and the second type of audio media item; and based on determining that the parallel audio standard set is not met: stopping playing the first type of audio media item; and playing the second type of audio media item.
[0018] An exemplary computer system is configured to communicate with a display generation component and one or more input devices and includes: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: when a user interface including a group of user interface objects is displayed via the display generation component, receiving a first voice input associated with a first predetermined action via one or more input devices; and in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing a first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0019] An exemplary computer system is configured to communicate with a display generating component and one or more input devices, comprising: one or more processors; and a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for the following operations: executing a sound registration process, the sound registration process comprising: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input meets a sound input standard; based on determining that the group of one or more sound inputs meets the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not meet the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0020] An exemplary computer system is configured to communicate with one or more input devices and includes: a device for receiving a request to play a second type of audio media item via one or more input devices while playing a first type of audio media item; a device for playing the following items in parallel based on determining that a parallel audio standard set is met: the first type of audio media item; and the second type of audio media item; and a device for performing the following operations based on determining that the parallel audio standard set is not met: stopping playing the first type of audio media item; and playing the second type of audio media item.
[0021] An exemplary computer system is configured to communicate with a display generation component and one or more input devices and includes: a device for receiving a first voice input associated with a first predetermined action via one or more input devices when a user interface including a group of user interface objects is displayed via the display generation component; and a device for performing the following operations in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0022] An exemplary computer system is configured to communicate with a display generating component and one or more input devices and includes: a device for performing a sound registration process, the sound registration process including: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input satisfies a sound input standard; based on determining that the group of one or more sound inputs satisfies the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not satisfy the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0023] An exemplary computer program product is disclosed herein. An exemplary computer program product includes one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices, the one or more programs including instructions for the following operations: while playing an audio media item of a first type, receiving a request to play an audio media item of a second type via the one or more input devices; upon determining that a parallel audio standard set is met, playing in parallel: the audio media item of the first type; and the audio media item of the second type; and upon determining that the parallel audio standard set is not met: stopping playing the audio media item of the first type; and playing the audio media item of the second type.
[0024] An exemplary computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs including instructions for the following operations: when a user interface including a group of user interface objects is displayed via the display generating component, receiving a first voice input associated with a first predetermined action via one or more input devices; and in response to receiving the first voice input: based on determining that a first user interface object in the group of user interface objects is currently selected, performing a first predetermined action based on the first user interface object; and based on determining that a second user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the second user interface object.
[0025] An exemplary computer program product includes one or more programs configured to be executed by one or more processors of a computer system that communicates with a display generating component and one or more input devices, the one or more programs including instructions for the following operations: performing a sound registration process, the sound registration process including: receiving a group of one or more sound inputs including a first sound input via one or more input devices; indicating whether the first sound input meets a sound input standard; based on determining that the group of one or more sound inputs meets the sound registration standard set, generating a model for identifying a first type of sound corresponding to the group of one or more sound inputs; and based on determining that the group of one or more sound inputs does not meet the sound registration standard set, abandoning the generation of a model for identifying a first type of sound corresponding to the group of one or more sound inputs.
[0026] Executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are optionally included in a transient computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0027] Thus, a faster, more efficient method and interface for providing auditory features is provided for devices, thereby increasing the effectiveness, efficiency, and user satisfaction of such devices.Such methods and interfaces may supplement or replace other methods for providing auditory features. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] For a better understanding of the various described embodiments, reference should be made to the following detailed description taken in conjunction with the following drawings, wherein like reference numerals designate corresponding parts throughout the several views.
[0029] Figure 1A is a block diagram illustrating a portable multifunction device with a touch-sensitive display according to some embodiments.
[0030] Figure 1B is a block diagram illustrating example components for event processing according to some embodiments.
[0031] Figure 2 A portable multifunction device with a touch screen according to some embodiments is shown.
[0032] Figure 3 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments.
[0033] Figure 4A An exemplary user interface for a menu of applications on a portable multifunction device is shown according to some embodiments.
[0034] Figure 4B An exemplary user interface is shown for a multifunction device having a touch-sensitive surface that is separate from the display according to some embodiments.
[0035] Figure 5A A personal electronic device according to some embodiments is shown.
[0036] Figure 5B is a block diagram illustrating a personal electronic device according to some embodiments.
[0037] FIG. 6A to FIG. 6M An exemplary user interface for providing background sound according to some embodiments is shown.
[0038] Figure 7 is a flowchart of a process for providing background sound according to some embodiments.
[0039] Figures 8A to 8V An exemplary user interface for providing auditory control according to some embodiments is shown.
[0040] Fig. 9 is a flowchart of a method for providing auditory control according to some embodiments.
[0041] Figure 10A to Figure 10V An exemplary user interface for providing notifications according to some embodiments is shown.
[0042] Fig.11 is a flowchart of a process for providing notifications according to some embodiments. DETAILED DESCRIPTION
[0043] The following description sets forth exemplary methods, parameters, etc. However, it should be appreciated that such description is not intended to limit the scope of the present disclosure, but is provided as a description of exemplary embodiments.
[0044] There is a need for electronic devices that provide efficient systems, methods, and interfaces for auditory features. Such techniques can reduce the cognitive burden on users who utilize auditory features, thereby increasing productivity. In addition, such techniques can reduce processor power and battery power that would otherwise be wasted on redundant or unnecessary user input.
[0045] under Figure 1A to Figure 1B , Figure 2 , Figure 3 , FIG. 4A to FIG. 4B , FIG. 5A to FIG. 5B A description of an exemplary device for performing techniques for providing auditory features is provided. FIG. 6A to FIG. 6M An exemplary user interface for providing background sound is shown. Figure 7 is a flowchart illustrating a method for providing background sound according to some embodiments. FIG. 6A to FIG. 6M The user interface in is used to illustrate the processes described below, which include Figure 7 process. Figures 8A to 8V An exemplary user interface for providing auditory control is shown. Fig. 9 is a flowchart illustrating a method for providing auditory control according to some embodiments. Figures 8A to 8V The user interface in is used to illustrate the processes described below, which include Fig. 9 process. Figure 10A to Figure 10V An exemplary user interface for providing auditory control is shown. Fig.11 is a flowchart illustrating a method for providing notifications according to some embodiments. Figure 10A to Figure 10V The user interface in is used to illustrate the processes described below, which include Fig.11 process.
[0046] The processes described below enhance the operability of the device and make the user-device interface more efficient (e.g., by helping the user provide appropriate input and reducing user errors when operating / interacting with the device) through various techniques, including by providing improved visual feedback to the user, reducing the number of inputs required to perform an operation, providing additional control options without cluttering the user interface with additional display controls, performing an operation without further user input when a set of conditions have been met, and / or additional techniques. These techniques also reduce power usage and extend the battery life of the device by enabling the user to use the device more quickly and more efficiently.
[0047] In addition, in the method described herein where one or more steps depend on one or more conditions being met, it should be understood that the method can be repeated in multiple repetitions so that in the process of repetition, all conditions for determining the steps in the method have been met in different repetitions of the method. For example, if the method needs to perform the first step (if the condition is met), and perform the second step (if the condition is not met), then the ordinary technician will know that the steps stated are repeated until both the condition is met and the condition is not met (in no particular order). Therefore, the method described as having one or more steps depending on one or more conditions being met can be rewritten as a method of repeating until each condition described in the method is met. However, this does not require the system or computer-readable medium to declare that the system or computer-readable medium contains instructions for performing a contingent operation based on the satisfaction of the corresponding one or more conditions, and is therefore able to determine whether the possible situation has been met without explicitly repeating the steps of the method until all conditions for determining the steps in the method have been met. It will also be understood by ordinary technicians in the art that, similar to the method with the contingent step, the system or computer-readable storage medium can repeat the steps of the method as many times as needed to ensure that all the contingent steps have been performed.
[0048] Although the following description uses the terms "first", "second", etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another element. For example, a first touch can be named a second touch and similarly a second touch can be named a first touch without departing from the scope of the various described embodiments. Both the first touch and the second touch are touches, but they are not the same touch.
[0049] The terms used in the description of various described embodiments in this article are only for the purpose of describing specific embodiments, and are not intended to be limited. As used in the description of various described embodiments and in the appended claims, the singular forms "one" and "the" are intended to also include plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more items in the associated listed items. It will also be understood that the terms "including" and / or "comprising" are used in this specification to specify the presence of stated features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or their grouping.
[0050] The term "if" is optionally interpreted to mean "when," "upon," or "in response to determining," or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined that ..." or "if [a stated condition or event] is detected" are optionally interpreted to mean "upon determining that ..." or "in response to determining that ..." or "upon detecting [a stated condition or event]," or "in response to detecting [a stated condition or event]," depending on the context.
[0051] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described herein. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as a PDA and / or a music player function. Exemplary embodiments of portable multifunction devices include, but are not limited to, the Apple Watch from Apple Inc. (Cupertino, California). Devices, iPod Equipment, and Device. Optionally use other portable electronic devices, such as laptops or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads). It should also be understood that in some embodiments, the device is not a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or touch pads). In some embodiments, the electronic device is a computer system that communicates with the display generation component (e.g., via wireless communication, via wired communication). The display generation component is configured to provide visual output, such as display via a CRT display, display via an LED display, or display via image projection. In some embodiments, the display generation component is integrated with the computer system. In some embodiments, the display generation component is separated from the computer system. As used herein, "display" content includes transmitting data (e.g., image data or video data) to an integrated or external display generation component via a wired or wireless connection to visually generate content to display content (e.g., video data rendered or decoded by display controller 156).
[0052] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, mouse and / or joystick.
[0053] The device typically supports a variety of applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk editing application, a spreadsheet application, a gaming application, a telephony application, a video conferencing application, an email application, an instant messaging application, a fitness support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.
[0054] Various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed on the device are optionally adjusted and / or varied for different applications, and / or adjusted and / or varied within the respective applications. In this way, the common physical architecture of the device (such as a touch-sensitive surface) optionally supports a variety of applications with a user interface that is intuitive and clear to the user.
[0055] Attention is now turned to embodiments of portable devices having touch-sensitive displays. Figure 1A is a block diagram showing a portable multifunction device 100 with a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 is sometimes referred to as a "touch screen" for convenience, and is sometimes referred to as or referred to as a "touch-sensitive display system." The device 100 includes a memory 102 (which optionally includes one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. The device 100 optionally includes one or more optical sensors 164. The device 100 optionally includes one or more contact force sensors 165 for detecting the intensity of contact on the device 100 (e.g., a touch-sensitive surface, such as the touch-sensitive display system 112 of the device 100). Device 100 optionally includes one or more tactile output generators 167 for generating tactile output on device 100 (e.g., generating tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touch pad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.
[0056] As used in this specification and claims, the term "intensity" of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or to a surrogate (surrogate) of the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values that includes at least four different values and more typically includes hundreds of different values (e.g., at least 256). The intensity of a contact is optionally determined (or measured) using various methods and various sensors or combinations of sensors. For example, one or more force sensors below or adjacent to the touch-sensitive surface are optionally used to measure the force at different points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted average) to determine an estimated contact force. Similarly, the pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size of the contact area detected on the touch-sensitive surface and / or its change, the capacitance of the touch-sensitive surface near the contact and / or its change, and / or the resistance of the touch-sensitive surface near the contact and / or its change are optionally used as a substitute for the force or pressure of the contact on the touch-sensitive surface. In some embodiments, the substitute measurement of the contact force or pressure is used directly to determine whether the intensity threshold has been exceeded (e.g., the intensity threshold is described in units corresponding to the substitute measurement). In some embodiments, the substitute measurement of the contact force or pressure is converted into an estimated force or pressure, and the estimated force or pressure is used to determine whether the intensity threshold has been exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of the contact as an attribute of the user input allows the user to access additional device functions that would otherwise be inaccessible to the user on a smaller device with limited real estate, which is used to display an enable indication (e.g., on a touch-sensitive display) and / or receive user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls, such as knobs or buttons).
[0057] As used in this specification and claims, the term "tactile output" refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, which will be detected by a user using the user's sense of touch. For example, in the case where a device or a component of the device is in contact with a user's touch-sensitive surface (e.g., a finger, palm, or other part of the user's hand), the tactile output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical characteristic of the device or component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a "press click" or "release click" to a physical actuation button. In some cases, the user will feel a tactile sensation, such as a "press click" or "release click", even when the physical actuation button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's movement does not move. As another example, even when there is no change in the smoothness of the touch-sensitive surface, movement of the touch-sensitive surface may optionally be interpreted or sensed by the user as "roughness" of the touch-sensitive surface. Although such a user's interpretation of touch will be limited by the user's individualized sensory perceptions, many sensory perceptions of touch are common to most users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "press click," "release click," "roughness"), unless otherwise stated, the generated tactile output corresponds to a physical displacement of the device or a component thereof that would generate that sensory perception for a typical (or average) user.
[0058] It should be understood that device 100 is merely one example of a portable multifunction device, and that device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of the components. Figure 1A The various components shown in the EMBODIMENTS 2000 are implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0059] Memory 102 optionally includes high-speed random access memory, and optionally also includes non-volatile memory, such as one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls other components of device 100 to access memory 102.
[0060] The peripheral device interface 118 can be used to couple the input peripheral devices and output peripheral devices of the device to the CPU 120 and the memory 102. The one or more processors 120 run or execute various software programs (such as computer programs (e.g., including instructions)) and / or instruction sets stored in the memory 102 to perform various functions of the device 100 and process data. In some embodiments, the peripheral device interface 118, the CPU 120, and the memory controller 122 are optionally implemented on a single chip such as the chip 104. In some other embodiments, they are optionally implemented on separate chips.
[0061] RF (radio frequency) circuit 108 receives and sends RF signals, also referred to as electromagnetic signals. RF circuit 108 converts electrical signals into / converts electromagnetic signals into electrical signals, and communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 108 optionally includes well-known circuits for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, user identity modules (SIM) cards, memories, and the like. RF circuit 108 optionally communicates with networks and other devices via wireless communications, such as the Internet (also referred to as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). RF circuit 108 optionally includes well-known circuits for detecting near field communications (NFC) fields, such as through short-range communications radio components. Wireless communication optionally uses any of a variety of communication standards, protocols and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High Speed Downlink Packet Access (HSDPA), High Speed Uplink Packet Access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual Cell HSPA (DC-HSPDA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11d), IEEE 802.11e, IEEE 802.11f, IEEE 802.11g, IEEE 802.11g). 802.11n and / or IEEE 802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, email protocols (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Utilizing Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols that have not been developed as of the filing date of this document.
[0062] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between a user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts the audio data into electrical signals, and transmits the electrical signals to the speaker 111. The speaker 111 converts the electrical signals into sound waves audible to humans. The audio circuit 110 also receives electrical signals converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signals into audio data and transmits the audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., Figure 2 The headset jack provides an interface between the audio circuit 110 and a removable audio input / output peripheral device, such as an output-only headset or a headset having both output (e.g., a single or dual-ear headset) and input (e.g., a microphone).
[0063] The I / O subsystem 106 couples input / output peripherals on the device 100, such as a touch screen 112 and other input control devices 116, to a peripheral device interface 118. The I / O subsystem 106 optionally includes a display controller 156, an optical sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a tactile feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive / send electrical signals from other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some embodiments, the input controller 160 is optionally coupled to any of the following (or none of the following): a keyboard, an infrared port, a USB port, and a pointing device such as a mouse. One or more buttons (e.g., Figure 2 208) optionally includes an increase / decrease button for volume control of the speaker 111 and / or the microphone 113. The one or more buttons optionally include a push button (e.g., Figure 2206 in). In some embodiments, the electronic device is a computer system that communicates with one or more input devices (e.g., via wireless communication, via wired communication). In some embodiments, the one or more input devices include a touch-sensitive surface (e.g., a touchpad as part of a touch-sensitive display). In some embodiments, the one or more input devices include one or more camera sensors (e.g., one or more optical sensors 164 and / or one or more depth camera sensors 175), such as for tracking a user's gestures (e.g., hand gestures) as input. In some embodiments, the one or more input devices are integrated with the computer system. In some embodiments, the one or more input devices are separate from the computer system.
[0064] A quick press of the push button optionally disengages the lock of the touch screen 112 or optionally begins the process of unlocking the device using gestures on the touch screen, as described in U.S. Patent Application 11 / 322,549, filed on December 23, 2005, entitled "Unlocking a Device by Performing Gestures on an Unlock Image" (i.e., U.S. Patent No. 7,657,849), which is hereby incorporated by reference in its entirety. A long press of the push button (e.g., 206) optionally turns the device 100 on or off. The functions of one or more buttons are optionally user-customizable. The touch screen 112 is used to implement virtual buttons or soft buttons and one or more soft keyboards.
[0065] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch screen 112 and / or sends electrical signals to the touch screen 112. The touch screen 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, videos, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.
[0066] The touch screen 112 has a touch-sensitive surface, sensor, or sensor group that accepts input from a user based on tactile and / or haptic contact. The touch screen 112 and display controller 156 (together with any associated modules and / or instruction sets in memory 102) detect contact on the touch screen 112 (and any movement or interruption of the contact) and convert the detected contact into interaction with a user interface object (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to a finger of the user.
[0067] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although other display technologies are used in other embodiments. The touch screen 112 and display controller 156 optionally use any of a variety of touch sensing technologies now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112 to detect contact and any movement or interruption thereof. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as in the Apple ® from Apple Inc. (Cupertino, California). and iPod The technology used in.
[0068] The touch-sensitive display in some embodiments of touch screen 112 is optionally similar to the multi-touch-sensitive touchpad described in the following U.S. Patents: 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman et al.), and / or U.S. Patent Publication 2002 / 0015024A1, each of which is hereby incorporated by reference in its entirety. However, touch screen 112 displays visual output from device 100, while a touch-sensitive touchpad does not provide visual output.
[0069] The touch-sensitive display in some embodiments of the touch screen 112 is described in the following applications: (1) U.S. patent application No. 11 / 381,313, filed on May 2, 2006, "Multipoint Touch Surface Controller"; (2) U.S. patent application No. 10 / 840,862, filed on May 6, 2004, "Multipoint Touchscreen"; (3) U.S. patent application No. 10 / 903,964, filed on July 30, 2004, "Gestures For Touch Sensitive Input Devices"; (4) U.S. patent application No. 11 / 048,264, filed on January 31, 2005, "Gestures For Touch Sensitive Input Devices"; (5) U.S. patent application No. 11 / 038,590, filed on January 18, 2005, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices”; (6) U.S. patent application No. 11 / 228,758, filed on September 16, 2005, “Virtual Input Device Placement On A Touch Screen User Interface”; (7) U.S. patent application No. 11 / 228,700, filed on September 16, 2005, “Operation Of A Computer With A Touch Screen Interface”; (8) U.S. patent application No. 11 / 228,737, filed on September 16, 2005, “Activating Virtual Keys Of A Touch-Screen Virtual Keyboard”; and (9) U.S. patent application No. 11 / 367,749, filed on March 3, 2006, “Multi-Functional Hand-Held Device”. All of these applications are incorporated herein by reference in their entirety.
[0070] The touch screen 112 optionally has a video resolution of more than 100 dpi. In some embodiments, the touch screen has a video resolution of about 160 dpi. The user optionally uses any suitable object or attachment such as a stylus, a finger, etc. to contact the touch screen 112. In some embodiments, the user interface is designed to work primarily through finger-based contacts and gestures, which may not be as accurate as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts rough finger-based input into precise pointer / cursor positions or commands for performing the actions desired by the user.
[0071] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touchpad for activating or deactivating specific functions. In some embodiments, the touchpad is a touch-sensitive area of the device that, unlike the touch screen, does not display visual output. The touchpad is optionally a touch-sensitive surface that is separate from the touch screen 112 or is an extension of the touch-sensitive surface formed by the touch screen.
[0072] The device 100 also includes a power system 162 for powering the various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator (e.g., a light emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in a portable device.
[0073] Device 100 optionally also includes one or more optical sensors 164 . Figure 1AAn optical sensor coupled to an optical sensor controller 158 in the I / O subsystem 106 is shown. The optical sensor 164 optionally includes a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light projected through one or more lenses from the environment and converts the light into data representing an image. In conjunction with the imaging module 143 (also called a camera module), the optical sensor 164 optionally captures a static image or video. In some embodiments, the optical sensor is located on the rear of the device 100, opposite to the touch screen display 112 on the front of the device, so that the touch screen display can be used as a viewfinder for static images and / or video image acquisition. In some embodiments, the optical sensor is located on the front of the device so that the user's image is optionally obtained for video conferencing while the user views other video conference participants on the touch screen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) so that a single optical sensor 164 is used with the touch screen display for both video conferencing and static image and / or video image acquisition.
[0074] Device 100 optionally also includes one or more depth camera sensors 175 . Figure 1A A depth camera sensor coupled to a depth camera controller 169 in the I / O subsystem 106 is shown. The depth camera sensor 175 receives data from the environment to create a three-dimensional model of an object (e.g., a face) within the scene from a viewpoint (e.g., a depth camera sensor). In some embodiments, in conjunction with the imaging module 143 (also referred to as a camera module), the depth camera sensor 175 is optionally used to determine a depth map of different portions of an image captured by the imaging module 143. In some embodiments, the depth camera sensor is located at the front of the device 100, so that an image of the user with depth information is optionally obtained for video conferencing while the user is viewing other video conference participants on the touch screen display, and a selfie with depth map data is captured. In some embodiments, the depth camera sensor 175 is located at the rear of the device, or at both the rear and front of the device 100. In some embodiments, the position of the depth camera sensor 175 can be changed by the user (e.g., by rotating the lens and sensor in the device housing) so that the depth camera sensor 175 is used with the touch screen display for both video conferencing and still image and / or video image acquisition.
[0075] Device 100 optionally also includes one or more contact intensity sensors 165 . Figure 1AA contact force sensor is shown coupled to a force sensor controller 159 in the I / O subsystem 106. Contact force sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electrical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other force sensors (e.g., sensors for measuring the force (or pressure) of contact on a touch-sensitive surface). Contact force sensor 165 receives contact force information (e.g., pressure information or a surrogate for pressure information) from the environment. In some embodiments, at least one contact force sensor is juxtaposed or adjacent to a touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact force sensor is located on the back of device 100, opposite to touch screen display 112 located on the front of device 100.
[0076] Device 100 optionally also includes one or more proximity sensors 166 . Figure 1A A proximity sensor 166 is shown coupled to the peripherals interface 118. Alternatively, the proximity sensor 166 is optionally coupled to the input controller 160 in the I / O subsystem 106. The proximity sensor 166 is optionally implemented as described in the following U.S. patent application numbers: 11 / 241,839, entitled "Proximity Detector In Handheld Device"; 11 / 240,788, entitled "Proximity Detector In Handheld Device"; 11 / 620,702, entitled "Using Ambient Light Sensor To Augment Proximity Sensor Output"; 11 / 586,862, entitled "Automated Response To And Sensing Of User Activity In Portable Devices"; and 11 / 638,251, entitled "Methods And Systems For Automatic Configuration Of Peripherals", which are hereby incorporated by reference in their entirety. In some embodiments, when the multifunction device is placed near the user's ear (e.g., when the user is on a phone call), the proximity sensor turns off and disables the touch screen 112.
[0077] Device 100 optionally also includes one or more tactile output generators 167 . Figure 1AA tactile output generator is shown coupled to a tactile feedback controller 161 in the I / O subsystem 106. The tactile output generator 167 optionally includes one or more electroacoustic devices such as a speaker or other audio component; and / or an electromechanical device for converting energy into linear motion such as a motor, solenoid, electroactive polymer, piezoelectric actuator, electrostatic actuator, or other tactile output generating component (e.g., a component for converting an electrical signal into a tactile output on the device). The contact force sensor 165 receives tactile feedback generation instructions from the tactile feedback module 133 and generates a tactile output on the device 100 that can be felt by a user of the device 100. In some embodiments, at least one tactile output generator is arranged in parallel or adjacent to a touch-sensitive surface (e.g., a touch-sensitive display system 112) and optionally generates a tactile output by moving the touch-sensitive surface vertically (e.g., inward / outward of the surface of the device 100) or laterally (e.g., backward and forward in the same plane as the surface of the device 100). In some embodiments, at least one tactile output generator sensor is located on the back of the device 100, opposite the touch screen display 112 located on the front of the device 100.
[0078] Device 100 optionally also includes one or more accelerometers 168 . Figure 1A An accelerometer 168 is shown coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to the input controller 160 in the I / O subsystem 106. The accelerometer 168 is optionally implemented as described in the following U.S. Patent Publication Nos.: 20050190059, entitled "Acceleration-based Theft Detection System for Portable Electronic Devices" and 20060017692, entitled "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", both of which are incorporated herein by reference in their entirety. In some embodiments, information is displayed in a portrait view or a landscape view on the touch screen display based on analysis of data received from one or more accelerometers. The device 100 optionally includes a magnetometer and a GPS (or GLONASS or other global navigation system) receiver in addition to the accelerometer 168 for obtaining information about the position and orientation (e.g., portrait or landscape) of the device 100.
[0079] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a contact / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a global positioning system (GPS) module (or instruction set) 135, and an application (or instruction set) 136. In addition, in some embodiments, memory 102 ( Figure 1A ) or 370( Figure 3 ) storage device / global internal state 157, such as Figure 1A and Figure 3 . The device / global internal state 157 includes one or more of the following: active application state, which indicates which applications (if any) are currently active; display state, which indicates what applications, views, or other information occupy various areas of the touch screen display 112; sensor state, including information obtained from the device's various sensors and input control devices 116; and position information related to the device's position and / or posture.
[0080] The operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, storage device control, power management, etc.), and facilitating communication between various hardware components and software components.
[0081] The communication module 128 facilitates communication with other devices through one or more external ports 124, and also includes various software components for processing data received by the RF circuit 108 and / or the external port 124. The external port 124 (e.g., Universal Serial Bus (USB), FireWire, etc.) is suitable for coupling directly to other devices, or indirectly through a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is connected to (trademark of Apple Inc.) devices.
[0082] The contact / motion module 130 optionally detects contact with the touch screen 112 (in conjunction with the display controller 156) and other touch-sensitive devices (e.g., a touchpad or physical click wheel). The contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact has occurred (e.g., detecting a finger press event), determining the contact strength (e.g., the force or pressure of the contact, or a substitute for the force or pressure of the contact), determining whether there is movement of the contact and tracking the movement on the touch-sensitive surface (e.g., detecting one or more finger drag events), and determining whether the contact has stopped (e.g., detecting a finger lift event or contact disconnection). The contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction) and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of the contact point being represented by a series of contact data. These operations are optionally applied to a single point of contact (e.g., a single finger contact) or multiple points of simultaneous contact (e.g., "multi-touch" / multiple finger contacts). In some embodiments, the contact / motion module 130 and display controller 156 detect contact on the touch pad.
[0083] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an operation has been performed by a user (e.g., to determine whether a user has "clicked" an icon). In some embodiments, at least a subset of the intensity thresholds are determined based on software parameters (e.g., the intensity thresholds are not determined by the activation thresholds of specific physical actuators and can be adjusted without changing the physical hardware of the device 100). For example, without changing the touchpad or touchscreen display hardware, the mouse "click" threshold of a touchpad or touchscreen can be set to any one of a large range of predefined thresholds. In addition, in some specific implementations, a software setting is provided to the user of the device for adjusting one or more intensity thresholds in a set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by utilizing a system-level click on an "intensity" parameter to adjust multiple intensity thresholds at once).
[0084] The contact / motion module 130 optionally detects gesture input by the user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different motions, timings, and / or intensities of the detected contacts). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a finger tap gesture includes detecting a finger press event, and then detecting a finger lift (lift-off) event at the same location (or substantially the same location) as the finger press event (e.g., at the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger press event, then detecting one or more finger drag events, and then detecting a finger lift (lift-off) event.
[0085] The graphics module 132 includes various known software components for rendering and displaying graphics on the touch screen 112 or other display, including components for changing the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual attributes) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including but not limited to text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, etc.
[0086] In some embodiments, the graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. The graphics module 132 receives one or more codes for specifying the graphics to be displayed from an application program, etc., and also receives coordinate data and other graphic attribute data together if necessary, and then generates screen image data to be output to the display controller 156.
[0087] Haptic feedback module 133 includes various software components for generating instructions used by tactile output generator 167 to produce tactile output at one or more locations on device 100 in response to user interaction with device 100 .
[0088] Text input module 134, which is optionally a component of graphics module 132, provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).
[0089] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for use in location-based dialing; to the camera 143 as picture / video metadata; and to applications that provide location-based services, such as a weather widget, a local yellow pages widget, and a map / navigation widget).
[0090] The application 136 optionally includes the following modules (or instruction sets) or a subset or superset thereof:
[0091] Contacts module 137 (sometimes called address book or contact list);
[0092] Telephone module 138;
[0093] Video conferencing module 139;
[0094] Email client module 140;
[0095] Instant messaging (IM) module 141;
[0096] Fitness support module 142;
[0097] A camera module 143 for still images and / or video images;
[0098] Image management module 144;
[0099] Video player module;
[0100] Music player module;
[0101] Browser module 147;
[0102] Calendar module 148;
[0103] A widget module 149, which optionally includes one or more of the following: a weather widget 149-1, a stock market widget 149-2, a calculator widget 149-3, an alarm widget 149-4, a dictionary widget 149-5, and other widgets acquired by the user, and a user-created widget 149-6;
[0104] A widget creator module 150 for forming a user-created widget 149 - 6 ;
[0105] Search module 151;
[0106] Video and music player module 152, which merges the video player module and the music player module;
[0107] Note module 153;
[0108] Map module 154; and / or
[0109] Online video module 155.
[0110] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, rendering applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice replication.
[0111] In conjunction with the touch screen 112, display controller 156, touch / motion module 130, graphics module 132, and text input module 134, the contact module 137 is optionally used to manage an address book or contact list (e.g., stored in the application internal state 192 of the contact module 137 in memory 102 or memory 370), including: adding one or more names to the address book; deleting names from the address book; associating phone numbers, email addresses, physical addresses, or other information with names; associating images with names; categorizing and classifying names; providing phone numbers or email addresses to initiate and / or facilitate communications via telephone 138, video conferencing module 139, email 140, or IM 141; and the like.
[0112] In conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, phone module 138 is optionally used to enter a sequence of characters corresponding to a phone number, access one or more phone numbers in contact module 137, modify an entered phone number, dial a corresponding phone number, conduct a conversation, and disconnect or hang up when the conversation is complete. As described above, wireless communication optionally uses any of a variety of communication standards, protocols, and technologies.
[0113] In combination with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the optical sensor 164, the optical sensor controller 158, the contact / motion module 130, the graphics module 132, the text input module 134, the contact module 137 and the telephone module 138, the video conferencing module 139 includes executable instructions for initiating, conducting and terminating a video conference between a user and one or more other participants in accordance with user instructions.
[0114] In conjunction with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, email client module 140 includes executable instructions for creating, sending, receiving, and managing emails in response to user instructions. In conjunction with image management module 144, email client module 140 makes it very easy to create and send emails with still images or video images captured by camera module 143.
[0115] In conjunction with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, instant messaging module 141 includes executable instructions for the following operations: inputting a character sequence corresponding to an instant message, modifying previously input characters, transmitting a corresponding instant message (e.g., using a short message service (SMS) or multimedia message service (MMS) protocol for phone-based instant messaging or using XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, the transmitted and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments supported in MMS and / or enhanced messaging services (EMS). As used herein, "instant messaging" refers to both phone-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).
[0116] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, the GPS module 135, the map module 154 and the music player module, the fitness support module 142 includes executable instructions for creating a fitness (e.g., with time, distance and / or calorie burn goals); communicating with fitness sensors (sports equipment); receiving fitness sensor data; calibrating sensors for monitoring fitness; selecting and playing music for fitness; and displaying, storing and transmitting fitness data.
[0117] In conjunction with the touch screen 112, display controller 156, optical sensor 164, optical sensor controller 158, contact / motion module 130, graphics module 132 and image management module 144, the camera module 143 includes executable instructions for the following operations: capturing still images or videos (including video streams) and storing them in the memory 102, modifying the characteristics of the still images or videos, or deleting the still images or videos from the memory 102.
[0118] In conjunction with touch screen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and camera module 143, image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still images and / or video images.
[0119] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132 and the text input module 134, the browser module 147 includes executable instructions for browsing the Internet in accordance with user instructions, including searching, linking to, receiving and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.
[0120] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132, the text input module 134, the email client module 140 and the browser module 147, the calendar module 148 includes executable instructions for creating, displaying, modifying and storing calendars and data associated with the calendar (e.g., calendar entries, to-do items, etc.) in accordance with user instructions.
[0121] In conjunction with RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and browser module 147, widget module 149 is a mini-application that is optionally downloaded and used by a user (e.g., weather widget 149-1, stock market widget 149-2, calculator widget 149-3, alarm widget 149-4, and dictionary widget 149-5) or a mini-application created by a user (e.g., user-created widget 149-6). In some embodiments, the widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).
[0122] In combination with the RF circuit 108, touch screen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134 and browser module 147, the widget creator module 150 is optionally used by a user to create widgets (e.g., converting a user-specified portion of a web page into a widget).
[0123] In combination with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132 and text input module 134, the search module 151 includes executable instructions for searching the memory 102 for text, music, sound, images, videos and / or other files that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.
[0124] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, video and music player module 152 includes executable instructions that allow a user to download and play back recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing back video (e.g., on touch screen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (trademark of Apple Inc.).
[0125] In conjunction with the touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, the note module 153 includes executable instructions for creating and managing notes, to-do lists, etc. according to user instructions.
[0126] In combination with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphics module 132, the text input module 134, the GPS module 135 and the browser module 147, the map module 154 is optionally used to receive, display, modify and store maps and data associated with the maps (e.g., driving directions, data relating to stores and other points of interest at or near a particular location, and other location-based data) in accordance with user instructions.
[0127] In conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuit 110, speaker 111, RF circuit 108, text input module 134, email client module 140, and browser module 147, online video module 155 includes instructions for performing the following operations: allowing a user to access, browse, receive (e.g., by streaming and / or downloading), playback (e.g., on the touch screen or on an external display connected via external port 124), send an email with a link to a specific online video, and otherwise manage online videos in one or more file formats such as H.264. In some embodiments, instant messaging module 141 is used instead of email client module 140 to send a link to a specific online video. Additional descriptions of online video applications can be found in U.S. Provisional Patent Application No. 60 / 936,562, filed on June 20, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” and U.S. Patent Application No. 11 / 968,067, filed on December 31, 2007, entitled “Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos,” the contents of which are hereby incorporated by reference in their entirety.
[0128] Each of the modules and applications described above corresponds to an executable instruction set for performing one or more of the functions described above and the methods described in this patent application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) do not have to be implemented as independent software programs (such as computer programs (e.g., including instructions)), processes, or modules, so various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. For example, a video player module is optionally combined with a music player module into a single module (e.g., Figure 1A In some embodiments, the memory 102 optionally stores a subset of the above modules and data structures. In addition, the memory 102 optionally stores additional modules and data structures not described above.
[0129] In some embodiments, the device 100 is a device where the operation of a predefined set of functions on the device is performed exclusively through a touch screen and / or a touch pad. By using a touch screen and / or a touch pad as the primary input control device for operating the device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on the device 100 is optionally reduced.
[0130] A predefined set of functions that are performed exclusively through a touch screen and / or a touch pad optionally includes navigation between user interfaces. In some embodiments, the touch pad, when touched by a user, navigates the device 100 from any user interface displayed on the device 100 to a main menu, a main desktop menu, or a root menu. In such embodiments, a touch pad is used to implement a "menu button." In some other embodiments, the menu button is a physical push button or other physical input control device, rather than a touch pad.
[0131] Figure 1B 1 is a block diagram illustrating exemplary components for event processing according to some embodiments. In some embodiments, memory 102 ( Figure 1A ) or memory 370( Figure 3 ) includes an event classifier 170 (e.g., in the operating system 126) and a corresponding application 136-1 (e.g., any one of the aforementioned applications 137 to 151, 155, 380 to 390).
[0132] Event classifier 170 receives event information and determines the application 136-1 and the application view 191 of application 136-1 to which the event information is to be delivered. Event classifier 170 includes event monitor 171 and event distributor module 174. In some embodiments, application 136-1 includes application internal state 192, which indicates one or more current application views displayed on touch-sensitive display 112 when the application is active or executing. In some embodiments, device / global internal state 157 is used by event classifier 170 to determine which application(s) is currently active, and application internal state 192 is used by event classifier 170 to determine the application view 191 to which the event information is to be delivered.
[0133] In some embodiments, the application internal state 192 includes additional information, such as one or more of the following: resumption information to be used when the application 136-1 resumes execution, user interface state information indicating that information is being displayed or is ready to be displayed by the application 136-1, a state queue for enabling a user to return to a previous state or view of the application 136-1, and a repeat / undo queue of previous actions taken by the user.
[0134] Event monitor 171 receives event information from peripherals interface 118. Event information includes information about sub-events (e.g., user touches on touch-sensitive display 112 as part of a multi-touch gesture). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, one or more accelerometers 168, and / or microphone 113 (through audio circuit 110). The information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display 112 or a touch-sensitive surface.
[0135] In some embodiments, event monitor 171 sends requests to peripheral device interface 118 at predetermined intervals. In response, peripheral device interface 118 transmits event information. In other embodiments, peripheral device interface 118 transmits event information only when there is a significant event (e.g., receiving an input above a predetermined noise threshold and / or receiving an input for more than a predetermined duration).
[0136] In some embodiments, the event classifier 170 also includes a hit view determination module 172 and / or an active event identifier determination module 173.
[0137] When the touch-sensitive display 112 displays more than one view, the hit view determination module 172 provides software procedures for determining where within one or more views a sub-event has occurred. A view consists of controls and other elements that a user can see on the display.
[0138] Another aspect of the user interface associated with an application is a set of views, sometimes also referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application views (of the respective application) in which a touch is detected optionally correspond to a programmatic level within the programmatic or view hierarchy of the application. For example, the lowest level view in which a touch is detected is optionally referred to as a hit view, and the set of events that are recognized as correct input is optionally determined at least in part based on the hit view of the initial touch that started the touch-based gesture.
[0139] Hit view determination module 172 receives information related to sub-events of touch-based gestures. When an application has multiple views organized in a hierarchy, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy where the sub-events should be processed. In most cases, the hit view is the lowest level view in which the initiating sub-event (e.g., the first sub-event in a sequence of sub-events that form an event or potential event) occurs. Once a hit view is identified by hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source for which it was identified as the hit view.
[0140] Active event recognizer determination module 173 determines which view or views within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that include the physical location of the sub-event are actively participating views, and therefore determines that all actively participating views should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely confined to an area associated with one particular view, higher views in the hierarchy will still remain as actively participating views.
[0141] Event distributor module 174 distributes event information to event recognizers (e.g., event recognizers 180). In embodiments including active event recognizer determination module 173, event distributor module 174 delivers the event information to the event recognizers determined by active event recognizer determination module 173. In some embodiments, event distributor module 174 stores the event information in an event queue, which is retrieved by corresponding event receivers 182.
[0142] In some embodiments, operating system 126 includes event classifier 170. Alternatively, application 136-1 includes event classifier 170. In yet another embodiment, event classifier 170 is a standalone module or is part of another module stored in memory 102, such as contact / motion module 130.
[0143] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each of which includes instructions for handling touch events occurring within a corresponding view of the user interface of the application. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, the corresponding application view 191 includes multiple event recognizers 180. In other embodiments, one or more event recognizers in event recognizers 180 are part of an independent module, which is a higher-level object such as a user interface toolkit or application 136-1 from which methods and other properties are inherited. In some embodiments, the corresponding event handler 190 includes one or more of the following: data updater 176, object updater 177, GUI updater 178, and / or event data 179 received from event classifier 170. Event handler 190 optionally utilizes or calls data updater 176, object updater 177, or GUI updater 178 to update application internal state 192. Alternatively, one or more of the application views in application view 191 include one or more corresponding event handlers 190. In addition, in some embodiments, one or more of data updater 176, object updater 177, and GUI updater 178 are included in the corresponding application view 191.
[0144] The corresponding event identifier 180 receives event information (e.g., event data 179) from the event classifier 170 and identifies the event based on the event information. The event identifier 180 includes an event receiver 182 and an event comparator 184. In some embodiments, the event identifier 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (which optionally include sub-event delivery instructions).
[0145] Event receiver 182 receives event information from event classifier 170. Event information includes information about sub-events such as touch or touch movement. Depending on the sub-event, the event information also includes additional information, such as the location of the sub-event. When the sub-event involves the movement of the touch, the event information optionally also includes the speed and direction of the sub-event. In some embodiments, the event includes the device rotating from one orientation to another orientation (for example, from a longitudinal orientation to a transverse orientation, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the device posture).
[0146] Event comparator 184 compares event information with predefined event or sub-event definitions, and determines an event or sub-event based on the comparison, or determines or updates the state of an event or sub-event. In some embodiments, event comparator 184 includes event definition 186. Event definition 186 includes the definition of an event (e.g., a predefined sequence of sub-events), such as event 1 (187-1), event 2 (187-2), and others. In some embodiments, sub-events in event (187) include, for example, touch start, touch end, touch move, touch cancel, and multi-touch. In one example, the definition of event 1 (187-1) is a double-click on a displayed object. For example, a double-click includes a first touch (touch start) of a predetermined duration on a displayed object, a first lift-off (touch end) of a predetermined duration, a second touch (touch start) of a predetermined duration on a displayed object, and a second lift-off (touch end) of a predetermined duration. In another example, the definition of event 2 (187-2) is a drag on a displayed object. For example, dragging includes a touch (or contact) of a predetermined duration on a displayed object, movement of the touch on the touch-sensitive display 112, and lifting of the touch (touch end). In some embodiments, the event also includes information for one or more associated event handlers 190.
[0147] In some embodiments, event definition 187 includes definitions of events for corresponding user interface objects. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with a sub-event. For example, in an application view displaying three user interface objects on touch-sensitive display 112, when a touch is detected on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a corresponding event handler 190, the event comparator uses the result of the hit test to determine which event handler 190 should be activated. For example, event comparator 184 selects an event handler associated with a sub-event and the object that triggered the hit test.
[0148] In some embodiments, the definition of the corresponding event (187) also includes a delay action that delays the delivery of the event information until it has been determined that the sub-event sequence does or does not correspond to the event type of the event identifier.
[0149] When a corresponding event recognizer 180 determines that a sequence of sub-events does not match any event in event definition 186, the corresponding event recognizer 180 enters an event impossible, event failed, or event ended state, after which subsequent sub-events of the touch-based gesture are ignored. In this case, other event recognizers (if any) that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.
[0150] In some embodiments, the corresponding event recognizers 180 include metadata 183 with configurable properties, flags, and / or lists that indicate how the event delivery system should perform sub-event delivery to actively participating event recognizers. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact or can interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate whether sub-events are delivered to different levels in a view or programmatic hierarchy.
[0151] In some embodiments, when one or more specific sub-events of an event are identified, the corresponding event recognizer 180 activates an event handler 190 associated with the event. In some embodiments, the corresponding event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is different from sending (and deferred sending) the sub-events to the corresponding hit view. In some embodiments, the event recognizer 180 throws a tag associated with the identified event, and the event handler 190 associated with the tag obtains the tag and executes a predefined process.
[0152] In some embodiments, the event delivery instructions 188 include a sub-event delivery instruction that delivers event information about a sub-event without activating an event handler. Instead, the sub-event delivery instruction delivers the event information to an event handler associated with a sub-event sequence or to an actively participating view. The event handler associated with the sub-event sequence or with an actively participating view receives the event information and executes a predetermined process.
[0153] In some embodiments, data updater 176 creates and updates data used in application 136-1. For example, data updater 176 updates phone numbers used in contact module 137 or stores video files used in video player module. In some embodiments, object updater 177 creates and updates objects used in application 136-1. For example, object updater 177 creates new user interface objects or updates the position of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on a touch-sensitive display.
[0154] In some embodiments, event handler 190 includes or has access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the corresponding application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.
[0155] It should be understood that the above discussion of event processing for user touches on a touch-sensitive display also applies to other forms of user input that utilize input devices to operate the multifunction device 100, and not all user input is initiated on the touch screen. For example, mouse movement and mouse button presses, optionally in conjunction with single or multiple keyboard presses or hold; contact movement on a touch pad, such as tapping, dragging, scrolling, etc.; stylus input; movement of the device; verbal commands; detected eye movement; biometric input; and / or any combination thereof are optionally used as input corresponding to sub-events that define the event to be distinguished.
[0156] Figure 2A portable multifunction device 100 with a touch screen 112 according to some embodiments is shown. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment and other embodiments described below, a user can select one or more of these graphics by, for example, making gestures on the graphics using one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, when the user interrupts contact with one or more graphics, selection of one or more graphics will occur. In some embodiments, gestures optionally include one or more taps, one or more swipes (from left to right, from right to left, up and / or down) and / or rolling of fingers that have been in contact with the device 100 (from right to left, from left to right, up and / or down). In some specific implementations or in some cases, inadvertent contact with a graphic will not select the graphic. For example, when the gesture corresponding to the selection is a tap, a swipe gesture that sweeps over an application icon optionally does not select the corresponding application.
[0157] The device 100 optionally also includes one or more physical buttons, such as a "home" or menu button 204. As previously described, the menu button 204 is optionally used to navigate to any application 136 in a set of applications that are optionally executed on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key in a GUI displayed on the touch screen 112.
[0158] In some embodiments, the device 100 includes a touch screen 112, a menu button 204, a push button 206 for turning the device on / off and for locking the device, one or more volume adjustment buttons 208, a user identity module (SIM) card slot 210, an earphone jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and keeping the button in a pressed state for a predefined time interval; lock the device by pressing the button and releasing the button before the predefined time interval passes; and / or unlock the device or initiate an unlocking process. In an alternative embodiment, the device 100 also accepts voice input for activating or deactivating certain functions through a microphone 113. The device 100 also optionally includes one or more contact strength sensors 165 for detecting the strength of contact on the touch screen 112, and / or one or more tactile output generators 167 for generating tactile output for a user of the device 100.
[0159] Figure 3300 is a block diagram of an exemplary multifunction device with a display and a touch-sensitive surface according to some embodiments. The device 300 does not have to be portable. In some embodiments, the device 300 is a laptop, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a children's learning toy), a game system, or a control device (e.g., a home controller or an industrial controller). The device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, a memory 370, and one or more communication buses 320 for interconnecting these components. The communication bus 320 optionally includes circuits (sometimes referred to as a chipset) that interconnect system components and control communications between system components. The device 300 includes an input / output (I / O) interface 330 with a display 340, which is typically a touch screen display. The I / O interface 330 also optionally includes a keyboard and / or a mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 for generating tactile output on the device 300 (e.g., similar to the above reference Figure 1A The tactile output generator 167 described above), sensor 359 (e.g., optical sensor, acceleration sensor, proximity sensor, touch sensor and / or contact intensity sensor (similar to the above reference Figure 1A The memory 370 may include a high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and may optionally include a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 370 may optionally include one or more storage devices located away from the CPU 310. In some embodiments, the memory 370 stores data related to the portable multifunction device 100 ( Figure 1A ) or a subset thereof. In addition, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk editing module 388, and / or a spreadsheet module 390, while portable multifunction device 100 ( Figure 1A )'s memory 102 optionally does not store these modules.
[0160] Figure 3Each element in the above-mentioned elements in is optionally stored in one or more memory devices of the previously mentioned memory device. Each module in the above-mentioned modules corresponds to an instruction set for performing the above-mentioned functions. The above-mentioned modules or computer programs (for example, instruction sets or including instructions) do not have to be implemented with separate software programs (such as computer programs (for example, including instructions)), processes or modules, and therefore the various subsets of these modules are optionally combined or otherwise rearranged in various embodiments. In some embodiments, memory 370 optionally stores a subset of the above-mentioned modules and data structures. In addition, memory 370 optionally stores additional modules and data structures not described above.
[0161] Attention is now turned to an embodiment of a user interface that is optionally implemented on, for example, portable multifunction device 100.
[0162] Figure 4A An exemplary user interface of an application menu on portable multifunction device 100 according to some embodiments is shown. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof:
[0163] Signal strength indicators 402 for wireless communications such as cellular signals and Wi-Fi signals;
[0164] Time 404;
[0165] Bluetooth indicator 405;
[0166] Battery status indicator 406;
[0167] A tray 408 with icons for commonly used applications, such as:
[0168] o An icon 416 of the phone module 138 labeled "Phone," which optionally includes an indicator 414 of the number of missed calls or voicemails;
[0169] o an icon 418 of the email client module 140 labeled “Mail”, the icon 418 optionally including an indicator 410 of the number of unread emails;
[0170] o An icon 420 labeled "Browser" of the browser module 147; and
[0171] o An icon 422 labeled “iPod” for a video and music player module 152 (also referred to as an iPod (trademark of Apple Inc.) module 152); and
[0172] Icons of other applications, such as:
[0173] o Icon 424 labeled "Message" of IM module 141;
[0174] o An icon 426 labeled “Calendar” of the calendar module 148;
[0175] o Icon 428 labeled “Photos” of the image management module 144;
[0176] o An icon 430 labeled “Camera” of the camera module 143;
[0177] ○ Icon 432 labeled “Online Video” of the online video module 155;
[0178] ○ Icon 434 labeled “Stock Market” of the Stock Market Widget 149 - 2 ;
[0179] o An icon 436 labeled “Map” of the map module 154;
[0180] ○ Icon 438 labeled “Weather” of the weather widget 149-1;
[0181] ○ Icon 440 labeled as “Clock” of the alarm clock widget 149 - 4 ;
[0182] o An icon 442 labeled “Fitness Support” of the fitness support module 142;
[0183] o An icon 444 labeled "Notes" of the notes module 153; and
[0184] o An icon 446 labeled “Settings” for a settings application or module that provides access to settings for the device 100 and its various applications 136 .
[0185] It should be noted that Figure 4A The icon labels shown in are exemplary only. For example, icon 422 of video and music player module 152 is labeled "Music" or "Music Player". Other labels are optionally used for various application icons. In some embodiments, the label of a corresponding application icon includes the name of the application corresponding to the corresponding application icon. In some embodiments, the label of a particular application icon is different from the name of the application corresponding to the particular application icon.
[0186] Figure 4B A touch-sensitive surface 451 (eg, touch screen display 112) is shown having a touch-sensitive surface 451 (eg, Figure 3 tablet computer or touch pad 355) device (e.g., Figure 3Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) for detecting intensity of contacts on touch-sensitive surface 451 and / or one or more tactile output generators 357 for generating tactile output for a user of device 300.
[0187] Although some of the examples below are given with reference to input on a touch screen display 112 (where a touch-sensitive surface and a display are combined), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, such as Figure 4B In some embodiments, the touch-sensitive surface (e.g., Figure 4B 451) has a main axis (e.g., Figure 4B 453) corresponding to the principal axis (for example, Figure 4B According to these embodiments, the device detects a position corresponding to a corresponding position on the display (e.g., Figure 4B 460 corresponds to 468 and 462 corresponds to 470) in contact with touch-sensitive surface 451 (e.g., Figure 4B Thus, when the touch-sensitive surface (e.g., Figure 4B 451) and a display of a multi-function device (e.g., Figure 4B When the user input detected by the device on the touch-sensitive surface (e.g., contacts 460 and 462 and their movement) is separated, the device is used to manipulate the user interface on the display. It should be understood that similar methods are optionally used for other user interfaces described herein.
[0188] In addition, although the following examples are primarily given with reference to finger inputs (e.g., finger contacts, single-finger tap gestures, finger swipe gestures), it should be understood that in some embodiments, one or more of these finger inputs are replaced by input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of contact), followed by movement of the cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click when the cursor is over the location of the tap gesture (e.g., instead of detecting contact, followed by ceasing to detect contact). Similarly, when multiple user inputs are detected simultaneously, it should be understood that multiple computer mice are optionally used simultaneously, or a mouse and finger contact are optionally used simultaneously.
[0189] Figure 5AAn exemplary personal electronic device 500 is shown. Device 500 includes a body 502. In some embodiments, device 500 may include a body 502 relative to devices 100 and 300 (e.g., Figures 1A to 4B ) some or all of the features described in the foregoing. In some embodiments, device 500 has a touch-sensitive display screen 504, referred to hereinafter as touch screen 504. As an alternative or in addition to touch screen 504, device 500 has a display and a touch-sensitive surface. As with devices 100 and 300, in some embodiments, touch screen 504 (or touch-sensitive surface) optionally includes one or more strength sensors for detecting the intensity of contact (e.g., touch) applied. One or more strength sensors of touch screen 504 (or touch-sensitive surface) can provide output data representing the intensity of touch. The user interface of device 500 can respond to touch based on the intensity of the touch, which means that touches of different intensities can invoke different user interface operations on device 500.
[0190] Exemplary techniques for detecting and processing touch intensity are found, for example, in the following related patent applications: International patent application serial number PCT / US2013 / 040061, entitled “Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application,” filed on May 8, 2013, published as WIPO patent publication number WO / 2013 / 169849; and International patent application serial number PCT / US2013 / 069483, entitled “Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships,” filed on November 11, 2013, published as WIPO patent publication number WO / 2014 / 105276, each of which is hereby incorporated by reference in its entirety.
[0191] In some embodiments, the device 500 has one or more input mechanisms 506 and 508. The input mechanisms 506 and 508 (if included) can be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, the device 500 has one or more attachment mechanisms. Such attachment mechanisms (if included) can allow the device 500 to be attached to, for example, hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watchbands, bracelets, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow the user to wear the device 500.
[0192] Figure 5B An exemplary personal electronic device 500 is depicted. In some embodiments, the device 500 may include a reference Figure 1A , Figure 1B and Figure 3 Some or all of the components described. Device 500 has a bus 512 that operatively couples an I / O portion 514 to one or more computer processors 516 and a memory 518. I / O portion 514 may be connected to a display 504, which may have a touch-sensitive component 522 and optionally a strength sensor 524 (e.g., a contact strength sensor). In addition, I / O portion 514 may be connected to a communication unit 530 for receiving application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular and / or other wireless communication technologies. Device 500 may include input mechanisms 506 and / or 508. For example, input mechanism 506 is optionally a rotatable input device or a depressible input device and a rotatable input device. In some examples, input mechanism 508 is optionally a button.
[0193] In some examples, input mechanism 508 is optionally a microphone. Personal electronic device 500 optionally includes various sensors, such as GPS sensor 532, accelerometer 534, orientation sensor 540 (e.g., compass), gyroscope 536, motion sensor 538, and / or combinations thereof, all of which are operably connected to I / O portion 514.
[0194] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, may cause the computer processors to perform the techniques described below, including processes 700, 900, and 1100 ( Figure 7 , Fig. 9 , Fig.11). Computer-readable storage media can be any medium that can tangibly contain or store computer-executable instructions for use by or in conjunction with instruction execution systems, devices, and apparatuses. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like. Personal electronic device 500 is not limited to Figure 5B components and configurations, but may include other components or additional components in a variety of configurations.
[0195] As used herein, the term "indicative representation" refers to an optional feature of the apparatus 100, 300, and / or 500 ( Figure 1A , Figure 3 and FIG. 5A to FIG. 5B ) is a user-interactive graphical user interface object displayed on a display screen of a computer. For example, an image (e.g., an icon), a button, and text (e.g., a hyperlink) optionally each constitute an affordance.
[0196] As used herein, the term "focus selector" refers to an input element used to indicate the current portion of a user interface that a user is interacting with. In some implementations that include a cursor or other position marker, the cursor acts as a "focus selector" such that when the cursor is over a particular user interface element (e.g., a button, window, slider, or other user interface element), a focus selector is displayed on a touch-sensitive surface (e.g., Figure 3 Touchpad 355 or Figure 4B In the event that an input (e.g., a press input) is detected on the touch-sensitive surface 451 in the display, the particular user interface element is adjusted according to the detected input. In the case that a touch-screen display (e.g., Figure 1A A touch-sensitive display system 112 or Figure 4AIn some implementations of the touch screen 112 in FIG. 1 , a contact detected on the touch screen acts as a “focus selector” such that when an input (e.g., a press input by the contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted in accordance with the detected input. In some implementations, the focus moves from one area of the user interface to another area of the user interface without corresponding movement of a cursor or movement of the contact on the touch screen display (e.g., by using a tab key or arrow keys to move the focus from one button to another); in these implementations, the focus selector moves in accordance with the movement of the focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is typically a user interface element (or contact on the touch screen display) that is controlled by the user to deliver the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface with which the user desires to interact). For example, when a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of a focus selector (e.g., a cursor, contact, or selection box) over a corresponding button will indicate that the user intends to activate the corresponding button (rather than other user interface elements shown on the device display).
[0197] As used in the specification and claims, the term "characteristic intensity" of a contact refers to a characteristic of a contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on multiple intensity samples. The characteristic intensity is optionally based on a predefined number of intensity samples or a set of intensity samples collected during a predetermined time period (e.g., 0.05 seconds, 0.1 seconds, 0.2 seconds, 0.5 seconds, 1 second, 2 seconds, 5 seconds, 10 seconds) relative to a predefined event (e.g., after contact is detected, before contact is detected to be lifted off, before or after contact starts to move, before contact ends, before or after contact intensity is detected to increase, and / or before or after contact intensity is detected to decrease). The characteristic intensity of a contact is optionally based on one or more of the following: the maximum value of the intensity of the contact, the mean value of the intensity of the contact, the average value of the intensity of the contact, the value at the top 10% of the intensity of the contact, the half-maximum value of the intensity of the contact, the 90% maximum value of the intensity of the contact, etc. In some embodiments, the duration of the contact is used when determining the characteristic intensity (e.g., when the characteristic intensity is the average value of the intensity of the contact over time). In some embodiments, the feature strength is compared to a set of one or more strength thresholds to determine whether the user has performed an operation. For example, the set of one or more strength thresholds optionally includes a first strength threshold and a second strength threshold. In this example, a contact whose feature strength does not exceed the first threshold results in a first operation, a contact whose feature strength exceeds the first strength threshold but does not exceed the second strength threshold results in a second operation, and a contact whose feature strength exceeds the second threshold results in a third operation. In some embodiments, a comparison between the feature strength and one or more thresholds is used to determine whether to perform one or more operations (e.g., whether to perform the corresponding operation or to abandon the corresponding operation) rather than to determine whether to perform the first operation or the second operation.
[0198] Attention is now turned to embodiments of a user interface ("UI") and associated processes implemented on an electronic device, such as portable multifunction device 100, device 300, or device 500.
[0199] FIG. 6A to FIG. 6M An exemplary user interface for providing background sound according to some embodiments is shown. The user interface in these figures is used to illustrate the following description including Figure 7 The process of the process.
[0200] It should be understood that for many users, everyday sounds may be distracting, uncomfortable or too loud. Therefore, exemplary user interfaces such as those described herein can be used to provide (e.g., play) background sounds to help minimize distractions and help such users concentrate, stay calm or rest. In some embodiments, balanced noise, bright noise or dark noise and ocean, rain or stream background sounds can be played (e.g., continuously) in the background of the user's hearing to cover up unwanted environmental or external noise. In addition, such sounds can be mixed into other audio and system sounds or attenuated for other audio and system sounds.
[0201] Fig. 6A An electronic device 600 is shown. Fig. 6A In the embodiment, electronic device 600 is a portable multifunction device and has one or more components described above with respect to one or more of devices 100, 300 and 500.
[0202] exist Fig. 6A , device 600 displays an audio settings interface 604 on display 602. Audio settings interface 604 includes feature enable indications 606a-606e. In some embodiments, feature enable indications 606a-606e correspond to various audio-related functions of device 600. For example, feature enable indication 606a corresponds to a headset adaptation function, feature enable indication 606b corresponds to a background sound function, feature enable indication 606c corresponds to a mono audio function, feature enable indication 606d corresponds to a phone noise cancellation function, and feature enable indication 606e corresponds to a headset notification function.
[0203] In some embodiments, in response to selection of feature affordance 606b, device 600 displays a background sound interface. Fig. 6A , when audio settings interface 604 is displayed, device 600 detects selection of feature affordance 606b. Fig. 6A , the selection is a tap gesture 610 on the feature affordance 606b. Figure 6B As shown in , in response to detecting the tap gesture 610 , the device 600 displays a background sound interface 620 .
[0204] Background sound interface 620 includes option 622, sound affordance 624, background sound volume control 626, and parallel volume control 628. In some embodiments, option 622 is used to switch the use of background sound on device 600. For example, in response to selection of option 622, device 600 switches the state of option 622 (e.g., deactivates the option if activated, activates the option if deactivated). Figure 6BIn response to detecting a tap gesture 630 at option 622, device 600 deactivates the use of background sounds on device 600 and modifies the display of option 622 to indicate that the use of background sounds has been deactivated, such as Figure 6C as shown in .
[0205] In some embodiments, sound enable indication 624 is used to select the type of background sound. For example, in response to selecting sound type enable indication 624, device 600 displays a sound selection interface. Figure 6B , when background sound interface 620 is displayed, device 600 detects selection of sound affordance 624. Figure 6B In the example, the selection is a tap gesture 632 on the sound enabling representation 624. Fig.6D As shown in , in response to detecting tap gesture 632 , device 600 displays sound selection interface 634 .
[0206] exist Fig.6D , the sound selection interface 634 includes candidate background sound indications 636a-636f, each of which corresponds to a corresponding type of background sound. Candidate background sound indication 636a corresponds to a "balanced noise" background sound, candidate background sound indication 636b corresponds to a "bright noise" background sound, candidate background sound indication 636c corresponds to a "dark noise" background sound, candidate background sound indication 636d corresponds to an "ocean" background sound, candidate background sound indication 636e corresponds to a "rain" background sound, and candidate background sound indication 636f corresponds to a "stream" background sound.
[0207] exist Fig.6D , indicator 638 indicates that the "bright noise" background sound is currently selected on device 600. However, the user can select a different background sound type by providing user input corresponding to a selection of candidate background sound affordances 636a-636f. Fig.6D , when sound selection interface 634 is displayed, device 600 detects user input 640 (e.g., a tap) corresponding to selection of candidate background sound affordance 636d. In response to user input 640, device 600 moves indicator 638 from candidate background sound affordance 636b to candidate background sound affordance 636d. As a result, the "ocean" background sound is selected on device 600, and the bright noise background sound is not selected.
[0208] Return to reference Figure 6B, background sound volume control 626 includes a slider 626a for adjusting the volume level of the background sound provided by device 600. Parallel volume control 628 includes option 628a and slider 628b. In some embodiments, option 628a is used to switch whether the background sound is played in parallel with other audio provided by device 600, and slider 628b is used to adjust the volume level of the background sound when provided in parallel with other audio.
[0209] In some examples, when background sounds and other audio are provided in parallel, the volume level of the background sounds is adjusted according to both slider 626a and slider 628b. Figure 6B Slider 626a is shown at a position corresponding to volume level "50" and slider 628a is shown at a position corresponding to volume level "50", which corresponds to a volume level of 50% of the maximum volume level. Thus, during concurrent playback of background and other audio, the volume level of the background sounds is first adjusted (e.g., reduced) by 50% based on the position of slider 626a, and then adjusted a second time by 50% based on the position of slider 628b, resulting in a volume level corresponding to "25" or 25% of the maximum volume level of the background sounds.
[0210] In some embodiments, the volume level of background sounds is adjusted based on contextual information of the device 600. Contextual information includes user-specific data stored on the device 600 and / or accessible by the device, as well as any information describing the operating state of the device 600 (time, day, week, whether the device 600 is charging, whether the device 600 is locked) or the environment (e.g., location, ambient noise). For example, the device 600 may play background sounds at different volumes based on the time of day. As another example, the device 600 may play background sounds at a volume level commensurate with the volume level of ambient noise (e.g., external noise detected by the device 600). In some embodiments, contextual information is used to adjust the volume level of the background on the device 600 and / or the volume level of the background sounds on the device 600 during parallel playback with other audio.
[0211] In some embodiments, when option 628a is activated, device 600 may play background sounds in parallel with other audio provided by device 600. Fig. 6F Operations are shown in which device 600 is playing background sound when option 628a is activated. While the background sound is playing, device 600 detects user input 644 (e.g., a tap) at play affordance 642. In response to user input 644, device 600 determines that concurrent playback of background sound and other audio is activated (via option 628a being activated) and initiates concurrent playback of background sound and audiobook 646, as shown in FIG. Figure 6G as shown in .
[0212] In some embodiments, when option 628a is deactivated, device 600 is unable to play background sounds in parallel with other audio provided by device 600. Figure 6H 6 shows an operation in which device 600 is playing background sound when option 628a is deactivated. While the background sound is playing, device 600 detects user input 648 (e.g., a tap) at play affordance 642. In response to user input 648, device 600 determines that concurrent playback of the background sound and other audio is not allowed (by virtue of option 628a being activated) and therefore terminates (e.g., stops, pauses) concurrent playback of the background sound and audiobook 646, as shown. Fig.6I as shown in .
[0213] As described, device 600 plays background sound when background sound option is enabled. In some embodiments, device 600 is configured to play background sound only when one or more other conditions are met. For example, device 600 may be configured to play background sound only when it is determined that the user of device 600 is wearing headphones. For another example, device 600 may be configured to abandon playing background sound when a sound corresponding to an alert or alarm is detected near the user.
[0214] refer to Figure 6J In some embodiments, the slider 628b is located at position 650 so that when playing with other media, the corresponding volume level of the background sound is 0 (e.g., the slider 628b has a set value of 0). In some embodiments, when the slider 628b is located at position 650 (e.g., the volume level of the background sound is 0 during parallel playback), the device 600 does not play the background sound in parallel with the other media (the volume is 0), but pauses the playback of the background sound. When the other media finishes playing, the device 600 resumes the playback of the background sound, for example, from the point where the background sound was paused.
[0215] In some embodiments, device 600 provides (e.g., generates) background sound in a random manner. For example, device 600 can provide randomly selected audio clips as background sound. For example, when playing the background sound of the "ocean" type, device 600 can provide (e.g., continuously provide) randomly selected ocean background sound clips (thus minimizing any perceived repetition of the user to the background). For example, device 600 can provide randomly arranged audio clips as background sound. For example, when playing the background sound of the "balanced noise" type, device 600 can provide balanced noise audio with one or more randomly generated audio characteristics (e.g., the randomized magnitude of one or more frequencies). In some embodiments, even when background sound is randomly provided, device 600 also pauses and resumes the playback of background sound (e.g., when slider 628a is at position 650).
[0216] Figure 6K An exemplary control center interface 660 is shown. In some embodiments, control center interface 660 is a user interface that can be displayed in response to a predetermined gesture (e.g., a swipe from the upper right edge of the display) received when most other user interfaces are being displayed. Control center interface 660 includes various enable representations for controlling corresponding functions and / or components of device 600, including audio enable representation 662. In some examples, in response to selection of audio enable representation 662, device 600 displays an audio interface. Figure 6K , when control center interface 660 is displayed, device 600 detects selection of audio affordance 662, i.e., tap gesture 664 on audio affordance 662. Figure 6L As shown in , in response to detecting tap gesture 664 , device 600 displays audio interface 670 .
[0217] Audio interface 670 includes region 672. Region 672 includes background sound indicator 674 and volume control 676. Background sound indicator 674 indicates whether background sound is currently provided by device 600, and if so, the type of background sound provided (e.g., rain, ocean, stream). Figure 6L As shown in FIG. 6 , the background sound indicator 674 indicates that the device 600 is currently playing the "ocean" background sound. The volume control 676 includes a volume slider 676a that can be used to adjust the volume level of the background sound provided by the device 600.
[0218] Audio interface 670 also includes status indicator 678 and option 680. Status indicator 678 indicates whether background sound is currently playing on device 600. Option 680 is used to switch the playback of background sound on device 600. For example, in response to selection of option 680, device 600 switches the state of option 680 (e.g., deactivates the option if activated, and activates the option if deactivated). Figure 6M As shown in , for example, in response to detecting a tap gesture 682 at option 680, device 600 deactivates playback of background sounds.
[0219] In further response to detecting tap gesture 682, device 600 modifies one or more elements of audio interface 670 to indicate that background sound has been deactivated. For example, device 600 replaces indicator 674 with indicator 684, indicating that background sound is currently "off". For another example, device 600 replaces indicator 678 with indicator 686, indicating that background sound is currently "off". For another example, device 600 modifies the display of option 680 (e.g., removes bold focus).
[0220] Figure 7700 is a flow chart illustrating a method for providing background sound using a computer system according to some embodiments. Method 700 is performed at a computer system (e.g., 100, 300, 500, 600) that communicates with one or more input devices (e.g., 602). Some operations in method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0221] As described below, method 700 provides an intuitive way to provide background sound. The method reduces the cognitive burden on the user caused by, for example, selectively playing background sound during playback of other audio, thereby more efficiently utilizing the computer system (e.g., computer system 100, 300, 500, 600). For battery-powered computer systems, enabling users to provide background sound more efficiently saves power and increases the time interval between battery charges.
[0222] When playing the first type of audio media item (e.g., corresponding to the sound of 636a-636f), the computer system receives (702) a request (e.g., 644) to play the second type of audio media item (e.g., 646) via one or more input devices. In some embodiments, the first type of audio media item is an audio media item without musical instruments or vocal audio elements, such as noise (e.g., white noise) or natural sounds (e.g., ocean, rain, stream). In some embodiments, the first type of audio media item includes randomly selected and / or arranged audio clips. In some embodiments, the first type of audio media item is played at a volume level corresponding to the value of the background sound volume feature. In some embodiments, the value of the background sound volume feature can be manually adjusted so that the user can selectively adjust the value of the background sound volume feature, for example, by adjusting the volume slider of the background sound user interface. In some embodiments, the value of the background sound volume feature is adjusted based on the context of the computer system (e.g., ambient noise, user-specific data stored on the computer system (e.g., calendar, message), proximity to other devices). In some embodiments, the second type of audio media item is an audio media item that includes one or more instrumental and / or vocal audio elements, such as music, a soundtrack to a video, an audiobook, or a podcast.
[0223] When playing the audio media item of the first type and according to determining that the parallel audio standard set is satisfied, the computer system plays (704) the audio media item of the first type (e.g., the sound corresponding to 636a-636f) and the audio media item of the second type (e.g., 646) in parallel (e.g., simultaneously, consistently). In some embodiments, the parallel audio standard includes a requirement to enable a parallel playback feature (e.g., "play with media") on the computer system. In some embodiments, the parallel playback feature is manually adjustable so that the user can selectively enable the parallel playback feature, for example, by switching the affordance of the background sound user interface. In some embodiments, the parallel audio standard includes a requirement that the value of the parallel volume feature exceeds a threshold value (e.g., 0). In some embodiments, the value of the parallel volume feature is manually adjustable so that the user can selectively adjust the value of the parallel volume feature, for example, by adjusting the volume slider of the background sound interface. In some embodiments, the volume level of the media item of the first type is determined according to both the magnitude of the background sound volume feature and the magnitude of the parallel volume feature. In some embodiments, the value of the concurrent sound volume feature is adjusted based on the context of the computer system (e.g., ambient noise, user-specific data stored on the computer system (e.g., calendar, message), proximity to other devices). In some embodiments, the computer system maintains playback of the audio media item of the first type. In some embodiments, the computer system restarts playback of the audio media item of the first type. In some embodiments, when playing the audio media item of the second type, the computer system adjusts the volume of the audio media item of the first type, for example, based on the magnitude of the background sound volume feature and / or the magnitude of the concurrent volume feature.
[0224] While playing the first type of audio media item and based on (706) determining that the parallel audio criteria set is not met (e.g., the parallel playback feature is disabled and / or the value of the parallel volume feature does not exceed a threshold amount), the computer system terminates (708) (e.g., pauses) playing of the first type of audio media item.
[0225] While playing the audio media item of the first type and according to (706) determining that the parallel audio standard set is not met, the computer system plays (710) the audio media item of the second type. In some embodiments, when the computer system stops playing the audio media item of the second type, the computer system resumes playing the media item of the first type. Playing the media item of the first type and the media item of the second type in parallel when the parallel playback feature is enabled, and stopping playing the audio media of the first type when playing the audio media of the second type when the parallel playback feature is disabled, allows the user to quickly and efficiently control the manner in which the parallel playback of the media items is achieved, which reduces the number of inputs required to perform the operation.
[0226] In some embodiments, the first type of audio media items include audio selected from the group consisting of ambient sounds (e.g., sounds corresponding to 636d-636f) (e.g., natural sounds (e.g., ocean, rain and / or stream sounds); non-artificial sounds), irregular noise (e.g., sounds corresponding to 636a-636c) (e.g., random noise (white noise, bright noise, balanced noise, dark noise)), and combinations thereof. In some embodiments, the first type of media items do not include speech / vocal or instrumental music / audio.
[0227] In some embodiments, the audio media item of the first type includes audio selected from the group consisting of randomly selected audio segments, randomly arranged audio segments, and combinations thereof. In some embodiments, the audio media item of the first type or background sound includes randomly selected and / or arranged audio segments so that the audio is non-repeatable or unpredictable for the user. In some embodiments, one or more auditory characteristics of the audio media item are adjusted so that the transition between the audio segments appears seamless.
[0228] In some embodiments, the parallel audio criteria set includes criteria that are met when it is determined that a parallel playback feature (e.g., 628a) (e.g., a parallel playback setting that enables parallel playback of a first type of audio and a second type of audio) is active (e.g., enabled, activated). In some embodiments, the parallel playback feature is activated in response to switching a corresponding option or an affordance of the parallel audio interface to an "on" state. In some embodiments, the parallel audio criteria are met when it is further determined that the parallel volume feature has a magnitude that exceeds a predetermined threshold (e.g., threshold 0). Including criteria that are met when it is determined that the parallel media feature is active in the parallel audio criteria set allows for selectively playing a first type of media item and a second type of media item in parallel based on whether the parallel media feature is enabled, which performs an operation without further user input when a set of conditions have been met.
[0229] In some embodiments, the parallel playback feature (e.g., 628a) is manually configurable. In some embodiments, the parallel playback feature can be activated (e.g., enabled) or deactivated (e.g., disabled) by the user input that switches the corresponding option or affordance of the background sound interface to the "on" state or the "off" state respectively. In some embodiments, the parallel playback feature can be activated or deactivated by adjusting the magnitude of the parallel volume feature. In some embodiments, the parallel playback feature is activated when the magnitude of the parallel volume feature exceeds a threshold magnitude (e.g., 0) and is deactivated when the magnitude of the parallel volume feature does not exceed the threshold magnitude. Using the playback feature that can be manually configured allows the user to quickly and efficiently control the mode of realizing the parallel playback of the media item, which reduces the number of inputs required for the execution of the operation.
[0230] In some embodiments, playing the first type of audio media item and the second type of audio media item in parallel includes adjusting the magnitude of the volume level of the first type of audio media item from an initial volume to an adjusted volume based on the magnitude of a parallel volume feature (e.g., 628b) (e.g., a parallel volume setting that adjusts the volume of the audio media item at the first time when parallel playback occurs). In some embodiments, the magnitude of the parallel volume feature is adjusted using a volume slider displayed in the background sound interface.
[0231] In some embodiments, the parallel audio criteria set includes a criterion that is met when it is determined that the magnitude of a second parallel volume feature (e.g., 628b) (e.g., a parallel volume setting that adjusts the volume of the audio media item at the first time when parallel playback occurs; a feature that is the same or different from the parallel volume feature) exceeds a threshold magnitude (e.g., 650). In some embodiments, parallel playback of the first media item and the second media item occurs when the magnitude of the parallel volume feature exceeds the threshold magnitude (e.g., 0) and is deactivated when the magnitude of the parallel volume feature does not exceed the threshold magnitude. In some embodiments, when the magnitude of the parallel volume feature does not exceed the threshold magnitude, playback of the first type of media item is paused during playback of the second type of media item, and optionally resumed when playback of the second type of media item is stopped.
[0232] In some embodiments, the computer system adjusts the background sound volume feature (e.g., 626a) for adjusting the volume level of the first type of media item and the third parallel volume feature (e.g., 628b) for adjusting the volume level of the first type of media item when the first type of media item and the second type of media item are played in parallel based on context information. In some embodiments, the volume level of the first type of audio media item is adjusted based on the context information of the computer system, for example. The context information includes user-specific data (e.g., user calendar) stored on the computer system and / or accessible by the computer system, the location of the computer system, the current time (e.g., time, day, week, month). In some embodiments, the context information also includes the proximity of the ambient noise detected by the computer system and / or the computer system to one or more other computer systems and / or devices. In some embodiments, additionally or alternatively, the volume level of the second media item is adjusted based on context information. In some embodiments, because the volume level of the first media item is adjusted according to the background sound volume feature and / or the parallel volume feature, the value of the background sound volume feature is adjusted based on the context information and the value of the parallel volume feature is adjusted based on the context information. Adjusting the magnitude of the volume level based on contextual information allows for intuitive and efficient automatic adjustment of the volume, which operates when a set of conditions have been met without further user input.
[0233] In some embodiments, playing the first type of audio media item includes: based on determining that the computer system is in a first context state (e.g., based on context information about the computer system), playing the first type of audio media item with the first audio content (e.g., corresponding to the sound of 636a-636f); and based on determining that the computer system is in the first context state, playing the first type of audio media item with the second audio content different from the first audio content. In some embodiments, the first audio content and / or the second audio content are selected based on the context information about the computer system. In some embodiments, the content of the first audio media is generated based on the time of the day; for example, the tone of the first audio media item can be adjusted based on the time of the day. In some embodiments, the content of the first audio media item is adjusted based on the proximity of the computer system to one or more other computer systems and / or devices. Playing audio media items based on the context state of the computer system can implement an improved technique for selecting background sound content, which performs operations when a set of conditions have been met without further user input.
[0234] In some embodiments, the computer system also communicates with the display generation component (e.g., 602). In some embodiments, the computer system displays the first user interface (e.g., 620) via the display generation component. In some embodiments, the first user interface includes a background sound user interactive graphical user interface object (e.g., 622) (e.g., affordance), which enables the playback of the audio media item of the first type when it is determined to meet the playback standard set when selected. In some embodiments, the background sound affordance is used to selectively enable the playback of the media item of the first type on the computer system. In some embodiments, the background sound interface includes a setting for adjusting the volume of the media item of the first type. In some embodiments, the first user interface includes an option for selectively enabling the parallel playback of the media item of the first type and the media item of the second type. In some embodiments, the first user interface includes a setting for adjusting the playback volume of the media item of the first type when playing in parallel with the media item of the second type. In some embodiments, the playback standard set includes the standard met when an external speaker (e.g., wired or wireless headphones) is connected to the computer system. In some embodiments, the first user interface (e.g., 620) includes a first audio content user interactive graphical user interface object (e.g., 624) that causes a second user interface (e.g., 634) to be displayed when selected. In some embodiments, the second user interface includes: a second audio content user interactive graphical user interface object (e.g., 636a-636f) that causes a third audio content (e.g., background noise; ocean sound; rain sound; stream sound) to be included in the audio media item of the first type during playback when selected; and a third audio content user interactive graphical user interface object (e.g., 636a-636f) that causes a fourth audio content (e.g., background noise; ocean sound; rain sound; stream sound; content) different from the third audio content to be included in the audio media item of the first type during playback when selected.
[0235] In some embodiments, the computer system detects a first input (e.g., 630) (e.g., a tap, a mouse click, a key press) corresponding to the background sound user-interactive graphical user interface object (e.g., 622) via one or more input devices. In some embodiments, in response to detecting the first input, the computer system enables (e.g., modifies the settings of the computer system) playback of the first type of audio media item when determining that the playback criteria set is met.
[0236] In some embodiments, when the computer system is enabled to playback the audio media item of the first type when it is determined that the set of playback criteria is met, the computer system plays the audio media item of the first type.
[0237] In some embodiments, when the second user interface is displayed, the computer system detects a second input (e.g., 640) (e.g., a tap, a mouse click, a key) via one or more input devices. In some embodiments, when the second user interface is displayed, in response to detecting the second input and according to determining that the second input corresponds to a second audio content user-interactive graphical user interface object (e.g., 636a-636f), the computer system configures the computer system to include the third audio content in the audio media item of the first type during playback (e.g., subsequent playback).
[0238] In some embodiments, when the second user interface is displayed, in response to detecting the second input and based on determining that the second input corresponds to a third audio content user-interactive graphical user interface object (e.g., 636a-636f), the computer system configures the computer system to include the fourth audio content in the audio media item of the first type during playback (e.g., subsequent playback).
[0239] In some embodiments, the computer system displays a fourth user interface (e.g., 670) via a display generation component, and the fourth user interface includes a background sound status indicator (e.g., 622, 674), a volume user interactive graphical user interface object (e.g., 626a, 676a), and a background sound enabled user interactive graphical user interface object (e.g., 622, 680). In some embodiments, the computer system detects a third input (e.g., 630, 682) (e.g., a tap, a mouse click, a key, a swipe gesture) via one or more input devices. In some embodiments, based on determining that the third input corresponds to a swipe gesture (e.g., a left swipe gesture, a right swipe gesture) at a position corresponding to a volume user interactive graphical user interface object (e.g., a volume slider), the computer system adjusts the volume level of the first type of media item from the second initial volume to the second adjusted volume based on the direction and magnitude of the swipe gesture. In some embodiments, the control center user interface includes a volume slider that can be used to adjust the volume of the first type of media item. In some embodiments, adjusting the volume level in this way is similar to adjusting the volume level of the background audio volume feature of the background sound interface of the settings menu. In some embodiments, based on determining that the third input corresponds to a selection of a user-interactive graphical user interface object (e.g., 630, 682) to enable background sound, the computer system selectively activates (e.g., activates or deactivates) the background sound feature. In some embodiments, based on determining that the third input corresponds to a selection of a user-interactive graphical user interface object to enable background sound, the computer system modifies the visual characteristics of the background sound status indicator (e.g., 622, 680). In some embodiments, when the background sound feature is deactivated, the background sound status indicator is modified to indicate that the background sound is "off". In some embodiments, when the background sound feature is activated, the background sound status indicator is modified to indicate that the background sound is enabled. In some embodiments, the background sound status indicator is modified to indicate the type of background sound.
[0240] It should be noted that the process described above with respect to method 700 (eg Figure 7 ) also applies in a similar manner to the methods described below. For example, methods 900 and 1100 optionally include one or more features of the various methods described above with reference to method 700. For example, when playing an audio media item of a first type (e.g., background sound), a user may use an auditory control to initiate playback of an audio media item of a second type, at which point the computer system determines whether to play these media items in parallel. For the sake of brevity, these details are not repeated below.
[0241] Figures 8A to 8VAn exemplary user interface for providing auditory control according to some embodiments is shown. The user interface in these figures is used to illustrate the following description including Fig. 9 The process of the process.
[0242] In some implementations, an exemplary user interface is used to provide voice actions for toggle controls. This can allow users (such as non-speaking and / or users with limited mobility) to invoke actions by providing voice input including sounds corresponding to actions (e.g., clicks, pops, or "ee" sounds) rather than using physical buttons, switches, and verbal commands.
[0243] exist Fig. 8A , device 600 displays an auxiliary function interface 804 on display 602. Accessibility interface 804 includes various affordances for controlling various auxiliary function functions of device 800, including a switch control affordance 806.
[0244] In some examples, in response to selecting the handover control affordance 806, the device 600 displays a handover control interface. Fig. 8A In the example, when auxiliary function interface 804 is displayed, device 600 detects selection of switch control enable indication 806, that is, tap gesture 808 on switch control enable indication 806. Figure 8B As shown in , in response to detecting the tap gesture 808 , the device 600 displays the switch control interface 810 .
[0245] The switch control interface 810 includes options 812 and switch settings enable representation 814. In some embodiments, the options 812 are used to switch the use of the switch control on the device 800. In some embodiments, when the switch control is implemented on the device 800, the switch control enables the use of any number of switches on the device 800 to control one or more functions. A "switch" corresponds to a specific type of input and an action performed by the device 600 in response to receiving the input. The switch can be used to associate an action with any number of input types, including but not limited to input from an external device (e.g., a Bluetooth device), touch input (e.g., a single tap, double tap, triple tap), gestures (hand gestures, head gestures), sound, etc.
[0246] In response to selection of option 812, device 600 switches the state of option 812 (e.g., deactivates the option if activated, activates the option if deactivated). Figure 8B In response to detecting a tap gesture 816 at option 812, device 600 activates a toggle control on device 600 and modifies the display of option 812 to indicate that the toggle control has been activated, such as Figure 8C as shown in .
[0247] exist Figure 8C In the example, the switch setting enable indication 814 includes an indicator 814a indicating the number of switches configured on the device 800. In response to the selection of the switch setting enable indication 814, the device 600 displays the switch setting interface (e.g., Fig.8D For example, when toggle control interface 810 is displayed, device 600 detects a selection of toggle settings enable representation 814. In some examples, the selection is a tap gesture 816 on toggle settings enable representation 814. Fig.8D As shown in , in response to detecting the tap gesture 808 , the device 600 displays the switch settings interface 820 .
[0248] The switch setting interface 820 includes switch enable indications 821a-821c and an add switch enable indication 822. Each switch enable indication 821a-821c corresponds to a corresponding switch configured on the device 800, and indicates the pairing between the action and the input type accordingly. For example, the switch enable indication 821a corresponds to a switch for the sound "oo" and the "move right" action (e.g., voice input), the switch enable indication 821b corresponds to a switch for the screen tap and the "move right" action, and the switch enable indication 821c corresponds to a switch for the sound "eh" and the "move left" action (e.g., voice input).
[0249] In some embodiments, in response to a selection to add a switch enable indication 822, device 600 initiates a process for adding a switch to device 800 (e.g., configuring a switch on the device). For example, when displaying switch settings interface 820, device 600 detects a selection to add a switch enable indication 822. Fig.8D In the example, select Yes to add a tap gesture 824 on the toggle affordance 822. Fig. 8E As shown in , in response to detecting tap gesture 824 , device 600 displays source selection interface 826 .
[0250] Source selection interface 826 includes source affordances 828a-828e. Each of source affordances 828a-828e corresponds to a set of input types that can be associated with a corresponding action. For example, source affordance 828a corresponds to external input (e.g., input provided by another device), source affordance 828b corresponds to screen input (e.g., tap input on display 802), source affordance 828c corresponds to camera input (e.g., recognition of a particular user gesture by a camera of device 600), source affordance 828d corresponds to back tap input (e.g., tap input on the back of the chassis of device 800), and source affordance 828e corresponds to sound (e.g., voiced, phoneme).
[0251] In some embodiments, the user may select source affordances 828a-828e and then select an input type for switching. For example, in response to selection of source affordances 828a-828e, device 600 displays an input selection interface. For example, when source selection interface 826 is displayed, device 600 detects selection of source affordance 828e. Fig. 8E In the example, the selection is a tap gesture 830 on the source affordance 828e. Figure 8F As shown in , in response to detecting tap gesture 830 , device 600 displays input selection interface 832 .
[0252] Input selection interface 832 includes multiple candidate sounds 834 (recall that selected source affordance 828e corresponds to a sound input), such as candidate sound 834a. When input selection interface 820 is displayed, device 600 detects selection of candidate sound 834a (corresponding to the "ah" sound). Figure 8F In the example, the selection is a tap gesture 835 on candidate sound 834a. Figure 8G As shown in , in response to detecting tap gesture 835 , device 600 displays action selection interface 836 .
[0253] Action selection interface 836 includes candidate actions 837, such as candidate action 837a. When action selection interface 836 is displayed, device 600 detects selection of candidate sound 837a (corresponding to the "select item" action). Figure 8G , the selection is a tap gesture 838 on candidate sound 837a.
[0254] In response to selecting tap gesture 838, device 600 provides candidate sound 834a ("ah") associated with candidate action 837a ("select item") to the switch. Thus, device 600 is configured such that in response to detecting input including the sound "ah", device 600 performs the action "select item". Further in response to detecting tap gesture 838, device 600 displays switch control interface 810. Figure 8H As shown in , indicator 814a is modified to reflect the updated number of switches configured on device 800.
[0255] exist Figure 8I In some embodiments, in response to selecting the exercise enable indication 839, the device 600 displays the sound practice interface. For example, when the input selection interface 820 is displayed, the device 600 detects the selection of the exercise enable indication 840. Figure 8I In the example, the selection is to practice the tap gesture 840 on the affordance 839. Figure 8JAs shown in , in response to detecting the tap gesture 840 , the device 600 displays a sound practice interface 841 .
[0256] The voice practice interface 840 includes a real-time preview 846 and candidate voices 848. In operation, when the practice interface 840 is displayed, the device 600 receives voice input from the user using an audio input device (e.g., a microphone) of the device 800. In some embodiments, when the voice input is received, the device 600 provides a real-time preview of the voice input, such as the real-time preview 846. Figure 8K As shown in , real-time preview 846 is a visual waveform indicating one or more auditory characteristics of the voice input. In some embodiments, device 600 prompts the user to provide voice input (e.g., "listen to the sound").
[0257] In some embodiments, the sound practice interface 840 can be used to help users practice the pronunciation of various sounds. Figure 8K When voice input 850 (“oo”) is received, device 600 determines whether voice input 850 includes a sound that corresponds to (e.g., matches) candidate sound 848. If it is determined that voice input 850 corresponds to candidate sound 848, device 600 indicates that a matching sound is received and optionally highlights the matching candidate sound. Figure 8L As shown in, for example, device 600 replaces the display of real-time preview 846 with an indicator 851 indicating that a match has been found ("Great"), and further bolds the matching sound (e.g., bolds candidate sound 848g corresponding to the sound "oo"). If it is determined that the voice input does not include a sound that matches candidate sound 848, device 800 optionally provides a notification that the voice input does not include a sound that matches the candidate sound.
[0258] In some embodiments, the user can select a specific candidate sound to practice the pronunciation of the candidate sound. Figure 8M , when the practice interface 840 is displayed, the device 600 detects the selection of the candidate sound 848a ("eh"), that is, the tap gesture 852 on the candidate sound input 848a. Figure 8N As shown in FIG. 8 , in response to detecting a tap gesture 852, device 600 displays a learning affordance 854, which may be selected to allow a user to practice pronunciation of candidate sound 848a, such as with respect to FIG. 8O to FIG. 8Q In some embodiments, selection of the candidate sound 848a causes the device 600 to audibly provide (eg, output) the candidate sound 848a.
[0259] refer to Figure 8N , when displaying practice sound interface 840, device 600 detects selection of learning affordance 854, i.e., tap gesture 856 on learning affordance 854. Fig.8OAs shown in , in response to detecting the tap gesture 856 , the device 600 displays a learning interface 860 of the candidate sounds 848 a .
[0260] Learning interface 860 includes real-time preview 862, sound indicator 864, sound description 866, and play affordance 868. Sound indicator 864 indicates the current sound of the learning affordance (e.g., the sound "eh" corresponding to selected candidate sound 848a). Selection of play affordance 868 causes device 600 to provide (e.g., output) the sound indicated by sound indicator 864 (e.g., "eh").
[0261] In some embodiments, the device 600 receives voice input from the user while displaying the learning interface 860. In some examples, when the voice input is received, the device 600 provides a real-time preview 862 of the voice input. In some embodiments, in response to receiving the voice input, the device 600 determines and indicates whether the voice input includes a sound that matches the sound of the learning interface 860.
[0262] In some embodiments, the learning interface 860 provides information about the ways in which one or more sounds can be made. The sound description 866 includes, for example, information describing the ways in which a user can make the "eh" sound. In some embodiments, the sound description 866 exceeds the display area of the device 800, and additional portions of the sound input description 866 are displayed in response to one or more inputs (e.g., a swipe gesture 860), such as Figure 8O to Figure 8P as shown in .
[0263] In some embodiments, the user may wish to learn a different voice than the current voice of the learning interface 860, and, for example, not necessarily use the practice voice interface 840 to select a different voice. Thus, in some embodiments, the device 600 changes the current voice of the learning interface 860 in response to an input (e.g., a swipe gesture). Figure 8Q , when displaying learning interface 860, device 600 detects a swipe gesture (e.g., a horizontal swipe gesture), such as swipe gesture 872. In response to swipe gesture 872, device 600 changes the current sound of learning interface 860 from "eh" to "E sound". In some embodiments, changing the sound in this manner includes replacing indicator 874 and sound description 876 with indicator 864 and sound description 866, respectively, such as Figure 8R Indicator 874 indicates that the current sound is "E sound", and sound description 876 describes the way in which the user can make the "E sound".
[0264] Figure 8S to Figure 8V FIG. 6 shows an exemplary operation of a device 600 using switching control. Figure 8S, device 600 displays home screen interface 880. Home screen interface 880 includes various application enable representations 882 (e.g., application enable representations 882a-882c), which, when selected, cause device 600 to execute (e.g., initiate execution, continue execution) an application corresponding to the selected application enable representation. Figure 8S As shown in , the focus of the home screen interface 880 is located on the application enable representation 882a, as indicated by the bold surrounding the application enable representation 882a.
[0265] When the device 600 displays the home screen interface 880 with the focus on the application enable representation 882a, the device 600 receives the voice input 884 including the sound "oo". Figure 8T As shown in , in response to receiving voice input 884, device 600 moves the focus of home screen interface 880 from application enable representation 882a to application enable representation 882b (recall that device 600 switches the sound "oo" to correspond to the "move right" action).
[0266] Thereafter, device 600 detects an input on display 802, such as a tap input 886. Figure 8U As shown in FIG. 8 , in response to receiving tap input 886, device 600 moves the focus of home screen interface 880 from application enable representation 882 b to application enable representation 882 c (recall that in FIG. 8 Fig.8D , switching of device 600 will correspond the tap input to the "move right" action).
[0267] Thereafter, device 600 receives input, such as voice input 888 ("eh"). Figure 8V As shown in FIG. 8 , in response to receiving voice input 888, device 600 moves the focus of home screen interface 880 from application enable representation 882c to application enable representation 882b (recall that in FIG. Fig.8D , the switching of device 600 corresponds the sound "eh" to the action "move left").
[0268] Fig. 9 900 is a flow chart illustrating a method for providing auditory control using a computer system according to some embodiments. Method 900 is performed at a computer system (e.g., 100, 300, 500, 600) in communication with a display generation component and one or more input devices. Some operations in method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0269] As described below, method 900 provides an intuitive way to provide auditory control. The method reduces the cognitive burden on the user of providing auditory control, thereby creating a more efficient human-computer interface. For battery-powered computing devices, enabling users to provide auditory control more quickly and more efficiently saves power and increases the time interval between battery charges.
[0270] When a user interface (e.g., 880) including a group of user interface objects (e.g., 882) is displayed via a display generation component (e.g., 602), a computer system (e.g., 800) receives (902) a first voice input (e.g., 884) associated with a first predetermined action via one or more input devices. In some embodiments, the computer system displays a user interface including the group of user interface objects. In some embodiments, the group of user interface objects includes one or more affordances and / or other objects that can be used to navigate the computer system and / or interact with the computer system. In some embodiments, the voice input is a phrase or word. In some embodiments, the voice input is a recognized voice sound, and is optionally voiced ("eh", "ah"). In some embodiments, a user can select a sound from a list of candidate predetermined sounds to associate the selected sound with a particular action. In some embodiments, other types of input (non-voice input, such as touch input or gesture input) may also be associated with an action. In some embodiments, the predetermined actions include actions for navigating the computer system and / or interacting with the computer system; for example, the predetermined actions include actions for navigating a user interface displayed by the computing system, such as "select", "move to next", "move to previous", "cancel selection"; for another example, the predetermined actions include actions for controlling various system features, such as "increase volume", "decrease volume", or "go to settings".
[0271] In response to (904) receiving the first voice input and based on determining that a first user interface object (e.g., 882a) in the group of user interface objects is currently selected (e.g., highlighted, in focus), the computer system performs (906) a first predetermined action based on the first user interface object. In some embodiments, only a single user interface object in the group of user interface objects is selected at any given time. In some embodiments, performing an action based on the user interface object includes performing a predetermined action based on the relative position of the object in the user interface; if, for example, the predetermined action is "move to next", the computing system deselects the first user interface object and selects the user interface object determined to be "next". In some embodiments, performing an action based on the user interface object includes performing an action on the first user interface object (e.g., deleting the object in response to a "delete" request).
[0272] In response to receiving the first voice input (904) and based on determining that a second user interface object (e.g., 882b) in the set of user interface objects is currently selected (and optionally, the first user interface object is not selected), the computer system performs (908) a first predetermined action based on the second user interface object. In some embodiments, the first object is not considered based on the first object. Performing the predetermined action based on the first user interface object and performing the predetermined action based on the second user interface object allow the user to quickly and efficiently control the manner in which the computer system operates, which reduces the number of inputs required to perform the operation.
[0273] In some embodiments, the voice input (e.g., 884) is a first sound (e.g., a sound corresponding to 848g) of a plurality of predetermined sounds (e.g., sounds corresponding to 848a-848h). In some embodiments, the voice input is a phrase or word. In some embodiments, the voice input is a recognized sound and / or may be a voiced sound ("eh", "ah"). In some embodiments, the sound is selected from a list of candidate predetermined sounds and associated with a particular action such that subsequent detection of the sound causes the computer system to perform the particular action.
[0274] In some embodiments, the plurality of predetermined sounds include a second sound (e.g., 848a) associated with a second predetermined action different from the first predetermined action. In some embodiments, when the user interface (e.g., 880) is displayed via the display generation component, the computer system receives a second voice input (e.g., 888) determined to include (e.g., determined to be) a second sound (e.g., a sound corresponding to 848a) via one or more input devices. In some embodiments, the voice input is a phrase or word. In some embodiments, the voice input is a recognized voice sound and can be a voiced sound ("eh", "ah"). In some embodiments, the user selects a sound from a list of candidate predetermined sounds to associate the selected sound with a specific action. In some embodiments, other types of input (non-voice input, such as touch input or gesture input) are also associated with the action. In some embodiments, the predetermined action includes an action for navigating the computer system and / or interacting with the computer system; for example, the predetermined action includes an action for navigating the user interface displayed by the computing system, such as "select", "move to next", "move to previous", "cancel selection"; for example, the predetermined action includes an action for controlling various system features, such as "increase volume", "decrease volume" or "go to settings".
[0275] In some embodiments, in response to the second voice input and based on determining that the first user interface object in the group of user interface objects is currently selected, the computer system performs a second predetermined action based on the first user interface object. In some embodiments, performing an action based on the user interface object includes performing a predetermined action based on the relative position of the object in the user interface; if, for example, the predetermined action is "move to next", the computing system deselects the first user interface object and selects the user interface object determined to be "next". In some embodiments, performing an action based on the user interface object includes performing an action on the first user interface object (e.g., deleting the object in response to a "delete" request).
[0276] In some embodiments, in response to receiving the second voice input and based on determining that a second user interface object in the set of user interface objects is currently selected (in some embodiments, and the first user interface object is not selected), the computer system performs a second predetermined action based on the second user interface object. In some embodiments, the first object is not considered based on the first object. Performing the predetermined action based on the first user interface object and performing the predetermined action based on the second user interface object enable further improved control over the manner in which the computer system operates, which reduces the number of inputs required to perform the operation.
[0277] In some embodiments, the computer system displays a second user interface (e.g., 820) (e.g., a switching interface including a switching affordance representation showing an association between an input type and an action) via a display generation component, the second user interface including a first action graphical user interface object (e.g., 821a) indicating an association between a first sound and a first predetermined action. In some embodiments, the switching interface includes one or more indicators showing an association between a sound type input and an action. In some embodiments, the first action graphical user interface object is a user interactive object (e.g., affordance representation) that can be selected to modify one or more settings associated with the first sound and / or the first predetermined action.
[0278] In some embodiments, the computer system displays a second user interface (e.g., 820) (e.g., a switch interface including a switch affordance showing the association between the input type and the action) via a display generation component, the second user interface including a second action graphical user interface object (e.g., 821c) indicating the association between the second sound and the second predetermined action. In some embodiments, the computer system displays an affordance interface including each "switch" configured on the computer system. In some embodiments, a switch represents a pairing of a specific type of input and an action that can be performed by the computer system. In some embodiments, each switch corresponds an input type to an action, so that providing an input type of input causes the computer system to perform a corresponding action. In some embodiments, the same input type cannot correspond to multiple switches. In some embodiments, the input assigned to the action is selected from a predetermined action set. In some embodiments, the predetermined action set includes a predetermined sound set. In some embodiments, once the input type is assigned to the action, the number of switches in the switch control interface is updated to reflect the current total number of switches identified by the computer system. Displaying a user interface including multiple action graphical user interface objects indicating the association between the sound and the predetermined action enables the user to quickly and effectively observe the association between the sound and the predetermined action, which provides improved visual feedback.
[0279] In some embodiments, when a user interface (e.g., 880) is displayed via a display generation component, the computer system receives a first non-voice input (e.g., 886) associated with a first predetermined action via one or more input devices. In some embodiments, non-voice input is any user input that does not include voice input. In some embodiments, non-voice input includes external input (e.g., input received from other computer systems and / or devices), screen input (e.g., taps on display generation components), camera input (e.g., detection of predetermined user movement by a camera of a computer system), and tap input on a predetermined portion of a computer system (e.g., tap input on the back of a computer system). In some embodiments, the predetermined action includes actions for navigating a computer system and / or interacting with a computer system; for example, the predetermined action includes actions for navigating a user interface displayed by a computing system, such as "select", "move to next", "move to previous", "cancel selection"; for example, the predetermined action includes actions for controlling various system features, such as "increase volume", "decrease volume" or "go to settings".
[0280] In some embodiments, in response to receiving a first non-voice input and based on determining that a first user interface object in the group of user interface objects is currently selected (e.g., 882b), the computer system performs a first predetermined action based on the first user interface object.
[0281] In some embodiments, in response to receiving the first non-voice input and based on determining that a second user interface object in the group of user interface objects is currently selected (e.g., 882c), the computer system performs a first predetermined action based on the second user interface object. Performing predetermined actions in response to voice input and performing predetermined actions in response to non-voice input allows for improved control over the manner in which the computer system operates, which reduces the number of inputs required to perform operations.
[0282] In some embodiments, the second user interface (e.g., 820) also includes a third action user interactive graphical user interface object (e.g., 821b) indicating an association between the first non-voice input and the first predetermined action. In some embodiments, voice input and non-voice input may be associated with the same action. Displaying a user interface including an action user interactive graphical user interface object indicating an association between a non-voice input and a predetermined action enables a user to quickly and efficiently observe the association between a non-voice input type and a corresponding action, which provides improved visual feedback.
[0283] In some embodiments, the computer system displays a third user interface (e.g., 840) (e.g., a practice sound interface that can be used to practice the pronunciation of various predetermined voice inputs), which includes a first group of practice sound user interactive graphical user interface objects (e.g., 848a-848h) (e.g., a group of affordances). In some embodiments, the practice sound interface prompts the user to provide a voice input corresponding to any one of a group of predetermined sounds. In some embodiments, the practice sound user interactive graphical user interface object includes a first practice sound user interactive graphical user interface object (e.g., 848g) (e.g., affordance; an object associated with a first sound that can be associated with a predetermined action) (e.g., affordance including a text representation of a voice input associated with a third sound (e.g., "eh")). In some embodiments, the affordance is selected to play the sound corresponding to the affordance. In some embodiments, selecting the affordance is to place the focus on the affordance. In some embodiments, after placing the focus on the affordance, the computer system displays a learning affordance that can be selected to access a submenu for voice input. In some embodiments, the submenu includes instructions for pronouncing the speech input and allows further practice of the sound. In some embodiments, the third sound is the same as the first sound. In some embodiments, the practice sound user interactive graphic the set of user interface objects includes a second practice sound user interactive graphic user interface object (e.g., 848b) associated with the fourth sound.
[0284] In some embodiments, when the third user interface is displayed, the computer system receives a first user input via one or more input devices (e.g., 850). In some embodiments, the user input is a voice input. In some embodiments, when the computer system receives the voice input, the computer system provides a real-time preview of the voice input. In some embodiments, the real-time preview is a dynamic waveform provided based on the auditory characteristics of the voice input. In some embodiments, the user input is a non-voice input, such as a tap input or a swipe input.
[0285] In some embodiments, in response to receiving a first user input (e.g., 850, 856) and based on determining that the first user input includes voice input, based on determining that the first user input includes a third sound, the computer system provides a first notification that the first user input corresponds to the third sound. In some embodiments, the first notification includes modifying a visual characteristic (e.g., highlighting) of a first practice sound user-interactive graphical user interface object (e.g., 848g) (e.g., if the received voice input matches the sound of the practice sound enable representation, the computer system highlights the practice sound enable representation). In some embodiments, highlighting the enable representation includes visually modifying the enable representation; for example, the computer system bolds the edges of the enable representation.
[0286] In some embodiments, in response to receiving a first user input (e.g., 850, 856) and based on determining that the first user input includes a voice input and based on determining that the first user input includes a fourth sound, the computer system provides a second notification that the first user input corresponds to the fourth sound. In some embodiments, the second notification includes modifying the visual characteristics (e.g., highlighting) of the second practice sound user interactive graphical user interface object (e.g., 848g). In some embodiments, if the voice input does not match any member of the first set of practice sound user interactive graphical user interface objects, no notification is provided (e.g., if the received voice input does not match the sound of the practice sound enabling representation, the computer system does not highlight the enabling representation). In some embodiments, the practice sound interface includes multiple practice sound enabling representations and if the voice input matches the voice input of different practice sound enabling representations, the computer system highlights the matching practice sound enabling representation.
[0287] In some embodiments, in response to receiving a first user input (e.g., 850, 856) and based on determining that the first user input does not include voice input (e.g., is non-voice input (e.g., contact on a touch-sensitive surface (e.g., tap); mouse click; key)), based on determining that the first user input corresponds to a first practice sound user-interactive graphical user interface object (e.g., 848a), the computer system outputs (e.g., via one or more speakers) a third sound. In some embodiments, in response to receiving a first user input and based on determining that the first user input does not include voice input (e.g., is non-voice input (e.g., contact on a touch-sensitive surface (e.g., tap); mouse click; key)), and based on determining that the first user input corresponds to a second practice sound user-interactive graphical user interface object, the computer system outputs (e.g., plays via one or more speakers) a fourth sound. In some embodiments, if the computer system detects a selection of an input (e.g., a tap input) to select a practice sound affordance, the computer system provides (e.g., plays) a sound output using an audio output device of the computer system. In some embodiments, the computer system also displays a learning affordance in response to selection of the practice sound affordance. Modifying visual characteristics of a user-interactive graphical user interface object in response to voice input corresponding to a sound associated with the user-interactive graphical user interface object allows a user to quickly and efficiently identify whether the voice input was provided correctly, which provides improved visual feedback.
[0288] In some embodiments, the third user interface (e.g., 840) (e.g., an exercise sound interface) includes a learning user-interactive graphical user interface object associated with the third sound (e.g., 854) (e.g., displayed after selecting the first exercise sound user-interactive graphical user interface object). In some embodiments, the learning affordance is displayed in response to selection of the exercise sound affordance.
[0289] In some embodiments, based on determining that the first user input corresponds to a selection (e.g., 856) of a learning user interactive graphical user interface object (e.g., 854), the computer system displays a fourth user interface (e.g., 860) (e.g., a learning interface for the sound selected in the practice sound interface), the fourth user interface including instructions for providing a voice input corresponding to a third sound (e.g., 866, 876). In some embodiments, the learning interface includes an instruction set indicating the manner in which the user can make the selected sound. In some embodiments, the learning interface includes an affordance for playing a sound corresponding to the selected sound.
[0290] In some embodiments, when displaying a fourth user interface (e.g., 860) (e.g., a learning interface corresponding to a selected sound), the computer system receives a third voice input. In some embodiments, in response to receiving the third voice input and according to determining that the third voice input corresponds to a third sound, the computer system provides a third notification that the third voice input corresponds to the third sound. In some embodiments, if the computer system receives a voice input of a sound matching the learning interface, the computer system provides a notification indicating that a match has been detected. In some embodiments, the notification is a displayed indicator, such as a check mark or a word ("Great"). In some embodiments, in response to receiving the third voice input and according to determining that the third voice input does not correspond to the third sound, the computer system abandons providing the third notification. In some embodiments, in response to the third voice input and according to determining that the third voice input does not correspond to the third sound, the computer system provides a fourth notification that the first voice input does not correspond to the third sound. In some embodiments, if the computer system receives a voice input of a sound that does not match the learning interface, the computer system provides a notification indicating that the voice input is improper.
[0291] In some embodiments, when the fourth user interface (e.g., 860) is displayed, the computer system receives a second user input (e.g., 870, 872). In some embodiments, the second user input is a swipe gesture. In some embodiments, the computer system detects a swipe gesture on the learning interface.
[0292] In some embodiments, in response to user input, the computer system replaces the instructions (e.g., 866, 876) for providing voice input corresponding to the third sound with instructions for providing voice input corresponding to the fifth sound. In some embodiments, in response to a swipe gesture, the computer system switches the learning interface to a different sound. In some embodiments, this function allows the user to navigate between various sounds (and view the instructions for making each sound) when the learning interface is displayed. In some embodiments, the instructions of the learning interface are not fully displayed within the display portion of the learning interface, so that some of the instructions are "hidden". In some embodiments, the computer system detects a scrolling input (e.g., a vertical scrolling input) and displays at least a portion of the hidden instructions in response to the scrolling input.
[0293] Note that the above description is relative to method 900 (eg, Fig. 9) also apply in a similar manner to the methods described below / above. For example, methods 700, 1100 optionally include one or more of the features of the various methods described above with reference to method 1100. For example, one or more sound actions configured by a user (e.g., as described with reference to method 900) can be used to register sounds, as described with reference to method 1100. For the sake of brevity, these details are not repeated below.
[0294] Figure 10A to Figure 10V An exemplary user interface for providing notifications according to some embodiments is shown. The user interface in some figures is used to illustrate the processes described below, which include Fig.11 process.
[0295] In some implementations, the exemplary user interface can be used to configure the device to provide a notification in response to the detection of a particular sound, such as an alarm. Thus, a user, such as a user with limited hearing, can receive a notification about the occurrence of a sound, even if the user would not otherwise hear such a sound.
[0296] exist Fig. 10A , device 600 displays auxiliary function interface 1004 on display 602. Accessibility interface 804 includes various affordances for controlling various auxiliary function functions of device 1000, including voice recognition affordance 1006.
[0297] In some examples, in response to selecting voice recognition affordance 1006, device 600 displays a voice recognition interface. For example, when auxiliary function interface 1004 is displayed, device 600 detects selection of voice recognition affordance 1006, i.e., tap gesture 1008 on voice recognition affordance 1006. Fig. 10B As shown in , in response to detecting the tap gesture 1008 , the electronic device 600 displays a sound recognition interface 1010 .
[0298] The voice recognition interface 1010 includes an option 1012. In some embodiments, the option 1012 is used to switch the use of a voice recognition on the device 1000. As will be described in detail below, in some embodiments, the device 600 can be configured to detect (e.g., recognize) one or more specific sounds and, in response to detecting the sound, provide a notification alerting the detection of the sound.
[0299] In response to selection of option 1012, device 600 switches the state of option 1012 (e.g., deactivates the option if activated, activates the option if deactivated). Fig. 10BIn response to detecting a tap gesture 1014 at option 1012, device 600 activates voice recognition on device 600 and modifies the display of option 1012 to indicate that voice recognition has been activated, such as Fig. 10C as shown in .
[0300] In further response to the selection of option 1012 (e.g., in response to device 600 activating voice recognition), device 600 displays voice setting enable indication 1016 in voice recognition interface 1010. In response to the selection of voice setting enable indication 1016, device 600 displays the voice setting interface. For example, when voice recognition interface 1010 is displayed, device 600 detects the selection of voice setting enable indication 1016, i.e., tap gesture 1018 on voice setting enable indication 1016. Fig. 10D As shown in , in response to detecting the tap gesture 1018 , the electronic device 600 displays a sound setting interface 1020 .
[0301] Sound settings interface 1020 includes various candidate sound affordances, each of which corresponds to a sound type. In some embodiments, the candidate sound affordances can be used to configure device 600 to detect sounds corresponding to the candidate sound affordances. For example, a "doorbell" candidate sound affordance can be used to configure device 600 to detect a "doorbell" sound.
[0302] However, in some cases, the user may wish to configure (e.g., train) device 600 to detect custom sounds (e.g., sounds that do not correspond to affordances in sound settings interface 1020). Accordingly, sound settings interface 1020 includes custom sound affordances 1022. In some embodiments, in response to selection of custom sound affordances 1022, device 600 initiates a sound enrollment process in which device 600 is configured to detect custom sounds.
[0303] Generally speaking, during the sound registration process, the device 600 receives a group of sound inputs corresponding to the custom sound, and obtains (e.g., generates) a model that can be used to detect the custom sound based on the group of sound inputs. As will be described in detail below, during the sound registration process, the device 600 continuously receives each sound input in the group of sound inputs. For each sound input, the device 600 determines whether the sound input meets the sound input standard. Based on the determination of each sound input, the device 600 determines whether the group of inputs meets the sound registration standard. If so, the device 600 obtains the model.
[0304] exist Fig. 10D In the example, when sound setting interface 1020 is displayed, device 600 detects selection of custom sound enabling representation 1022, i.e., tap gesture 1024 on custom sound enabling representation 1022. Fig.10E As shown in , in response to detecting tap gesture 1024 , device 600 displays sound tag interface 1026 .
[0305] Sound label interface 1026 includes a name field 1028 (which can be used to enter the name of the custom sound) and a continue enable indication 1030. In response to the selection of the continue enable indication 1016, device 600 displays the sound registration interface. Fig. 10E In the example, when the sound recognition interface 1026 is displayed, the device 600 detects the selection of the continuation enable indication 1030, that is, the tap gesture 1032 on the continuation enable indication 1030. Fig.10F As shown in , in response to detecting tap gesture 1032 , device 600 displays sound registration interface 1034 .
[0306] The sound registration interface 1034 includes an indicator 1036, an instruction 1038, a progress area 1040, a real-time preview 1042, a status indicator 1044, and a postponement affordance 1046. The indicator 1034 indicates the name of the custom sound entered in the name field 1028. The instruction 1036 includes instructions for performing the sound registration process. As shown, the instruction 1038 includes instructions for providing a predetermined number (e.g., five) of sound inputs including the custom sound. The status indicator 1044 indicates the state of the device 600 during the sound registration process. The status indicator 1044 indicates that the device 600 is waiting to receive the sound input.
[0307] In operation, when the voice registration interface 1034 is displayed, the device 600 receives voice input using an audio input device (e.g., a microphone) of the device 1000. In some embodiments, when the voice input is received, the device 600 provides a real-time preview of the custom voice input, such as the real-time preview 1042. Figure 10G As shown in , real-time preview 1042 is a visual waveform indicating one or more auditory characteristics of the received sound input. In some examples, when the sound input is received, device 600 stops displaying status indicator 1044.
[0308] After receiving the voice input, the device 600 determines (e.g., analyzes the voice input to determine) whether the voice input meets the voice input criteria. In some embodiments, the device 600 provides an indication when determining whether the voice input meets the voice input criteria. Fig. 10H As shown in , for example, when determining whether the sound input satisfies the sound input standard, the device 600 modifies the display of the progress indicator 1040a of the progress area 1040.
[0309] If the device 600 determines that the sound input meets the sound input criteria, the device 600 saves (e.g., stores) the sound input and modifies the display of the progress indicator 1040a to indicate that the sound input meets the sound input criteria (e.g., the device 600 modifies the progress indicator 1040a to a check mark). Optionally, the device 600 replaces the real-time preview 1042 with a check mark 1050 to indicate that the custom sound input meets the sound input criteria.
[0310] Thereafter, device 600 iteratively receives the remaining sound inputs at device 1000 until all sound inputs necessary to complete the sound enrollment process have been received. As described, for each received sound input, device 600 determines whether the sound input meets the sound input criteria and indicates whether the sound input meets the sound input criteria. Fig.10J An example is shown in which device 600 has determined (and indicated) that all sound inputs satisfy the sound input criteria.
[0311] In some embodiments, the sound input criteria include criteria that are met when the sound input includes a specific type of sound. For example, when the sound input includes a sound corresponding to an alarm or a sound generated by an electronic device, the sound input meets the sound input criteria. In this way, the device 600 verifies that the sound is of a type that can be detected by the device 600 (e.g., the sound has sufficient discrimination). In some embodiments, the sound input criteria include criteria that are met when the sound input is similar enough to other sound inputs (e.g., the sound input includes sufficiently similar sounds). In this way, the device 600 verifies that the changes between the sound inputs are small enough so that a reliable model can be generated using the sound input.
[0312] Once the sound enrollment process is successfully completed, the device 600 determines whether the set of inputs meets the sound enrollment criteria. In some embodiments, the sound enrollment criteria include a criterion that is met when a threshold number of sound inputs in the set of sound inputs meet the sound input criteria.
[0313] If the set of inputs meets the sound registration criteria, the device 600 is configured to detect the custom sound. For example, if the set of inputs meets the sound registration criteria, the device 600 obtains a model for detecting the custom sound to be generated based on the set of inputs.
[0314] In some embodiments, the device 600 further adds a candidate sound enable representation 1048 corresponding to a custom sound (e.g., a toaster) to the sound setting interface 1020, such as Figure 10K In some embodiments, the candidate sound affordances 1048 are used to configure one or more features of a custom sound.
[0315] For example, when sound settings interface 1020 is displayed, device 600 detects selection of candidate sound enabling indication 1048, that is, tap gesture 1050 on candidate sound enabling indication 1048. Fig.10L As shown in , in response to detecting tap gesture 1050 , device 600 displays sound configuration interface 1052 .
[0316] Sound configuration interface 1052 includes options 1054, audio enable representation 1056, model indicator 1058 and return enable representation 1062. Audio enable representation 1056 indicates the audio output provided by device 600 in response to detecting a custom sound (e.g., "tritone"), and can further be used (e.g., selected to) select different audio outputs. For example, in response to the selection of audio enable representation 1056, device 600 displays an audio selection interface (not shown) in which the user can select an audio output.
[0317] In some embodiments, the sound configuration interface 1052 is displayed while the device 600 is obtaining a model of a custom sound. Fig.10L , model indicator 1058 indicates that device 600 is in the process of obtaining a model of a custom sound. In some embodiments, once device 1000 obtains the model, device 600 removes the display of model indicator 1058, such as Figure 10M as shown in .
[0318] In some embodiments, in response to selection of option 1054, device 600 switches voice recognition to a custom voice. Figure 10M In response to detecting a tap gesture 1060 at option 1054, device 600 activates voice recognition for custom voice input on device 600 and modifies the display of option 1054 to indicate that voice recognition for custom voice input has been activated, such as Fig.10N as shown in .
[0319] Selecting return affordance 1062 displays (eg, restores display of) sound settings interface 1020. When sound input configuration interface 1052 is displayed, device 600 detects selection of return affordance 1062, i.e., tap gesture 1064 on return affordance 1064. Fig.10O In response to detecting tap gesture 1064, device 600 displays sound settings interface 1020. As shown, because sound recognition for custom sounds has been activated, candidate sound affordances 1048 indicate that sound recognition for custom sounds is activated (e.g., “on”).
[0320] Once voice recognition for a custom sound is activated (e.g., by performing a voice registration process for a custom sound and / or activating an option for a custom sound on device 600), device 600 may detect subsequent occurrences of the custom sound and provide a notification indicating the custom sound in response to detecting the custom sound. Figure 10P As shown in , device 600 displays home screen interface 1070. Home screen interface 1070 includes various application affordances, which, when selected, cause device 600 to execute (e.g., initiate execution, continue execution) an application corresponding to the selected application affordance. While home screen 1070 is displayed, device 600 detects the occurrence of a custom sound. In response, device 600 displays notification 1072 and optionally provides audio output 1074 (e.g., "tritone"), thereby indicating that a custom sound has been detected (e.g., "toaster detected").
[0321] In some cases, the voice registration process may not be completed successfully. For example, the user may choose to complete the voice registration process at a later time. Figure 10Q In response to the selection of the postponement indication 1046 of the sound registration interface 1034, the device 600 saves (e.g., stores) the current state of the sound registration process and terminates the sound registration process. Thereafter, the sound registration process may be resumed. Figure 10R As shown in , the sound setting interface 1020 includes a resume enable indication 1080. The indicator 1082 of the resume enable indication 1080 indicates the number of sound inputs provided during the sound registration process, and optionally indicates the number of sound inputs remaining in the sound registration process. In some embodiments, in response to the selection of the resume enable indication 1080, the device 600 resumes the sound registration process.
[0322] As another example, during the voice registration process, the device 600 may determine that the voice input does not meet the voice input criteria. Fig. 10S As shown in , in the case where the sound input does not meet the sound input standard, the device 600 displays the sound registration interface 1034. In response to determining that the sound input does not meet the sound input standard, the device 600 modifies the display of the progress indicator 1040a to indicate that the sound input has failed to meet the sound input standard (for example, the progress indicator 1040a is changed to an exclamation mark), and optionally displays a failure indicator 1090 to indicate the situation. Further in response to determining that the sound input does not meet the sound input standard, the device 600 displays a restart enable indication 1082, a continue enable indication 1084, and a learning enable indication 1086. In response to the selection of the restart enable indication 1082, the device 600 restarts the sound registration process. In response to the selection of the continue enable indication 1084, the device 600 proceeds to the next iteration of the sound registration process.
[0323] In response to selection of learn affordance 1086, device 600 displays a compatible voice interface. Fig. 10S , when sound enrollment interface 1034 is displayed, device 600 detects selection of learning affordance 1086, i.e., tap gesture 1088 on learning affordance 1086. Figure 10T As shown in , in response to detecting tap gesture 1088 , device 600 displays a compatible sound interface 1090 .
[0324] The compatible sound interface includes information 1092 and a continue enable indication 1094. Information 1092 includes information about which sounds are suitable for recognition by device 1000 and which sounds meet the sound input criteria accordingly. In response to the selection of the continue enable indication 1094, device 600 displays (e.g., resumes displaying) the sound registration interface 1034. For example, when displaying compatible sound interface 1090, device 600 detects the selection of the continue enable indication 1094. In some examples, the selection is a tap gesture 1096 on the continue enable indication 1094. In response to detecting the tap gesture 1096, device 600 displays the sound registration interface 1034.
[0325] In some embodiments, one or more sound affordances may be removed (e.g., deleted) from the sound settings interface. Figure 10U As shown in, for example, device 600 displays a sound settings interface 1020 including custom sound enable representation 1048 and edit enable representation 1098. In response to the selection of edit enable representation 1098, device 600 displays a delete enable representation for the custom sound input enable representation. For example, when displaying sound settings interface 1020, device 600 detects the selection of edit enable representation 1098. In some examples, the selection is a tap gesture 1002a on edit enable representation 1098. Figure 10V As shown in , in response to detecting tap gesture 1002a, device 600 displays delete enable indication 1004a. In response to the selection of delete enable indication 1004a, device 600 removes custom sound enable indication 1048 from sound settings interface 1020, and is optionally configured so that device 600 no longer detects the custom sound and / or no longer provides notification about the detection of the custom sound.
[0326] Fig.111 is a flowchart illustrating a method for providing notifications using a computer system according to some embodiments. Method 1100 is performed at a computer system (e.g., a smart phone, a tablet computer, a personal computer) (e.g., 100, 300, 500, 600) that communicates with a display generation component (e.g., a television, a display controller, an internal or external touch-sensitive display system) and one or more input devices (e.g., a touch-sensitive surface, a hardware button, a microphone). Some operations in method 1100 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.
[0327] As described below, method 1100 provides an intuitive way to provide notifications. The method reduces the cognitive burden of providing notifications on the user, thereby creating a more effective human-computer interface. For battery-powered computing devices, enabling users to provide notifications faster and more efficiently saves power and increases the time interval between battery charges.
[0328] The computer system (e.g., 600) performs (1102) a sound registration process. In some embodiments, the sound registration process is a process in which the computing system learns to recognize a particular sound. In some embodiments, before performing the sound registration process, the computer system displays a label interface that can be used to name the sound to be registered.
[0329] When performing the sound registration process, a group of one or more sound inputs including a first sound input (e.g., an input represented by 1042) is received via one or more input devices. In some embodiments, during the sound registration process, one or more sound inputs are provided to the computing system by a user. In some embodiments, during the sound registration process, the computer system displays a sound registration interface that tracks the progress of the sound registration process. In some embodiments, the sound registration interface includes a real-time preview of the sound input received by the computer system. In some embodiments, the real-time preview is a dynamic waveform provided based on the auditory characteristics of the voice input. In some embodiments, the sound registration interface includes a postponed enable indication that can be selected to save the state of the sound registration process so that the sound registration process can be restored and / or completed at a later time. In some embodiments, the sound registration process can be restored by selecting the enable indication of the sound settings interface. In some embodiments, the enable indication of the sound settings interface indicates the degree of completion of the sound registration process (e.g., a score indicating the number of completed samples).
[0330] When the sound registration process is performed (in some embodiments, and in response to receiving the group of one or more sound inputs), it is indicated whether the first sound input meets the sound input standard (e.g., using indicator 1040a). In some embodiments, the sound input standard includes the requirement that the sound input is a non-verbal sound. In some embodiments, the sound input standard includes the requirement that the sound input is a sound and / or alarm provided by an electronic device (e.g., a device separate and / or different from the computing system). In some embodiments, the sound input standard includes the requirement that each sound input is a sound input of the same type and / or sufficiently similar (e.g., each sound input is the same alarm). In some embodiments, indicating whether the sound input meets the sound input standard includes providing an output indicating the situation. In some embodiments, the computing system indicates whether the sound input meets the sound input standard by providing a visual indicator (e.g., a green check mark if the sound input meets the standard, or a red exclamation mark if the sound input does not meet the standard). In some embodiments, an indication is provided for each sound in the group of one or more sound inputs.
[0331] While performing the sound enrollment process and based on determining that the group of one or more sound inputs meets the sound enrollment criteria set (in some embodiments, and in response to receiving the group of one or more sound inputs), the computer system causes (1104) to generate a model (e.g., a classification model) for identifying a first type of sound corresponding to the group of one or more sound inputs. In some embodiments, the sound enrollment criteria set includes criteria that are met when each sound input in the group of one or more sound inputs satisfies the sound input criteria. In some embodiments, even if each sound input in the group of sound inputs does not meet the sound input criteria, the sound enrollment criteria set is met. In some embodiments, if each sound input meets the sound input criteria, the computing system generates a model for identifying whether future sound inputs match the sound inputs provided during the sound enrollment process. In some embodiments, if a match is detected, the computing system provides a notification indicating that a match has been detected. In some embodiments, the model is a machine learning model, such as a classification network (e.g., a neural network). In some embodiments, the generated model is adjusted to take into account auditory hallucinations. In some embodiments, the model is generated by the computer system. In some embodiments, after the computer system sends data corresponding to the group of one or more sound inputs, the model is generated by one or more remote computer systems in communication with the computer system.
[0332] While performing the sound enrollment process and based on determining that the set of one or more sound inputs does not satisfy the sound enrollment criteria set (in some embodiments, and in response to receiving the set of one or more sound inputs), the computer system abandons (1106) causing generation of a model for identifying a first type of sound corresponding to the set of one or more sound inputs. Indicating whether the sound input satisfies the sound input criteria allows a user to quickly and efficiently identify whether the sound input is usable for generating a model for identifying a sound, which provides improved visual feedback.
[0333] In some embodiments, the computer system determines whether the group of one or more sound inputs meets the sound registration standard set. In some embodiments, the computer system determines whether the group of one or more sound inputs meets the sound registration standard set. In some embodiments, the computer system provides each one or more sound inputs to an external server, which in turn determines whether the group of one or more sound inputs meets the sound registration standard set. In some embodiments, the computer system displays a sound registration interface that tracks the progress of the sound registration process. In some embodiments, when determining whether the group of one or more sound inputs meets the sound registration standard set, the computer system provides (e.g., displays) a notification indicating that the computer system is determining whether the group of one or more sound inputs meets the sound registration standard set. In some embodiments, the notification is an animation indicating that the computer system is analyzing one or more sound inputs.
[0334] In some embodiments, the sound input criteria include criteria that are met when the first sound input is determined to be a non-verbal sound (e.g., as indicated at 1092). In some embodiments, the sound registration criteria include a requirement that one or more sound inputs are non-verbal sounds, such as an alarm provided by an electronic device. Including criteria that are met when the first sound input is determined to be a non-verbal sound in the sound input criteria reduces the complexity of configuring sound recognition, and thus allows a user to quickly and efficiently configure a computer system to recognize non-verbal sounds, such as an alarm, which reduces the number of inputs required to perform an operation.
[0335] In some embodiments, the sound registration criteria set includes criteria that are met when each sound input in the group of one or more sound inputs is determined to be a first type of sound input. In some embodiments, the sound registration criteria include a requirement that each sound input is of the same or similar type (e.g., an alarm provided by the same device). In some embodiments, the computer system analyzes each sound input to determine whether the sound inputs have sufficient similarity. Including in the sound registration criteria criteria that are met when each sound input in the group of one or more sound inputs is determined to be the same type of sound input improves the accuracy and efficacy of sound recognition, allowing for more reliable and efficient notification of detected sounds, which provides improved visual and / or auditory feedback.
[0336] In some embodiments, the group of one or more sound inputs includes a second sound input. In some embodiments, the computer system indicates whether the second sound input meets the sound input criteria. In some embodiments, during the sound registration process, the computer system provides an indication (e.g., a notification) for each sound input indicating whether the sound input meets the sound registration criteria. In some embodiments, indicating whether the first sound input meets the sound registration criteria includes displaying a notification (e.g., 1040a) indicating that the first sound input meets the sound input criteria via a display generation component based on determining that the first sound input meets the sound input criteria. In some embodiments, the notification is an agreement or affirmative notification, such as a check mark. In some embodiments, the notification is displayed with a color associated with a successful result (e.g., green).
[0337] In some embodiments, indicating whether the first sound input meets the sound registration standard includes displaying a notification (e.g., 1040a) indicating that the first sound input does not meet the sound input standard via a display generation component based on determining that the first sound input does not meet the sound input standard. In some embodiments, the notification is a negative notification, such as an exclamation mark. In some embodiments, the notification is displayed with a color associated with an unsuccessful result (e.g., red). In some embodiments, the notification is a text string (e.g., "hearing a sound but cannot be recognized as an alarm"). In some embodiments, in response to determining that the sound input does not meet the sound registration standard, the computer system displays a restart enable indication in the sound registration interface. In some embodiments, the restart enable indication restarts the sound registration process when selected. In some embodiments, in response to determining that the sound input does not meet the sound registration standard, the computer system displays a continue enable indication in the sound registration interface. In some embodiments, when the continue enable indication is selected, the registration process proceeds to the next sampling step of the sound registration process despite the sound input not meeting the sound registration standard. In some embodiments, when the continue enable indication is selected, the registration process repeats the failed sampling step of the sound registration process. In some embodiments, a threshold number of sound inputs must satisfy the sound registration criteria or the sound registration process is terminated. In some embodiments, in response to determining that the sound input does not satisfy the sound registration criteria, the computer system displays a learning affordance in the sound registration interface, which, when the learning affordance is selected, causes the computer system to display an interface describing sounds that are compatible with the sound registration process (e.g., sounds that satisfy the sound registration process).
[0338] In some embodiments, the sound registration process also includes providing a prompt (e.g., 1044) (e.g., an audio prompt; a visual prompt) to the user for providing the set of one or more sound inputs before receiving the set of one or more sound inputs. In some embodiments, in response to initiating the sound registration process, the computer system displays a sound registration interface that includes a prompt (e.g., "Teach the phone by playing the sound 5 times") to inform the user to provide a set of sound inputs for registering the sound for subsequent detection by the computer system.
[0339] In some embodiments, the model is a machine learning model generated based on the group of one or more sound inputs. In some embodiments, the model is generated by a computer system. In some embodiments, the group of one or more sound inputs are provided to a remote server (e.g., a device), and the remote server generates the model. In some embodiments, the model includes one or more adjustments for considering auditory hallucinations. In some embodiments, the computer system adjusts one or more parameters of the model so that the model better considers auditory hallucinations. In some embodiments, adjusting for auditory hallucinations reduces the number of false positive detections of sounds. Including adjustments in the model to consider auditory hallucinations improves model accuracy, thereby reducing the possibility of false positives and allowing more reliable and effective notification of detected sounds, which provides improved visual and / or auditory feedback.
[0340] In some embodiments, after generating the model, the computer system receives (e.g., detects) subsequent sound input. In some embodiments, in response to receiving the subsequent sound input and determining that the subsequent sound input is a first type of sound input based on the model, the computer system provides (e.g., displays) a notification indicating that the first type of sound input has been received (e.g., detected) (e.g., 1072). In some embodiments, once a sound is registered on the computer system using a sound registration process, the computer system then receives the sound input and uses the model generated during the sound registration process to determine whether the received sound input matches the registered type. In some embodiments, if a match is detected, the computer system provides a notification alerting the user that the registered sound has been detected.
[0341] In some embodiments, in response to receiving a subsequent sound input, based on a model-based determination that the subsequent sound input is not the first type of sound input, the computer system forgoes providing a notification indicating receipt of the first type of sound input.
[0342] It should be noted that the above description with respect to method 1100 (eg, Fig.11) also apply in a similar manner to the methods described above. For example, methods 700, 900 optionally include one or more of the features of the various methods described above with reference to method 1100. For example, a voice action (such as those described with reference to method 1100) can be used to adjust the volume settings of a parallel audio feature (such as those described with reference to method 700). For the sake of brevity, these details are not repeated below.
[0343] For the purpose of explanation, the preceding description is described by reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the invention to the disclosed precise form. According to the above teachings, many modifications and variations are possible. These embodiments are selected and described in order to best explain the principles of these technologies and their practical applications. Others skilled in the art can thus best utilize these technologies and various embodiments with various modifications suitable for the intended specific use.
[0344] Although the disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. It should be understood that such changes and modifications are considered to be included within the scope of the disclosure and examples defined by the claims.
[0345] As described above, one aspect of the present technology is to collect and use data from various sources to provide auditory features to users. The present disclosure contemplates that, in some instances, these collected data may include personal information data that uniquely identifies or can be used to contact or locate a specific person. Such personal information data may include demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records related to the user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying or personal information.
[0346] The present disclosure recognizes that the use of such personal information data in the present technology can be used to benefit users. For example, personal information data can be used to more accurately identify user behavior and / or input (e.g., voice input). Therefore, more reliable device operation is achieved using such personal information data. In addition, the present disclosure also anticipates other uses of personal information data that benefit users. For example, health and fitness data can be used to provide insights into the user's overall health status, or can be used as positive feedback to individuals who use technology to pursue health goals.
[0347] This disclosure envisions that entities responsible for collecting, analyzing, disclosing, transmitting, storing or otherwise using such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. Such policies should be easily accessible to users and should be updated as the collection and / or use of data changes. Personal information from users should be collected for legitimate and reasonable purposes of the entity and not shared or sold outside of these legitimate uses. In addition, such collection / sharing should be performed after receiving informed consent from the user. In addition, such entities should consider taking any necessary steps to safeguard and protect access to such personal information data and ensure that others who have access to personal information data comply with their privacy policies and processes. In addition, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices. In addition, policies and practices should be adjusted to specific types of personal information data collected and / or accessed, and to applicable laws and standards including specific considerations of jurisdiction. For example, in the United States, the collection or access of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA), while health data in other countries may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy practices should be maintained in each country for different types of personal data.
[0348] Regardless of the foregoing, the present disclosure also contemplates implementation schemes in which users selectively block the use or access of personal information data. That is, the present disclosure contemplates that hardware elements and / or software elements may be provided to prevent or block access to such personal information data. For example, where an auditory feature is provided, the technology of the present invention may be configured to allow a user to choose to "opt in" or "opt out" to participate in the collection of personal information data during registration for a service or at any time thereafter. In another example, a user may choose not to provide data indicating the use of background sounds and / or auditory controls. In addition to providing "opt-in" and "opt-out" options, the present disclosure also contemplates providing notifications related to access or use of personal information. For example, a user may be notified that their personal information data will be accessed when downloading an application, and then reminded again just before the personal information data is accessed by the application.
[0349] In addition, it is an object of the present disclosure that personal information data should be managed and processed to minimize the risk of unintentional or unauthorized access or use. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. In addition, and when applicable, including in certain health-related applications, data de-identification can be used to protect the privacy of users. De-identification can be facilitated by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than at the address level), controlling how data is stored (e.g., aggregating data between users), and / or other methods when appropriate.
[0350] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may also be implemented without access to such personal information data. That is, various embodiments of the present technology will not fail to function properly due to the lack of all or a portion of such personal information data. For example, auditory control may be implemented based on non-personal information data or an absolute minimum amount of personal information (such as auditory characteristics of sounds provided by a user or publicly available information).
Claims
1. A method comprising: At a computer system in communication with one or more input devices: While playing the audio media item of the first type, receiving, via the one or more input devices, a request to play an audio media item of a second type; Based on determining that a parallel audio standard set is satisfied, playing the following items in parallel, wherein the parallel audio standard set includes a requirement to enable a parallel playback feature on the computer system: the audio media item of the first type; and the audio media item of the second type; and Based on the determination that the parallel audio standard set is not met: stopping playing of the audio media item of the first type; and The audio media item of the second type is played.
2. The method according to claim 1, wherein: The first type of audio media item includes audio selected from the group consisting of ambient sound, irregular noise, and combinations thereof.
3. The method according to any one of claims 1 to 2, wherein: The first type of audio media item includes audio selected from the group consisting of randomly selected audio clips, randomly arranged audio clips, and combinations thereof.
4. The method according to any one of claims 1 to 2, wherein: The parallel playback feature is manually configurable.
5. The method according to any one of claims 1 to 2, wherein: Playing the audio media item of the first type and the audio media item of the second type in parallel comprises: The magnitude of the volume level of the first type of audio media item is adjusted from an initial volume to an adjusted volume based on the magnitude of the parallel volume feature.
6. The method according to any one of claims 1 to 2, wherein: The set of parallel audio criteria includes criteria that are satisfied when a magnitude of a second parallel volume feature is determined to exceed a threshold magnitude.
7. The method of any one of claims 1 to 2, further comprising adjusting a magnitude of at least one of the following based on the context information: a background sound volume characteristic for adjusting a volume level of an audio media item of the first type; and A third parallel volume feature is used to adjust the volume level of the first type of audio media item when the first type of audio media item is played in parallel with the second type of audio media item.
8. The method according to any one of claims 1 to 2, wherein: Playing the audio media item of the first type includes: In response to determining that the computer system is in a first context state, playing an audio media item of the first type having first audio content; and Based on determining that the computer system is in a first context state, an audio media item of the first type having second audio content different from the first audio content is played.
9. The method according to any one of claims 1 to 2, wherein: The computer system is also in communication with a display generation component, and the method further comprises: Displaying a first user interface via the display generation component, the first user interface comprising: a background sound user-interactive graphical user interface object that, when selected, enables playback of audio media items of the first type upon determining that a set of playback criteria is satisfied; and a first audio content user interactive graphical user interface object that, when selected, causes display of a second user interface, the second user interface comprising: a second audio content user interactive graphical user interface object that, when selected, causes third audio content is included in the audio media item of the first type during playback; and A third audio content user interactive graphical user interface object that, when selected, causes fourth audio content, different from the third audio content, to be included in the audio media item of the first type during playback.
10. The method according to claim 9, further comprising: detecting, via the one or more input devices, a first input corresponding to the background sound user-interactive graphical user interface object; In response to detecting the first input, enabling playback of the first type of audio media item upon determining that the set of playback criteria is satisfied; as well as When the computer system is enabled to play back the audio media item of the first type when it is determined that the set of playback criteria is met: Based on determining that the set of playback criteria is met, an audio media item of the first type is played.
11. The method according to claim 9, further comprising: When the second user interface is displayed: detecting a second input via the one or more input devices; as well as In response to detecting the second input: Based on determining that the second input corresponds to the second audio content user-interactive graphical user interface object, configuring the computer system to include the third audio content in the audio media item of the first type during playback; as well as Based on determining that the second input corresponds to the third audio content user-interactive graphical user interface object, the computer system is configured to include the fourth audio content in the audio media item of the first type during playback.
12. The method according to any one of claims 1 to 2, wherein: The computer system is also in communication with a display generation component, and the method further comprises: displaying, via the display generation component, a fourth user interface, the fourth user interface comprising a background sound status indicator, a volume user interactive graphical user interface object, and a background sound enable user interactive graphical user interface object; detecting a third input via the one or more input devices; Based on determining that the third input corresponds to a swipe gesture at a location corresponding to the volume user-interactive graphical user interface object, adjusting the volume level of the first type of audio media item from a second initial volume to a second adjusted volume based on a direction and a magnitude of the swipe gesture; and Based on determining that the third input corresponds to selection of the background sound enabling user-interactive graphical user interface object: selectively activating background sound features; and A visual characteristic of the background sound status indicator is modified.
13. A computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 12.
14. A computer system configured to communicate with one or more input devices, the computer system comprising: one or more processors; as well as A memory storing one or more programs configured to be executed by the one or more processors, the one or more programs comprising instructions for executing the method according to any one of claims 1 to 12.
15. A computer system configured to communicate with one or more input devices, comprising: Memory; as well as Device for carrying out the method according to any one of claims 1 to 12.
16. A computer program product comprising one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices, the one or more programs comprising instructions for performing the method according to any one of claims 1 to 12.
17. A method comprising: At a computer system in communication with a display generating component and one or more input devices: receiving, via the one or more input devices, a first voice input associated with a first predetermined action while displaying, via the display generation component, a user interface including a set of user interface objects; as well as In response to receiving the first voice input: In response to determining that a first user interface object in the group of user interface objects is currently selected, performing the first predetermined action based on the first user interface object; as well as According to determining that a second user interface object in the group of user interface objects is currently selected, the first predetermined action is performed based on the second user interface object.
18. A method comprising: At a computer system in communication with a display generating component and one or more input devices: Performing a sound registration process, the sound registration process comprising: receiving, via the one or more input devices, a set of one or more sound inputs including a first sound input; indicating whether the first sound input meets a sound input standard; Based on determining that the set of one or more sound inputs satisfies the sound enrollment criteria set, causing generation of a model for identifying a first type of sound corresponding to the set of one or more sound inputs; and Based on determining that the set of one or more sound inputs does not satisfy the set of sound enrollment criteria, causing generation of the model for identifying the first type of sound corresponding to the set of one or more sound inputs to be abandoned.
Citation Information
Patent Citations
Method and apparatus for integrating manual input
US20020015024A1
Gestures for touch sensitive input devices
US20060026521A1
Gestures for touch sensitive input devices
US20060026536A1
Virtual input device placement on a touch screen user interface
US20060033724A1
Multipoint touchscreen
US20060097991A1