Spatial management of audio

Efficient spatial audio management techniques using visual and audio transitions on electronic devices with touch-sensitive surfaces improve user interaction and conserve power, addressing inefficiencies in existing methods.

JP2025163072APending Publication Date: 2025-10-28APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025123348
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-09-26
Filing Date
2025-07-23
Publication Date
2025-10-28

Smart Images

  • Figure 2025163072000001_ABST
    Figure 2025163072000001_ABST
Patent Text Reader

Abstract

To provide a user interface for managing a spatial audio.SOLUTION: A method includes: a user interface for making a transition between visual elements; a user interface for previewing an audio; a user interface for finding a music; a user interface for managing transmittivity of a headphone; and a user interface for operating a plurality of audio streams of an audio source.SELECTED DRAWING: Figure 6A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 736,990, entitled "SPATIAL MANAGEMENT OF AUDIO," filed September 26, 2018, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to computer user interfaces, and more particularly to techniques for managing spatial audio. [Background technology]

[0003] Humans can localize sound in three dimensions (up and down, front and back, and left and right). Various techniques can be used to modify the audio produced by a device so that the listener perceives it as coming from a particular point in space. Summary of the Invention

[0004] However, some techniques for managing spatial audio using electronic devices are generally cumbersome and inefficient. For example, some techniques do not provide users with contextual awareness of the state of the electronic device through spatial management of audio. In another example, some existing techniques use complex and time-consuming user interfaces that may involve multiple key presses or keystrokes. Existing techniques take more time than necessary, wasting user time and device energy. The latter problem is particularly acute in battery-operated devices. Furthermore, existing audio techniques do not adequately assist users in navigating graphical user interfaces.

[0005] The present techniques thus provide electronic devices with faster, more efficient methods and interfaces for managing spatial audio. Such methods and interfaces optionally complement or replace other methods for managing spatial audio. Such methods and interfaces reduce the cognitive burden on users and create more efficient human-machine interfaces. For battery-operated computing devices, such methods and interfaces conserve power and extend the time between battery charges.

[0006] According to some embodiments, a method is described that is performed on an electronic device having a display, the electronic device being operatively connected to two or more speakers, the method including: displaying a first visual element at a first position on the display; accessing first audio corresponding to the first visual element; generating audio on the two or more speakers using the first audio in a first mode while displaying the first visual element at the first position on the display and receiving a first user input; transitioning the display of the first visual element from the first position on the display to a first visual element that is not displayed on the display in response to receiving the first user input; and generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that audio generated in the second mode is perceived by a user to be generated from a direction away from the display.

[0007] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display, the electronic device being operatively connected to two or more speakers, the one or more programs including instructions for displaying a first visual element at a first position on the display, accessing first audio corresponding to the first visual element, generating audio on the two or more speakers using the first audio in a first mode while displaying the first visual element at the first position on the display, receiving a first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not displayed on the display in response to receiving the first user input, and generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display.

[0008] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display, the electronic device being operatively connected to two or more speakers, the one or more programs including instructions for displaying a first visual element at a first position on the display, accessing first audio corresponding to the first visual element, generating audio on the two or more speakers using the first audio in a first mode while displaying the first visual element at the first position on the display, receiving a first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not displayed on the display in response to receiving the first user input, and generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display.

[0009] According to some embodiments, an electronic device is described that includes a display and one or more processors and a memory storing one or more programs configured to be executed by the one or more processors, the electronic device being operably connected to two or more speakers, the one or more programs including instructions for displaying a first visual element at a first position on the display, accessing first audio corresponding to the first visual element, generating audio on the two or more speakers using the first audio in a first mode while displaying the first visual element at the first position on the display, receiving a first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not displayed on the display in response to receiving the first user input, and generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display.

[0010] According to some embodiments, an electronic device is described that includes a display operatively connected to two or more speakers, means for displaying a first visual element at a first position on the display, means for accessing first audio corresponding to the first visual element, means for generating audio on the two or more speakers using the first audio in a first mode while displaying the first visual element at the first position on the display and receiving a first user input, and means for transitioning the display of the first visual element from the first position on the display to a first visual element that is not displayed on the display in response to receiving the first user input, and means for generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that audio generated in the second mode is perceived by a user to be generated from a direction away from the display.

[0011] According to some embodiments, a method is described that is performed in an electronic device that includes a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers.The method includes displaying, on a display, a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file; detecting, using the touch-sensitive surface, a user contact at a location corresponding to a first media element; in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, generating audio using a first audio file corresponding to the first media element using two or more speakers without exceeding a predetermined audio playback period; continuing to generate audio using the first audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the location corresponding to the first media element; stopping generating audio using the first audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; and using the touch-sensitive surface to generate audio from the location corresponding to the first media element. detecting movement of a user contact from the first audio file to a position corresponding to a second media element; generating audio using a second audio file corresponding to the second media element using two or more speakers without exceeding a predetermined audio playback period in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input; continuing to generate audio using the second audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the position corresponding to the second media element; stopping generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; detecting lift-off of the user contact using the touch-sensitive surface; and stopping generating audio using the first audio file or the second audio file using the two or more speakers in response to detecting lift-off of the user contact.

[0012] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file, detecting, using the touch-sensitive surface, a user contact at a location corresponding to a first media element, generating, using the two or more speakers, audio using a first audio file corresponding to the first media element in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, without exceeding a predetermined audio playback period, and continuing to generate, using the two or more speakers, audio using the first audio file in accordance with not exceeding the predetermined audio playback period while the user contact remains at the location corresponding to the first media element. and stopping generating audio using the two or more speakers in accordance with exceeding a predetermined audio playback period; detecting, using the touch-sensitive surface, movement of a user contact from a position corresponding to the first media element to a position corresponding to a second media element; in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input, generating audio using a second audio file corresponding to the second media element using the two or more speakers without exceeding the predetermined audio playback period; continuing to generate audio using the second audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the position corresponding to the second media element; stopping generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; and detecting, using the touch-sensitive surface, lift-off of the user contact;In response to detecting lift-off of the user contact, stop generating audio using the two or more speakers using the first audio file or the second audio file.

[0013] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file, detecting, using the touch-sensitive surface, a user contact at a location corresponding to a first media element, generating audio using a first audio file corresponding to the first media element using the two or more speakers in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, without exceeding a predetermined audio playback period, and continuing to generate audio using the first audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the location corresponding to the first media element. stop generating audio using the first audio file using the two or more speakers in accordance with exceeding a predetermined audio playback period; detect, using the touch-sensitive surface, movement of a user contact from a position corresponding to the first media element to a position corresponding to a second media element; in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input, generate audio using a second audio file corresponding to the second media element using the two or more speakers without exceeding the predetermined audio playback period; continue generating audio using the second audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the position corresponding to the second media element; stop generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; detect lift-off of the user contact using the touch-sensitive surface;In response to detecting lift-off of the user contact, stop generating audio using the two or more speakers using the first audio file or the second audio file.

[0014] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface, one or more processors, and memory storing one or more programs configured to be executed by the one or more processors, the one or more programs operatively connected to two or more speakers, the one or more programs displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file, detecting, using the touch-sensitive surface, a user contact at a location corresponding to a first media element, generating, using the two or more speakers, audio using a first audio file corresponding to the first media element in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, without exceeding a predetermined audio playback period, and continuing to generate, using the two or more speakers, audio using the first audio file in accordance with not exceeding the predetermined audio playback period while the user contact remains at the location corresponding to the first media element. stop generating audio using the first audio file using the two or more speakers in accordance with exceeding a predetermined audio playback period; detect, using the touch-sensitive surface, movement of a user contact from a position corresponding to the first media element to a position corresponding to a second media element; in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input, generate audio using a second audio file corresponding to the second media element using the two or more speakers without exceeding the predetermined audio playback period; continue generating audio using the second audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the position corresponding to the second media element; stop generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; detect lift-off of the user contact using the touch-sensitive surface;In response to detecting lift-off of the user contact, stop generating audio using the two or more speakers using the first audio file or the second audio file.

[0015] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface operatively connected to two or more speakers, and means for displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file, means for detecting, using the touch-sensitive surface, a user contact at a location corresponding to a first media element, and means for generating, in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, audio using a first audio file corresponding to the first media element, using the two or more speakers without exceeding a predetermined audio playback period, continuing to generate audio using the first audio file, using the two or more speakers, while the user contact remains at the location corresponding to the first media element, in accordance with not exceeding the predetermined audio playback period, and playing audio using the first audio file, using the two or more speakers, in accordance with exceeding the predetermined audio playback period. means for detecting, using the touch-sensitive surface, movement of a user contact from a position corresponding to a first media element to a position corresponding to a second media element; means for generating audio using a second audio file corresponding to the second media element using two or more speakers without exceeding a predetermined audio playback period in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input; means for continuing to generate audio using the second audio file using the two or more speakers in accordance with not exceeding the predetermined audio playback period while the user contact remains at the position corresponding to the second media element, and stopping generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period; means for detecting, using the touch-sensitive surface, lift-off of the user contact; and means for generating audio using the two or more speakers in response to detecting lift-off of the user contact.and means for stopping generating audio using the first audio file or the second audio file.

[0016] According to some embodiments, a method is described that is performed on an electronic device with a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the method including: detecting a first user input to activate a discovery mode; and, in response to detecting the first user input to activate the discovery mode, using the two or more speakers to generate audio from a first audio source in a first mode, the first mode being configured such that audio generated using the first mode is perceived by a user as being generated from a first point in space moving over time in a first direction along a predetermined path at a first speed; and a second audio source in a second mode, the second mode being configured such that audio generated using the second mode is generated from a first point in space moving over time in a first direction along a predetermined path at a second speed. and simultaneously generating audio using a second audio source configured to be perceived by a user as being generated from a second point in space moving over time in a first direction along a predetermined path at a third speed; and a third audio source in a third mode configured to be perceived by a user as being generated from a third point in space moving over time in the first direction along a predetermined path at a third speed, wherein the first point, the second point, and the third point are different points in space.

[0017] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs configured to detect a first user input to activate a discovery mode, and in response to detecting the first user input to activate the discovery mode, use the two or more speakers to generate audio from a first audio source in a first mode, the first mode being configured such that audio generated using the first mode is perceived by a user as being generated from a first point in space moving over time along a predefined path at a first speed and in a first direction. and a third audio source in a third mode, wherein the third mode is configured such that audio generated using the third mode is perceived by the user as being generated from a third point in space moving over time in the first direction along a predetermined path at a third speed, wherein the first point, the second point, and the third point are different points in space.

[0018] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs configured to detect a first user input to activate a discovery mode and, in response to detecting the first user input to activate the discovery mode, use the two or more speakers to generate audio from a first audio source in a first mode, the first mode being such that audio generated using the first mode is perceived by a user as being generated from a first point in space moving over time along a predefined path in a first direction at a first speed; and a third audio source in a third mode, wherein the third mode is configured such that audio generated using the third mode is perceived by the user as being generated from a third point in space moving over time in the first direction along a predetermined path at a third speed, wherein the first point, the second point, and the third point are different points in space.

[0019] According to some embodiments, an electronic device is described, the electronic device comprising: a display; a touch-sensitive surface; one or more processors; and memory storing one or more programs configured to be executed by the one or more processors, the electronic device operatively connected to two or more speakers, the one or more programs configured to detect a first user input to activate a discovery mode and, in response to detecting the first user input to activate the discovery mode, use the two or more speakers to generate audio from a first audio source in a first mode, the first mode being perceived by a user as being generated from a first point in space moving over time along a predetermined path in a first direction at a first speed; The method includes instructions for simultaneously generating audio using a second audio source in a second mode, the second mode configured such that audio generated using the second mode is perceived by a user as being generated from a second point in space moving over time in a first direction along a predetermined path at a second speed, and a third audio source in a third mode, the third mode configured such that audio generated using the third mode is perceived by a user as being generated from a third point in space moving over time in the first direction along a predetermined path at a third speed, wherein the first point, the second point, and the third point are different points in space.

[0020] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface to which the electronic device is operatively connected, means for detecting a first user input to activate a discovery mode, and means for, in response to detecting the first user input to activate the discovery mode, using the two or more speakers, transmitting a first audio source in a first mode configured such that audio generated using the first mode is perceived by a user as being generated from a first point in space moving over time in a first direction along a predetermined path at a first speed, and a second audio source in a second mode configured to generate audio using the first mode. and means for simultaneously generating audio using a second audio source configured to cause audio generated using the third mode to be perceived by a user as being generated from a second point in space moving over time in a first direction along a predetermined path at a second speed; and a third audio source in a third mode configured to cause audio generated using the third mode to be perceived by a user as being generated from a third point in space moving over time in a first direction along a predetermined path at a third speed, wherein the first point, the second point, and the third point are different points in space.

[0021] According to some embodiments, a method is described that is performed in an electronic device with a display and a touch-sensitive surface, the electronic device operatively connected to two or more speakers. The method includes displaying a user-movable affordance at a first location on the display, operating the electronic device in a first state of ambient sound transparency while the user-movable affordance is displayed at the first location, generating audio using an audio source in a first mode using the two or more speakers, detecting user input using the touch-sensitive surface, and, in response to detecting the user input, operating the electronic device in a second state of ambient sound transparency that is different from the first state of ambient sound transparency in accordance with a set of one or more conditions being met, the first condition including a first condition that is met if the user input is a touch-and-drag motion on the user-movable affordance, transitioning the generation of audio using the audio source from the first mode to a second mode that is different from the first mode, and maintaining the electronic device in the first state of ambient sound transparency and maintaining the generation of audio using the audio source in the first mode in accordance with the set of one or more conditions not being met.

[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs including instructions to: display a user-movable affordance at a first position on the display; operate the electronic device in a first state of ambient sound transparency while the user-movable affordance is displayed at the first position; generate audio using an audio source in a first mode using the two or more speakers; detect user input using the touch-sensitive surface; and, in response to detecting the user input, operate the electronic device in a second state of ambient sound transparency that is different from the first state of ambient sound transparency in accordance with a set of one or more conditions being met, the first condition including a first condition that is met if the user input is a touch-and-drag motion on the user-movable affordance; transition the generation of audio using the audio source from the first mode to a second mode that is different from the first mode; and maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode in accordance with a set of one or more conditions not being met.

[0023] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs including instructions to: display a user-movable affordance at a first position on the display; operate the electronic device in a first state of ambient sound transparency while the user-movable affordance is displayed at the first position; generate audio using an audio source in a first mode using the two or more speakers; detect user input using the touch-sensitive surface; and, in response to detecting the user input, operate the electronic device in a second state of ambient sound transparency that is different from the first state of ambient sound transparency in accordance with a set of one or more conditions being met, the first condition including a first condition that is met if the user input is a touch-and-drag motion on the user-movable affordance; transition the generation of audio using the audio source from the first mode to a second mode that is different from the first mode; and maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode in accordance with a set of one or more conditions not being met.

[0024] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface, one or more processors, and memory that stores one or more programs configured to be executed by the one or more processors, the electronic device being operatively connected to two or more speakers, the one or more programs displaying a user-movable affordance at a first position on the display, operating the electronic device in an ambient sound transparent first state while the user-movable affordance is displayed at the first position, generating audio using an audio source in a first mode using the two or more speakers, detecting user input using the touch-sensitive surface, and In response to detecting a user input, the electronic device includes instructions to operate the electronic device in a second state of ambient sound transparency that is different from the first state of ambient sound transparency and transition generating audio using the audio source from the first mode to the second mode that is different from the first mode in accordance with a set of one or more conditions being satisfied, the set including a first condition that is satisfied if the user input is a touch-and-drag action on the user-movable affordance; and to maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode in accordance with the set of one or more conditions not being satisfied.

[0025] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface to which the electronic device is operatively connected, and means for displaying a user-movable affordance at a first location on the display, means for operating the electronic device in a first state of ambient sound transparency and generating audio using an audio source in a first mode using the two or more speakers and detecting user input using the touch-sensitive surface while the user-movable affordance is displayed at the first location, and means for operating the electronic device in a second state of ambient sound transparency that is different from the first state of ambient sound transparency and transitioning generation of audio using the audio source from the first mode to a second mode that is different from the first mode in response to detecting the user input according to a set of one or more conditions being met, the first condition including a first condition that is met if the user input is a touch-and-drag motion on the user-movable affordance, and means for maintaining the electronic device in the first state of ambient sound transparency and maintaining generation of audio using the audio source in the first mode according to a set of one or more conditions not being met.

[0026] According to some embodiments, a method is described that is performed in an electronic device that includes a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, including a first speaker and a second speaker. The method includes generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams, including a first audio stream and a second audio stream; detecting a first user input using the touch-sensitive surface; in response to detecting the first user input, simultaneously transitioning, using the two or more speakers, generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode; transitioning, using the two or more speakers, generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode; displaying, on the display, a first visual representation of the first audio stream of the audio source; and displaying, on the display, a second visual representation of the second audio stream of the audio source, wherein the first visual representation is different from the second visual representation.

[0027] According to some embodiments, a non-transitory computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device that includes a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, including a first speaker and a second speaker, the one or more programs for generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams, including a first audio stream and a second audio stream, and detecting a first user input using the touch-sensitive surface and generating audio using the touch-sensitive surface. In response to detecting a user input, simultaneously transitioning, using two or more speakers, generation of a first audio stream of an audio source from a first mode to a second mode different from the first mode, transitioning, using two or more speakers, generation of a second audio stream of an audio source from the first mode to a third mode different from the first mode and the second mode, displaying on the display a first visual representation of the first audio stream of the audio source, and displaying on the display a second visual representation of the second audio stream of the audio source, wherein the first visual representation is different from the second visual representation.

[0028] According to some embodiments, a temporary computer-readable storage medium is described that stores one or more programs configured to be executed by one or more processors of an electronic device that includes a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, including a first speaker and a second speaker, the one or more programs for generating audio using an audio source in a first mode, using the two or more speakers, the audio source including a plurality of audio streams, including a first audio stream and a second audio stream, and detecting a first user input using the touch-sensitive surface and generating audio using the touch-sensitive surface. In response to detecting the user input, simultaneously transition, using two or more speakers, generation of a first audio stream of an audio source from a first mode to a second mode different from the first mode, transition, using two or more speakers, generation of a second audio stream of an audio source from the first mode to a third mode different from the first mode and the second mode, display a first visual representation of the first audio stream of the audio source on the display, and display a second visual representation of the second audio stream of the audio source on the display, wherein the first visual representation is different from the second visual representation.

[0029] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface, one or more processors, and memory that stores one or more programs configured to be executed by the one or more processors, the electronic device being operatively connected to two or more speakers, including a first speaker and a second speaker, the one or more programs being configured to generate audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams, including a first audio stream and a second audio stream, and to detect a first user input using the touch-sensitive surface and In response to detecting the user input, simultaneously transition, using two or more speakers, generation of a first audio stream of an audio source from a first mode to a second mode different from the first mode, transition, using two or more speakers, generation of a second audio stream of an audio source from the first mode to a third mode different from the first mode and the second mode, display a first visual representation of the first audio stream of the audio source on the display, and display a second visual representation of the second audio stream of the audio source on the display, wherein the first visual representation is different from the second visual representation.

[0030] According to some embodiments, an electronic device is described that includes a display, a touch-sensitive surface, the electronic device operatively connected to two or more speakers, including a first speaker and a second speaker, means for generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams, including a first audio stream and a second audio stream, means for detecting a first user input using the touch-sensitive surface, and in response to detecting the first user input, simultaneously transitioning, using the two or more speakers, the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode, and transitioning, using the two or more speakers, the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode, and displaying, on the display, a first visual representation of the first audio stream of the audio source;

[0031] and means for displaying, on the display, a second visual representation of a second audio stream of the audio source, wherein the first visual representation is different from the second visual representation.

[0032] Executable instructions to perform these functions are optionally contained in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions to perform these functions are optionally contained in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0033] This provides devices with faster, more efficient methods and interfaces for managing spatial audio, thereby increasing the effectiveness, efficiency, and user satisfaction of such devices. Such methods and interfaces can complement or replace other methods for managing spatial audio. [Brief explanation of the drawings]

[0034] For a better understanding of the various described embodiments, reference should be made to the following Detailed Description of the Invention in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout:

[0035] [Figure 1A] 1 is a block diagram illustrating a portable multifunction device having a touch-sensitive display in accordance with some embodiments.

[0036] [Figure 1B] FIG. 2 is a block diagram illustrating exemplary components for event processing according to some embodiments.

[0037] [Figure 2] FIG. 1 illustrates a portable multifunction device with a touch screen in accordance with some embodiments.

[0038] [Figure 3] FIG. 1 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface in accordance with some embodiments.

[0039] [Figure 4A] 1 illustrates an exemplary user interface for a menu of applications on a portable multifunction device in accordance with some embodiments.

[0040] [Figure 4B] 1A-1C illustrate exemplary user interfaces for a multifunction device having a touch-sensitive surface separate from a display in accordance with some embodiments.

[0041] [Figure 5A] FIG. 1 illustrates a personal electronic device according to some embodiments.

[0042] [Figure 5B] FIG. 1 is a block diagram illustrating a personal electronic device according to some embodiments.

[0043] [Figure 5C] 1 illustrates exemplary components of a personal electronic device having a touch-sensitive display and intensity sensor in accordance with some embodiments. [Figure 5D] 1 illustrates exemplary components of a personal electronic device having a touch-sensitive display and intensity sensor in accordance with some embodiments.

[0044] [Figure 5E] 1 illustrates exemplary components and a user interface of a personal electronic device according to some embodiments. [Figure 5F] 1 illustrates exemplary components and a user interface of a personal electronic device according to some embodiments. [Figure 5G] 1 illustrates exemplary components and a user interface of a personal electronic device according to some embodiments. [Figure 5H] 1 illustrates exemplary components and a user interface of a personal electronic device according to some embodiments.

[0045] [Figure 6A] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6B] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6C] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6D] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6E] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6F]1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6G] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6H] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6I] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6J] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6K] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6L] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6M] 1 illustrates an example technique for transitioning between visual elements according to some embodiments. [Figure 6N] 1 illustrates an example technique for transitioning between visual elements according to some embodiments.

[0046] [Figure 7A] FIG. 1 is a flow diagram illustrating a method for transitioning between visual elements using an electronic device, according to some embodiments. [Figure 7B] FIG. 1 is a flow diagram illustrating a method for transitioning between visual elements using an electronic device, according to some embodiments. [Figure 7C] FIG. 1 is a flow diagram illustrating a method for transitioning between visual elements using an electronic device, according to some embodiments.

[0047] [Figure 8A] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8B] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8C] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8D] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8E] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8F] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8G] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8H] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8I] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8J] 1 illustrates an exemplary technique for previewing audio, according to some embodiments. [Figure 8K] 1 illustrates an exemplary technique for previewing audio, according to some embodiments.

[0048] [Figure 9A] FIG. 1 is a flow diagram illustrating a method for previewing audio using an electronic device according to some embodiments. [Figure 9B] FIG. 1 is a flow diagram illustrating a method for previewing audio using an electronic device according to some embodiments. [Figure 9C] FIG. 1 is a flow diagram illustrating a method for previewing audio using an electronic device according to some embodiments.

[0049] [Figure 10A] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10B]1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10C] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10D] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10E] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10F] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10G] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10H] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10I] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10J] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 10K] 1 illustrates an exemplary technique for discovering music, according to some embodiments.

[0050] [Figure 11A] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11B] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11C] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11D] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11E] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11F] 1 illustrates an exemplary technique for discovering music, according to some embodiments. [Figure 11G]1 illustrates an exemplary technique for discovering music, according to some embodiments.

[0051] [Figure 12A] FIG. 1 is a flow diagram illustrating a method for discovering music using an electronic device, according to some embodiments. [Figure 12B] FIG. 1 is a flow diagram illustrating a method for discovering music using an electronic device, according to some embodiments.

[0052] [Figure 13A] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments. [Figure 13B] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments. [Figure 13C] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments. [Figure 13D] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments. [Figure 13E] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments. [Figure 13F] 1 illustrates an exemplary technique for managing headphone transparency, according to some embodiments.

[0053] [Figure 13G] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13H] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13I] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13J]1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13K] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13L] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments. [Figure 13M] 1 illustrates an exemplary technique for manipulating multiple audio streams of an audio source, according to some embodiments.

[0054] [Figure 14A] FIG. 1 is a flow diagram illustrating a method for managing headphone transparency using an electronic device, according to some embodiments. [Figure 14B] FIG. 1 is a flow diagram illustrating a method for managing headphone transparency using an electronic device, according to some embodiments.

[0055] [Figure 15] FIG. 1 is a flow diagram illustrating a method for manipulating multiple audio streams of an audio source using an electronic device, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0056] The following description sets forth example methods, parameters, etc. However, it should be recognized that such description is not intended as a limitation on the scope of the present disclosure, but rather is provided as a description of example embodiments.

[0057] There is a need for electronic devices that provide efficient methods and interfaces for managing spatial audio. For example, spatial audio can provide a user with contextual awareness of the state of the electronic device. Such techniques can reduce the cognitive burden on users using electronic devices, thereby increasing their productivity. Furthermore, such techniques can reduce processor and battery power that would otherwise be wasted on redundant user input.

[0058] Below, FIGS. 1A-1B, 2, 3, 4A-4B, and 5A-5H provide descriptions of example devices for performing techniques for managing event notifications.

[0059] 6A-6N illustrate example techniques for transitioning between visual elements, according to some embodiments. 7A-7C are flow diagrams illustrating methods for transitioning between visual elements using an electronic device, according to some embodiments. The user interfaces of FIGS. 6A-6N are used to illustrate processes described below, including the processes of FIGS. 7A-7C.

[0060] 8A-8K illustrate exemplary techniques for previewing audio, according to some embodiments. 9A-9C are flow diagrams illustrating methods for previewing audio using an electronic device, according to some embodiments. The user interfaces of FIGS. 8A-8K are used to illustrate processes described below, including the processes of FIGS. 9A-9C.

[0061] 10A-10K illustrate an exemplary technique for discovering music, according to some embodiments. 11A-11G illustrate an exemplary technique for discovering music, according to some embodiments. 12A-12B are flow diagrams illustrating a method for discovering music using an electronic device, according to some embodiments. The user interfaces of FIGS. 10A-10K and 11A-11G are used to illustrate processes described below, including the process of FIGS. 12A-12B.

[0062] 13A-13F illustrate an exemplary technique for managing headphone transparency, according to some embodiments. 14A-14B are flow diagrams illustrating a method for managing headphone transparency using an electronic device, according to some embodiments. The user interfaces in FIGS. 13A-13F are used to illustrate processes described below, including the processes in FIGS. 14A-14B.

[0063] 13G-13M illustrate exemplary techniques for manipulating multiple audio streams of an audio source, according to some embodiments. FIG. 15 is a flow diagram illustrating a method for manipulating multiple audio streams of an audio source using an electronic device, according to some embodiments. The user interfaces in FIGS. 13G-13M are used to illustrate processes described below, including the process of FIG. 15.

[0064] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first touch can be referred to as a second touch, and similarly, a second touch can be referred to as a first touch, without departing from the scope of the various embodiments described. Although a first touch and a second touch are both touches, they are not the same touch.

[0065] The terminology used in the description of the various embodiments set forth herein is for the purpose of describing particular embodiments only and is not intended to be limiting. In the description of the various embodiments set forth and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" should be understood to refer to and include any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms "includes," "including," "comprises," and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0066] The term "if" is interpreted, optionally, depending on the context, to mean "when" or "upon," or "in response to determining" or "in response to detecting." Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" are interpreted, optionally, depending on the context, to mean "upon determining" or "in response to determining," or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."

[0067] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as PDA and / or music player functions. Exemplary embodiments of portable multifunction devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Optionally, other portable electronic devices, such as laptops or tablet computers having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad), are also used. It should also be understood that in some embodiments, the device is not a portable communication device, but rather a desktop computer having a touch-sensitive surface (e.g., a touchscreen display and / or touchpad).

[0068] In the following discussion, electronic devices are described that include a display and a touch-sensitive surface, however, it should be understood that the electronic device optionally includes one or more other physical user-interface devices, such as a physical keyboard, a mouse, and / or a joystick.

[0069] The device typically supports a variety of applications such as one or more of a drawing application, a presentation application, a word processing application, a website creation application, a disc authoring application, a spreadsheet application, a gaming application, a telephone application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0070] Various applications running on the device optionally use at least one common physical user-interface device, such as a touch-sensitive surface. One or more features of the touch-sensitive surface and corresponding information displayed on the device are optionally adjusted and / or changed for each application and / or within each application. In this way, the common physical architecture of the device (such as the touch-sensitive surface) optionally supports various applications with user interfaces that are intuitive and transparent to the user.

[0071] Attention now turns to embodiments of portable devices with touch-sensitive displays. FIG. 1A is a block diagram illustrating portable multifunction device 100 having touch-sensitive display system 112, according to some embodiments. Touch-sensitive display 112 may conveniently be referred to as a "touch screen" and may also be known or referred to as a "touch-sensitive display system." Device 100 includes memory 102 (optionally including one or more computer-readable storage media), memory controller 122, one or more processing units (CPUs) 120, peripherals interface 118, RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, input / output (I / O) subsystem 106, other input control devices 116, and external port 124. Device 100 optionally includes one or more optical sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 that detect the intensity of a contact on device 100 (e.g., a touch-sensitive surface, such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more tactile output generators 167 that generate tactile output on device 100 (e.g., generate tactile output on a touch-sensitive surface such as touch-sensitive display system 112 of device 100 or touchpad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0072] As used herein and in the claims, the term “intensity” of a contact on a touch-sensitive surface refers to the force or pressure (force per unit area) of a contact (e.g., a finger contact) on the touch-sensitive surface, or a proxy for the force or pressure of a contact on the touch-sensitive surface. The intensity of a contact has a range of values ​​that includes at least four distinct values ​​and more typically includes hundreds (e.g., at least 256) distinct values. The intensity of a contact is optionally determined (or measured) using various techniques and various sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch-sensitive surface are optionally used to measure force at various points on the touch-sensitive surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine an estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch-sensitive surface. Alternatively, the size and / or change in the contact area detected on the touch-sensitive surface, the capacitance and / or change in the capacitance of the touch-sensitive surface proximate the contact, and / or the resistance and / or change in the capacitance of the touch-sensitive surface proximate the contact are optionally used as a surrogate for the force or pressure of the contact on the touch-sensitive surface. In some implementations, the surrogate measure of the force or pressure of the contact is used directly to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is described in units corresponding to the surrogate measure). In some implementations, the surrogate measure of the contact force or pressure is converted to an estimate of the force or pressure, and the estimate of the force or pressure is used to determine whether an intensity threshold is exceeded (e.g., the intensity threshold is a pressure threshold measured in units of pressure). Using the intensity of contact as an attribute of user input may, in some circumstances, enable a user to access additional device functionality that would not otherwise be accessible to a user on a reduced-size device with a limited area for displaying affordances (e.g., on a touch-sensitive display) and / or receiving user input (e.g., via a touch-sensitive display, touch-sensitive surface, or physical / mechanical controls such as knobs or buttons).

[0073] As used herein and in the claims, the term “tactile output” refers to a physical displacement of a device relative to a previous position of the device, a physical displacement of a component of the device (e.g., a touch-sensitive surface) relative to another component of the device (e.g., a housing), or a displacement of a component relative to the center of mass of the device, that will be detected by a user with the user's sense of touch. For example, in a situation where a device or a component of a device is in contact with a touch-sensitive surface of a user (e.g., the fingers, palm, or other part of the user's hand), the tactile output produced by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in a physical property of the device or a component of the device. For example, movement of a touch-sensitive surface (e.g., a touch-sensitive display or trackpad) is optionally interpreted by the user as a “downclick” or “upclick” of a physical actuator button. In some cases, a user feels a tactile sensation such as a “downclick” or “upclick” even when there is no movement of a physical actuator button associated with the touch-sensitive surface that is physically pressed (e.g., displaced) by the user's action. As another example, movement of a touch-sensitive surface is optionally interpreted or perceived by a user as "roughness" of the touch-sensitive surface, even when there is no change in the smoothness of the touch-sensitive surface. While such user interpretation of touch depends on the user's personal sensory perception, there are many sensory perceptions of touch that are common to the majority of users. Thus, when a tactile output is described as corresponding to a particular sensory perception of a user (e.g., "upclick," "downclick," "roughness"), unless otherwise specified, the generated tactile output corresponds to a physical displacement of the device, or a component of the device, that produces the described sensory perception for a typical (or average) user.

[0074] It should be understood that device 100 is only one example of a portable multifunction device, and that device 100 optionally has more or fewer components than those shown, optionally combines two or more components, or optionally has a different configuration or arrangement of its components. The various components shown in Figure 1A are implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0075] Memory 102 optionally includes high-speed random access memory, and optionally includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0076] Peripheral interface 118 may be used to couple input and output peripherals of the device to CPU 120 and memory 102. One or more processors 120 operate or execute various software programs and / or instruction sets stored in memory 102 to perform various functions and process data for device 100. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0077] RF (radio frequency) circuitry 108 transmits and receives RF signals, also called electromagnetic signals. RF circuitry 108 converts electrical signals to electromagnetic signals and electromagnetic signals to communicate with communication networks and other communication devices via electromagnetic signals. RF circuitry 108 optionally includes well-known circuitry for performing these functions, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. RF circuitry 108 optionally communicates via wireless communication with networks, such as the Internet, also known as the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and with other devices. RF circuitry 108 optionally includes well-known circuitry for detecting near field communication (NFC) fields, such as by short-range radios. Wireless communication optionally includes, but is not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), Long Term Evolution (LTE), and other standards.evolution (LTE), near field communications (NFC), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, IEEE 802.11n, and / or IEEE 802.11ac), voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet message access protocol (IMAP) and / or post office protocol (POP)), instant messaging (e.g., extensible messaging and presence protocol), The present invention may use any of a number of communication standards, protocols, and technologies, including the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (XMPP), the Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), the Instant Messaging and Presence Service (IMPS), and / or the Short Message Service (SMS), or any other suitable communication protocol, including communication protocols not yet developed as of the filing date of this application.

[0078] Audio circuit 110, speaker 111, and microphone 113 provide an audio interface between a user and device 100. Audio circuit 110 receives audio data from peripherals interface 118, converts the audio data into electrical signals, and transmits the electrical signals to speaker 111. Speaker 111 converts the electrical signals into sound waves audible to humans. Audio circuit 110 also receives electrical signals converted from sound waves by microphone 113. Audio circuit 110 converts the electrical signals into audio data and transmits the audio data to peripherals interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to memory 102 and / or RF circuit 108 by peripherals interface 118. In some embodiments, audio circuit 110 also includes a headset jack (e.g., 212 in FIG. 2 ). The headset jack provides an interface between audio circuitry 110 and a detachable audio input / output peripheral, such as an output-only headphone or a headset with both an output (e.g., mono or binaural headphones) and an input (e.g., a microphone).

[0079] I / O subsystem 106 couples input / output peripherals on device 100, such as touchscreen 112 and other input control devices 116, to peripheral interface 118. I / O subsystem 106 optionally includes display controller 156, light sensor controller 158, intensity sensor controller 159, haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. One or more input controllers 160 receive / send electrical signals from / to other input control devices 116. Other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, input controller 160 is optionally coupled to any (or none) of a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 208 in FIG. 2) optionally include up / down buttons for volume control of speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., 206 in FIG. 2).

[0080] As described in U.S. Patent Application No. 11 / 322,549, filed December 23, 2005, "Unlocking a Device by Performing Gestures on an Unlock Image," U.S. Patent No. 7,657,849, which is incorporated herein by reference in its entirety, a quick press of a push button optionally unlocks touchscreen 112 or, optionally, initiates the process of unlocking the device using gestures on the touchscreen. A longer press of a push button (e.g., 206) optionally turns power on or off to device 100. The functionality of one or more of the buttons is optionally customizable by the user. Touchscreen 112 is used to implement virtual or soft buttons and one or more soft keyboards.

[0081] Touch-sensitive display 112 provides an input and output interface between the device and a user. Display controller 156 receives and / or sends electrical signals to touchscreen 112. Touchscreen 112 displays visual output to the user. This visual output optionally includes graphics, text, icons, animation, and any combination thereof (collectively "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0082] Touchscreen 112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touchscreen 112 and display controller 156 (along with any associated modules and / or instruction sets in memory 102) detects contacts (and any movement or cessation of contact) on touchscreen 112 and translates the detected contacts into interactions with user interface objects (e.g., one or more softkeys, icons, web pages, or images) displayed on touchscreen 112. In an exemplary embodiment, the point of contact between touchscreen 112 and the user corresponds to the user's finger.

[0083] Touchscreen 112 optionally uses LCD (liquid crystal display), LPD (light emitting polymer display), or LED (light emitting diode) technology, although other display technologies are used in other embodiments. Touchscreen 112 and display controller 156 optionally use any of a number of now known or later developed touch sensing technologies to detect contact and any movement or disruption thereof, including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more points of contact with touchscreen 112. In an exemplary embodiment, projected mutual capacitance sensing technology is used, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California.

[0084] The touch-sensitive display in some embodiments of touchscreen 112 is optionally similar to the multi-touch-sensing touchpad described in U.S. Patent Nos. 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman), and / or U.S. Patent Publication No. 2002 / 0015024 A1, each of which is incorporated by reference herein in its entirety. However, touchscreen 112 displays visual output from device 100, whereas touch-sensitive touchpads do not provide visual output.

[0085] Touch-sensitive displays in some embodiments of touchscreen 112 are described in the following applications: (1) U.S. patent application Ser. No. 11 / 381,313, filed May 2, 2006, entitled "Multipoint Touch Surface Controller"; (2) U.S. patent application Ser. No. 10 / 840,862, filed May 6, 2004, entitled "Multipoint Touchscreen"; (3) U.S. patent application Ser. No. 10 / 903,964, filed July 30, 2004, entitled "Gestures For Touch Sensitive Input Devices"; (4) U.S. patent application Ser. No. 11 / 048,264, filed January 31, 2005, entitled "Gestures For Touch Sensitive Input Devices"; (5) U.S. patent application Ser. No. 11 / 038,590, filed January 18, 2005, entitled "Mode-Based Graphical User Interfaces For Touch Sensitive Input"; No. 11 / 228,758, filed September 16, 2005, entitled "Virtual Input Device Placement On A Touch Screen User Interface," (7) U.S. Patent Application No. 11 / 228,700, filed September 16, 2005, entitled "Operation Of A Computer With A Touch Screen Interface," (8) U.S. Patent Application No. 11 / 228,737, filed September 16, 2005, entitled "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard," and (9) U.S. Patent Application No. 11 / 367,749, filed March 3, 2006, entitled "Multi-Functional Hand-Held Device," all of which are incorporated herein by reference in their entireties.

[0086] Touchscreen 112 optionally has a video resolution greater than 100 dpi. In some embodiments, the touchscreen has a video resolution of approximately 160 dpi. A user optionally contacts touchscreen 112 using any suitable object or accessory, such as a stylus, a finger, or the like. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of ​​a finger on the touchscreen. In some embodiments, the device translates the coarse finger input into precise pointer / cursor positions or commands to perform the action desired by the user.

[0087] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad (not shown) for activating or deactivating certain functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is optionally a touch-sensitive surface separate from touchscreen 112 or an extension of the touch-sensitive surface formed by the touchscreen.

[0088] Device 100 also includes a power system 162 that provides power to the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, power failure detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of electrical power within a portable device.

[0089] Device 100 also optionally includes one or more light sensors 164. FIG. 1A shows a light sensor coupled to light sensor controller 158 in I / O subsystem 106. Light sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Light sensor 164 receives light from the environment projected through one or more lenses and converts the light into data representing an image. Light sensor 164 optionally works in conjunction with imaging module 143 (also called a camera module) to capture still images or video. In some embodiments, the light sensor is located on the back side of device 100 opposite touchscreen display 112 on the front of the device, so that the touchscreen display can be used as a viewfinder for capturing still images and / or video. In some embodiments, the light sensor is located on the front of the device so that an image of a user is optionally acquired for a videoconference while the user views other videoconference participants on the touchscreen display. In some embodiments, the position of the light sensor 164 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), so that a single light sensor 164 is used for both video conferencing and capturing still images and / or video, along with a touchscreen display.

[0090] Device 100 also optionally includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled to intensity sensor controller 159 in I / O subsystem 106. Contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of a contact on a touch-sensitive surface). Contact intensity sensor 165 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with or proximate to the touch-sensitive surface (e.g., touch-sensitive display system 112). In some embodiments, at least one contact intensity sensor is located on the back of device 100, opposite touchscreen display 112, which is located on the front of device 100.

[0091] Device 100 also optionally includes one or more proximity sensors 166. Figure 1A shows proximity sensor 166 coupled to peripheral interface 118. Alternatively, proximity sensor 166 is optionally coupled to input controller 160 within I / O subsystem 106. Proximity sensor 166 optionally functions as described in U.S. patent application Ser. Nos. 11 / 241,839, "Proximity Detector In Handheld Device," 11 / 240,788, "Proximity Detector In Handheld Device," 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output," 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices," and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals," which are incorporated herein by reference in their entireties. In some embodiments, when the multifunction device is placed near the user's ear (eg, when the user is making a phone call), the proximity sensor turns off and disables touchscreen 112.

[0092] Device 100 also optionally includes one or more tactile output generators 167. FIG. 1A shows tactile output generators coupled to haptic feedback controller 161 in I / O subsystem 106. Tactile output generator 167 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile output on the device). Contact intensity sensor 165 receives tactile feedback generation instructions from haptic feedback module 133 and generates a tactile output on device 100 that can be sensed by a user of device 100. In some embodiments, at least one tactile output generator is juxtaposed with or proximate to a touch-sensitive surface (e.g., touch-sensitive display system 112) and generates a tactile output, optionally by moving the touch-sensitive surface vertically (e.g., in / out of the surface of device 100) or horizontally (e.g., back and forth in the same plane as the surface of device 100). In some embodiments, at least one tactile output generator sensor is located on the back of device 100, opposite touchscreen display 112, which is located on the front of device 100.

[0093] Device 100 also optionally includes one or more accelerometers 168. FIG. 1A shows accelerometer 168 coupled to peripherals interface 118. Alternatively, accelerometer 168 is optionally coupled to input controller 160 in I / O subsystem 106. Accelerometer 168 optionally functions as described in U.S. Patent Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices," and U.S. Patent Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer," both of which are incorporated by reference herein in their entireties. In some embodiments, information is displayed on the touchscreen display in portrait or landscape orientation based on an analysis of data received from the one or more accelerometers. Device 100 optionally includes, in addition to accelerometer(s) 168, a magnetometer (not shown), and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the location and orientation (e.g., portrait or landscape) of device 100.

[0094] In some embodiments, software components stored in memory 102 include operating system 126, communications module (or instruction set) 128, touch / motion module (or instruction set) 130, graphics module (or instruction set) 132, text input module (or instruction set) 134, Global Positioning System (GPS) module (or instruction set) 135, and applications (or instruction sets) 136. Additionally, in some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) stores device / global internal state 157, as shown in FIGS. 1A and 3. Device / global internal state 157 includes one or more of: active application state indicating which applications, if any, are currently active; display state indicating which applications, views, or other information occupy various regions of touchscreen display 112; sensor state including information obtained from the device's various sensors and input control devices 116; and location information regarding the device's position and / or orientation.

[0095] Operating system 126 (e.g., Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or an embedded operating system such as VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage control, power management, etc.) and facilitate communication between various hardware and software components.

[0096] Communications module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by RF circuitry 108 and / or external port 124. External port 124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as, similar to, and / or compatible with the 30-pin connector used on iPod® (trademark of Apple Inc.) devices.

[0097] Contact / motion module 130, optionally in cooperation with display controller 156, detects contact with touchscreen 112 and other touch-sensing devices (e.g., a touchpad or physical click wheel). Contact / motion module 130 includes various software components for performing various operations related to contact detection, such as determining whether contact occurs (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a surrogate for the force or pressure of the contact), determining whether there is contact movement and tracking the movement across the touch-sensitive surface (e.g., detecting one or more finger-drag events), and determining whether the contact has stopped (e.g., detecting a finger-up event or an interruption of the contact). Contact / motion module 130 receives contact data from the touch-sensitive surface. Determining the movement of the contact point, as represented by the series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These actions are optionally applied to a single contact (e.g., a single finger contact) or multiple simultaneous contacts (e.g., "multi-touch" / multiple finger contacts). In some embodiments, contact / motion module 130 and display controller 156 detect contacts on the touchpad.

[0098] In some embodiments, contact / motion module 130 uses one or more sets of intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has “clicked” on an icon). In some embodiments, at least a subset of the intensity thresholds are determined according to software parameters (e.g., the intensity thresholds are not determined by the activation threshold of a particular physical actuator, but can be adjusted without modifying the physical hardware of device 100). For example, the mouse “click” threshold of a trackpad or touchscreen display can be set to any of a wide range of pre-defined thresholds without modifying the trackpad or touchscreen display hardware. Additionally, in some implementations, a user of the device is provided with a software setting to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once via a system-level click “intensity” parameter).

[0099] Contact / motion module 130 optionally detects gesture input by a user. Different gestures on the touch-sensitive surface have different contact patterns (e.g., different movements, timing, and / or intensities of detected contacts). Thus, gestures are optionally detected by detecting particular contact patterns. For example, detecting a finger tap gesture includes detecting a finger down event, followed by detecting a finger up (lift off) event at the same location (or substantially the same location) as the finger down event (e.g., the location of an icon). As another example, detecting a finger swipe gesture on the touch-sensitive surface includes detecting a finger down event, followed by one or more finger drag events, followed by detecting a finger up (lift off) event.

[0100] Graphics module 132 includes various known software components that render and display graphics on touchscreen 112 or other display, including components that vary the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual characteristics) of the displayed graphics. As used herein, the term "graphic" includes any object that can be displayed to a user, including, but not limited to, text, web pages, icons (such as user interface objects including soft keys), digital images, video, animation, etc.

[0101] In some embodiments, graphics module 132 stores data representing graphics to be used. Each graphic is optionally assigned a corresponding code. Graphics module 132 receives one or more codes specifying the graphics to be displayed, including coordinate data and other graphic characteristic data, as needed, from an application or the like, and then generates screen image data to output to display controller 156.

[0102] The tactile feedback module 133 includes various software components for generating instructions used by the tactile output generator 167, which generates tactile outputs at one or more locations on the device 100 in response to a user's interaction with the device 100.

[0103] Text input module 134 is optionally a component of graphics module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).

[0104] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for use in location-based dialing, to the camera 143 as photo / video metadata, and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and map / navigation widgets).

[0105] Application 136 optionally includes the following modules (or sets of instructions), or a subset or superset thereof: • a contacts module 137 (sometimes called an address book or contact list); ●Telephone module 138, ●Videoconferencing module 139, ● an email client module 140; ● Instant messaging (IM) module 141, ●Training support module 142, camera module 143 for still images and / or video; ● Image management module 144; ●Video player module, ●Music player module, ● Browser module 147, ● Calendar module 148, • A widget module 149 optionally including one or more of a weather widget 149-1, a stock price widget 149-2, a calculator widget 149-3, an alarm clock widget 149-4, a dictionary widget 149-5, and other widgets obtained by the user, and a user-created widget 149-6; a widget creator module 150 for creating user-created widgets 149-6; ● Search module 151, A video and music player module 152 that integrates a video player module and a music player module; ● Memo module 153, Map module 154, and / or ●Online video module 155.

[0106] Examples of other applications 136 optionally stored in memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-enabled applications, encryption, digital rights management, voice recognition, and voice duplication.

[0107] Contacts module 137, in conjunction with touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, is optionally used to manage an address book or contact list (e.g., stored in memory 102 or in the application internal state 192 of contacts module 137 in memory 370). These include adding a name to an address book, removing a name(s) from an address book, associating phone number(s), email address(es), physical address(es), or other information with a name, associating an image with a name, categorizing and sorting names, providing phone numbers or email addresses to initiate and / or facilitate communication by phone 138, video conferencing module 139, email 140, or IM 141, and the like.

[0108] Telephone module 138, in conjunction with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, is optionally used to enter character sequences corresponding to telephone numbers, access one or more telephone numbers in contacts module 137, modify entered telephone numbers, dial respective telephone numbers, place calls, and disconnect and hang up when the call is completed. As previously mentioned, wireless communication optionally uses any of a number of communication standards, protocols, and technologies.

[0109] Videoconferencing module 139 includes executable instructions to cooperate with RF circuitry 108, audio circuitry 110, speaker 111, microphone 113, touchscreen 112, display controller 156, light sensor 164, light sensor controller 158, contact / motion module 130, graphics module 132, text input module 134, contact module 137, and telephone module 138 to initiate, conduct, and end a videoconference between a user and one or more other participants according to the user's instructions.

[0110] Email client module 140, in conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, contains executable instructions for composing, sending, receiving, and managing emails in response to user instructions. In conjunction with image management module 144, email client module 140 greatly facilitates the creation and sending of emails with still or video images captured by camera module 143.

[0111] Instant messaging module 141 includes executable instructions, in cooperation with RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, for entering character sequences corresponding to instant messages, modifying previously entered characters, sending respective instant messages (e.g., using Short Message Service (SMS) or Multimedia Message Service (MMS) protocols for telephony-based instant messaging, or XMPP, SIMPLE, or IMPS for Internet-based instant messaging), receiving instant messages, and viewing received instant messages. In some embodiments, sent and / or received instant messages optionally include graphics, photos, audio files, video files, and / or other attachments, such as those supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephony-based messages (e.g., messages sent using SMS or MMS) and Internet-based messages (e.g., messages sent using XMPP, SIMPLE, or IMPS).

[0112] The training support module 142 includes executable instructions to cooperate with the RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, map module 154, and music player module to create workouts (e.g., with time, distance, and / or calorie burn goals), communicate with training sensors (sports devices), receive training sensor data, calibrate sensors used to monitor workouts, select and play music for workouts, and display, store, and transmit workout data.

[0113] Camera module 143, in conjunction with touchscreen 112, display controller 156, light sensor 164, light sensor controller 158, contact / motion module 130, graphics module 132, and image management module 144, contains executable instructions for capturing and storing still images or video (including video streams) in memory 102, modifying the characteristics of the still images or video, or deleting the still images or video from memory 102.

[0114] Image management module 144 includes executable instructions for arranging, modifying (e.g., editing), or otherwise manipulating, labeling, deleting, presenting (e.g., in a digital slideshow or album), and storing still and / or video images in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, and camera module 143.

[0115] Browser module 147, in conjunction with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, and text input module 134, contains executable instructions for browsing the Internet according to user directions, including retrieving, linking to, receiving, and displaying web pages or portions thereof, as well as attachments and other files linked to web pages.

[0116] Calendar module 148 includes executable instructions to cooperate with RF circuitry 108, touch screen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, email client module 140, and browser module 147 to create, display, modify, and store calendars and data associated with the calendars (e.g., calendar items, to-do lists, etc.) according to user instructions.

[0117] Widget module 149, in conjunction with RF circuitry 108, touchscreen 112, display controller 156, touch / motion module 130, graphics module 132, text input module 134, and browser module 147, optionally provides mini-applications (e.g., weather widget 149-1, stock quotes widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) downloaded and used by a user, or mini-applications created by a user (e.g., user-created widget 149-6). In some embodiments, a widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, a widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! Widgets).

[0118] The widget creator module 150, in conjunction with the RF circuitry 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, the text input module 134, and the browser module 147, is optionally used by a user to create a widget (e.g., turn a user-specified portion of a web page into a widget).

[0119] The search module 151 includes executable instructions for working in conjunction with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134 to search for text, music, sound, images, video, and / or other files in the memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) in accordance with user instructions.

[0120] Video and music player module 152 includes executable instructions that, in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, and browser module 147, enable a user to download and play pre-recorded music and other sound files stored in one or more file formats, such as MP3 or AAC files, as well as executable instructions for displaying, presenting, or otherwise playing videos (e.g., on touchscreen 112 or on an external display connected via external port 124). In some embodiments, device 100 optionally includes the functionality of an MP3 player, such as an iPod (a trademark of Apple Inc.).

[0121] The notes module 153 includes executable instructions for working with the touch screen 112, the display controller 156, the contact / motion module 130, the graphics module 132, and the text input module 134 to create and manage notes, to-do lists, and the like as directed by a user.

[0122] Map module 154, in conjunction with RF circuitry 108, touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, text input module 134, GPS module 135, and browser module 147, is used to receive, display, modify, and store maps and data associated with maps (e.g., driving directions, data about businesses and other points of interest at or near a particular location, and other location-based data), optionally in accordance with user instructions.

[0123] Online video module 155, in conjunction with touchscreen 112, display controller 156, contact / motion module 130, graphics module 132, audio circuitry 110, speaker 111, RF circuitry 108, text input module 134, email client module 140, and browser module 147, contains instructions that enable a user to access, browse for, receive (e.g., by streaming and / or downloading), and play (e.g., on the touchscreen or on an external display connected via external port 124) particular online videos, send emails with links to particular online videos, and otherwise manage online videos in one or more file formats, such as H.264. In some embodiments, instant messaging module 141 is used to send links to particular online videos, rather than email client module 140. For additional description of online video applications, see U.S. Provisional Patent Application No. 60 / 936,562, filed June 20, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," and U.S. Patent Application No. 11 / 968,067, filed December 31, 2007, entitled "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos," the contents of which are incorporated herein by reference in their entireties.

[0124] The above-identified modules and applications each correspond to sets of executable instructions that perform one or more of the functions previously described and methods described herein (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. For example, a video player module is optionally combined with a music player module into a single module (e.g., video and music player module 152 of FIG. 1A). In some embodiments, memory 102 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 102 optionally stores additional modules and data structures not described above.

[0125] In some embodiments, device 100 is a device in which operation of a predetermined set of functions on the device is performed solely via a touchscreen and / or touchpad. Using the touchscreen and / or touchpad as the primary input control device for operation of device 100 optionally reduces the number of physical input control devices (push buttons, dials, etc.) on device 100.

[0126] The set of predefined functions performed only through the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 to a main menu, home menu, or root menu from any user interface displayed on device 100. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touchpad.

[0127] 1B is a block diagram illustrating exemplary components for event processing, according to some embodiments. In some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) includes event sorter 170 (e.g., within operating system 126) and a respective application 136-1 (e.g., any of applications 137-151, 155, 380-390 described above).

[0128] Event sorter 170 receives the event information and determines which application 136-1 to deliver the event information to and application view 191 for application 136-1. Event sorter 170 includes event monitor 171 and event dispatcher module 174. In some embodiments, application 136-1 includes application internal state 192 that indicates the current application view that is displayed on touch-sensitive display 112 when the application is active or running. In some embodiments, device / global internal state 157 is used by event sorter 170 to determine which application(s) is currently active, and application internal state 192 is used by event sorter 170 to determine which application(s) is / are currently active, and application internal state 192 is used by event sorter 170 to determine which application view 191 to deliver the event information to.

[0129] In some embodiments, application internal state 192 includes additional information such as one or more of resume information to be used if application 136-1 resumes execution, user interface state information indicating or ready to display information being displayed by application 136-1, state cues that allow the user to return to a previous state or view of application 136-1, and redo / undo cues of previous actions taken by the user.

[0130] Event monitor 171 receives event information from peripherals interface 118. The event information includes information about sub-events (e.g., a user touch as part of a multi-touch gesture on touch-sensitive display 112). Peripherals interface 118 transmits information it receives from I / O subsystem 106 or sensors such as proximity sensor 166, accelerometer(s) 168, and / or microphone 113 (via audio circuitry 110). Information that peripherals interface 118 receives from I / O subsystem 106 includes information from touch-sensitive display 112 or a touch-sensitive surface.

[0131] In some embodiments, event monitor 171 sends requests to peripherals interface 118 at predetermined intervals. In response, peripherals interface 118 transmits event information. In other embodiments, peripherals interface 118 transmits event information only when there is a significant event (e.g., receipt of an input above a predetermined noise threshold and / or for more than a predetermined duration).

[0132] In some embodiments, event sorter 170 also includes a hit view determination module 172 and / or an active event recognizer determination module 173 .

[0133] Hit view determination module 172 provides software procedures that determine where a sub-event occurred within one or more views when touch-sensitive display 112 displays more than one view. A view consists of the controls and other elements that a user can see on the display.

[0134] Another aspect of a user interface associated with an application is the set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (of the respective application) in which the touch is detected optionally corresponds to a programmatic level within the application's programmatic or view hierarchy. For example, the lowest-level view in which the touch is detected is optionally referred to as the hit view, and the set of events that are recognized as appropriate inputs is optionally determined based at least in part on the hit view of the initial touch that initiates the touch gesture.

[0135] Hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has multiple views organized hierarchically, hit view determination module 172 identifies the hit view as the lowest view in the hierarchy that should process the sub-events. In most situations, the hit view is the lowest-level view in which an initiating sub-event occurs (e.g., the first sub-event in a sequence of sub-events that form an event or potential event). Once a hit view is identified by hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source as the touch or input source identified as the hit view.

[0136] Active event recognizer determination module 173 determines which view(s) in the view hierarchy should receive the particular sequence of sub-events. In some embodiments, active event recognizer determination module 173 determines that only the hit view should receive the particular sequence of sub-events. In other embodiments, active event recognizer determination module 173 determines that all views that contain the physical location of the sub-event are actively participating views, and therefore all actively participating views should receive the particular sequence of sub-events. In other embodiments, even if the touch sub-event is completely confined to the area associated with one particular view, views higher in the hierarchy still remain actively participating views.

[0137] Event dispatcher module 174 dispatches event information to event recognizers (e.g., event recognizer 180). In embodiments that include active event recognizer determination module 173, event dispatcher module 174 delivers the event information to the event recognizers determined by active event recognizer determination module 173. In some embodiments, event dispatcher module 174 stores event information obtained by each event receiver 182 in an event queue.

[0138] In some embodiments, operating system 126 includes event sorter 170. Alternatively, application 136-1 includes event sorter 170. In still other embodiments, event sorter 170 is a stand-alone module or is part of another module stored in memory 102, such as contact / motion module 130.

[0139] In some embodiments, application 136-1 includes multiple event handlers 190 and one or more application views 191, each containing instructions for processing touch events that occur within a respective view of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognizers 180. Typically, each application view 191 includes multiple event recognizers 180. In other embodiments, one or more of event recognizers 180 are part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 136-1 inherits methods and other properties. In some embodiments, each event handler 190 includes one or more of data updater 176, object updater 177, GUI updater 178, and / or event data 179 received from event sorter 170. Event handler 190 optionally utilizes or invokes data updater 176, object updater 177, or GUI updater 178 to update application internal state 192. Alternatively, one or more of the application views 191 include one or more respective event handlers 190. Also, in some embodiments, one or more of the data updater 176, the object updater 177, and the GUI updater 178 are included in each application view 191.

[0140] Each event recognizer 180 receives event information (e.g., event data 179) from event sorter 170 and identifies an event from the event information. Event recognizer 180 includes event receiver 182 and event comparator 184. In some embodiments, event recognizer 180 also includes metadata 183 and at least a subset of event delivery instructions 188 (optionally including sub-event delivery instructions).

[0141] Event receiver 182 receives event information from event sorter 170. The event information includes information about a sub-event, e.g., a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information, such as the position of the sub-event. When the sub-event involves a movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (e.g., from portrait to landscape or vice versa), and the event information includes corresponding information about the current orientation of the device (also called the device's attitude).

[0142] The event comparator 184 compares the event information with predefined event or sub-event definitions and determines the event or sub-event, or determines or updates the state of the event or sub-event, based on the comparison. In some embodiments, the event comparator 184 includes an event definition 186. The event definition 186 includes definitions of events (e.g., a predefined sequence of sub-events), such as Event 1 (187-1) and Event 2 (187-2). In some embodiments, sub-events within Event 1 (187) include, for example, touch start, touch end, touch movement, touch cancellation, and multiple touches. In one example, the definition for Event 1 (187-1) is a double tap on a displayed object. The double tap includes, for example, a first touch on a displayed object relative to a predetermined stage (touch start), a first lift-off (touch end) relative to the predetermined stage, a second touch on a displayed object relative to the predetermined stage (touch start), and a second lift-off (touch end) relative to the predetermined stage. In another example, a definition of event 2 (187-2) is a drag on a displayed object. Drag includes, for example, a touch (or contact) on the displayed object to a predetermined stage, a movement of the touch across the touch-sensitive display 112, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event handlers 190.

[0143] In some embodiments, event definition 187 includes a definition of the event for each user interface object. In some embodiments, event comparator 184 performs a hit test to determine which user interface object is associated with the sub-event. For example, if a touch is detected on touch-sensitive display 112 in an application view in which three user interface objects are displayed on touch-sensitive display 112, event comparator 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub-event). If each displayed object is associated with a respective event handler 190, event comparator 184 uses the results of the hit test to determine which event handler 190 to activate. For example, event comparator 184 selects the event handler associated with the sub-event and object that triggers the hit test.

[0144] In some embodiments, each event 187 definition also includes a delay action that delays transmission of the event information until it is determined whether the sequence of sub-events corresponds to the event type of the event recognizer.

[0145] If the respective event recognizer 180 determines that the sequence of sub-events does not match any of the events in the event definition 186, the respective event recognizer 180 enters an event-disabled, event-failed, or event-ended state and thereafter ignores the next sub-event of the touch-based gesture. In this situation, any other event recognizers that remain active for the hit view continue to track and process sub-events of the ongoing touch-based gesture.

[0146] In some embodiments, each event recognizer 180 includes metadata 183 with configurable properties, flags, and / or lists that indicate to actively participating event recognizers how the event delivery system should perform sub-event delivery. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how event recognizers interact with each other or how event recognizers are allowed to interact with each other. In some embodiments, metadata 183 includes configurable properties, flags, and / or lists that indicate how sub-events are delivered to various levels in the view or programmatic hierarchy.

[0147] In some embodiments, each event recognizer 180 activates an event handler 190 associated with an event when one or more specific sub-events of the event are recognized. In some embodiments, each event recognizer 180 delivers event information associated with the event to the event handler 190. Activating the event handler 190 is separate from sending (and postponing sending) sub-events to the respective hit view. In some embodiments, the event recognizer 180 pops a flag associated with the recognized event, and the event handler 190 associated with the flag captures the flag and performs a predetermined process.

[0148] In some embodiments, the event delivery instructions 188 include sub-event delivery instructions that deliver event information about a sub-event without activating an event handler. Instead, the sub-event delivery instructions deliver the event information to an event handler associated with a set of sub-events or to an actively participating view. The event handler associated with the set of sub-events or the actively participating view receives the event information and performs predetermined processing.

[0149] In some embodiments, data updater 176 creates and updates data used by application 136-1. For example, data updater 176 updates phone numbers used by contacts module 137 or stores video files used by a video player module. In some embodiments, object updater 177 creates and updates objects used by application 136-1. For example, object updater 177 creates new user interface objects or updates the positions of user interface objects. GUI updater 178 updates the GUI. For example, GUI updater 178 prepares display information and sends the display information to graphics module 132 for display on the touch-sensitive display.

[0150] In some embodiments, event handler(s) 190 include or have access to data updater 176, object updater 177, and GUI updater 178. In some embodiments, data updater 176, object updater 177, and GUI updater 178 are included in a single module of the respective application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.

[0151] It should be understood that the foregoing description of event processing of a user's touch on a touch-sensitive display also applies to other forms of user input for operating multifunction device 100 using input devices, although not all of them are initiated on a touchscreen. For example, mouse movements and mouse button presses, contact movements such as tapping, dragging, scrolling on a touchpad, optionally coordinated with single or multiple keyboard presses or holds, pen stylus input, device movement, verbal commands, detected eye movement, biometric input, and / or any combination thereof, are optionally utilized as inputs corresponding to sub-events that define the recognized event.

[0152] FIG. 2 illustrates portable multifunction device 100 having touchscreen 112, according to some embodiments. The touchscreen optionally displays one or more graphics within user interface (UI) 200. In this embodiment, as well as other embodiments described below, a user may select one or more of the graphics by performing a gesture on the graphics, for example, using one or more fingers 202 (not drawn to scale) or one or more styluses 203 (not drawn to scale). In some embodiments, selection of one or more graphics is performed when the user breaks contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward and / or downward), and / or rolling (right to left, left to right, upward and / or downward) of a finger in contact with device 100. In some implementations or situations, accidental contact with a graphic does not select the graphic, for example, if the gesture corresponding to selection is a tap, a swipe gesture sweeping over an application icon optionally does not select the corresponding application.

[0153] Device 100 also optionally includes one or more physical buttons, such as a "home" button or menu button 204. As previously mentioned, menu button 204 is optionally used to navigate to any application 136 within a set of applications running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on touchscreen 112.

[0154] In some embodiments, device 100 includes touchscreen 112, menu button 204, pushbuttons 206 for powering the device on / off and locking the device, volume control buttons 208, subscriber identity module (SIM) card slot 210, headset jack 212, and external docking / charging port 124. Pushbutton 206 is optionally used to power the device on / off by pressing and holding the button down for a predetermined period of time, to lock the device by pressing and releasing the button before the predetermined time has elapsed, and / or to unlock the device or initiate the unlocking process. In alternative embodiments, device 100 also accepts verbal input via microphone 113 to activate or deactivate certain functions. Device 100 also optionally includes one or more contact intensity sensors 165 for detecting the intensity of a contact on touchscreen 112 and / or one or more tactile output generators 167 for generating a tactile output for a user of device 100.

[0155] FIG. 3 is a block diagram of an exemplary multifunction device having a display and a touch-sensitive surface, according to some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, a desktop computer, a tablet computer, a multimedia player device, a navigation device, an educational device (such as a child's learning toy), a gaming system, or a control device (e.g., a home or commercial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more network or other communication interfaces 360, memory 370, and one or more communication buses 320 interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communication between system components. Device 300 includes input / output (I / O) interface 330, including display 340, which is typically a touchscreen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350 and a touchpad 355, a tactile output generator 357 (e.g., similar to tactile output generator 167 described above with reference to FIG. 1A ) that generates tactile output on device 300, and sensors 359 (e.g., light, acceleration, proximity, touch-sensing, and / or contact intensity sensors similar to contact intensity sensor 165 described above with reference to FIG. 1A ). Memory 370 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 370 optionally includes one or more storage devices located remotely from CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or a subset of, programs, modules, and data structures stored in memory 102 of portable multifunction device 100 (FIG. 1A). Additionally, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunction device 100. For example, memory 370 of device 300 optionally stores drawing module 380, presentation module 382, ​​word processing module 384, website creation module 386, disc authoring module 388, and / or spreadsheet module 390, whereas memory 102 of portable multifunction device 100 (FIG. 1A) optionally does not store these modules.

[0156] Each of the above-identified elements of FIG. 3 is optionally stored in one or more of the memory devices mentioned above. Each of the above-identified modules corresponds to an instruction set that performs the function described above. The above-identified modules or programs (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules; thus, in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. In some embodiments, memory 370 optionally stores a subset of the above-identified modules and data structures. Additionally, memory 370 optionally stores additional modules and data structures not described above.

[0157] Attention is now optionally directed to user interface embodiments, for example, as implemented on portable multifunction device 100.

[0158] 4A shows an exemplary user interface for a menu of applications on portable multifunction device 100, according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, user interface 400 includes the following elements, or a subset or superset thereof: signal strength indicator(s) 402 for wireless communication(s), such as cellular and Wi-Fi signals; ●Time 404, ●Bluetooth indicator 405, ● Battery status indicator 406, Tray 408 with icons of frequently used applications, such as: An icon 416 for the phone module 138, labeled "Phone," optionally including an indicator 414 of the number of missed calls or voicemail messages; icon 418 of the email client module 140, labeled "Mail," optionally including an indicator 410 of the number of unread emails; ○ An icon 420 for the browser module 147, labeled "Browser"; and ○ An icon 422 for the video and music player module 152, also called the iPod (trademark of Apple Inc.) module 152, labeled "iPod"; and ● Icons of other applications, such as: ○ Icon 424 of IM module 141, labeled "Messages"; icon 426 of the calendar module 148, labeled "Calendar"; ○ Icon 428 of the image management module 144, labeled "Photos" ○ An icon 430 of the camera module 143, labeled "camera"; ○ Icon 432 of the online video module 155, labeled "Online Video"; Icon 434 of Stock Price Widget 149-2, labeled "Stock Price" ○ Icon 436 of the map module 154, labeled "Map" Icon 438 of weather widget 149-1, labeled "Weather" ○ Icon 440 of alarm clock widget 149-4, labeled "Clock" ○ Icon 442 of Training Support Module 142, labeled "Training Support"; ○ An icon 444 of the Notes module 153 labeled "Notes," and A settings application or module icon 446 labeled "Settings" that provides access to settings for the device 100 and its various applications 136.

[0159] 4A are merely exemplary. For example, icon 422 of video and music player module 152 is labeled "music" or "music player," although other labels are optionally used for various application icons. In some embodiments, the label for each application icon includes the name of the application corresponding to the respective application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to that particular application icon.

[0160] 4B shows an example user interface on a device (e.g., device 300 of FIG. 3 ) that has touch-sensitive surface 451 (e.g., tablet or touchpad 355 of FIG. 3 ) that is separate from display 450 (e.g., touchscreen display 112). Device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of sensors 359) that detect the intensity of a contact on touch-sensitive surface 451, and / or one or more tactile output generators 357 that generate a tactile output for a user of device 300.

[0161] Although some of the following examples are given with reference to input on touchscreen display 112 (which combines a touch-sensitive surface and a display), in some embodiments, the device detects input on a touch-sensitive surface that is separate from the display, as shown in FIG. 4B . In some embodiments, the touch-sensitive surface (e.g., 451 in FIG. 4B ) has a primary axis (e.g., 452 in FIG. 4B ) that corresponds to a primary axis (e.g., 453 in FIG. 4B ) on the display (e.g., 450). According to these embodiments, the device detects contact with touch-sensitive surface 451 (e.g., 460 and 462 in FIG. 4B ) at locations that correspond to respective locations on the display (e.g., in FIG. 4B , 460 corresponds to 468 and 462 corresponds to 470). In this way, user input (e.g., contacts 460 and 462 and their movement) detected by the device on the touch-sensitive surface (e.g., 451 in FIG. 4B ) is used by the device to operate a user interface on the display (e.g., 450 in FIG. 4B ) of the multifunction device when the touch-sensitive surface is separate from the display. It should be understood that similar methods are optionally used for the other user interfaces described herein.

[0162] Additionally, while the following examples are given primarily with reference to finger input (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs are replaced with input from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture is optionally replaced by a mouse click (e.g., instead of a contact) followed by movement of a cursor along the path of the swipe (e.g., instead of movement of the contact). As another example, a tap gesture is optionally replaced by a mouse click (e.g., instead of detecting a contact and then ceasing contact detection) while the cursor is located over the location of the tap gesture. Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice are optionally used simultaneously, or a mouse and finger contacts are optionally used simultaneously.

[0163] FIG. 5A shows an exemplary personal electronic device 500. Device 500 includes a main body 502. In some embodiments, device 500 can include some or all of the features described with respect to devices 100 and 300 (e.g., FIGS. 1A-4B ). In some embodiments, device 500 has a touch-sensitive display screen 504, hereafter touch screen 504. Alternatively, or in addition to touch screen 504, device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, touch screen 504 (or the touch-sensitive surface) optionally includes one or more intensity sensors that detect the intensity of contact (e.g., touches) being applied. The one or more intensity sensors of touch screen 504 (or the touch-sensitive surface) can provide output data representing the intensity of the touch. The user interface of device 500 can respond to touches based on their intensity, meaning that touches of different intensities can invoke different user interface actions on device 500.

[0164] For exemplary techniques for detecting and processing touch intensity, see, for example, related applications International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, entitled "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," published as International Patent Application No. WO / 2013 / 169849, and International Patent Application No. PCT / US2013 / 069483, filed November 11, 2013, entitled "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," published as International Patent Application No. WO / 2014 / 105276, each of which is incorporated herein by reference in its entirety.

[0165] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508, if included, may be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms, if included, may allow device 500 to be attached to, for example, hats, eyewear, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, etc. These attachment mechanisms allow device 500 to be worn by a user.

[0166] FIG. 5B illustrates an exemplary personal electronic device 500. In some embodiments, device 500 can include some or all of the components described with respect to FIGS. 1A, 1B, and 3. Device 500 has a bus 512 operably coupling an I / O section 514 to one or more computer processors 516 and memory 518. I / O section 514 can be connected to a display 504, which can have touch-sensing components 522 and, optionally, an intensity sensor 524 (e.g., a contact intensity sensor). Additionally, I / O section 514 can be connected to a communication unit 530 that receives application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication techniques. Device 500 can include input mechanisms 506 and / or 508. Input mechanism 506 is optionally a rotatable input device or a depressible and rotatable input device, for example. In some examples, input mechanism 508 is optionally a button.

[0167] In some examples, the input mechanism 508 is optionally a microphone. The personal electronic device 500 optionally includes various sensors, such as a GPS sensor 532, an accelerometer 534, an orientation sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which may be operably connected to the I / O section 514.

[0168] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions that, when executed by one or more computer processors 516, can cause the computer processors to perform the techniques described below, including, for example, process 700 (FIGS. 7A-7C), process 900 (FIGS. 9A-9C), process 1200 (FIGS. 12A-12B), process 1400 (FIGS. 14A-14B), and process 1500 (FIG. 15). A computer-readable storage medium may be any medium that can tangibly contain or store computer-executable instructions used by or in connection with an instruction execution system, apparatus, or device. In some embodiments, the storage medium is a transient computer-readable storage medium. In some embodiments, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic, optical, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, optical disks based on CD, DVD, or Blu-ray technology, and resident solid-state memory such as flash, solid-state drives, etc. Personal electronic device 500 is not limited to the components and configuration of Figure 5B and may include other or additional components in multiple configurations.

[0169] As used herein, the term "affordance" optionally refers to a user-interactive graphical user interface object displayed on a display screen of device 100, 300, and / or 500 (FIGS. 1A, 3, and 5A-5B). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each, optionally, constitute an affordance.

[0170] As used herein, the term “focus selector” refers to an input element that indicates the current portion of the user interface with which the user is interacting. In some implementations involving a cursor or other position marker, the cursor acts as the “focus selector,” such that when input (e.g., a press input) is detected on a touch-sensitive surface (e.g., touchpad 355 of FIG. 3 or touch-sensitive surface 451 of FIG. 4B) while the cursor is positioned over a particular user interface element (e.g., a button, window, slider, or other user interface element), the particular user interface element is adjusted according to the detected input. In some implementations involving a touchscreen display (e.g., touch-sensitive display system 112 of FIG. 1A or touchscreen 112 of FIG. 4A) that allows direct interaction with user interface elements on the touchscreen display, a detected contact on the touchscreen acts as the “focus selector,” such that when input (e.g., a press input by contact) is detected at the location of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touchscreen display, the particular user interface element is adjusted according to the detected input. In some implementations, focus is moved from one region of the user interface to another region of the user interface without a corresponding cursor movement or contact movement on the touchscreen display (e.g., by using the tab key or arrow keys to move focus from one button to another), and in these implementations, the focus selector moves to follow the movement of focus between various regions of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or contact on a touchscreen display) that is controlled by the user to communicate the user's intended interaction with the user interface (e.g., by indicating to the device the element of the user interface through which the user intends to interact).For example, while a press input is detected on a touch-sensitive surface (e.g., a touchpad or touchscreen), the position of a focus selector (e.g., a cursor, touch, or selection box) over a corresponding button indicates that the user intends to activate that corresponding button (and not other user interface elements shown on the device's display).

[0171] As used herein and in the claims, the term "characteristic intensity" of a contact refers to a characteristic of that contact based on one or more intensities of the contact. In some embodiments, the characteristic intensity is based on a plurality of intensity samples. The characteristic intensity is optionally based on a predetermined number of intensity samples, i.e., a set of intensity samples collected during a predetermined time period (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined event (e.g., after detecting the contact, before detecting lift-off of the contact, before or after detecting the start of contact movement, before detecting the end of the contact, before or after detecting an increase in the intensity of the contact, and / or before or after detecting a decrease in the intensity of the contact). The characteristic intensity of the contact is optionally based on one or more of the maximum intensity of the contact, the mean intensity of the contact, the average intensity of the contact, the top 10 percentile intensity of the contact, half the maximum intensity of the contact, 90 percent of the maximum intensity of the contact, etc. In some embodiments, the duration of the contact is used in determining the characteristic intensity (e.g., where the characteristic intensity is an average of the intensity of the contact over time). In some embodiments, the characteristic intensity is compared to a set of one or more intensity thresholds to determine whether an action is performed by the user. For example, the set of one or more intensity thresholds optionally includes a first intensity threshold and a second intensity threshold. In this example, a contact having a characteristic intensity that does not exceed the first threshold results in a first action, a contact having a characteristic intensity that exceeds the first intensity threshold but not the second intensity threshold results in a second action, and a contact having a characteristic intensity that exceeds the second threshold results in a third action. In some embodiments, the comparison between the characteristic intensity and the one or more thresholds is not used to determine whether the first action or the second action should be performed, but rather is used to determine whether one or more actions should be performed (e.g., whether to perform the respective action or to forgo performing the respective action).

[0172] FIG. 5C illustrates the detection of multiple contacts 552A-552E on the touch-sensitive display screen 504 by multiple intensity sensors 524A-524D. FIG. 5C additionally includes an intensity diagram illustrating the current intensity measurements of intensity sensors 524A-524D relative to intensity units. In this example, intensity sensors 524A and 524D each measure 9 intensity units, and intensity sensors 524B and 524C each measure 7 intensity units. In some implementations, the aggregate intensity is the sum of the intensity measurements of multiple intensity sensors 524A-524D, which in this example is 32 intensity units. In some embodiments, each contact is assigned a respective intensity that is a fraction of the aggregate intensity. FIG. 5D illustrates assigning aggregate intensities to contacts 552A-552E based on their distance from the center of force 554. In this example, contacts 552A, 552B, and 552E are each assigned a contact intensity of 8 intensity units of aggregate intensity, and contacts 552C and 552D are each assigned a contact intensity of 4 intensity units of aggregate intensity. More generally, in some implementations, each contact j is assigned a respective intensity Ij, which is a fraction of a total intensity A, according to a predetermined mathematical function Ij=A·(Dj / ΣDi), where Dj is the distance from the center of force to the respective contact j, and ΣDi is the sum of the distances from the center of force to all respective contacts (e.g., from i=1 to the end). The operations described with reference to FIGS. 5C-5D can be performed using electronic devices similar to or identical to device 100, 300, or 500. In some embodiments, the characteristic intensity of a contact is based on one or more intensities of the contact. In some embodiments, an intensity sensor is used to determine a single characteristic intensity (e.g., a single characteristic intensity of a single contact). Note that the intensity diagrams are not part of the displayed user interface, but are included in FIGS. 5C-5D as an aid to the reader.

[0173] In some embodiments, a portion of the gesture is identified for purposes of determining the characteristic intensity. For example, the touch-sensitive surface optionally receives successive swipe contacts that transition from a start position to an end position, where the intensity of the contact increases. In this example, the characteristic intensity of the contact at the end position is optionally based on only a portion of the successive swipe contacts (e.g., only the portion of the swipe contact at the end position), rather than the entire swipe contact. In some embodiments, a smoothing algorithm is optionally applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of an unweighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or dips in the intensity of the swipe contact for purposes of determining the characteristic intensity.

[0174] The intensity of a contact on the touch-sensitive surface is optionally characterized with respect to one or more intensity thresholds, such as a contact-detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold corresponds to an intensity at which the device performs an action normally associated with clicking a physical mouse button or trackpad. In some embodiments, the deep press intensity threshold corresponds to an intensity at which the device performs an action different from an action normally associated with clicking a physical mouse button or trackpad. In some embodiments, when a contact is detected having a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact-detection intensity threshold below which the contact is not detected), the device follows the movement of the contact on the touch-sensitive surface and moves the focus selector without performing an action associated with the light press intensity threshold or the deep press intensity threshold. In general, unless otherwise specified, these intensity thresholds are consistent across various sets of values ​​for a user interface.

[0175] An increase in the characteristic intensity of a contact from an intensity below the light press intensity threshold to an intensity between the light press intensity threshold and the deep press intensity threshold may be referred to as inputting a "light press." An increase in the characteristic intensity of a contact from an intensity below the deep press intensity threshold to an intensity above the deep press intensity threshold may be referred to as inputting a "deep press." An increase in the characteristic intensity of a contact from an intensity below the contact-detection intensity threshold to an intensity between the contact-detection intensity threshold and the light press intensity threshold may be referred to as detecting a contact on the touch surface. A decrease in the characteristic intensity of a contact from an intensity above the contact-detection intensity threshold to an intensity below the contact-detection intensity threshold may be referred to as detecting a lift-off of the contact from the touch surface. In some embodiments, the contact-detection intensity threshold is zero. In some embodiments, the contact-detection intensity threshold is greater than zero.

[0176] In some embodiments described herein, one or more actions are performed in response to detecting a gesture including a respective press input or in response to detecting a respective press input performed by a respective contact (or multiple contacts), where the respective press inputs are detected based at least in part on detecting an increase in intensity of the contact (or multiple contacts) above a press input intensity threshold. In some embodiments, the respective actions are performed in response to detecting an increase in intensity of the respective contact above the press input intensity threshold (e.g., a “downstroke” of the respective press input). In some embodiments, the press input includes an increase in intensity of the respective contact above the press input intensity threshold followed by a decrease in intensity of the contact below the press input intensity threshold, and the respective actions are performed in response to detecting a subsequent decrease in intensity of the respective contact below the press input threshold (e.g., an “upstroke” of the respective press input).

[0177] Figures 5E-5H show the light press intensity thresholds (e.g., "IT L ") to the deep press intensity threshold (e.g., "IT D5 illustrates the detection of a gesture including a press input corresponding to an increase in the intensity of contact 562 to an intensity above a deep press intensity threshold (e.g., "IT 1"). The gesture performed by contact 562 is detected on touch-sensitive surface 560, and cursor 576 is displayed over application icon 572B corresponding to app2 on display user interface 570, which includes application icons 572A-572D displayed within predetermined region 574. In some embodiments, the gesture is detected on touch-sensitive display 504. An intensity sensor detects the intensity of the contact on touch-sensitive surface 560. The device detects when the intensity of contact 562 exceeds a deep press intensity threshold (e.g., "IT 1"). D ") and reaches a peak. Contact 562 is maintained on touch-sensitive surface 560. In response to detecting the gesture, a deep press intensity threshold (e.g., "IT D "), reduced-scale representations 578A-578C (e.g., thumbnails) of recently opened documents are displayed for App2, as shown in FIGS. 5F-5H. In some embodiments, this intensity, which is compared to one or more intensity thresholds, is the characteristic intensity of the contact. Note that the intensity diagrams for contact 562 are not part of the displayed user interface, but are included in FIGS. 5E-5H as an aid to the reader.

[0178] In some embodiments, the display of representations 578A-578C includes animation. For example, as shown in FIG. 5F, representation 578A is first displayed adjacent to application icon 572B. As the animation progresses, as shown in FIG. 5G, representation 578A moves upward and representation 578B is displayed adjacent to application icon 572B. Then, as shown in FIG. 5H, representation 578A moves upward and representation 578B moves upward toward representation 578A, and representation 578C is displayed adjacent to application icon 572B. Representations 578A-578C form an array above icon 572B. In some embodiments, the animation progresses according to the intensity of contact 562, as shown in FIGS. 5F-5G, as the intensity of contact 562 exceeds a deep press intensity threshold (e.g., "ITD "), representations 578A-578C appear and move upward. In some embodiments, the intensity on which the animation progression is based is a characteristic intensity of the contact. The operations described with reference to FIGS. 5E-5H can be performed using electronic devices similar to or identical to device 100, 300, or 500.

[0179] In some embodiments, the device employs intensity hysteresis to avoid accidental input, sometimes referred to as “jitter,” and the device defines or selects a hysteresis intensity threshold that has a predetermined relationship to the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units below the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Thus, in some embodiments, the press input includes an increase in the intensity of each contact above the press input intensity threshold followed by a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to the press input intensity threshold, and a respective action is performed in response to detecting a subsequent decrease in the intensity of each contact below the hysteresis intensity threshold (e.g., an “upstroke” of each press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of the contact from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, and optionally a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and a respective action is performed in response to detecting the press input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, as the case may be).

[0180] For ease of explanation, descriptions of operations performed in response to a press input associated with a press input intensity threshold, or a gesture including a press input, are optionally triggered in response to detecting any of: an increase in the intensity of the contact above the press input intensity threshold; an increase in the intensity of the contact from an intensity below a hysteresis intensity threshold to an intensity above the press input intensity threshold; a decrease in the intensity of the contact below the press input intensity threshold; and / or a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to the press input intensity threshold. Further, in examples where an operation is described as being performed in response to detecting a decrease in the intensity of the contact below a press input intensity threshold, the operation is optionally performed in response to detecting a decrease in the intensity of the contact below a hysteresis intensity threshold corresponding to and lower than the press input intensity threshold.

[0181] As used herein, an "installed application" refers to a software application that has been downloaded onto an electronic device (e.g., device 100, 300, and / or 500) and is ready to run (e.g., opened) on the device. In some embodiments, a downloaded application becomes an installed application by an installation program that extracts program portions from a downloaded package and integrates the extracted portions with the computer system's operating system.

[0182] As used herein, the terms "open application" or "running application" refer to a software application that has retained state information (e.g., as part of device / global internal state 157 and / or application internal state 192). An open or running application is, optionally, any one of the following types of application: ● the active application currently displayed on the display screen of the device on which the application is being used; background applications (or background processes) that are not currently displayed, but for which one or more processes are being processed by one or more processors; and A suspended or hibernated application that is not running but has state information stored in memory (volatile and non-volatile, respectively) and that can be used to resume execution of the application.

[0183] As used herein, the term "closed application" refers to a software application that does not have retained state information (e.g., state information for a closed application is not stored in the device's memory). Thus, closing an application includes stopping and / or removing the application process for the application and removing state information for the application from the device's memory. Generally, opening a second application during a first application does not close the first application. When the second application is displayed and the first application is terminated, the first application becomes a background application.

[0184] Spatial management of audio includes techniques for modifying the characteristics of sound (e.g., by applying filters) so that a listener perceives the sound as emanating from a particular location in space (e.g., three-dimensional (3D) space). Such techniques can be achieved using speakers, such as headphones, earbuds, or loudspeakers. In some embodiments, for example, when a listener is using headphones, binaural simulation is used to recreate binaural cues that give the listener the illusion that sound is coming from a particular location in space. For example, the listener perceives a sound source as coming from the left of the listener. In another embodiment, the listener perceives a sound source as passing in front of the listener from left to right. This effect can be enhanced by using head tracking to create the illusion that the location of the sound source remains stationary in space, even when the listener's head moves or rotates. In some embodiments, for example, when a listener is using loudspeakers, a similar effect can be achieved using crosstalk cancellation to give the listener the illusion that sound is coming from a particular location in space.

[0185] A head-related transfer function (HRTF) characterizes how the human ear receives sound from various points in space. HRTFs can be based on one or more of the direction, altitude, and distance of the sound. By using HRTFs, a device (e.g., device 100) applies different functions to audio to recreate the directional pattern of the human ear. In some embodiments, a pair of HRTFs for two ears can be used to synthesize binaural sound that the listener perceives as coming from a specific point in space relative to the listener, such as above, below, in front, behind, left, or right of the user, or a combination thereof. A personalized HRTF provides better results for the listener for whom the HRTF is personalized compared to a generic HRTF. In some embodiments, the HRTF is applied to a listener using a listening device such as headphones, earphones, and earbuds.

[0186] In another example, when a device (e.g., device 100) generates sound using two or more loudspeakers, the sound from each loudspeaker is heard through the listener's respective nearest ear, but also through the opposite ear, resulting in crosstalk. Effectively managing the cancellation of this unintended crosstalk helps modify the sound so that the listener perceives it as emanating from a particular location in space.

[0187] A device (e.g., device 100, 300, 500) can also modify the characteristics of multiple audio sources simultaneously (e.g., by applying different filters to each source) to give a listener the illusion that sounds from different audio sources are coming from different locations in space. Such techniques can be achieved using headphones or loudspeakers.

[0188] In some embodiments, modifying a stereo sound source so that a listener perceives the sound as emanating from a particular location in space (e.g., 3D space) includes generating a mono sound from the stereo sound. For example, the stereo sound includes a left audio channel and a right audio channel. The left audio channel, for example, includes a first device without including a second device. The second audio channel, for example, includes a second device without including the first device.

[0189] When a stereo sound source is placed in the space, the device optionally combines the left and right audio channels to form composite channel audio and then applies interaural time difference to the composite channel audio. Further, the device optionally (or alternatively) also applies HRTFs and / or cross-cancellation to the composite channel audio before generating it on the different speakers.

[0190] If stereo sound sources are not positioned in the space, the device optionally does not combine the left and right audio channels and does not apply interaural time difference, HRTF, or cross-cancellation. Instead, the device generates stereo sound by using the device's left loudspeaker to generate the left audio channel and the device's right loudspeaker to generate the right channel. As a result, the device produces sound in stereo, and the listener perceives audio in stereo.

[0191] Many of the techniques described below use various processes to modify the sound so that the listener perceives it as coming from a particular location in space.

[0192] Attention is now directed to embodiments of user interfaces (“UIs”) and related processes implemented on an electronic device such as portable multifunction device 100, device 300, or device 500.

[0193] 6A-6N illustrate example techniques for transitioning between visual elements according to some embodiments. The techniques in these figures are used to illustrate processes described below, including the processes in FIGS. 7A-7C.

[0194] 6A-6G show a user 606 sitting in front of a device 600 (e.g., a laptop computer) having a display 600a, left and right loudspeakers, and a touch-sensitive surface 600b (e.g., a touchpad). Throughout FIGS. 6A-6G and 6K-6N, an additional close-up of touch-sensitive surface 600b of device 600 is shown to the right of the user to provide the reader with a better understanding of the described techniques, particularly with respect to exemplary user input. Similarly, an overhead view 650, a visual representation of the spatial configuration of audio being generated by device 600, is shown throughout FIGS. 6A-6G to provide the reader with a better understanding of the techniques, particularly with respect to where user 606 perceives sound to be coming from (e.g., as a result of device 600 positioning the audio in space). Overhead view 650 is not part of the device's user interface. Similarly, visual elements displayed outside the display device, as represented by dotted outlines, are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the techniques. Similarly, visual elements displayed outside the device's display, as represented by dotted outlines, are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the techniques. Throughout Figures 6A-6G, device 600 is running (1) a web browser 604a that includes playback of a basketball game with game audio and game video, (2) a music player 604b that includes playback of music, and (3) a video player 604c that includes playback of show audio and show video. Audio element 654a corresponds to audio provided by web browser 604a, audio element 654b corresponds to audio provided by music player 604b, and audio element 654c corresponds to audio provided by video player 604c. Throughout Figures 6A-6G, device 600 simultaneously generates audio provided by each of web browser 604a, music player 604b, and video player 604c.

[0195] In FIG. 6A , device 600 displays web browser 604a on display 600a. While displaying web browser 604a, device 600 generates game audio provided by web browser 604a using left and right loudspeakers. In FIG. 6A , while displaying web browser 604a, device 600 does not spatially position the game audio of web browser 604a (e.g., device 600 does not apply interaural time difference, HRTF, or cross-cancellation). As a result, user 606 perceives the audio as being in front of user 606 in stereo, as represented in FIG. 6A by the position of audio element 654a relative to user 606 in overhead view 650. For example, the left audio channel of the game audio includes a sports commentator and does not include crowd noise, while the right audio channel includes crowd noise and does not include the sports commentator. While device 600a displays web browser 604a, device 600 uses game audio to generate stereo audio by generating the left audio channel using device 600a's left loudspeaker and the right channel using device 600a's right loudspeaker. As a result, user 606 perceives audio as being in front of user 606 in stereo, as represented in FIG. 6A by the position of audio element 654a relative to user 606 in overhead view 650. In some examples, if web browser 604a was part of the currently accessed desktop, the device generates game audio for web browser 604a in the same manner, even if web browser 604a was not actively displayed (e.g., when a different visual element, such as a word processing application, is displayed on top of web browser 604a, thereby blocking the display of web browser 604a on display 600a). Optionally, audio for various applications follows curved path 650a when device 600 repositions the audio in space. In some embodiments, device 600a positions various sound sources equidistant in space from adjacent sound sources.In some embodiments, device 600a moves various sound sources in only two axes (eg, left-right and front-back, but not up-down).

[0196] In FIG. 6A , device 600 does not display music player 604b. While not displaying music player 604b, device 600 generates music provided by music player 604b using left and right loudspeakers. In FIG. 6A , device 600 spatially positions the music provided by music player 604b (e.g., device 600 applies interaural time difference, HRTF, and / or cross-cancellation to the music). Device 600 positions the music so that the user perceives it as coming from a location in space to the right of display 600a (and user 606), as indicated by audio element 654b. This allows the user to recognize that music player 604b is running and generating audio, even when music player 604b is not displayed. Furthermore, the spatial placement of music helps the user recognize how to access the display of music player 604c, as described in FIGS. 6B-6G .

[0197] In FIG. 6A , device 600 is not displaying video player 604c. While not displaying video player 604c, device 600 uses left and right loudspeakers to generate audio provided by video player 604c. In FIG. 6A , device 600 spatially positions show audio provided by video player 604b (e.g., device 600 applies interaural time difference, HRTF, and / or cross-cancellation to the music). Device 600 positions the show audio so that the user perceives it as coming from a location in space further to the right of display 600a (and user 606) than the music from music player 604d. This allows the user to recognize that video player 604c is running and generating audio even when video player 604c is not displayed. Furthermore, the placement of show audio in space helps the user recognize how to access the display of video player 604c.

[0198] In addition, device 600 optionally applies a low-pass (or high-pass, or band-pass) filter to audio corresponding to applications that are not on display (e.g., not part of the currently accessed desktop), thereby attenuating (e.g., removing) audio above a certain frequency threshold before the audio is produced by the loudspeakers. As a result, the user perceives such audio as background noise compared to audio to which the low-pass filter is not applied. This allows the user to more easily hear audio from particular applications, such as applications that are not currently displayed. In some embodiments, the same low-pass filter is applied to all audio corresponding to applications that are not displayed. In some embodiments, different low-pass filters are applied to each piece of audio based on how far the corresponding application should be perceived by the user from the display. In some embodiments, device 600 optionally attenuates audio corresponding to applications that are not on display (e.g., across all frequencies of the audio).

[0199] 6A , for example, device 600 does not apply a low-pass filter to game audio provided by web browser 604 a, device 600 applies a first low-pass filter having a first cutoff frequency to music provided by music player 604 b, and device 600 applies a second low-pass filter having a second cutoff frequency (lower than the first cutoff frequency) to show audio provided by video player 604 c. Thus, in FIG. 6A , device 600 simultaneously generates audio from each of web browser 604 a, music player 604 b, and video player 604 c.

[0200] 6B-6C, device 600 receives left swipe user input 610a on touch-sensitive surface 600b. In response to receiving left swipe user input 610a, device 600 transitions the display of web browser 604a away from display 600a by sliding web browser 604a to the left, and transitions the display of music player 604b onto display 600a by sliding music player 604b to the left, as shown in Figures 6B-6C. Furthermore, in response to receiving left swipe user input 610a, device 600 changes the location in space at which the user perceives audio from the corresponding application, as shown in overhead view 650 of Figures 6B-6C.

[0201] 6D , device 600 is not displaying web browser 604a. While not displaying web browser 604a, device 600 uses left and right loudspeakers to position the audio provided by web browser 604a in space (e.g., device 600 applies interaural time difference, HRTF, and / or cross-cancellation to the audio) so that the user perceives the game audio as coming from a location in space to the left of display 600a (and user 606). This allows the user to recognize that web browser 604a is running and generating audio, even when web browser 604a is not displayed. Furthermore, the positioning of the game audio in space helps the user recognize how to access the display of web browser 604a (e.g., using a swipe-right user input).

[0202] In Figure 6D, device 600 displays music player 604b on display 600a. While displaying music player 604b, device 600 generates audio provided by music player 604b using left and right loudspeakers. In Figure 6D, while displaying music player 604b, device 600 does not spatially position the music provided by music player 604b (e.g., device 600 does not apply interaural time difference, HRTF, or cross-cancellation). As a result, user 606 perceives the audio as being in front of user 606 in stereo, as represented in overhead view 650 of Figure 6D by the position of audio element 654c relative to user 606.

[0203] In FIG. 6D , device 600 is not displaying video player 604c. While not displaying video player 604c, device 600 uses left and right loudspeakers to position the audio provided by video player 604c in space (e.g., device 600 applies interaural time difference, HRTF, and / or cross-cancellation to the audio) so that the user perceives the audio provided by video player 604c as coming from a location in space to the right of display 600a (and user 606), not as far to the right as previously perceived by the user in FIG. 6A . This allows the user to recognize that video player 604c is running and generating audio even when video player 604c is not displayed. Furthermore, the positioning of the show audio in space helps the user recognize how to access the display of video player 604c (e.g., using a swipe-left user input).

[0204] 6D , for example, device 600 does not apply a low-pass filter to music provided by music player 604 b, device 600 applies a first low-pass filter with a first cutoff frequency to game audio provided by web browser 604 a, and device 600 applies a first low-pass filter with a first cutoff frequency to show audio provided by video player 604 c. Thus, in FIG. 6D , device 600 simultaneously generates audio from each of web browser 604 a, music player 604 b, and video player 604 c.

[0205] 6E-6F, device 600 receives right swipe user input 610b on touch-sensitive surface 600b. In response to receiving right swipe user input 610b, device 600 transitions the display of web browser 604a onto display 600a by sliding web browser 604a to the right, and transitions the display of music player 604b away from display 600a by sliding music player 604b to the right, as shown in FIGS. 6E-6F. Furthermore, in response to receiving right swipe user input 610b, device 600 changes the location in space at which the user perceives audio from the corresponding application, as shown in overhead view 650 of FIGS. 6E-6F. In this example, the device modifies the audio so that the user perceives the audio as described in FIG. 6A.

[0206] FIG. 6G shows an example corresponding to FIG. 6A. In FIG. 6G, a user is listening to audio generated by device 600a using headphones. As a result, rather than perceiving the game audio of web browser 604a as being in front of the user, the user perceives the game audio as being generated within the user's head. Optionally, the audio of various applications follows a linear path 650b when device 600 repositions the audio in space. In some examples, device 600a positions various sound sources equidistant in space from adjacent sound sources. In some examples, device 600a moves various sound sources in only two axes (e.g., left-right and front-back, but not up-down).

[0207] 6H-6J show device 660 (e.g., a mobile phone) having display 660a (e.g., a touchscreen), touch-sensitive surface 660b (e.g., a portion of the touchscreen), and connected (wirelessly, wired) to headphones. In this example, user 606 listens to device 660 using the headphones.

[0208] An overhead view 670, a visual representation of the spatial configuration of the audio being generated by device 660, is shown throughout FIGS. 6H-6J to provide the reader with a better understanding of the technique, particularly with respect to where user 606 perceives sound as coming from (e.g., as a result of device 660 positioning the audio in space). The overhead view 670 is not part of the user interface of device 660. Similarly, visual elements displayed outside of the display device, as represented by dotted outlines, are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the technique. Throughout FIGS. 6H-6J, device 660 is running a music player that includes playing music with corresponding album art.

[0209] Audio element 674a corresponds to the audio for track 1 provided by the music player, audio element 674b corresponds to the audio for track 2 provided by the music player, audio element 674c corresponds to the audio for track 3 provided by the music player, and audio element 674d corresponds to the audio for track 4 provided by the music player.

[0210] In FIG. 6H , device 660 does not display album art 664a for track 1 on display 660a. While not displaying album art 664a for track 1, device 660 generates audio for track 1 by positioning the audio for track 1 in space using the left and right speakers of the headphones (e.g., device 660 applies interaural time difference, HRTF, and / or cross-cancellation), so that the user perceives the audio as coming from a first point in space (e.g., to the left of the user, to the left of the device), as illustrated in FIG. 6H by the position of audio element 670a relative to user 606 in overhead view 670. In this example, device 660 additionally modifies the audio for track 1 by attenuating the audio and / or applying a low-pass (or high-pass, or band-pass) filter to the audio. In some examples, device 660 does not generate audio for track 1 if corresponding album art is not on display.

[0211] In FIG. 6H , device 660 displays album art 664b for track 2 on display 660a. While displaying album art 664a for track 2, device 600 generates the audio for track 2 using left and right headphones by positioning the audio for track 2 in space (e.g., device 600 applies interaural time difference, HRTF, or cross-cancellation), so that the user perceives the audio as coming from a second point in space (e.g., different from the first point in space, at a position corresponding to device 660, and to the right of the first point in space in front of the user), as shown in FIG. 6H by the position of audio element 670b relative to user 606 in overhead view 670. In this example, device 660 does not modify the audio for track 2 by attenuating the audio or applying a low-pass (or high-pass, or band-pass) filter to the audio.

[0212] 6H, device 660 does not display album art 664c for track 3 on display 660a. Device 660 also does not produce audio for track 3 using the left or right speaker of the headphones.

[0213] 6I-6J , device 660 receives left swipe user input 666 on touch-sensitive surface 660a. In response to receiving left swipe user input 666, device 660 transitions the display of album art 664b away from display 660a by sliding album art 664b to the left, and transitions the display of album art 664c onto display 660a by sliding album art 664c to the left, as shown in FIGS. 6I-6J . Further, in response to receiving left swipe user input 666, device 660 begins generating audio for track 3 (simultaneously with track 2), changing the location in space at which the user perceives audio from tracks 2 and 3, as represented in overhead view 670 of FIGS. 6I-6J . In some examples, generating audio for track 3 in response to receiving left swipe user input 666 includes skipping a predetermined amount of audio (e.g., the first 0.5 seconds of track 2). This provides the user with a sense that track 3 was previously playing, even though device 660 was not previously producing audio for track 3.

[0214] In FIG. 6J , device 660 stops generating audio for track 1. Device 660 generates audio for track 2 by positioning the audio for track 2 in space using the left and right speakers of the headphones (e.g., device 660 applies interaural time difference, HRTF, and / or cross-cancellation), so that the user perceives the audio as coming from a first point in space (e.g., to the left of the user, to the left of the device), as shown in FIG. 6J by the position of audio element 670b relative to user 606 in overhead view 670. In this example, device 660 additionally modifies the audio for track 2 by attenuating the audio and / or applying a low-pass (or high-pass, or band-pass) filter to the audio. In some examples, device 660 fades out the audio for track 2 (by stopping generating the audio) as the corresponding album art moves away from the display.

[0215] In FIG. 6J , device 660 displays album art 664c for track 3 on display 660a. While displaying album art 664b for track 3, device 600 generates the audio for track 3 using left and right headphones by positioning the audio for track 3 in space (e.g., device 600 applies interaural time difference, HRTF, or cross-cancellation), so that the user perceives the audio as coming from a second point in space (e.g., different from the first point in space, at a position corresponding to device 660, and to the right of the first point in space in front of the user), as shown in FIG. 6H by the position of audio element 670c relative to user 606 in overhead view 670. In this example, device 660 does not modify the audio for track 3 by attenuating the audio or applying a low-pass (or high-pass, or band-pass) filter to the audio.

[0216] As a result, the user 606 perceives the music passing in front of them as they swipe through various album art.

[0217] 6K-6N show a user 606 sitting in front of a device 600 (e.g., a laptop computer) having a display 600a, left and right loudspeakers, and a touch-sensitive surface 600b (e.g., a touchpad). Throughout FIGS. 6K-6N, an additional close-up of touch-sensitive surface 600b of device 600 is shown to the right of the user to provide the reader with a better understanding of the described techniques, particularly with respect to user input (or lack thereof). Similarly, an overhead view 680, a visual representation of the spatial configuration of audio being generated by device 600, is shown throughout FIGS. 6K-6N to provide the reader with a better understanding of the techniques, particularly with respect to where user 606 perceives sound to be coming from (e.g., as a result of device 600 positioning the audio in space). Overhead view 680 is not part of the user interface of device 600. Similarly, visual elements displayed outside the display device, as represented by dotted outlines, are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the techniques.

[0218] In FIG. 6K, music player 604b is playing music. While displaying music player 604a, device 600 uses left and right speakers to generate the audio provided by music player 604a. For example, device 600 does not spatially position the audio provided by music player 604a (e.g., device 600 does not apply interaural time difference, HRTF, or cross-cancellation). As a result, user 606 perceives stereo music as being in front of user 606, as represented in FIG. 6K by the position of audio element 680a relative to user 606 in overhead view 680. Throughout FIGS. 6K-6N, audio element 680a corresponds to the audio being provided by music player 604b.

[0219] In FIG. 6L , device 600 receives a notification (e.g., a message notification received via a network connection). In response to receiving the notification (and without receiving user input), device 600 transitions to generating music provided by music player 604a using left and right speakers by spatially positioning the music (e.g., device 600 applies interaural time difference, HRTF, or cross-cancellation) so that the user perceives the audio as coming from a point in space to the left of device 600, as illustrated in FIGS. 6L-6M by the position of audio element 680a relative to user 606 in overhead view 680. Device 600 maintains a display of music player 604a on display 600a. In some embodiments, device 600 also displays the notification in response to receiving the notification. Throughout FIGS. 6K-6N , audio element 680b corresponds to the audio being provided by the notification.

[0220] Further in response to receiving the notification (and without receiving user input), device 600 transitions from not generating notification audio to generating the notification audio by spatially positioning the notification audio using the left and right speakers (e.g., device 600 applies interaural time difference, HRTF, or cross-cancellation) so that the user perceives the notification audio as coming from a point in space to the right of device 600, as illustrated in FIG. 6L by the position of audio element 680b relative to user 606 in overhead view 680. In some examples, device 600 emphasizes the notification audio by ducking the music provided by music player 604a. For example, device 600 attenuates the music provided by music player 604a while generating the notification audio. Next, device 600 transitions to generating the audio of the notification without positioning the audio in space (e.g., device 600 does not apply interaural time difference, HRTF, or cross-cancellation), as shown in FIG. 6M by the position of audio element 680b relative to user 606 in overhead view 680.

[0221] Device 600 then uses the left and right speakers to generate the audio provided by music player 604a without spatially placing the audio provided by music player 604a, as shown in FIG. 6N (e.g., device 600 does not apply interaural time difference, HRTF, or cross-cancellation).

[0222] In some embodiments, devices 600 and 660 include digital assistants that generate audio feedback, such as returning the results of a query by speaking the results. In some embodiments, devices 600 and 660 generate the audio of the digital assistant by positioning the audio for the digital assistant in space (e.g., over the user's right shoulder), so that the user perceives the digital assistant as remaining stationary in space even when other audio moves in space. In some embodiments, devices 600 and 660 emphasize the audio of the digital assistant by ducking one or more (or all) other audio.

[0223] 7A-7C are flow diagrams illustrating a method for transitioning between visual elements using an electronic device, according to some embodiments. Method 700 is performed on a device (e.g., 100, 300, 500, 600, 660) having a display, the electronic device operatively connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earbuds, left and right earbuds). Some operations of method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0224] As described below, method 700 provides an intuitive way to transition between visual elements. This method reduces the cognitive burden on a user to transition between visual elements, thereby creating a more efficient human-machine interface. For battery-operated computing devices, allowing users to transition between visual elements more quickly and efficiently conserves power and extends the time between battery charges.

[0225] The electronic device displays (702) a first visual element (e.g., a first application 604a, a video playback window, album art) at a first location on the display.

[0226] The electronic device accesses (704) first audio (e.g., 654a from the first source audio) corresponding to a first visual element (e.g., audio of a video in a playback window, audio of a song from an album corresponding to album art, audio generated by or from a first application).

[0227] According to some embodiments, the electronic device accesses (706) second audio (eg, 654b from a second source audio) corresponding to the second visual element (eg, 604b).

[0228] While displaying (708) a first visual element (e.g., 604a) in a first position (e.g., a position on the display that is substantially horizontally centered on the display), the electronic device generates audio in a first mode (e.g., without modifying the first audio (e.g., 654a) shown in FIG. 6A and without applying any interaural time difference, HRTF, or cross-cancellation to 654a) on two or more speakers using the first audio (e.g., 654a).

[0229] According to some embodiments, the first mode is configured (712) such that audio produced using the first mode is perceived by the user as originating from a first direction (and optionally a position) corresponding to (e.g., aligned with) the display.

[0230] While displaying a first visual element (e.g., 604a) at a first position on the display (e.g., a position on the display substantially horizontally centered on the display) (708), the electronic device generates audio (e.g., a discrete audio output or a combined audio output including a component based on the second audio) in two or more speakers using the second audio (e.g., 654b) in a third mode different from the first and second modes (e.g., by applying interaural time difference, HRTF, and / or cross-cancellation to 654b shown in FIG. 6A ). The third mode is configured such that the audio generated in the third mode is perceived by the user as originating from a direction (and optionally a position) away from the display (e.g., to the right of the display that is not aligned) (e.g., modifying the source audio before generating the audio such that the audio is perceived by the user as originating from a direction to the right of the display or the user). In some embodiments, a first mode places the audio in front of the user or display, a second mode places the audio to the left of the user or display, and a third mode places the audio to the right of the user or display.

[0231] While displaying a first visual element (e.g., 604a) on the display at a first position (e.g., a position on the display that is substantially horizontally centered on the display) (708), the electronic device refrains from displaying a second visual element (e.g., 604b in FIG. 6A , second video playback window of the second application, second album art) on the display corresponding to second audio (e.g., 654b from a second source audio) (e.g., audio of a video in a playback window, audio of a song from an album corresponding to album art, audio generated by, or audio received from, a second application) (716).

[0232] While displaying a first visual element (e.g., 604a) at a first position on the display (e.g., a position on the display that is substantially horizontally centered on the display) (708), the electronic device receives a first user input (e.g., 610a, a swipe input on touch-sensitive surface 600b) (718).

[0233] In response to receiving (720) a first user input (e.g., 610a), the electronic device transitions (e.g., by sliding) the display of a first visual element (e.g., 604a in FIG. 6B) from a first position on the display to a first visual element (e.g., 604a in FIG. 6C) that is not displayed on the display (e.g., by sliding off an edge (e.g., left edge) of the display).

[0234] Further, in response to receiving a first user input (e.g., 610a) (720), while not displaying a first visual element (e.g., 604a in FIG. 6C ) on the display, the electronic device generates audio on two or more speakers using a first audio (e.g., 654a) in a second mode different from the first mode (724). The second mode is configured such that the audio generated in the second mode is perceived by the user as originating from a direction (and optionally a position) away from the display (e.g., to the left of an unaligned location) (e.g., applying interaural time difference, HRTF, and / or cross-cancellation to modify the source audio before generating the audio, such that the audio is perceived by the user as originating from a direction away from the display in a direction corresponding to the last displayed position of the first visual element (e.g., to the left of the display or the user) as shown in FIG. 6C ). Generating audio associated with content having varying characteristics allows a user to visualize where the content is located relative to the user without requiring the content to be displayed. This allows the user to quickly and easily recognize the input needed to access the content (e.g., to cause the content to be displayed). Generating audio with varying characteristics also provides the user with contextual feedback regarding different content locations. Providing improved audio feedback to the user enhances device operability (e.g., by assisting the user in providing appropriate inputs) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0235] In some embodiments, the volume of the audio produced using the first mode is greater than the volume of the audio produced using the second mode or the volume of the audio produced using the third mode, and in some embodiments, the volume of the audio produced using the second mode is the same as the volume of the audio produced using the first mode.

[0236] In some examples, a frequency filter, such as a low-pass filter, a high-pass filter, or a band-pass filter, is not applied to audio generated in the first mode (e.g., when the audio is perceived as being centered, such as in front of the user). In some examples, a first frequency filter, such as a low-pass filter, a high-pass filter, or a band-pass filter, is applied to audio generated using the first mode. In some examples, a second frequency filter, such as a low-pass filter, a high-pass filter, or a band-pass filter, is applied to audio generated using the second mode. In some examples, a third frequency filter, such as a low-pass filter, a high-pass filter, or a band-pass filter, is applied to audio generated using the third mode. In some embodiments, the second frequency filter is the same as the third frequency filter. In some examples, a frequency filter is applied to audio generated using the second mode and the third mode, but not to audio generated using the first mode.

[0237] According to some embodiments, while displaying a first visual element (e.g., 604a in FIG. 6A , a video playback window for a first application, album art) in a first position on the display, the electronic device refrains from displaying a second visual element (e.g., 654b in FIG. 6A , a second video playback window for a second application, second album art) on the display corresponding to a second audio (e.g., 654b from a second source audio) (e.g., audio from a video in the playback window, audio of a song from an album corresponding to the album art, audio generated by a second application, or audio received from a second application). Not displaying the content conserves display space and allows the device to provide other visual feedback to the user on the display. Providing improved visual feedback to the user improves device usability and makes the user-device interface more efficient (e.g., by assisting the user in providing appropriate inputs when operating / interacting with the device and reducing user errors), as well as reducing power usage and improving the device's battery life by allowing the user to use the device more quickly and efficiently.

[0238] According to some embodiments, further in response to receiving (720) a first user input (e.g., 610a), the electronic device transitions (726) the display of the second visual element (e.g., 604b) from not being displayed on the display to a fourth position on the display (e.g., 604b in FIG. 6C , which is the same as the first position) (e.g., by sliding) the second visual element onto the display from an edge (e.g., right edge) of the display.

[0239] According to some embodiments, further in response to receiving (720) a first user input (e.g., 610a), the electronic device generates (728) audio on two or more speakers using a second audio (e.g., 654b of FIG. 6C) in the first mode (e.g., without modifying the second audio) while simultaneously generating audio using the first audio (e.g., 654a of FIG. 6C) in the second mode. In other examples, the second audio is generated in a mode different from the first mode.

[0240] According to some embodiments, the second mode is configured such that audio produced using the second mode is perceived by a user as being produced from a second direction (e.g., different from the first direction). According to some embodiments, the third mode is configured such that audio produced using the third mode is perceived by a user as being produced from a third direction different from the second direction (and optionally different from the first direction). Thus, optionally, the perceived location of the source of audio produced using the various modes is different.

[0241] According to some embodiments, following displaying a first visual element (e.g., 604a in FIG. 6A ) at a first position on the display (e.g., a position on the display that is substantially horizontally centered on the display), the electronic device displays the first visual element (e.g., 604a in FIG. 6B ) at a second position on the display (e.g., a position on the display to the left of the first position on the display, a position on the display that is not substantially horizontally centered on the display, a position on the display that is adjacent to an edge (e.g., the left edge) of the display) before the first visual element is no longer displayed on the display (e.g., 604a in FIG. 6C ). According to some embodiments, following displaying a first visual element (e.g., 604a in FIG. 6A ) at a first position on the display (e.g., a position on the display substantially centered horizontally on the display), and before the first visual element is no longer displayed on the display (e.g., 604a in FIG. 6C ), while displaying the first visual element (e.g., 604a in FIG. 6B ) at a second position on the display, the electronic device generates audio using a first audio (e.g., 654a in FIG. 6B ) in two or more speakers in a fourth mode different from the first, second, and third modes (e.g., modifying the source audio before generating the audio so that the audio is perceived by the user as coming from a direction different from the location from which the audio would be perceived if generated in the first mode, such as from or near the left side of the display).

[0242] According to some embodiments, the first mode does not include correction of interaural time difference of the audio. In some examples, the audio generated in the first mode (e.g., 654a in FIG. 6A ) is not corrected using any of HRTFs, interaural time difference of the audio, or cross-cancellation. In some examples, while the first visual element is displayed, the electronic device plays the first audio using an HRTF configured to cause the first audio to be perceived as being in front of the user, such as within 10 degrees left and 10 degrees right of the direction the user is facing. In some examples, while displaying the first visual element on the display, the electronic device generates the first audio using an HRTF that corrects the interaural time difference of the first audio by less than a first predetermined amount (e.g., minimal change).

[0243] According to some embodiments, the second mode includes modification of the audio interaural time difference. In some examples, the second mode includes modification of the audio interaural time difference by a first degree, and the second mode includes modification of the audio interaural time difference by a second degree greater than the first degree. In some examples, the third mode includes modification of the audio interaural time difference. In some examples, the audio generated in the second mode (e.g., 654b in FIG. 6C ) is modified using one or more of HRTFs, audio interaural time difference, and cross-cancellation. In some examples, the audio generated in the third mode (e.g., 654b in FIG. 6A ) is modified using one or more of HRTFs, audio interaural time difference, and cross-cancellation. Generating audio associated with content having varying characteristics allows a user to visualize where the content is located relative to the user without requiring the content to be displayed. This allows a user to quickly and easily recognize the input required to access the content (e.g., to trigger the display of the content). Generating audio with varying characteristics also provides the user with contextual feedback regarding the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by assisting the user in providing appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0244] According to some embodiments, correcting the interaural time difference of the audio includes combining a first channel audio (e.g., right channel) of the audio and a second channel audio (e.g., left channel) of the audio to form composite channel audio, updating the second channel audio to include the composite channel audio at a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio at a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount).

[0245] In some embodiments, correcting the interaural time difference of the audio includes introducing a time delay into a first channel audio (e.g., a right channel) without introducing a time delay into a second channel audio (e.g., a left channel) that is different from the first audio channel.

[0246] According to some embodiments, in response to beginning to receive (e.g., begin detecting) a first user input (e.g., a swipe input), the electronic device transitions from not generating audio using second audio (e.g., from a second source audio) corresponding to a second visual element (e.g., a second video playback window, second album art of a second application) to generating audio on two or more speakers using the second audio corresponding to the second visual element. For example, before the first input is received, no audio is generated on the speakers using the second audio. Once the start of the first input is detected (or after a portion of the first input is detected, or after the first input is detected), the device generates audio on two or more speakers using the second audio corresponding to the second visual element. In some embodiments, the secondary audio is an audio file (e.g., a song), and transitioning to generate the audio using the secondary audio includes forgoing generating the audio using a first predetermined portion (e.g., the first 0.1 seconds) of the secondary audio, for example, to provide the effect that the audio was being played before the audio was heard by the user.

[0247] According to some embodiments, generating audio using the second mode (and optionally the third mode) includes one or more of attenuating the audio, applying a high-pass filter to the audio, applying a low-pass filter to the audio, and altering the volume balance between two or more speakers. Optionally, generating the audio uses one or more of attenuating the audio, applying a high-pass filter to the audio, applying a low-pass filter to the audio, and altering the volume balance between two or more speakers. Generating audio associated with content with varying characteristics allows a user to visualize where the content is located relative to the user without requiring the content to be displayed. This allows a user to quickly and easily recognize the input needed to access the content (e.g., to cause the content to be displayed). Generating audio with varying characteristics also provides a user with contextual feedback regarding the placement of different content. Providing improved audio feedback to the user enhances the usability of the device (e.g., by assisting the user in providing appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.

[0248] In some embodiments, generating audio using the second mode (or the third and fourth modes) includes applying crosstalk cancellation techniques to the audio such that the audio is configured to be perceived by a user as coming from a particular direction. In some embodiments, generating audio using the third mode includes applying crosstalk cancellation techniques to the audio such that the audio is configured to be perceived as coming from a particular direction. In some embodiments, generating audio using the first mode does not include applying crosstalk cancellation techniques to the audio.

[0249] It should be noted that the details of the processes described above with respect to method 700 (e.g., FIGS. 7A-7C) are also applicable in an analogous manner to the methods described below. For example, methods 900, 1200, 1400, and 1500 optionally include one or more of the characteristics of the various methods described above with respect to method 700. For example, the same or similar techniques are used to place audio within a space. In other embodiments, the same audio source may be used in the various techniques. In yet other embodiments, the currently playing audio in each of the various methods may be manipulated using techniques described in the other methods. For the sake of brevity, these details will not be repeated below.

[0250] 8A-8K show exemplary techniques for previewing audio according to some embodiments. The techniques in these figures are used to illustrate processes described below, including the processes of FIGS. 9A-9C.

[0251] 8A-8K show a device 800 (e.g., a mobile phone) having a display, a touch-sensitive surface, and left and right speakers (e.g., headphones). An overhead view 850, a visual representation of the spatial configuration of audio being generated by device 800, is shown throughout FIGS. 8A-8K to provide the reader with a better understanding of the techniques, particularly with respect to where a user 856 perceives sound to be coming from (e.g., as a result of device 800 positioning the audio in space). Overhead view 850 is not part of the user interface of device 800. In some examples, the techniques described below enable a user to more easily and efficiently preview and select music for playback.

[0252] 8A , device 800 displays a music player 802 for playing music. Music player 802 includes multiple affordances 804, each corresponding to a different song that can be played using device 800. As shown in overhead view 850, device 800 is not playing any audio.

[0253] 8B, device 800 detects a tap-and-hold input 810a on affordance 804a of track 3. As shown in overhead view 850, device 800 is not playing any audio.

[0254] In Figure 8C, in response to detecting input 810a on affordance 804a, and while continuing to detect input 810a on affordance 804a, device 800 updates the plurality of affordances 804 to distinguish affordance 804a for track 3. In the example of Figure 8C, device 800 blurs the plurality of affordances 804 other than the selected affordance 804a for track 3. This indicates to the user that the song corresponding to affordance 804a will be previewed.

[0255] For example, the preview is limited to a predetermined audio playback period (e.g., shorter than the duration of the song). After the predetermined audio playback period is reached during playback of a song, the device stops playing that song. In some embodiments, after stopping playback of a song, the device proceeds to provide a preview of a different song. In some embodiments, after stopping playback of a song, the device proceeds to provide another preview of the same song (e.g., looping the same portion of the song).

[0256] As shown in overhead view 850 of FIG. 8C , in response to detecting input 810a on affordance 804a, and while continuing to detect input 810a on affordance 804a, device 800 generates a preview of the audio of track 3 by positioning the audio of track 3 in space using two or more speakers (e.g., device 800 applies interaural time difference, HRTF, and / or cross-cancellation to the music). Device 800 positions the audio of track 3 so that the user perceives the music as coming from a location in space to the left of user 856 (or device 800), as indicated by audio element 850a in overhead view 850 of FIG. 8C . In FIGS. 8C-8D , device 800 continues to detect input 810a on affordance 804a and updates the location in space of the audio of track 3 from which the user perceives the music so that the user perceives the music as moving toward the user.

[0257] In Figure 8D, device 800 continues to generate audio for track 3 but stops positioning the audio in space, resulting in the user perceiving the audio as being within the user's head, as indicated by audio element 850a in overhead view 850 of Figure 8D. For example, in Figure 8D, device 800 generates audio for track 3 using left and right speakers without positioning the audio (e.g., device 800 does not apply interaural time difference, HRTF, or cross-cancellation).

[0258] Device 800 continues to play a preview of track 3 for the user until (1) the predetermined audio playback duration is reached, (2) the device detects movement of input 810a on the touch-sensitive surface to an affordance that corresponds to a different song, or (3) the device detects lift-off of input 810a.

[0259] 8E, device 800 detects movement of input 810a from affordance 804a to affordance 804b (without detecting lift-off of input 810a). Note that device 800 does not scroll (or otherwise move) multiple affordances 804 in response to detecting movement of input 810a. In response to detecting input at affordance 804b, device 800 updates the visual aspects and audio spatialization of the device. Device 800 blurs track 3 affordance 804a and ceases blurring track 4 affordance 804b. Device 800 also transitions to producing the audio for track 3 by spatially positioning the audio using the left and right speakers (e.g., device 800 applies interaural time difference, HRTF, and / or cross-cancellation), so that the user perceives the audio as moving away from the user's head and to the user's right, as indicated by audio element 850a in overhead view 850 of Figures 8E-8F. Device 800 also, optionally, begins attenuating the audio for track 3, and then stops producing the audio for track 3, as indicated by audio element 850a no longer being shown in overhead view 850 of Figure 8G. Further, in response to detecting input at affordance 804b, device 800 generates a preview of the audio of track 4 by positioning the audio in space (e.g., device 800 applies interaural time difference, HRTF, and / or cross-cancellation) so that the user perceives the audio of track 4 as coming from the user's left and moving into the user's head, as illustrated by audio element 850b in overhead view 850 of FIGS. 8E-8G. As mentioned above, the preview is limited to a predetermined audio playback period (e.g., shorter than the duration of the song). After the predetermined audio playback period is reached, the device stops playing the song.

[0260] 8F-8G, device 800 generates audio for track 4 using left and right speakers without positioning the audio (e.g., device 800 does not apply interaural time difference, HRTF, or cross-cancellation). For example, as a result of not positioning the audio in space, a user of device 800 wearing headphones perceives the audio as being inside the user's head, as indicated by audio element 850b in overhead view 850 of FIGS. 8F-8G.

[0261] In FIG. 8H , device 800 detects lift-off of input 810a. In response to detecting lift-off of input 810a, device 800 transitions to generating audio for track 4 using the left and right speakers by spatially positioning the audio (e.g., device 800 applies interaural time difference, HRTF, and / or cross-cancellation), such that the user perceives the audio as moving away from the user's head and to the user's right, as indicated by audio element 850b in overhead view 850 of FIG. 8H . In response to detecting lift-off of input 810a, device 800 stops blurring multiple affordances, as shown in FIG. 8H . Further, in response to detecting lift-off of input 810a, device 800 optionally begins attenuating the audio for track 4 and then stops generating the audio for track 4, as indicated by audio element 850b, which is no longer shown in overhead view 850 of FIG. 8I .

[0262] In FIG. 8J , device 800 detects tap input 810b on affordance 804c for track 6. In FIG. 8K , as shown in overhead view 850, device 800 begins generating audio for track 6 using two or more speakers without positioning the audio for track 6 in space, as indicated by audio element 850c in overhead view 850 of FIG. 8J . As a result, the user perceives the audio as being within the user's head. For example, in FIG. 8J , device 800 generates stereo audio for track 6 using left and right speakers without positioning the audio (e.g., device 800 does not apply interaural time difference, HRTF, or cross-cancellation). In response to detecting tap input 810b, device 800 does not blur any of multiple affordances 804. Device 800 also, optionally, updates affordance 804c to include additional information about media controls and / or tracks. Track 6 continues to play until the end of the track is reached, without the playback being limited to a preview of the given audio playback duration.

[0263] In some examples, device 800 includes a digital assistant that generates audio feedback, such as returning the results of a query by speaking the results. In some examples, device 800 generates the digital assistant's audio by positioning the audio for the digital assistant at a location in space (e.g., over the user's right shoulder), so that the user perceives the digital assistant as remaining stationary in space even when other audio moves in space. In some examples, device 800 emphasizes the digital assistant's audio by ducking one or more (or all) other audio.

[0264] 9A-9C are flow diagrams illustrating a method for previewing audio using an electronic device, according to some embodiments. Method 900 is performed on a device (e.g., 100, 300, 500, 800) having a display and a touch-sensitive surface. The electronic device is operatively connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earbuds, left and right earbuds). Some operations of method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0265] As described below, method 900 provides an intuitive way to preview audio. This method reduces the cognitive burden on a user to preview audio, thereby creating a more efficient human-machine interface. For battery-operated computing devices, enabling users to preview audio more quickly and efficiently conserves power and extends the time between battery charges.

[0266] The electronic device displays (902) on a display a list (e.g., in a column) of a plurality of media elements (e.g., 804). Each media element (e.g., 804a-804c) (or at least two media elements) of the plurality of media elements corresponds to a respective media file (e.g., an audio file, a song, a video file). In some embodiments, the respective media files are all different from each other.

[0267] The electronic device uses the touch-sensitive surface to detect (904) a user contact (e.g., a touch-and-hold user input of 810a, a touch input that is sustained for longer than a predetermined period of time greater than 0 seconds) at a location corresponding to a first media element (e.g., 804a).

[0268] In response to detecting (906) a user contact (e.g., 810a in FIGS. 8B-8C) at a location corresponding to a first media element (e.g., 804a), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input (e.g., as determined by the electronic device), the electronic device generates audio (e.g., 850a) using a first audio file corresponding to the first media element using two or more speakers without exceeding a predetermined audio playback period. For example, the audio plays for a maximum of five seconds. For example, the device generates a preview of the audio. In some examples, the predetermined audio playback period is shorter than the duration of the audio file.

[0269] Further, in response to detecting 906 a user contact (e.g., 810a in FIGS. 8B-8C ) at a location corresponding to the first media element (e.g., 804a), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input (e.g., as determined by the electronic device), and while the user contact (e.g., 810a) remains (910) at the location corresponding to the first media element (e.g., 804a) (without a lift-off event) and in accordance with not exceeding a predetermined audio playback period, the electronic device continues 912 to generate audio (e.g., 850a) using the first audio file using two or more speakers.

[0270] Further, in response to detecting (906) a user contact (e.g., 810a in FIGS. 8B-8C ) at a location corresponding to the first media element (e.g., 804a), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input (e.g., as determined by the electronic device), and while the user contact (e.g., 810a) remains (910) at a location corresponding to the first media element (e.g., 804a) (without a lift-off event) and in accordance with exceeding a predetermined audio playback period, the electronic device stops (914) generating audio (e.g., 850a) using the first audio file using the two or more speakers (and, optionally, without starting to generate audio using a different audio file).

[0271] The electronic device uses the touch-sensitive surface to detect (916) movement of a user touch from a location corresponding to a first media element (e.g., 804a) to a location corresponding to a second media element (e.g., 804b) (e.g., 810a in Figures 8D-8E, toward the bottom of the display and away from the top of the display).

[0272] In response to detecting 918 a user contact (e.g., 810a in FIG. 8F ) at a location corresponding to the second media element, and in accordance with the user contact (e.g., 810a) including a touch-and-hold input, the electronic device generates 920 audio (e.g., 850b) using a second audio file (different from the first audio file) corresponding to the second media element (e.g., 804b) using two or more speakers without exceeding a predetermined audio playback period. For example, the audio plays for a maximum of five seconds (less than the entire duration of the second audio file).

[0273] Further, in response to detecting a user contact (e.g., 810a in FIG. 8F ) at a location corresponding to the second media element (918), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input, and while the user contact remains (922) at a location corresponding to the second media element (e.g., 804b) (without a lift-off event) and in accordance with not exceeding a predetermined audio playback period, the electronic device continues to generate audio (e.g., 850b) using the second audio file using two or more speakers (924).

[0274] Further, in response to detecting (918) a user contact (e.g., 810a in FIG. 8F ) at a location corresponding to the second media element, and in accordance with the user contact (e.g., 810a) including a touch-and-hold input (e.g., as determined by the electronic device), and while the user contact remains (without a lift-off event) at a location corresponding to the second media element (e.g., 804b) and in accordance with exceeding a predetermined audio playback period, the electronic device stops (926) generating audio (e.g., 850b) using the second audio file using the two or more speakers (and, optionally, without starting to generate audio using a different audio file).

[0275] The electronic device detects (928) lift-off of a user contact (eg, 810a in Figures 8G and 8H) using the touch-sensitive surface.

[0276] In response to detecting lift-off of the user contact (930), the electronic device stops generating audio (e.g., 850a and 850b in FIG. 8I) using the two or more speakers using the first audio file or the second audio file (and, optionally, stops generating any audio using audio files corresponding to the multiple media elements) (932), thus stopping playback of any song being previewed.

[0277] According to some embodiments, generating audio using the first audio file without exceeding a predetermined audio playback duration, further pursuant to user contact (e.g., as determined by the electronic device) including a touch-and-hold input (906), includes transitioning the audio between a plurality of modes, including a first mode, a second mode different from the first mode (e.g., 850a in FIG. 8D , 850b in FIG. 8F ), and a third mode different from the first and second modes. Generating audio with varying characteristics also provides the user with contextual feedback regarding the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by assisting the user in providing appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0278] According to some embodiments, the first mode is configured such that audio generated using the first mode is perceived by the user as originating from a first direction (e.g., 850a in FIG. 8C , 850b in FIG. 8E ) (and optionally a location). In some examples, the first direction is a direction originating from a location to the left of the user (e.g., not in front of the user). In some examples, the electronic device modifies the source audio before generating the audio such that the audio is perceived by the user as originating from a direction originating from the display or a location to the left of the user.

[0279] According to some embodiments, the second mode is configured such that audio produced using the second mode is perceived by the user as being in the user's head (e.g., non-point source stereo). For example, the second mode does not include applying any of HRTFs, interaural time differences of the audio, or cross-cancellation.

[0280] According to some embodiments, the third mode is configured such that audio generated using the third mode is perceived by the user as originating from a third direction (e.g., 850a in FIG. 8F or 850b in FIG. 8H) different from the first direction. In some examples, the third direction is a direction originating from a position to the right of the user (e.g., not in front of the user). In some examples, the electronic device modifies the source audio before generating the audio such that the audio is perceived by the user as originating from a direction originating from the display or a position to the right of the user.

[0281] According to some embodiments, the first mode includes a first modification of the audio interaural time difference, and the third mode includes a second modification of the audio interaural time difference that is different from the first modification. In some examples, the first mode includes a first degree of modification of the audio interaural time difference, and the third mode includes a second degree of modification of the audio interaural time difference that is greater than the first degree. In some examples, the first mode includes a modification of the audio interaural time difference such that the audio is perceived by the user as originating from a direction originating from the display or a position to the left of the user. In some examples, the third mode includes a modification of the audio interaural time difference such that the audio is perceived by the user as originating from a direction originating from the display or a position to the right of the user. Modifying the audio interaural time difference allows the user to perceive the audio as coming from a particular direction. This audio directional information provides the user with additional feedback about the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by assisting the user in providing appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.

[0282] In some embodiments, audio generated in a first mode (e.g., 850a in FIG. 8C ) is modified using one or more of HRTFs, interaural time differences of audio, and cross-cancellation. In some embodiments, audio generated in a second mode (e.g., 850a in FIG. 8D ) is not modified using any of HRTFs, interaural time differences of audio, or cross-cancellation. In some embodiments, audio generated in a third mode (e.g., 850a in FIG. 8F ) is modified using one or more of HRTFs, interaural time differences of audio, and cross-cancellation.

[0283] According to some embodiments, the second mode does not include correction of interaural time difference of the audio. In some examples, the audio generated in the second mode is not corrected using HRTFs. In some examples, the audio generated in the second mode (e.g., 850a of FIG. 8D) is not corrected using any of HRTFs, interaural time difference of the audio, or cross-cancellation.

[0284] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio (e.g., a right channel) of the audio and a second channel audio (e.g., a left channel) of the audio to form composite channel audio, updating the second channel audio to include the composite channel audio at a first delay (e.g., a 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio at a second delay different from the first delay (e.g., a 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay into the first channel audio (e.g., a right channel) without introducing a time delay into the second channel audio (e.g., a left channel) different from the first audio channel. Modifying the interaural time difference of the audio allows a user to perceive audio as coming from a particular direction. This directional information of the audio provides the user with additional feedback about the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by assisting the user in providing appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves the device's battery life by allowing the user to use the device more quickly and efficiently.

[0285] According to some embodiments, further in response to detecting (906) a user contact (e.g., 810a) at a location corresponding to the first media element (e.g., 804a), the electronic device changes a first visual characteristic (e.g., focus / blur level) of a displayed media element within the plurality of media elements (e.g., 804b of FIG. 8C) other than the first media element (e.g., 804a of FIG. 8C). In response to detecting a user contact (e.g., 810a) at a location corresponding to a second media element (e.g., 804ba), the electronic device reverts the change in the first visual characteristic (e.g., 804b of FIG. 8E) of the second media element and changes the first visual characteristic (e.g., 804a of FIG. 8E) of the first media element. In some examples, the electronic device fades inactive media elements. In some examples, the electronic device adds a blur effect to inactive media elements. In some examples, the electronic device changes the color of inactive media elements. Visually distinguishing playing content from non-playing content provides the user with feedback regarding the state of the device. Providing improved visual feedback to the user improves usability of the device, makes the user device interface more efficient (e.g., by assisting the user in providing appropriate input when operating / interacting with the device and reducing user errors), and also reduces power usage and improves the battery life of the device by allowing the user to use the device more quickly and efficiently.

[0286] According to some embodiments, in response to detecting a user touch at a location corresponding to a first media element, the electronic device changes the second visual characteristic of the first media element without changing the second visual characteristic of any displayed media element in the plurality of media elements other than the first media element. In response to detecting a user touch at a location corresponding to a second media element, the electronic device undoes the change in the second visual characteristic of the first media element and changes the second visual characteristic of the second media element. In some examples, the electronic device highlights the activated media element. In some examples, the electronic device brightens the activated media element. In some examples, the electronic device changes the color of the activated media element. Visually distinguishing playing content from non-playing content provides a user with feedback regarding the status of the device. Providing improved visual feedback to the user improves device usability and makes the user device interface more efficient (e.g., by assisting the user in providing appropriate input when operating / interacting with the device and reducing user errors), as well as allowing the user to use the device more quickly and efficiently, thereby reducing power usage and improving the device's battery life.

[0287] According to some embodiments, in response to detecting a user contact at a location corresponding to the first media element, and in accordance with the user contact including a tap input (e.g., as determined by the electronic device), the electronic device generates audio using a first audio file corresponding to the first media element without automatically ceasing to generate audio using the first audio file after a predetermined audio playback period and without automatically ceasing to generate audio using the first audio file upon detecting lift-off of the user contact, e.g., the audio plays for longer than a 5-second preview time.

[0288] According to some embodiments, generating audio using the first audio file in accordance with user contact including a tap input (e.g., as determined by the electronic device) includes generating audio using the first audio file in a second mode without transitioning the audio between a first mode and a third mode of the plurality of modes.

[0289] It should be noted that the details of the processes described above with respect to method 900 (e.g., FIGS. 9A-9C) are also applicable in an analogous manner to the methods described below and above. For example, methods 700, 1200, 1400, and 1500 optionally include one or more of the characteristics of the various methods described above with respect to method 900. For example, the same or similar techniques are used to place audio within a space. In other embodiments, the same audio source may be used in the various techniques. In yet other embodiments, the currently playing audio in each of the various methods may be manipulated using techniques described in the other methods. For the sake of brevity, these details will not be repeated below.

[0290] 10A-10K illustrate exemplary techniques for discovering music, according to some embodiments. The techniques in these figures are used to illustrate processes described below, including the processes of FIGS. 12A-12B.

[0291] 10A-10K illustrate a device 1000 (e.g., a mobile phone) having a display, a touch-sensitive surface, and left and right speakers (e.g., headphones). An overhead view 1050, a visual representation of the spatial configuration of audio being generated by device 1000, is shown throughout FIGS. 10A-10K to provide the reader with a better understanding of the techniques, particularly with respect to where a user 1056 perceives sound as coming from (e.g., as a result of device 1000 positioning the audio in space). The overhead view 1050 is not part of the user interface of device 1000. Similarly, visual elements displayed outside the display device, as represented by dotted outlines, are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the techniques. In some examples, the techniques described below enable users to more easily and efficiently discover new songs from a song repository.

[0292] In Figure 10A, device 1000 displays music player 1002 including affordance 1002a for discovering new audio content. In Figure 10A, device 1000 is not generating any audio, as indicated by the absence of audio elements in overhead view 1050 of Figure 10A. While device 1000 is not generating any audio, device 1000 detects tap input 1010a on affordance 1002a. In response to detecting tap input 1010a on affordance 1002a, device 1000 enters discovery mode.

[0293] 10B, device 1000 replaces the display of affordance 1002a with a display of album art 1004a. Album art 1004a corresponds to a first song (e.g., the album art is for the album to which the song belongs), and album art 1004b corresponds to a second song that is different from the first song. Audio element 1054a corresponds to the first song, and audio element 1054b corresponds to the second song.

[0294] In FIG. 10B , device 1000 generates audio for a first song and a second song by positioning the audio for the first song and the second song along path 1050a in space using left and right speakers (e.g., device 1000 applies interaural time difference, HRTF, and / or cross-cancellation). For example, path 1050a is a curved, fixed path along which the device positions the audio for various songs while in discovery mode. Device 1000 generates the audio for the first song and the second song and updates the positioning of the audio so that the user perceives the song audio as drifting in front of the user from left to right along path 1050a, as shown by audio elements 1054a and 1054b in overhead view 1050 of FIGS. 10B-10G . In some examples, the direction in which the song moves along path 1050a is not based on input provided by the user (e.g., not based on tap input 1010a).

[0295] In FIG. 10C, as the first and second songs progress along path 1050a, device 1000 begins to generate the audio of the third song by using the left and right speakers to position and transition the audio in space (e.g., device 1000 applies interaural time difference, HRTF, and / or cross-cancellation), thereby causing the user to perceive the audio as drifting in front of the user from left to right along path 1050a, as shown by audio element 1054c in the overhead view 1050 of FIGS. 10C-10G.

[0296] 10D-10F, device 1000 begins generating audio for additional songs (while continuing to generate audio for the first and second songs) by spatially positioning and transitioning the audio for each song using the left and right speakers (e.g., device 1000 applies interaural time difference, HRTFs, and / or cross-cancellation), causing the user to perceive the audio as drifting in front of the user from left to right along path 1050a, as illustrated by audio elements 1054a-1054e in overhead view 1050 of FIGS. 10C-10G. In some examples, device 1000 positions the songs at equal distances along path 1050a (e.g., perceived as 2-3 meters away from the user). In some embodiments, the distance between songs changes (e.g., increases) as the songs approach a point on path 1050a (e.g., a point in front of the user), and the distance between songs changes (e.g., decreases) as the songs move away from a point on path 1050a. In some embodiments, each song moves at the same speed along path 1050a. In some embodiments, device 1000 stops generating songs that reach a certain point along path 1050a (e.g., they are more than a threshold distance away from the user). In some embodiments, device 1000 attenuates songs based on position along path 1050a (e.g., songs farther from the user are attenuated).

[0297] As shown in Figures 10B-10I, album art 1004a-1004e corresponds to various songs. Device 1000 displays the respective album art for the songs located in front of the user. For example, device 1000 displays album art corresponding to songs located along a particular subset of path 1050a. The displayed album art moves in the same direction on the display as the various audio moves along path 1050a (e.g., from left to right).

[0298] The first song (visualized as audio element 1054a in overhead view 1050) corresponds to album art 1004a (e.g., the album art is for the album to which the song belongs). The second song (visualized as audio element 1054b in overhead view 1050) corresponds to album art 1004b. The third song (visualized as audio element 1054c in overhead view 1050) corresponds to album art 1004c. The fourth song (visualized as audio element 1054d in overhead view 1050) corresponds to album art 1004d. The fifth song (visualized as audio element 1054e in overhead view 1050) corresponds to album art 1004e.

[0299] 10G-10H, device 1000 detects left swipe input 1010b. In response to detecting left swipe input 1010b, device 1000 causes album art to move on the display and change direction to correspond to left swipe input 1010b, and various audio to move along path 1050a and change direction to correspond to left swipe input 1010b. In FIG. 10G, when the user places their finger on the touch-sensitive surface, album art 1004d and 1004c stop moving on the display, and audio corresponding to audio elements 1054a-1054e stops moving along path 1050a. 10H , when device 1000 detects right-to-left left swipe input 1010b, device 1000 updates the display such that album art 1004d and 1004c move from right to left on the display and audio corresponding to audio elements 1054a-1054e move from right to left along path 1050a. In some embodiments, the speed at which album art 1004a-1004e moves across the display and the speed at which audio corresponding to audio elements 1054a-1054e moves along path 1050a is based on one or more characteristics (e.g., length, speed, characteristic intensity) of left swipe input 1010b. In some embodiments, a faster swipe results in faster movement of album art on the display and audio along path 1050a. In some embodiments, a longer swipe results in faster movement of album art on the display and audio along path 1050a.

[0300] 10I, after device 1000 stops detecting left swipe input 1010b, device 1000 continues to display album art 1004a-1004e moving left to right across the display, and the audio corresponding to audio elements 1054a-1054e continues to move left to right along path 1050a. Thus, left swipe input 1010b changes the direction the user perceives the audio to be moving along path 1050a, changing the direction in which device 1000 moves the corresponding album art across the display.

[0301] In Figure 10J, device 1000 detects tap input 1010c on album art 1004b. In response to detecting tap input 1010c on album art 1004b, device 1000 transitions to generating audio for a second song corresponding to album art 1004b using the left and right speakers without positioning the audio (e.g., device 1000 does not apply interaural time difference, HRTF, or cross-cancellation), as shown in Figures 10J-10K. For example, as a result of not positioning the audio in space, a user of device 1000 wearing headphones perceives the audio for the second song as being within the user's head, as indicated by audio element 1054b in overhead view 1050 of Figure 10K. Further, in response to detecting tap input 1010c on album art 1004b, device 1000 moves the audio corresponding to audio elements 1054c and 1054a in opposite directions away from the user before ceasing to generate the audio corresponding to audio elements 1054c and 1054a, as shown in Figures 10J-10K. Thus, the user perceives the audio of the unselected songs as floating.

[0302] In some examples, device 1000 includes a digital assistant that generates audio feedback, such as returning the results of a query by speaking the results. In some examples, device 1000 generates the digital assistant's audio by positioning the audio for the digital assistant at a location in space (e.g., over the user's right shoulder), so that the user perceives the digital assistant as remaining stationary in space even when other audio moves in space. In some examples, device 1000 emphasizes the digital assistant's audio by ducking one or more (or all) other audio.

[0303] 11A-11G illustrate an exemplary technique for discovering music, according to some embodiments. FIGS. 11A-11G illustrate an exemplary user interface for display by a device (e.g., a laptop) having a display, a touch-sensitive surface (e.g., 1100), and left and right speakers (e.g., headphones). Touch-sensitive surface 1100 of the device is illustrated to provide the reader with a better understanding of the described technique, particularly with respect to exemplary user inputs. Touch-sensitive surface 1100 is not part of the device's displayed user interface. Overhead view 1150 illustrated throughout FIGS. 11A-11G is displayed by the device. Overhead view 1150 is also a visual representation of the spatial configuration of the audio being generated by the device, providing the reader with a better understanding of the technique, particularly with respect to where user 1106 perceives sound to be coming from (e.g., as a result of the device positioning the audio in space). The arrows on the audio elements 1150a-1150g indicate the direction and speed (including visual indications and as perceived by the user via speakers) that the audio elements 1150a-1150g are moving, providing the reader with a better understanding of the technique. In some embodiments, the displayed user interface does not include the arrows on the audio elements 1150a-1150g. Similarly, visual elements that are displayed outside of the device's display are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the technique. In some embodiments, the techniques described below enable users to more easily and efficiently discover new songs from a song repository.

[0304] 11A-11D show a device displaying a user's representation 1106 in a space. The device also illustrates audio elements 1150a-1150f corresponding to songs 1-6, respectively, in the same space. The device simultaneously generates audio for each of songs 1-6 using left and right speakers by positioning and transitioning the individual songs in the space (e.g., the device applies interaural time difference, HRTF, and / or cross-cancellation), thereby causing the user to perceive the song audio as drifting past the user, as illustrated by audio elements 1150a-1150f in the overhead view 1150 of FIGS. 11A-11D. In some examples, the audio elements are displayed as being equidistant from one another. In some examples, the generated audio is positioned so that the user perceives the song sources as being equidistant from one another. In Figure 6A, for example, the user perceives song 1 (corresponding to audio element 1150a) and song 2 (corresponding to audio element 1150b) as being substantially in front of the user, song 6 (corresponding to audio element 1150f) as being substantially to the right of the user, and song 4 (corresponding to audio element 1150d) as being substantially to the left of the user. As the device continues to generate songs, the user perceives the audio source as passing by. For example, in Figure 11D, the user perceives song 6 as being substantially behind the user. This allows the user to hear various songs simultaneously.

[0305] 11D-11E, a device receives swipe input 1110 on touch-sensitive surface 1100. In response to receiving swipe input 1110, the device updates the direction and / or velocity of displayed visual elements 1150a-1150h on the display (e.g., according to the direction and / or velocity of swipe input 1110), as shown in FIGS. 11E-11F. Further, in response to receiving swipe input 1110, the device transitions the audio generated for each song in space (e.g., the device applies interaural time difference, HRTF, and / or cross-cancellation), such that the user perceives the song audio as drifting past the user in the updated direction and / or velocity (e.g., according to the direction and / or velocity of swipe input 1110).

[0306] In some examples, the device includes a digital assistant that generates audio feedback, such as returning the results of a query by speaking the results. In some examples, the device generates the digital assistant's audio by placing the audio for the digital assistant at a position in space (e.g., over the user's right shoulder), so that the user perceives the digital assistant as remaining stationary in space even when other audio moves in space. In some examples, the device emphasizes the digital assistant's audio by ducking one or more (or all) other audio.

[0307] In some examples, the device detects a tap input on a displayed audio element (e.g., at a corresponding location on the touch-sensitive surface). In response to detecting the tap input on the audio element, the device transitions to generating audio for the respective song corresponding to the selected audio element using the left and right speakers without positioning the audio in space (e.g., device 1000 does not apply interaural time difference, HRTF, or cross-cancellation). For example, as a result of not positioning the audio in space, a user wearing headphones perceives the audio of the selected song as being within the user's head. Further, in response to detecting the tap input on the displayed audio element, the device stops generating audio for the remaining audio elements. Thus, the user is able to select individual songs to listen to.

[0308] 12A-12B are flow diagrams illustrating a method for discovering music using an electronic device, according to some embodiments. Method 1200 is performed on a device (e.g., 100, 300, 500, 1000) having a display and a touch-sensitive surface. The electronic device is operatively connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earbuds, left and right earbuds). Some operations of method 1200 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0309] As described below, method 1200 provides an intuitive way to discover music. This method reduces the cognitive burden on the user to discover music, thereby creating a more efficient human-machine interface. For battery-operated computing devices, enabling users to discover music more quickly and efficiently conserves power and extends the time between battery charges.

[0310] The electronic device detects (1202) a first user input (e.g., 1010a) to activate a discovery mode (e.g., by tapping the "Discover" affordance, providing a voice input such as "sample some songs" to the digital assistant).

[0311] In response to detecting (1204) a first user input (e.g., 1010a) to activate a discovery mode, the electronic device simultaneously generates (1204) audio (e.g., 1054a, 1054b, 1054c) using two or more speakers using a first audio source (e.g., audio file, music file, media file) in a first mode (1206), a second audio source (e.g., audio file, music file, media file) in a second mode (1210), and a third audio source (e.g., audio file, music file, media file) in a third mode (1214).

[0312] The first mode (1206) is configured such that audio produced using the first mode is perceived by a user as originating from a first point in space moving over time in a first direction (e.g., left to right) along a predetermined path (e.g., 1050a) at a first speed.

[0313] According to some embodiments, a first audio source (e.g., an audio file, music file, media file) corresponds to a first visual element (e.g., audio of a video in a playback window of 1004a, audio of a song from an album corresponding to album art, audio generated by or received from a first application) (1208).

[0314] The second mode (1210) is configured such that audio produced using the second mode is perceived by the user as originating from a second point in space moving over time in a first direction (e.g., left to right) along a predetermined path (e.g., 1050a) at a second speed.

[0315] According to some embodiments, a second audio source (e.g., an audio file, music file, media file) corresponds to a second visual element (e.g., audio of a video in a playback window of 1004b, audio of a song from an album corresponding to album art, audio generated by or received from the first application) (1212).

[0316] The third mode (1214) is configured such that audio produced using the third mode is perceived by the user as originating from a third point in space moving in a first direction (e.g., left to right) along a predetermined path (e.g., 1050a) at a third speed.

[0317] According to some embodiments, a third audio source (e.g., an audio file, music file, media file) corresponds to a third visual element (e.g., audio of a video in a playback window, audio of a song from an album corresponding to album art, audio generated by or received from the first application, of 1004c) (1216).

[0318] The first point (e.g., location 1054a in FIG. 10C ), the second point (e.g., location 1054b in FIG. 10C ), and the third point (e.g., location 1054c in FIG. 10C ) are different points in space (e.g., different points in space perceived by a user) (1218). Generating audio with varying characteristics also provides the user with contextual feedback regarding the placement of different content, allowing the user to quickly and easily recognize which input is needed to access the content (e.g., cause the display of the content). For example, the user can recognize that particular audio can access a particular input based on where in the space the user can perceive the audio to be coming from. Providing improved audio feedback to the user enhances device operability (e.g., by assisting the user in providing the appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0319] According to some embodiments, concurrently with generating audio using the first audio source, the second audio source, and the third audio source, the electronic device displays (1220) on the display the simultaneous movement of two or more of the first visual element (e.g., 1004a), the second visual element (e.g., 1004b), and the third visual element (e.g., 1004c) at a fourth rate. In some examples, the device displays the simultaneous movement of all of the first visual element (e.g., 1004a), the second visual element (e.g., 1004b), and the third visual element (e.g., 1004c) at the fourth rate.

[0320] According to some embodiments, the movement of two or more of the first visual element (e.g., 1004a), the second visual element (e.g., 1004b), and the third visual element (e.g., 1004c) is in a first direction (e.g., from left to right).

[0321] According to some embodiments, the predefined path (e.g., 1050a) varies along a first dimension (e.g., x-dimension, left / right dimension). The predefined path varies along a second dimension (e.g., z-dimension, near / far dimension) that is different from the first dimension. The predefined path does not vary along a third dimension (e.g., y-dimension, up / down dimension, height) that is different from the first and second dimensions. Thus, in some examples, the device generates audio along the path such that the user perceives the audio as moving from left to right and / or right to left, and from far to near or near to far, but does not perceive the audio as moving up and down or down and up.

[0322] According to some embodiments, simultaneously generating audio using two or more speakers (1204) further includes using a fourth audio source (e.g., an audio file, a music file, a media file) in a fourth mode. The fourth mode is configured such that audio generated using the fourth mode (e.g., 1054d) is perceived by a user as originating from a fourth point in space (e.g., location 1054a) moving over time in a first direction (e.g., from left to right) along a predetermined path (e.g., 1050a), the fourth point in space being farther from the user than the first, second, and third points. In some examples, visual elements of audio perceived as farther away (e.g., farther than a predetermined distance) are not displayed on the display (e.g., 1004a in FIG. 10E is not displayed on the display). While displaying the simultaneous movement of two or more of the first visual element (e.g., 1004a), the second visual element (e.g., 1004b), and the third visual element (e.g., 1004c) on the display, the electronic device refrains from displaying a fourth visual element (e.g., 1004d) on the display that corresponds to a fourth audio source.

[0323] According to some embodiments, the first speed, the second speed, and the third speed are the same. According to some embodiments, the first speed, the second speed, and the third speed are different.

[0324] According to some embodiments, while simultaneously generating audio using two or more speakers (e.g., as shown at 1050 in FIG. 10F), the electronic device detects a second user input (e.g., 1010b) in a second direction (e.g., a right-to-left direction) that is different from the first direction (e.g., a left-to-right direction). In response to detecting the second user input (e.g., 1010b) in the second direction, the electronic device updates the generation of audio for the first audio source, the second audio source, and the third audio source using a mode configured such that the audio sources are perceived by the user as moving over time in the second direction (e.g., a right-to-left direction as shown at 1050 in FIG. 10H) along a predetermined path (e.g., 1050a). Further, in response to detecting a second user input (e.g., 1010b) in a second direction, the electronic device updates the display of the movement of one or more (e.g., two or more, all) of the first visual element (e.g., 1004a), the second visual element (e.g., 1004a), and the third visual element (e.g., 1004a) so that the movement is in the second direction (e.g., from right to left as shown in Figures 10H-10I).

[0325] According to some embodiments, while simultaneously producing audio using two or more speakers, the electronic device detects a third user input (e.g., a first direction, left to right). In response to detecting the third user input, the electronic device updates the audio production of the first audio source, the second audio source, and the third audio source using a mode configured to cause the audio sources to be perceived by the user as moving over time in the first direction (e.g., right to left) along a predetermined path at a fifth speed that is faster than the first speed. Further, in response to detecting the third user input, the electronic device updates the representation of the simultaneous movement of the first visual element, the second visual element, and the third visual element on the display so that the movement is in the first direction (e.g., left to right) at a sixth speed that is faster than the fourth speed.

[0326] According to some embodiments, the electronic device detects a selection input (e.g., 1010c, tap input, tap-and-hold input) at a location corresponding to the second visual element (e.g., 1004b). In response to detecting the selection input (e.g., 1010c), the electronic device generates audio of the second audio file in a fifth mode (rather than the second mode) using two or more speakers. The audio generated using the fifth mode is not perceived by the user as being generated from a point in space that moves over time.

[0327] According to some embodiments, the fifth mode does not include correction of interaural time difference of audio. In some examples, the audio produced in the fifth mode (e.g., 1054b of FIG. 10K) is not corrected using any of HRTFs, interaural time difference of audio, or cross-cancellation.

[0328] According to some embodiments, the first mode, the second mode, and the third mode include interaural audio time difference modification. In some examples, the audio generated in the first mode is modified using one or more of HRTFs, interaural audio time difference, and cross cancellation. In some examples, the audio generated in the second mode is modified using one or more of HRTFs, interaural audio time difference, and cross cancellation. In some examples, the audio generated in the third mode is modified using one or more of HRTFs, interaural audio time difference, and cross cancellation. Modifying the interaural audio time difference allows a user to perceive audio as coming from a particular direction. This audio directional information provides the user with additional feedback about the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by assisting the user in providing appropriate inputs) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0329] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio (e.g., right channel) of the audio and a second channel audio (e.g., left channel) of the audio to form composite channel audio, updating the second channel audio to include the composite channel audio at a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio at a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay into the first channel audio (e.g., right channel) without introducing a time delay into the second channel audio (e.g., left channel) different from the first audio channel.

[0330] It should be noted that the details of the processes described above with respect to method 1200 (e.g., FIGS. 12A-12B) are also applicable in an analogous manner to the methods described below and above. For example, methods 700, 900, 1400, and 1500 optionally include one or more of the characteristics of the various methods described above with respect to method 1200. For example, the same or similar techniques are used to place audio within a space. In other embodiments, the same audio source may be used in the various techniques. In yet other embodiments, the currently playing audio in each of the various methods may be manipulated using techniques described in the other methods. For the sake of brevity, these details will not be repeated below.

[0331] 13A-13F illustrate exemplary techniques for managing headphone transparency, according to some embodiments. The techniques in these figures are used to illustrate processes described below, including the processes of FIGS. 14A-14B.

[0332] 13G-13M illustrate example techniques for manipulating multiple audio streams of an audio source, according to some embodiments. The techniques in these figures are used to illustrate processes described below, including the process of FIG.

[0333] 13A-13M show device 1300 (e.g., a mobile phone) having a touch-sensitive surface connected to a display and left and right speakers (e.g., headphones 1358). In some examples, the left and right speakers are operable to operate independently at a noise cancellation level (e.g., noise from outside the headphones is suppressed so the user hears less noise, thus low noise transparency), and at a full transparency level (e.g., noise from outside the headphones is passed completely to the user, or the headphones pass as much as possible so the user can hear those noises, thus high noise transparency). In FIGS. 13A-13C, device 1300 operates left speaker 1358a and right speaker 1358b at a noise cancellation level, as indicated by the filled-in speakers. However, for example, device 1300 could operate left speaker 1358a at a noise cancellation level while operating right speaker 1358b at a full transparency level.

[0334] An overhead view 1350 is a visual representation of the spatial configuration of audio being generated by device 1300 and is shown throughout FIGS. 13A-13M to provide the reader with a better understanding of the techniques, particularly with respect to where user 1356 perceives sound as coming from (e.g., as a result of device 1300 positioning the audio in space). Time 1350 is not part of the user interface of device 1300. In some examples, the techniques described below enable a user to more easily and efficiently listen to audio sources that are not generated by the device they are currently listening to. For example, this technique allows a user to more clearly hear someone speaking to them while listening to music using headphones. In some examples, the techniques described below enable a user to easily manipulate various audio streams.

[0335] 13A, device 1300 displays a music player 1304 including album art 1304a and, optionally, axes 1304b. The album art corresponds to the song being played by device 1300. In some embodiments, the song includes multiple audio streams. In this embodiment, the song includes five audio streams, each corresponding to a particular instrument.

[0336] In FIG. 13A , device 1300 uses left speaker 1358a and right speaker 1358b to generate audio for a song (including all five audio streams) without spatially positioning the audio (e.g., device 1300 does not apply interaural time difference, HRTF, or cross-cancellation). For example, this results in a user perceiving the audio as being within the user's 1356's head, as illustrated by audio element 1354 in overhead view 850 of FIGS. 8F-8G . Audio element 1354 corresponds to a song, including five audio streams. Device 1300 is configured so that user input received at affordance 1304c controls the music, such as by pausing, playing, fast-forwarding, and rewinding the music.

[0337] In some examples, user 1356 sees the person to their right speaking, and the user provides a drag input 1310a. In Figures 13B-13D, device 1300 detects drag input 1310a, which displaces album art 1304a. For example, device 1300 updates the display of the position of album art 1304a to correspond to the movement of drag input 1310a.

[0338] In Figures 13B and 13C, the displacement of the album art 1304a does not exceed a predetermined distance, and the device 1300 continues to generate the audio of the song (including all five audio streams) using the left speaker 1358a and the right speaker 1358b without placing the audio in space (e.g., the device 1300 does not apply interaural time difference, HRTF, or cross-cancellation) and while maintaining operation of the left speaker 1358a and the right speaker 1358b at noise cancellation levels.

[0339] In Figure 13D, device 1300 determines that the displacement of album art 1304a exceeds a predetermined distance (e.g., the user has moved album art 1304a far enough). In response to determining that the displacement of album art 1304a exceeds a predetermined distance, device 1300 transitions to generating audio for the song (including all five audio streams) by positioning the audio in space using left speaker 1358a and right speaker 1358b (e.g., device 1300 applies interaural time difference, HRTF, and / or cross-cancellation) so that the user perceives the audio as being located in space away from the person to the user's right (e.g., forward and to the left), as shown in overhead view 1350 of Figure 13D. In some embodiments, the audio is moved to a location in space based on the direction and / or distance of drag input 1310a. Further, in response to determining that the displacement of the album art 1304a exceeds a predetermined distance, the device 1300 maintains operation of the left speaker 1358a at a noise cancellation level and transitions operation of the right speaker 1358b to a fully transparent level, as illustrated in the overhead view 1350 of FIG. 13D by the left speaker 1358a being filled in and the right speaker 1358b not being filled in.

[0340] While device 1300 maintains the placement of audio in space (e.g., device 1300 applies interaural time difference, HRTF, and / or cross-cancellation) and operates left speaker 1358a at a noise cancellation level and right speaker 1358b at a full transparency level, in FIG. 13E device 1300 detects lift-off of drag input 1310a.

[0341] In response to detecting lift-off of the drag input 1310a, as shown in FIG. 13F, device 1300 returns the display of album art 1304a to its original pre-drag position and transitions to generating the song's audio (including all five audio streams) using left 1358a and right 1358b speakers without spatially positioning the audio (e.g., device 1300 does not apply interaural time difference, HRTF, and / or cross-cancellation), thereby causing the user to perceive the audio as being within the user's head. Further, in response to detecting lift-off of the drag input 1310a, device 1300 maintains operation of left speaker 1358a at a noise cancellation level and transitions operation of right speaker 1358b to a noise cancellation level, as shown in the overhead view 1350 of FIG. 13D by the left speaker 1358a and right speaker 1358b being filled in.

[0342] In some embodiments, in response to detecting lift-off of the drag input 1310a, the device 1300 maintains the placement of the audio in space and maintains operation of the right speaker 1358b at a fully transparent level.

[0343] In Figure 13G, device 1300 continues to generate the audio of the song (including all five audio streams) using left speaker 1358a and right speaker 1358b without spatially positioning the audio (e.g., device 1300 does not apply interaural time difference, HRTF, or cross-cancellation). In Figure 13G, device 1300 detects input 1310b.

[0344] In some embodiments, device 1300 determines whether the characteristic intensity of input 1310b exceeds the intensity threshold. In response to device 1300 determining that the characteristic intensity of input 1310b does not exceed the intensity threshold, device 1300 continues to generate the audio of the song (including all five audio streams) using left speaker 1358a and right speaker 1358b without placing the audio in space.

[0345] In response to device 1300 determining that the characteristic intensity of input 1310b exceeds the intensity threshold, device 1300 transitions to generating the audio of the various audio streams for the song (including all five audio streams) using left speaker 1358a and right speaker 1358b by spatially positioning the various audio, so that the user perceives the various audio streams as being in different directions away from the user's head, as illustrated by audio element 1354 being separated into audio elements 1354a-1354e in FIGS. 13G-13I. Further, in response to device 1300 determining that the characteristic intensity of input 1310b exceeds the intensity threshold, device 1300 updates the display of music player 1304 to show animation of stream affordances 1306a-1306e being displayed and spreading apart, as shown in FIGS. 13H-13I. For example, each of stream affordances 1306a-1306e corresponds to a respective audio stream of the song. For example, stream affordance 1306a corresponds to a first audio stream that includes a singer (but does not include a guitar, keyboard, drums, etc.). Similarly, audio element 1354a corresponds to the first audio stream. In another example, stream affordance 1306d corresponds to a second audio stream that includes a guitar (but does not include a singer, keyboard, drums, etc.). Similarly, audio element 1354d corresponds to the second audio stream.

[0346] As shown in Figures 13H-13I, device 1300 displays stream affordances 1306a-1306e at positions on the display that correspond to positions in space where device 1300 has placed the corresponding audio streams (e.g., relative to each other, relative to points on the display).

[0347] 13J-13K, device 1300 detects drag gesture 1310c on stream affordance 1306d. In response to detecting drag gesture 1310c, device 1300 transitions to generating audio for a second audio stream (corresponding to audio element 1354e) using left speaker 1358a and right speaker 1358b without positioning the audio in space, so that the user perceives the audio for the second audio stream as being within the user's head, and continues to generate various other audio streams for songs (e.g., corresponding to audio elements 1354a-1354d) at various locations in space, as shown by audio elements 1354a-1354e in FIG. 13K. Thus, the device moves the audio streams of individual songs into and out of the user's head and to different locations in space based on detecting drag inputs on the corresponding stream affordances 1306a-1306d. In some embodiments, the device positions different audio streams at positions in space based on detected drag input (e.g., direction of movement, lift-off placement).

[0348] In FIG. 13L , device 1300 detects tap input 1310d on stream affordance 1306a. In response to detecting tap input 1310d on stream affordance 1306a, device 1300 stops generating audio for the second audio stream, thus disabling the audio stream. Further, in response to detecting tap input 1310d on stream affordance 1306a, device 1300 updates one or more visual characteristics (e.g., shading, size, movement) of stream affordance 1306a to indicate that the corresponding audio stream has been disabled. In some examples, device 1300 receives input (e.g., a drag gesture on stream affordance 1306a) that moves the disabled audio stream to a new position in space without generating the disabled audio stream. In response to detecting an additional tap input on the stream affordance corresponding to the disabled audio stream, device 1300 enables the audio stream and begins generating audio for the enabled audio stream at the new position in space.

[0349] In some examples, device 1300 includes a digital assistant that generates audio feedback, such as returning the results of a query by speaking the results. In some examples, device 1300 generates the digital assistant's audio by positioning the audio for the digital assistant at a location in space (e.g., over the user's right shoulder), so that the user perceives the digital assistant as remaining stationary in space even when other audio moves in space. In some examples, device 1300 emphasizes the digital assistant's audio by ducking one or more (or all) other audio.

[0350] 14A-14B are flow diagrams illustrating a method for managing headphone transparency using an electronic device, according to some embodiments. Method 1400 is performed on a device (e.g., 100, 300, 500, 1300) having a display and a touch-sensitive surface. The electronic device is operatively connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earbuds, left and right earbuds). Some operations of method 1400 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0351] As described below, method 1400 provides an intuitive way to manage headphone transparency. This method reduces the cognitive burden on the user to manage headphone transparency, thereby creating a more efficient human-machine interface. For battery-operated computing devices, allowing the user to manage headphone transparency faster and more efficiently conserves power and extends the time between battery charges.

[0352] The electronic device displays (1402) a user-movable affordance (e.g., 1304a, album art) at a first location on the display.

[0353] While the user-movable affordance (e.g., 1304a) is displayed in a first position (e.g., 1304a as shown in FIG. 13A ) (1404), the electronic device operates the electronic device in a first state of ambient sound transparency (e.g., a state in which transparency is disabled for both a first speaker and a second speaker of the two or more speakers, 1358a and 1358b as shown in FIG. 13A ), and external noise is suppressed for a first speaker and a second speaker of the two or more speakers, 1358a and 1358b, as shown in FIG. 13A ) (1408).

[0354] While a user-movable affordance (e.g., 1304a) is displayed in a first position (e.g., 1304a as shown in FIG. 13A) (1404), the electronic device generates audio (e.g., 1354) using two or more speakers and an audio source (e.g., an audio file, a music file, a media file) in a first mode (e.g., 1354 in FIG. 13A, no HRTF, cross-cancellation, or interaural time difference applied, stereo mode) (1410).

[0355] While the user-movable affordance (e.g., 1304a) is displayed (1404) in a first position (e.g., 1304a as shown in FIG. 13A), the electronic device detects (1412) a user input (e.g., 1310a, a drag gesture) using the touch-sensitive surface.

[0356] The set of one or more conditions includes a first condition 1416 that is met if the user input (e.g., 1310a) is a touch-and-drag action on the user-movable affordance (e.g., 1304a). According to some embodiments, the set of one or more conditions further includes a second condition 1418 that is met if the user input causes a displacement of the user-movable affordance by at least a predetermined amount from a first position. For example, small movements of the movable affordance (e.g., 1304a) do not cause a change, and the movable affordance (e.g., 1304a) must move a minimum distance while the mode is changed from a first mode to a second mode and before the state is changed from a first state to a second state. This helps avoid inadvertent mode and state changes.

[0357] In response to detecting 1414 a user input (e.g., 1310a) and in response to one or more sets of conditions being met 1416, the electronic device operates 1420 the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency (e.g., a state in which transparency is enabled for a first speaker and disabled for a second speaker of two or more speakers, or a state in which external noise is suppressed for a first speaker and not suppressed for a second speaker of two or more speakers, 1358a and 1358b as shown in FIG. 13D ). Changing the ambient sound transparency state allows the user to better hear sounds from the user's environment, and in particular sounds from particular directions within the environment.

[0358] Further, in response to detecting 1414 a user input (e.g., 1310a) and in response to one or more sets of conditions being met 1416, the electronic device transitions 1422 the generation of audio using the audio source (e.g., 1354 shown in FIG. 13D ) from a first mode to a second mode different from the first mode (e.g., mono mode applying one or more of HRTF, cross-cancellation, and interaural time difference). Changing the mode in which audio is generated allows the user to better hear sounds from the user's environment, and in particular sounds from particular directions within the environment (e.g., particular directions away from the direction in which the audio generation is moved in space).

[0359] Further, in response to detecting 1414 a user input (e.g., 1310a) and in response to one or more sets of conditions not being met 1424, the electronic device maintains 1426 the electronic device in a first ambient sound transparency state (e.g., a state in which transparency is disabled for both the first speaker and the second speaker).

[0360] Further, in response to detecting 1414 a user input (e.g., 1310a) and in response to the set of one or more conditions not being met 1424, the electronic device continues to generate audio (e.g., 1354 as shown in FIG. 13C) using the audio source in the first mode.

[0361] In some embodiments, the speakers (e.g., headphones) are operable to operate at a full noise cancellation level (e.g., noises from outside the headphones are suppressed and the headphones are suppressed as much as possible so the user cannot hear them, thus low noise transparency), and a full transparency level (noises from outside the headphones are passed completely to the user, or passed as much as possible through the headphones so the user can hear them, thus high noise transparency).

[0362] Further, in response to detecting 1414 the user input (e.g., 1310a), the electronic device updates 1430 on the display the representation of the user-movable affordance (e.g., 1304a) from a first position on the display (e.g., position 1304a in FIG. 13A ) to a second position on the display (e.g., position 1304a in FIG. 13D ) according to movement of the user input (e.g., 1310a) on the touch-sensitive surface. In some examples, the display position of the user-movable affordance (e.g., 1304a) is based on the user input, such as by updating the representation of the user-movable affordance (e.g., 1304a) to correspond to the location of the contact of the user input. In this way, the farther the user input (e.g., 1310a) is moved, the farther the user-movable affordance (e.g., 1304a) moves on the display.

[0363] According to some embodiments, after (e.g., during) a set of one or more conditions is met, the electronic device detects an end of user input (e.g., by detecting liftoff of the user input contact on the touch-sensitive surface, 1310a). In response to detecting the end of user input, the electronic device updates, on the display, the representation of the user-movable affordance (e.g., 1304a in FIG. 13F) to a first position on the display. Further, in response to detecting the end of user input, the electronic device transitions the electronic device to operate in an ambient sound transparent first state (e.g., as shown by 1358a and 1358b in FIG. 13F). Further, in response to detecting the end of user input, the electronic device transitions audio generation from the second mode to the first mode using the audio source (e.g., as shown by the change in position of 1354 in FIGS. 13E-13F).

[0364] According to some embodiments, after (e.g., during) a set of one or more conditions is met, the electronic device detects an end of user input (e.g., by detecting lift-off of the user input contact on the touch-sensitive surface). In response to detecting the end of user input, the electronic device maintains, on the display, the representation of the user-movable affordance at a second position on the display. Further, in response to detecting the end of user input, the electronic device maintains operation of the electronic device in a second state of ambient sound transparency. Further, in response to detecting the end of user input, the electronic device maintains generating audio using the audio source in a second mode.

[0365] According to some embodiments, the user input (e.g., 1310a) includes a direction of movement. According to some embodiments, the second state of the ambient sound transparency (as shown by 1358a and 1358b in FIG. 13D ) is based on the direction of movement of the user input (e.g., 1310a). For example, the electronic device selects a particular speaker (e.g., 1358b) to change the transparency state based on the direction of movement of the user input (e.g., 1310a). According to some embodiments, the second mode is based on the direction of movement of the user input. For example, in the second mode, the electronic device generates audio such that the audio is perceived to be generated from a direction in space, with the direction of movement being based on the direction of movement of the first input (and / or based on the position of the user-movable affordance 1304a).

[0366] According to some embodiments, the second mode is configured such that audio produced using the second mode is perceived by the user as originating from a point in space corresponding to a second position of the displayed user-movable affordance.

[0367] According to some embodiments, the first mode does not include interaural time difference correction of the audio, and in some examples, the audio produced in the first mode is not corrected using any of HRTFs, cross-cancellation, or interaural time difference.

[0368] According to some embodiments, the second mode includes modifying the interaural time difference of the audio. In some examples, the second mode includes modifying the interaural time difference of the audio, such that the audio is perceived by the user as originating from a direction corresponding to the location of the display or the user-movable affordance, such as the left or right of the user. In some examples, the audio generated in the second mode is modified using one or more of HRTFs, cross-cancellation, and interaural time difference. Modifying the interaural time difference of the audio allows the user to perceive the audio as coming from a particular direction. This audio direction information provides the user with additional feedback about the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by helping the user provide appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0369] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio (e.g., right channel) of the audio and a second channel audio (e.g., left channel) of the audio to form composite channel audio, updating the second channel audio to include the composite channel audio at a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio at a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay into the first channel audio (e.g., right channel) without introducing a time delay into the second channel audio (e.g., left channel) different from the first audio channel.

[0370] It should be noted that the details of the processes described above with respect to method 1400 (e.g., FIGS. 14A-14B) are also applicable in an analogous manner to the methods described below and above. For example, methods 700, 900, 1200, and 1500 optionally include one or more of the characteristics of the various methods described above with respect to method 1400. For example, the same or similar techniques are used to place audio within a space. In other embodiments, the same audio source may be used in the various techniques. In yet other embodiments, the currently playing audio in each of the various methods may be manipulated using techniques described in the other methods. For the sake of brevity, these details will not be repeated below.

[0371] 15 is a flow diagram illustrating a method for manipulating multiple audio streams of an audio source using an electronic device, according to some embodiments. Method 1500 is performed on a device (e.g., 100, 300, 500, 1300) having a display and a touch-sensitive surface. The electronic device is operatively connected to two or more speakers, including a first (e.g., left, 1358a) speaker and a second (e.g., right, 1358b) speaker. For example, the two or more speakers are left and right speakers, left and right headphones, left and right earbuds, or left and right earbuds. Some operations of method 1500 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0372] As described below, method 1500 provides an intuitive way to manipulate multiple audio streams of an audio source. This method reduces the cognitive burden on a user to manipulate multiple audio streams of an audio source, thereby creating a more efficient human-machine interface. For battery-operated computing devices, allowing a user to manipulate multiple audio streams of an audio source faster and more efficiently conserves power and extends the time between battery charges.

[0373] The electronic device generates (1502) audio (e.g., 1354 in FIG. 13F ) using an audio source (e.g., an audio file, a music file, a media file) in a first mode (e.g., a stereo mode without applying HRTF, cross-cancellation, or interaural time difference) using two or more speakers. The audio source includes multiple (e.g., five) audio streams, including a first audio stream and a second audio stream. In some embodiments, each audio stream of the audio source is a stereo audio stream. In some embodiments, each audio stream is limited to a single respective device. In some embodiments, each audio stream is limited to the voice of a single respective singer. Thus, each audio stream of the audio source is generated in the first mode.

[0374] The electronic device detects (1504) a first user input (eg, 1310b, a tap on the affordance, an input having a characteristic intensity that exceeds an intensity threshold) using the touch-sensitive surface.

[0375] In response to detecting (1506) a first user input (e.g., 1310b), the electronic device transitions (1508) the generation of the first audio stream (e.g., 1354a in FIG. 13H) of the audio source from a first mode to a second mode, different from the first mode, using two or more speakers.

[0376] Further, in response to detecting 1506 a first user input (e.g., 1310b), the electronic device transitions 1510 the generation of a second audio stream (e.g., 1354b in FIG. 13H) of the audio source from the first mode to a third mode, different from the first mode and the second mode, using two or more speakers.

[0377] By placing various audio streams at different locations in space, a user can better distinguish between different audio streams. This allows a user to more quickly and efficiently turn off (and on) specific portions of audio (e.g., specific audio streams) that the user wants to exclude from their listening experience. Providing improved audio feedback to the user improves device usability and makes the user-device interface more efficient (e.g., by assisting the user in providing appropriate input when operating / interacting with the device and reducing user errors), as well as reducing power usage and improving the device's battery life by allowing the user to use the device more quickly and efficiently.

[0378] Further, in response to detecting (1506) a first user input (e.g., 1310b), the electronic device displays (1512) on the display a first visual representation (e.g., 1306a) of the first audio stream of the audio source.

[0379] Further, in response to detecting (1506) a first user input (e.g., 1310b), the electronic device displays (1514) on the display a second visual representation (e.g., 1306b) of a second audio stream of the audio source, wherein the first visual representation (e.g., 1306a) is different from the second visual representation (e.g., 1306b).

[0380] According to some embodiments, blocks 1508-1514 occur simultaneously.

[0381] According to some embodiments, the first mode does not include interaural time difference correction of the audio, and in some examples, the audio produced in the first mode is not corrected using any of HRTFs, cross-cancellation, or interaural time difference.

[0382] According to some embodiments, the second mode includes modifying the interaural time difference of the audio. In some examples, the second mode includes modifying the interaural time difference of the audio, such that the audio is perceived by the user as originating from a direction corresponding to the location of the corresponding visual representation. In some examples, the second mode includes applying one or more of HRTF, cross-cancellation, and interaural time difference. In some examples, the third mode includes modifying the interaural time difference of the audio. In some examples, the third mode includes applying one or more of HRTF, cross-cancellation, and interaural time difference. Modifying the interaural time difference of the audio allows the user to perceive the audio as coming from a particular direction. This audio direction information provides the user with additional feedback about the placement of different content. Providing improved audio feedback to the user enhances device usability (e.g., by helping the user provide appropriate input) and makes the user device interface more efficient, which in turn reduces power usage and improves device battery life by allowing the user to use the device more quickly and efficiently.

[0383] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio (e.g., right channel) of the audio stream and a second channel audio (e.g., left channel) of the audio stream to form composite channel audio, updating the second channel audio to include the composite channel audio at a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio at a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay into the first channel audio (e.g., right channel) without introducing a time delay into the second channel audio (e.g., left channel) different from the first audio channel.

[0384] According to some embodiments, the electronic device displays a first visual representation (e.g., 1306a) of a first audio stream of an audio source on a display by displaying a first visual representation at a first position and sliding the first visual representation in a first direction toward a second position (e.g., 1306a in FIG. 13I) on the display. According to some embodiments, the electronic device displays a second visual representation (e.g., 1306b) of a second audio stream of an audio source on a display by displaying a second visual representation at the first position and sliding the second visual representation in a second direction different from the first direction toward a third position (e.g., 1306b in FIG. 13I) different from the second position.

[0385] According to some embodiments, the electronic device detects a second user input (e.g., 1310c) using the touch-sensitive surface, where the second input (e.g., 1310c) begins at a position corresponding to the second position and ends at a position corresponding to the first position. In response to detecting the second user input, the electronic device slides the first visual representation on the display from the second position to the first position while maintaining the second visual representation at a third position on the display. Further, in response to detecting the second user input, the electronic device transitions audio generation using the first audio stream using two or more speakers from the second mode to the first mode while maintaining audio generation using the second audio stream in the third mode.

[0386] According to some embodiments, detecting a first user input (e.g., 1310a) includes accessing a characteristic intensity of the first user input and determining that the characteristic intensity of the first user input exceeds an intensity threshold.

[0387] According to some embodiments, while the electronic device generates audio using two or more speakers and a first audio stream of an audio source and a second audio stream of an audio source, the electronic device detects, using the touch-sensitive surface, a third user input (e.g., 1310d, a tap input) at a location on the touch-sensitive surface that corresponds to the first visual representation (e.g., 1306a). In response to detecting the third user input (e.g., 1310d), the electronic device stops generating audio (e.g., 1354a) using the first audio stream of the audio source and using the two or more speakers. Further, in response to detecting the third user input (e.g., 1310d), the electronic device continues generating audio using the second audio stream of the audio source and using the two or more speakers (e.g., 1354b). Thus, the electronic device detects a tap on a device affordance and disables the audio stream corresponding to the device.

[0388] According to some embodiments, while the electronic device does not generate audio using the first audio stream of the audio source using the two or more speakers and generates audio using the second audio stream of the audio source using the two or more speakers, the electronic device detects, using the touch-sensitive surface, a fourth user input (e.g., a tap input) at a location on the touch-sensitive surface that corresponds to the first visual representation. In response to detecting the fourth user input, the electronic device generates audio using the first audio stream of the audio source using the two or more speakers. Further, in response to detecting the fourth user input, the electronic device continues generating audio using the second audio stream of the audio source using the two or more speakers. Thus, the electronic device detects a tap on a device affordance and enables the audio stream corresponding to the device.

[0389] According to some embodiments, the electronic device detects a volume control input (e.g., a user input requesting a volume be turned down), and in response to detecting the volume control input, the electronic device modifies a volume of audio generated using two or more speakers for each audio stream in the plurality of audio streams.

[0390] It should be noted that the details of the processes (e.g., FIG. 15) described above with respect to method 1500 are also applicable in a similar manner to the methods described above. For example, methods 700, 900, 1200, and 1400 optionally include one or more of the characteristics of the various methods described above with respect to method 1500. For example, the same or similar techniques are used to place audio within a space. In other embodiments, the same audio source may be used in the various techniques. In yet other embodiments, the currently playing audio in each of the various methods may be manipulated using techniques described in the other methods.

[0391] The foregoing has been described with reference to specific embodiments for purposes of explanation. However, the exemplary discussion above is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described in order to best explain the principles of the present technique and their practical application. This will enable others skilled in the art to best utilize the present technique and various embodiments with various modifications as suited to the particular applications intended.

[0392] Although the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art, and such changes and modifications are to be understood as being included within the scope of the present disclosure and examples, as defined by the claims.

[0393] As mentioned above, one aspect of the present technology is the collection and use of data available from various sources to improve audio delivery. This disclosure contemplates that, in some examples, this collected data may include personal information data that uniquely identifies a particular person or that can be used to contact or locate a particular person. Such personal information data may include demographic data, location-based data, phone numbers, email addresses, Twitter IDs, home addresses, data or records regarding a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), birth date, or any other identifying or personal information.

[0394] This disclosure recognizes that the use of such personal information data in the present technology can be used to the benefit of the user. For example, the personal information data can be used to provide an improved audio experience. Additionally, other uses of personal information data that benefit the user are also contemplated by this disclosure. For example, health and fitness data can be used to provide insight into the user's overall wellness, or can be used as proactive feedback to individuals using the technology to pursue wellness goals.

[0395] This disclosure contemplates that entities involved in the collection, analysis, disclosure, transmission, storage, or other use of such personal information data adhere to robust privacy policies and / or privacy practices. Specifically, such entities should implement and consistently use privacy policies and practices that are generally recognized as meeting or exceeding industry or government requirements for maintaining the strict confidentiality of personal information data. Such policies should be easily accessible to users and should be updated as data collection and / or use changes. Personal information from users should be collected for the entity's lawful and legitimate use and should not be shared or sold except for those lawful uses. Furthermore, such collection / sharing should be carried out after the user's informed consent is obtained. Furthermore, such entities should consider taking all necessary measures to protect and secure access to such personal information data and to ensure that others with access to that personal information data comply with their privacy policies and procedures. Furthermore, such entities may be able to undergo third-party assessments to attest to their adherence to widely accepted privacy policies and practices. Furthermore, policies and practices should be tailored to the specific types of personal data collected and / or accessed and should comply with applicable laws and standards, including jurisdiction-specific considerations. For example, in the United States, the collection of or access to certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA), while health data in other countries may be subject to other regulations and policies and should be addressed accordingly. Therefore, different privacy practices should be maintained for different types of personal data in each country.

[0396] Notwithstanding the foregoing, the present disclosure also contemplates embodiments in which a user selectively blocks use of or access to personal information data. That is, the present disclosure contemplates that hardware and / or software elements may be provided to prevent or block access to such personal information data. For example, in the case of an audio distribution service, the present technology may be configured to allow a user to choose to "opt in" or "opt out" of participating in the collection of personal information data during registration for the service or at any time thereafter. In yet another example, a user may choose to limit the amount of time location data is collected or maintained. In addition to providing "opt-in" and "opt-out" options, the present disclosure contemplates providing notifications regarding access or use of personal information. For example, a user may be notified upon downloading an app that will access the user's personal information data, and then again immediately before the app accesses the user's personal information data.

[0397] Furthermore, it is the intent of this disclosure that personal information data should be managed and handled in a manner that minimizes the risk of unintentional or unauthorized access or use. Risk can be minimized by limiting data collection and deleting data when it is no longer needed. Furthermore, where applicable, de-identification of data can be used to protect user privacy in certain health-related applications. De-identification can be facilitated, where appropriate, by removing certain identifiers (e.g., date of birth, etc.), controlling the amount or specificity of data stored (e.g., collecting location data at a city level rather than a street address level), controlling how data is stored (e.g., aggregating data across users), and / or other methods.

[0398] Thus, while this disclosure broadly encompasses the use of personal information data to implement one or more various disclosed embodiments, this disclosure also contemplates that the various embodiments may be implemented without requiring access to such personal information data. That is, various embodiments of the present technology are not rendered inoperable by the absence of all or par...

Claims

1. 1. An electronic device comprising a display operatively connected to two or more speakers, displaying a first visual element at a first location on the display; accessing a first audio corresponding to the first visual element; While displaying the first visual element at the first location on the display, generating audio on the two or more speakers using the first audio in a first mode; Receiving a first user input; In response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to the first visual element not being displayed on the display; generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display; and A method comprising:

2. 2. The method of claim 1 , wherein the first mode is configured such that audio produced using the first mode is perceived by the user as originating from a first direction corresponding to the display.

3. accessing a second audio corresponding to a second visual element; While displaying the first visual element at the first position on the display and while not displaying the second visual element, generating audio in a third mode, different from the first mode and the second mode, in the two or more speakers using the second audio, wherein the third mode is configured such that the audio generated in the third mode is perceived by the user as being generated from a direction away from the display; and forgoing displaying a second visual element on the display corresponding to the second audio; and The method of claim 1 , further comprising:

4. While displaying the first visual element at the first location on the display, forgoing displaying a second visual element on the display corresponding to the second audio; and In response to receiving the first user input, transitioning the display of the second visual element that is not displayed on the display to a fourth position on the display; generating audio on the two or more speakers using the second audio in the first mode simultaneously with generating audio using the first audio in the second mode; The method of claim 1 , further comprising:

5. the second mode is configured such that audio produced using the second mode is perceived by the user as being produced from a second direction; the third mode is configured such that audio produced using the third mode is perceived by the user as being produced from a third direction different from the second direction.

5. The method according to any one of claims 1 to 4.

6. following displaying the first visual element at the first location on the display, and before the first visual element is no longer displayed on the display: displaying the first visual element at a second location on the display; and generating audio at the two or more speakers using the first audio in a fourth mode different from the first mode, the second mode, and the third mode while displaying the first visual element at the second location on the display; The method of claim 1 , further comprising:

7. The method of claim 1 , wherein the first mode does not include correction of interaural time differences of audio.

8. The method of claim 1 , wherein the second mode includes interaural time difference correction of audio.

9. Correcting the interaural time difference of the audio includes: combining the first channel audio of the audio and the second channel audio of the audio to form composite channel audio; updating the second channel audio to include the composite channel audio at a first delay; updating the first channel audio to include the composite channel audio at a second delay different from the first delay; 9. The method of claim 7, comprising:

10. In response to beginning to receive the first user input, transitioning from not generating audio using a second audio corresponding to a second visual element to generating audio using the second audio corresponding to the second visual element on the two or more speakers.

10. The method of claim 1, further comprising:

11. generating audio using the second mode includes: attenuating the audio; and applying a high pass filter to the audio; applying a low pass filter to the audio; Varying the volume balance between two or more speakers; 11. The method of claim 1, comprising one or more of:

12. 12. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display, the one or more programs including instructions for performing the method of any one of claims 1 to 11.

13. The display and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 12. An electronic device comprising: the one or more programs comprising instructions for performing the method of any one of claims 1 to 11, Electronic devices.

14. The display and means for carrying out the method according to any one of claims 1 to 11; , an electronic device.

15. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display, the electronic device being operatively connected to two or more speakers, the one or more programs comprising: displaying a first visual element at a first location on the display; accessing a first audio corresponding to the first visual element; While displaying the first visual element at the first location on the display, in a first mode, generating audio on the two or more speakers using the first audio; receiving a first user input; In response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to the first visual element not being displayed on the display; generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display. A non-transitory computer-readable storage medium containing instructions.

16. The display and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 1. An electronic device comprising: the electronic device operatively connected to two or more speakers; and the one or more programs: displaying a first visual element at a first location on the display; accessing a first audio corresponding to the first visual element; While displaying the first visual element at the first location on the display, in a first mode, generating audio on the two or more speakers using the first audio; receiving a first user input; In response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to the first visual element not being displayed on the display; generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display. An electronic device containing instructions.

17. 1. An electronic device comprising: a display to which the electronic device is operatively connected and two or more speakers; means for displaying a first visual element at a first location on the display; means for accessing a first audio corresponding to the first visual element; While displaying the first visual element at the first location on the display, in a first mode, generating audio on the two or more speakers using the first audio; means for receiving a first user input; In response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to the first visual element not being displayed on the display; generating audio on the two or more speakers using the first audio in a second mode different from the first mode while not displaying the first visual element on the display, the second mode being configured such that the audio generated in the second mode is perceived by a user as being generated from a direction away from the display. Means and , an electronic device.

18. 1. An electronic device comprising a display and a touch-sensitive surface operatively connected to two or more speakers, displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file; Detecting a user contact at a location corresponding to a first media element using the touch-sensitive surface; in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a first audio file corresponding to the first media element without exceeding a predetermined audio playback duration; while the user contact remains at the location corresponding to the first media element; continuing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; and ceasing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; using the touch-sensitive surface to detect movement of the user contact from the location corresponding to the first media element to a location corresponding to a second media element; in response to detecting the user contact at the location corresponding to the second media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a second audio file corresponding to the second media element without exceeding the predetermined audio playback duration; while the user contact remains at the location corresponding to the second media element; continuing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; and ceasing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; detecting lift-off of the user contact using the touch-sensitive surface; In response to detecting the lift-off of the user contact, ceasing to generate audio using the two or more speakers using the first audio file or the second audio file; and A method comprising:

19. further in response to said user contact including a touch and hold input; generating audio using the first audio file without exceeding a predetermined audio playback duration; a first mode configured such that audio produced using the first mode is perceived by a user as being produced from a first direction; and a second mode different from the first mode; and a third mode different from the first mode and the second mode, the third mode being configured such that audio produced using the third mode is perceived by the user as being produced from a third direction different from the first direction; and transitioning the audio between a plurality of modes, including 20. The method of claim 18, further comprising:

20. 20. The method of claim 18, wherein the first mode comprises a first modification of an audio interaural time difference, and the third mode comprises a second modification of the audio interaural time difference that is different from the first modification.

21. 21. The method of any one of claims 18 to 20, wherein the second mode does not include correction of interaural time differences of audio.

22. Correcting the interaural time difference of the audio includes: combining the first channel audio of the audio and the second channel audio of the audio to form composite channel audio; updating the second channel audio to include the composite channel audio at a first delay; updating the first channel audio to include the composite channel audio at a second delay different from the first delay; 22. The method of any one of claims 20 to 21, comprising:

23. In response to detecting the user contact at the location corresponding to the first media element, modifying a first visual characteristic of the displayed media elements in the plurality of media elements other than the first media element; In response to detecting the user contact at the location corresponding to the second media element, undoing the change in the first visual characteristic of the second media element; and modifying the first visual characteristic of the first media element; 23. The method of any one of claims 18 to 22, further comprising:

24. In response to detecting the user contact at the location corresponding to the first media element, modifying the second visual characteristic of the first media element without modifying the second visual characteristic of the displayed media elements in the plurality of media elements other than the first media element; In response to detecting the user contact at the location corresponding to the second media element, undoing the change in the second visual characteristic of the first media element; and modifying the second visual characteristic of the second media element; 24. The method of any one of claims 18 to 23, further comprising:

25. in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a tap input; generating audio using the first audio file corresponding to the first media element without automatically ceasing generating audio using the first audio file after the predetermined audio playback period and without automatically ceasing generating audio using the first audio file upon detecting lift-off of the user contact; 25. The method of any one of claims 18 to 24, further comprising:

26. In response to said user contact, including a tap input, generating audio using the first audio file, including generating audio using the first audio file in the second mode without transitioning the audio between the first mode and the third mode of the plurality of modes; 26. The method of any one of claims 18 to 25, further comprising:

27. 27. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display, the one or more programs including instructions for performing the method of any one of claims 18 to 26.

28. The display and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 27. An electronic device comprising: the one or more programs comprising instructions for performing the method of any one of claims 18 to 26, Electronic devices.

29. The display and means for carrying out the method according to any one of claims 18 to 26; , an electronic device.

30. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs comprising: displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file; Detecting a user contact at a location corresponding to a first media element using the touch-sensitive surface; in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a first audio file corresponding to the first media element without exceeding a predetermined audio playback duration; while the user contact remains at the location corresponding to the first media element; continuing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; using the touch-sensitive surface to detect movement of the user contact from the location corresponding to the first media element to a location corresponding to a second media element; in response to detecting the user contact at the location corresponding to the second media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a second audio file corresponding to the second media element without exceeding the predetermined audio playback duration; while the user contact remains at the location corresponding to the second media element; continuing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the second audio file using the two or more speakers in response to the predetermined audio playback duration being exceeded; Detecting lift-off of the user contact using the touch-sensitive surface; In response to detecting the lift-off of the user contact, ceasing to generate audio using the two or more speakers using the first audio file or the second audio file; A non-transitory computer-readable storage medium containing instructions.

31. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 1. An electronic device comprising: the electronic device operatively connected to two or more speakers; and the one or more programs: displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file; Detecting a user contact at a location corresponding to a first media element using the touch-sensitive surface; in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a first audio file corresponding to the first media element without exceeding a predetermined audio playback duration; while the user contact remains at the location corresponding to the first media element; continuing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; using the touch-sensitive surface to detect movement of the user contact from the location corresponding to the first media element to a location corresponding to a second media element; in response to detecting the user contact at the location corresponding to the second media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a second audio file corresponding to the second media element without exceeding the predetermined audio playback duration; while the user contact remains at the location corresponding to the second media element; continuing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the second audio file using the two or more speakers in response to the predetermined audio playback duration being exceeded; Detecting lift-off of the user contact using the touch-sensitive surface; In response to detecting the lift-off of the user contact, ceasing to generate audio using the two or more speakers using the first audio file or the second audio file; Including instructions, Electronic devices.

32. 1. An electronic device comprising: The display and a touch-sensitive surface of the electronic device operatively connected to two or more speakers; means for displaying on the display a list of a plurality of media elements, each media element of the plurality of media elements corresponding to a respective media file; means for detecting a user contact at a location corresponding to a first media element using the touch-sensitive surface; in response to detecting the user contact at the location corresponding to the first media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a first audio file corresponding to the first media element without exceeding a predetermined audio playback duration; while the user contact remains at the location corresponding to the first media element; continuing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the first audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; Means and means for detecting, using the touch-sensitive surface, movement of the user contact from the location corresponding to the first media element to a location corresponding to a second media element; in response to detecting the user contact at the location corresponding to the second media element and in accordance with the user contact including a touch and hold input; generating audio using the two or more speakers using a second audio file corresponding to the second media element without exceeding the predetermined audio playback duration; Means and while the user contact remains at the location corresponding to the second media element; continuing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration not being exceeded; ceasing to generate audio using the second audio file using the two or more speakers in accordance with the predetermined audio playback duration being exceeded; Means and means for detecting lift-off of the user contact using the touch-sensitive surface; In response to detecting the lift-off of the user contact, ceasing to generate audio using the two or more speakers using the first audio file or the second audio file; Means and , an electronic device.

33. 1. An electronic device comprising a display and a touch-sensitive surface operatively connected to two or more speakers, Detecting a first user input to activate a discovery mode; In response to detecting the first user input to activate the discovery mode, using the two or more speakers: a first audio source in a first mode configured such that audio produced using the first mode is perceived by a user as being produced from a first point in space moving over time in a first direction along a predetermined path at a first speed; a second audio source in a second mode configured such that audio produced using the second mode is perceived by the user as being produced from a second point in space moving over time in the first direction along the predetermined path at a second speed; and a third audio source in a third mode, the third mode configured such that audio produced using the third mode is perceived by the user as being produced from a third point in space moving over time in the first direction along the predetermined path at a third speed; and and simultaneously generating audio using the first point, the second point, and the third point are different points in space; method.

34. the first audio source corresponds to a first visual element; the second audio source corresponds to a second visual element; the third audio source corresponds to a third visual element; The method comprises: generating audio using the first audio source, the second audio source, and the third audio source while simultaneously displaying simultaneous movement of two or more of the first visual element, the second visual element, and the third visual element on the display at a fourth rate; Further comprising:

34. The method of claim 33.

35. 35. The method of claim 34, wherein the movement of the two or more of the first visual element, the second visual element, and the third visual element is in the first direction.

36. the predetermined path varies along a first dimension; the predetermined path varies along a second dimension different from the first dimension; the predetermined path does not vary along a third dimension different from the first dimension and the second dimension; 36. The method of any one of claims 33 to 35.

37. Simultaneously generating audio using two or more speakers includes: using a fourth audio source in a fourth mode, the fourth mode configured such that audio produced using the fourth mode is perceived by the user as being produced from a fourth point in space moving over time in the first direction along the predetermined path, the fourth point in space being farther from the user than the first point, the second point, and the third point; refraining from displaying a fourth visual element corresponding to the fourth audio source on the display while displaying the simultaneous movement of the two or more of the first visual element, the second visual element, and the third visual element on the display; 37. The method of any one of claims 33 to 36, further comprising:

38. 38. The method of any one of claims 33 to 37, wherein the first speed, the second speed, and the third speed are the same speed.

39. detecting a second user input in a second direction different from the first direction while simultaneously producing the audio using the two or more speakers; in response to detecting the second user input in the second direction; updating the audio production of the first audio source, the second audio source, and the third audio source using a mode configured such that the audio sources are perceived by the user as moving over time in the second direction along the predetermined path; updating, on the display, a representation of the movement of one or more of the first visual element, the second visual element, and the third visual element such that the movement is in the second direction; 39. The method of any one of claims 34 to 38, further comprising:

40. detecting a third user input while simultaneously producing the audio using the two or more speakers; In response to detecting the third user input, updating the audio production of the first audio source, the second audio source, and the third audio source using a mode configured such that the audio sources are perceived by the user as moving over time in the first direction along the predetermined path at a fifth speed that is faster than the first speed; updating on the display a representation of the simultaneous movement of the first visual element, the second visual element, and the third visual element such that the movement is in the first direction at a sixth rate that is faster than the fourth rate; 39. The method of any one of claims 34 to 38, further comprising:

41. detecting a selection input at a location corresponding to the second visual element; in response to detecting the selection input, generating audio of the second audio file in a fifth mode using the two or more speakers, wherein audio generated using the fifth mode is not perceived by the user as being generated from a point in space that moves over time; and 41. The method of any one of claims 34 to 40, further comprising:

42. 42. The method of claim 41, wherein the fifth mode does not include correction of interaural time differences in audio.

43. 43. The method of any one of claims 33 to 42, wherein the first, second, and third modes include correction of interaural time differences of audio.

44. Correcting the interaural time difference of the audio includes: combining the first channel audio of the audio and the second channel audio of the audio to form composite channel audio; updating the second channel audio to include the composite channel audio at a first delay; updating the first channel audio to include the composite channel audio at a second delay different from the first delay; 44. The method of any one of claims 42 to 43, comprising:

45. 45. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the one or more programs including instructions for performing the method of any one of claims 33 to 44.

46. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 45. An electronic device comprising: the one or more programs comprising instructions for performing the method of any one of claims 33 to 44, Electronic devices.

47. 1. An electronic device comprising: The display and a touch-sensitive surface; and means for carrying out the method of any one of claims 33 to 44; , an electronic device.

48. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs comprising: Detecting a first user input to activate a discovery mode; In response to detecting the first user input to activate the discovery mode, using the two or more speakers: a first audio source in a first mode configured such that audio produced using the first mode is perceived by a user as being produced from a first point in space moving over time in a first direction along a predetermined path at a first speed; a second audio source in a second mode configured such that audio produced using the second mode is perceived by the user as being produced from a second point in space moving over time in the first direction along the predetermined path at a second speed; and a third audio source in a third mode, the third mode configured such that audio produced using the third mode is perceived by the user as being produced from a third point in space moving over time in the first direction along the predetermined path at a third speed; and Generate audio simultaneously using Contains instructions, the first point, the second point, and the third point are different points in space; A non-transitory computer-readable storage medium.

49. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 1. An electronic device comprising: the electronic device operatively connected to two or more speakers; and the one or more programs: Detecting a first user input to activate a discovery mode; In response to detecting the first user input to activate the discovery mode, using the two or more speakers: a first audio source in a first mode configured such that audio produced using the first mode is perceived by a user as being produced from a first point in space moving over time in a first direction along a predetermined path at a first speed; a second audio source in a second mode configured such that audio produced using the second mode is perceived by the user as being produced from a second point in space moving over time in the first direction along the predetermined path at a second speed; and a third audio source in a third mode, the third mode configured such that audio produced using the third mode is perceived by the user as being produced from a third point in space moving over time in the first direction along the predetermined path at a third speed; and Generate audio simultaneously using Contains instructions, the first point, the second point, and the third point are different points in space; Electronic devices.

50. 1. An electronic device comprising: The display and a touch-sensitive surface of the electronic device operatively connected to two or more speakers; means for detecting a first user input for activating a discovery mode; In response to detecting the first user input to activate the discovery mode, using the two or more speakers: a first audio source in a first mode configured such that audio produced using the first mode is perceived by a user as being produced from a first point in space moving over time in a first direction along a predetermined path at a first speed; a second audio source in a second mode configured such that audio produced using the second mode is perceived by the user as being produced from a second point in space moving over time in the first direction along the predetermined path at a second speed; and a third audio source in a third mode, the third mode configured such that audio produced using the third mode is perceived by the user as being produced from a third point in space moving over time in the first direction along the predetermined path at a third speed; and Generate audio simultaneously using and means for the first point, the second point, and the third point are different points in space; Electronic devices.

51. 1. An electronic device comprising a display and a touch-sensitive surface operatively connected to two or more speakers, Displaying a user-movable affordance at a first location on the display; While the user-movable affordance is displayed at the first position, operating the electronic device in a first state that is transparent to ambient sound; generating audio using an audio source in a first mode using the two or more speakers; Detecting user input using the touch-sensitive surface; In response to detecting the user input, according to a set of one or more conditions being satisfied, including a first condition being satisfied if the user input is a touch-and-drag action on the user-movable affordance; operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency; transitioning audio generation using the audio source from the first mode to a second mode different from the first mode; According to said set of one or more conditions not being met, maintaining the electronic device in the first state transparent to ambient sound; maintaining generating audio using the audio source in the first mode; and A method comprising:

52. further in response to detecting the user input, updating, on the display, a representation of the user-movable affordance from the first position on the display to a second position on the display in accordance with movement of the user input on the touch-sensitive surface; 52. The method of claim 51, further comprising:

53. 52. The method of claim 51 , wherein the set of one or more conditions includes a second condition that is satisfied if the user input displaces the user-movable affordance from the first position by at least a predetermined amount.

54. detecting an end of the user input after the set of one or more conditions is met; and In response to detecting the end of the user input, updating on the display a representation of the user-movable affordance to the first location on the display; and transitioning the electronic device to operate in the first state that is transparent to ambient sound; transitioning audio generation using the audio source from the second mode to the first mode; 54. The method of any one of claims 52 to 53, further comprising:

55. detecting an end of the user input after the set of one or more conditions is met; and In response to detecting the end of the user input, maintaining on the display a representation of the user-movable affordance at the second position on the display; and maintaining operation of the electronic device in the second state of ambient sound transparency; and maintaining audio generation using the audio source in the second mode; and 54. The method of any one of claims 52 to 53, further comprising:

56. the user input includes a direction of movement; the second state of ambient sound transparency is based on the direction of movement of the user input; the second mode is based on the direction of movement of the user input.

56. The method of any one of claims 51 to 55.

57. 57. The method of any one of claims 52 to 56, wherein the second mode is configured such that audio produced using the second mode is perceived by a user as originating from a point in space corresponding to the second position of the displayed user-movable affordance.

58. 58. The method of any one of claims 51 to 57, wherein the first mode does not include correction of interaural time differences in audio.

59. 59. The method of any one of claims 51 to 58, wherein the second mode includes interaural time difference correction of audio.

60. Correcting the interaural time difference of the audio includes: combining the first channel audio of the audio and the second channel audio of the audio to form composite channel audio; updating the second channel audio to include the composite channel audio at a first delay; updating the first channel audio to include the composite channel audio at a second delay different from the first delay; 60. The method of any one of claims 58 to 59, comprising:

61. 61. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the one or more programs comprising instructions for performing the method of any one of claims 51 to 60.

62. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 61. An electronic device comprising: the one or more programs comprising instructions for performing the method of any one of claims 51 to 60, Electronic devices.

63. The display and a touch-sensitive surface; and means for carrying out the method of any one of claims 51 to 60; , an electronic device.

64. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, the one or more programs comprising: displaying a user-movable affordance at a first location on the display; While the user-movable affordance is displayed at the first position, operating the electronic device in a first state that is transparent to ambient sound; generating audio using an audio source in a first mode using the two or more speakers; Detecting user input using the touch-sensitive surface; In response to detecting the user input, according to a set of one or more conditions being satisfied, including a first condition being satisfied if the user input is a touch-and-drag action on the user-movable affordance; operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency; transitioning audio generation using the audio source from the first mode to a second mode different from the first mode; According to said set of one or more conditions not being met, maintaining the electronic device in the first state transparent to ambient sound; maintaining generating audio using the audio source in the first mode; Including instructions, A non-transitory computer-readable storage medium.

65. 1. An electronic device comprising: The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 1. An electronic device comprising: the electronic device operatively connected to two or more speakers; and the one or more programs: displaying a user-movable affordance at a first location on the display; While the user-movable affordance is displayed at the first position, operating the electronic device in a first state that is transparent to ambient sound; generating audio using an audio source in a first mode using the two or more speakers; Detecting user input using the touch-sensitive surface; In response to detecting the user input, according to a set of one or more conditions being satisfied, including a first condition being satisfied if the user input is a touch-and-drag action on the user-movable affordance; operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency; transitioning audio generation using the audio source from the first mode to a second mode different from the first mode; According to said set of one or more conditions not being met, maintaining the electronic device in the first state transparent to ambient sound; maintaining generating audio using the audio source in the first mode; Including instructions, Electronic devices.

66. 1. An electronic device comprising: The display and a touch-sensitive surface of the electronic device operatively connected to two or more speakers; means for displaying a user-movable affordance at a first location on the display; While the user-movable affordance is displayed at the first position, operating the electronic device in a first state that is transparent to ambient sound; generating audio using an audio source in a first mode using the two or more speakers; using the touch-sensitive surface to detect user input; Means and In response to detecting the user input, according to a set of one or more conditions being satisfied, including a first condition being satisfied if the user input is a touch-and-drag action on the user-movable affordance; operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency; transitioning audio generation using the audio source from the first mode to a second mode different from the first mode; According to said set of one or more conditions not being met, maintaining the electronic device in the first state transparent to ambient sound; maintaining generating audio using the audio source in the first mode; Means and , an electronic device.

67. 1. An electronic device comprising a display and a touch-sensitive surface operatively connected to two or more speakers, including a first speaker and a second speaker, generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams including a first audio stream and a second audio stream; Detecting a first user input using the touch-sensitive surface; simultaneously, in response to detecting the first user input; transitioning the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode using the two or more speakers; transitioning the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode using the two or more speakers; displaying on the display a first visual representation of the first audio stream of the audio source; displaying on the display a second visual representation of the second audio stream of the audio source, the first visual representation being different from the second visual representation; and A method comprising:

68. 68. The method of claim 67, wherein the first mode does not include correction of interaural time differences in audio.

69. 69. The method of any one of claims 67 to 68, wherein the second mode includes interaural time difference correction of audio.

70. Correcting the interaural time difference of the audio includes: combining the first channel audio of the audio stream with the second channel audio of the audio stream to form composite channel audio; updating the second channel audio to include the composite channel audio at a first delay; updating the first channel audio to include the composite channel audio at a second delay different from the first delay; 70. The method of any one of claims 68 to 69, comprising:

71. displaying the first visual representation of the first audio stream of the audio source on the display includes displaying the first visual representation at a first position and sliding the first visual representation in a first direction towards a second position on the display; displaying the second visual representation of the second audio stream of the audio source on the display includes displaying the second visual representation at the first position and sliding the second visual representation in a second direction different from the first direction toward a third position different from the second position.

71. The method of any one of claims 67 to 70.

72. detecting a second user input using the touch-sensitive surface, the second input starting at a location corresponding to the second position and ending at a location corresponding to the first position; In response to detecting the second user input, sliding the first visual representation from the second position to the first position on the display while maintaining the second visual representation at the third position on the display; transitioning audio generation using the first audio stream using the two or more speakers from the second mode to the first mode while maintaining audio generation using the second audio stream in the third mode; 72. The method of claim 71, further comprising:

73. Detecting the first user input includes: accessing a characteristic strength of the first user input; determining that the characteristic intensity of the first user input exceeds an intensity threshold; 73. The method of any one of claims 67 to 72, comprising:

74. while the electronic device is generating audio using the first audio stream of the audio source and the second audio stream of the audio source using the two or more speakers; using the touch-sensitive surface to detect a third user input at a location on the touch-sensitive surface that corresponds to the first visual representation; In response to detecting the third user input, ceasing to generate audio using the first audio stream of the audio source using the two or more speakers; maintaining audio generation using the second audio stream of the audio source using the two or more speakers; 74. The method of any one of claims 67 to 73, further comprising:

75. While the electronic device is not generating audio using the first audio stream of the audio source using the two or more speakers and is generating audio using the second audio stream of the audio source using the two or more speakers, using the touch-sensitive surface to detect a fourth user input at a location on the touch-sensitive surface that corresponds to the first visual representation; In response to detecting the fourth user input, generating audio using the first audio stream of the audio source using the two or more speakers; maintaining audio generation using the second audio stream of the audio source using the two or more speakers; 75. The method of claim 74, further comprising:

76. Detecting a volume control input; modifying the volume of audio produced using the two or more speakers for each audio stream in the plurality of audio streams in response to detecting the volume control input; 76. The method of any one of claims 67 to 75, further comprising:

77. 77. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the one or more programs including instructions for performing the method of any one of claims 67 to 76.

78. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 77. An electronic device comprising: the one or more programs comprising instructions for performing the method of any one of claims 67 to 76, Electronic devices.

79. The display and a touch-sensitive surface; and means for carrying out the method of any one of claims 67 to 76; , an electronic device.

80. 1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers, including a first speaker and a second speaker, the one or more programs comprising: generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams including a first audio stream and a second audio stream; Detecting a first user input using the touch-sensitive surface; simultaneously, in response to detecting the first user input; transitioning the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode using the two or more speakers; transitioning the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode using the two or more speakers; displaying on the display a first visual representation of the first audio stream of the audio source; displaying on the display a second visual representation of the second audio stream of the audio source, the first visual representation being different from the second visual representation; A non-transitory computer-readable storage medium containing instructions.

81. The display and a touch-sensitive surface; and one or more processors; a memory storing one or more programs configured to be executed by the one or more processors; 1. An electronic device comprising: the electronic device operatively connected to two or more speakers, including a first speaker and a second speaker; and the one or more programs comprising: generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams including a first audio stream and a second audio stream; Detecting a first user input using the touch-sensitive surface; simultaneously, in response to detecting the first user input; transitioning the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode using the two or more speakers; transitioning the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode using the two or more speakers; displaying on the display a first visual representation of the first audio stream of the audio source; displaying on the display a second visual representation of the second audio stream of the audio source, the first visual representation being different from the second visual representation; An electronic device containing instructions.

82. 1. An electronic device comprising: The display and the electronic device having a touch-sensitive surface operatively connected to two or more speakers, including a first speaker and a second speaker; means for generating audio using an audio source in a first mode using the two or more speakers, the audio source including a plurality of audio streams including a first audio stream and a second audio stream; means for detecting a first user input using the touch-sensitive surface; simultaneously, in response to detecting the first user input; transitioning the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode using the two or more speakers; transitioning the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode using the two or more speakers; displaying on the display a first visual representation of the first audio stream of the audio source; displaying on the display a second visual representation of the second audio stream of the audio source, the first visual representation being different from the second visual representation; Means and , an electronic device.