Audio Spatial Management

The described method enhances spatial audio management on electronic devices by using visual and audio transitions based on user inputs, addressing inefficiencies and power consumption issues in existing techniques.

JP7717923B2Active Publication Date: 2025-08-04APPLE INC
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024135700
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-09-26
Filing Date
2024-08-15
Publication Date
2025-08-04
Estimated Expiration
2039-08-29

AI Technical Summary

Technical Problem

Existing techniques for managing spatial audio on electronic devices are cumbersome, inefficient, and consume excessive time and energy, particularly in battery-operated devices, often requiring complex user interfaces and lacking context recognition.

Method used

A method and interface for managing spatial audio that involves displaying visual elements on a display, generating audio through multiple speakers, and transitioning audio modes based on user inputs, such as touch-and-hold or touch-and-drag actions, to enhance user interaction and conserve power.

Benefits of technology

The method improves user efficiency and reduces cognitive burden while conserving battery life by providing a faster and more intuitive interface for spatial audio management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717923000001
    Figure 0007717923000001
  • Figure 0007717923000002
    Figure 0007717923000002
  • Figure 0007717923000003
    Figure 0007717923000003
Patent Text Reader

Abstract

To provide a faster and more efficient method for managing spatial audio.SOLUTION: An electronic device including a display and a touch sensitive surface includes: being connected to two or more speakers in an operable manner to generate audio by using a single audio source in a first mode; including a plurality of audio streams including first and second audio streams; detecting a first user input by using the touch sensitive surface; transitioning from the first mode to a second mode different from the first mode; transitioning from the first mode to a third mode different from the first mode; displaying a first visual representation through the display and corresponding to a first position corresponding the first audio stream; and displaying a second visual representation of the second audio stream at a second position and simultaneously performing the first visual representation so as to correspond to the second position and to be different from the second visual representation.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 736,990, filed on September 26, 2018, titled "SPATIAL MANAGEMENT OF AUDIO", the entire content of which is incorporated herein by reference.

[0002] This disclosure generally relates to computer user interfaces, and more specifically to techniques for managing spatial audio.

Background Art

[0003] Humans can detect the position of three - dimensional (up - down, front - back, and left - right) sound. Various techniques can be used to modify audio so that a listener perceives the audio as arriving from a specific point in space, where the audio is generated by a device.

Summary of the Invention

[0004] However, some techniques for managing spatial audio using electronic devices are generally cumbersome and inefficient. For example, some techniques do not provide the user with context recognition of the state of the electronic device through spatial management of audio. In another example, some existing techniques use complex and time - consuming user interfaces that may involve multiple key presses or keystrokes. Existing techniques take more time than necessary, wasting the user's time and the device's energy. The latter problem is particularly critical in battery - operated devices. Further, existing audio techniques do not adequately assist the user in navigating a graphical user interface.

[0005] Accordingly, the present technique provides a faster and more efficient method and interface for managing spatial audio to an electronic device. Such a method and interface optionally complement or replace other methods for managing spatial audio. Such a method and interface reduce the cognitive burden on the user and create a more efficient human-machine interface. In the case of a battery-operated computing device, such a method and interface conserve power and increase the battery charging interval.

[0006] According to some embodiments, a method is described that is executed in an electronic device comprising a display, the electronic device being operably connected to two or more speakers. The method includes displaying a first visual element at a first position on the display, accessing a first audio corresponding to the first visual element, generating audio in two or more speakers using the first audio in a first mode while the first visual element is being displayed at the first position on the display, receiving a first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not being displayed on the display in response to receiving the first user input, and generating audio in two or more speakers using the first audio in a second mode different from the first mode while the first visual element is not being displayed on the display, wherein the second mode is configured such that the audio generated in the second mode is perceived by the user as being generated from a direction away from the display.

[0007] According to some embodiments, a non-transitory computer-readable storage medium is described. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device comprising a display, the electronic device being operably connected to two or more speakers, the one or more programs displaying a first visual element at a first position on the display, accessing a first audio corresponding to the first visual element, generating audio on two or more speakers using the first audio in a first mode while the first visual element is being displayed at the first position on the display, receiving a first user input, in response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not being displayed on the display, generating audio on two or more speakers using the first audio in a second mode different from the first mode while the first visual element is not being displayed on the display, the second mode being configured such that the audio generated in the second mode is perceived by the user as being generated from a direction away from the display, including instructions.

[0008] According to some embodiments, a non - persistent computer - readable storage medium is described. A non - persistent computer - readable storage medium that stores one or more programs configured to be executed by one or more processors of an electronic device comprising a display, wherein the electronic device is operably connected to two or more speakers, and the one or more programs cause the display to display a first visual element at a first position on the display, access a first audio corresponding to the first visual element, generate audio on the two or more speakers using the first audio in a first mode while the first visual element is being displayed at the first position on the display, receive a first user input, in response to receiving the first user input, transition the display of the first visual element from the first position on the display to a first visual element not being displayed on the display, generate audio on the two or more speakers using the first audio in a second mode different from the first mode while the first visual element is not being displayed on the display, and the second mode is configured such that the audio generated in the second mode is perceived by the user as being generated from a direction away from the display.

[0009] According to some embodiments, an electronic device is described. The electronic device includes a display, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The electronic device is operably connected to two or more speakers. The one or more programs display a first visual element at a first position on the display, access a first audio corresponding to the first visual element, and while displaying the first visual element at the first position on the display, generate audio at the two or more speakers using the first audio in a first mode, receive a first user input, and in response to receiving the first user input, transition the display of the first visual element from the first position on the display to a first visual element not displayed on the display. While the first visual element is not displayed on the display, generate audio at the two or more speakers using the first audio in a second mode different from the first mode. The second mode is configured such that the audio generated in the second mode is perceived by the user as being generated from a direction away from the display. The instructions include

[0010] According to some embodiments, an electronic device is described. The electronic device includes a display operably connected to two or more speakers, means for displaying a first visual element at a first position on the display, means for accessing a first audio corresponding to the first visual element, means for generating audio in two or more speakers using the first audio in a first mode while the first visual element is being displayed at the first position on the display, means for receiving a first user input, and in response to receiving the first user input, transitioning the display of the first visual element from the first position on the display to a first visual element not being displayed on the display, and while the first visual element is not being displayed on the display, generating audio in two or more speakers using the first audio in a second mode different from the first mode, wherein the second mode is configured such that the audio generated in the second mode is perceived by the user as being generated from a direction away from the display.

[0011] According to some embodiments, a method is described that is executed in an electronic device comprising a display and a touch sensing surface, the electronic device being operably connected to two or more speakers.This method involves displaying a list of multiple media elements on a display, where each of the multiple media elements corresponds to a respective media file, detecting user contact at a position corresponding to a first media element using a touch-sensitive surface, generating audio using two or more speakers without exceeding a predetermined audio playback period according to the detected user contact at the position corresponding to the first media element and according to a user contact including a touch-and-hold input, continuing to generate audio using the first audio file corresponding to the first media element using two or more speakers according to the fact that the user contact remains at the position corresponding to the first media element and does not exceed the predetermined audio playback period, stopping generating audio using the first audio file using two or more speakers according to the fact that it exceeds the predetermined audio playback period, detecting movement of the user contact from the position corresponding to the first media element to the position corresponding to a second media element using the touch-sensitive surface, generating audio using two or more speakers without exceeding a predetermined audio playback period according to the detected user contact at the position corresponding to the second media element and according to a user contact including a touch-and-hold input, continuing to generate audio using the second audio file corresponding to the second media element using two or more speakers according to the fact that the user contact remains at the position corresponding to the second media element and does not exceed the predetermined audio playback period, stopping generating audio using the second audio file using two or more speakers according to the fact that it exceeds the predetermined audio playback period, detecting lift-off of the user contact using the touch-sensitive surface, and stopping generating audio using the first audio file or the second audio file using two or more speakers according to the detected lift-off of the user contact.

[0012] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers, and the one or more programs display a list of a plurality of media elements on the display, each of the plurality of media elements corresponding to a respective media file, detect a user contact at a position corresponding to a first media element using the touch-sensing surface, in response to detecting the user contact at the position corresponding to the first media element and according to the user contact including a touch-and-hold input, generate audio using the two or more speakers without exceeding a predetermined audio playback period using a first audio file corresponding to the first media element, continue to generate audio using the two or more speakers using the first audio file as long as the user contact remains at the position corresponding to the first media element and does not exceed the predetermined audio playback period, stop generating audio using the two or more speakers using the first audio file when the predetermined audio playback period is exceeded, detect a movement of the user contact from the position corresponding to the first media element to the position corresponding to a second media element using the touch-sensing surface, in response to detecting the user contact at the position corresponding to the second media element and according to the user contact including a touch-and-hold input, generate audio using the two or more speakers without exceeding a predetermined audio playback period using a second audio file corresponding to the second media element, continue to generate audio using the two or more speakers using the second audio file as long as the user contact remains at the position corresponding to the second media element and does not exceed the predetermined audio playback period, stop generating audio using the two or more speakers using the second audio file when the predetermined audio playback period is exceeded, detect a lift-off of the user contact using the touch-sensing surface,In response to detecting a lift-off of user contact, it includes an instruction to stop generating audio using two or more speakers and using a first audio file or a second audio file.

[0013] According to some embodiments, a temporary computer-readable storage medium is described. The temporary computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensitive surface, the electronic device being operably connected to two or more speakers, and the one or more programs cause a list of a plurality of media elements to be displayed on the display, each of the plurality of media elements corresponding to a respective media file, detect a user contact at a position corresponding to a first media element using the touch-sensitive surface, in response to detecting the user contact at the position corresponding to the first media element and according to the user contact including a touch-and-hold input, generate audio using the first audio file corresponding to the first media element using two or more speakers without exceeding a predetermined audio playback period, continue to generate audio using the first audio file using two or more speakers while the user contact remains at the position corresponding to the first media element and according to not exceeding the predetermined audio playback period, stop generating audio using the first audio file using two or more speakers according to exceeding the predetermined audio playback period, detect a movement of the user contact from the position corresponding to the first media element to the position corresponding to a second media element using the touch-sensitive surface, in response to detecting the user contact at the position corresponding to the second media element and according to the user contact including a touch-and-hold input, generate audio using the second audio file corresponding to the second media element using two or more speakers without exceeding a predetermined audio playback period, continue to generate audio using the second audio file using two or more speakers while the user contact remains at the position corresponding to the second media element and according to not exceeding the predetermined audio playback period, stop generating audio using the second audio file using two or more speakers according to exceeding the predetermined audio playback period, detect a lift-off of the user contact using the touch-sensitive surface,In response to detecting a lift-off of user contact, stop generating audio using two or more speakers and using a first audio file or a second audio file, including an instruction.

[0014] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The electronic device is operably connected to two or more speakers. The one or more programs are configured to display a list of a plurality of media elements on the display, each media element of the plurality of media elements corresponding to a respective media file, detect a user contact at a position corresponding to a first media element using the touch sensing surface, in response to detecting the user contact at the position corresponding to the first media element and in accordance with the user contact including a touch-and-hold input, generate audio using the two or more speakers without exceeding a predetermined audio playback period using a first audio file corresponding to the first media element, continue to generate audio using the first audio file using the two or more speakers while the user contact remains at the position corresponding to the first media element and without exceeding the predetermined audio playback period, stop generating audio using the first audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period, detect a movement of the user contact from the position corresponding to the first media element to the position corresponding to a second media element using the touch sensing surface, in response to detecting the user contact at the position corresponding to the second media element and in accordance with the user contact including a touch-and-hold input, generate audio using the two or more speakers without exceeding a predetermined audio playback period using a second audio file corresponding to the second media element, continue to generate audio using the second audio file using the two or more speakers while the user contact remains at the position corresponding to the second media element and without exceeding the predetermined audio playback period, stop generating audio using the second audio file using the two or more speakers in accordance with exceeding the predetermined audio playback period, detect a lift-off of the user contact using the touch sensing surface,In response to detecting a lift-off of user contact, stop generating audio using two or more speakers and using a first audio file or a second audio file, including an instruction.

[0015] According to some embodiments, an electronic device is described. The present electronic device includes a display, a touch sensing surface to which the electronic device is operably connected to two or more speakers, means for displaying a list of a plurality of media elements on the display, wherein each media element of the plurality of media elements corresponds to a respective media file, means for detecting a user contact at a position corresponding to a first media element using the touch sensing surface, and in response to detecting a user contact at a position corresponding to the first media element and according to a user contact including a touch-and-hold input, using two or more speakers to generate audio using a first audio file corresponding to the first media element without exceeding a predetermined audio playback period, and continuing to generate audio using the first audio file using two or more speakers according to not exceeding a predetermined audio playback period while the user contact remains at a position corresponding to the first media element, and stopping generating audio using the first audio file using two or more speakers according to exceeding a predetermined audio playback period, means for detecting a movement of a user contact from a position corresponding to the first media element to a position corresponding to a second media element using the touch sensing surface, and in response to detecting a user contact at a position corresponding to the second media element and according to a user contact including a touch-and-hold input, using two or more speakers to generate audio using a second audio file corresponding to the second media element without exceeding a predetermined audio playback period, and continuing to generate audio using the second audio file using two or more speakers according to not exceeding a predetermined audio playback period while the user contact remains at a position corresponding to the second media element, and stopping generating audio using the second audio file using two or more speakers according to exceeding a predetermined audio playback period, means for detecting a lift-off of a user contact using the touch sensing surface, and in response to detecting a lift-off of a user contact, using two or more speakers toMeans for stopping generating audio using the first audio file or the second audio file.

[0016] According to some embodiments, a method is described that is executed in an electronic device comprising a display and a touch-sensitive surface, the electronic device being operatively connected to two or more speakers. The method includes detecting a first user input for activating a discovery mode, and in response to detecting the first user input for activating the discovery mode, using the two or more speakers to generate audio simultaneously using a first audio source of a first mode, the first mode being configured such that the user perceives that the audio generated using the first mode is generated from a first point in a space that moves in a first direction along a predefined path over time at a first speed, a second audio source of a second mode, the second mode being configured such that the user perceives that the audio generated using the second mode is generated from a second point in a space that moves in a first direction along a predefined path over time at a second speed, and a third audio source of a third mode, the third mode being configured such that the user perceives that the audio generated using the third mode is generated from a third point in a space that moves in a first direction along a predefined path over time at a third speed, wherein the first point, the second point, and the third point are different points in the space.

[0017] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers, and the one or more programs detect a first user input for activating a discovery mode, and in response to detecting the first user input for activating the discovery mode, use the two or more speakers to generate, simultaneously, a first audio source of a first mode, wherein the first mode is configured such that the user perceives that the audio generated using the first mode is generated from a first point in a space where the audio moves over time in a first direction along a predetermined path at a first speed, a second audio source of a second mode, wherein the second mode is configured such that the user perceives that the audio generated using the second mode is generated from a second point in a space where the audio moves over time in the first direction along the predetermined path at a second speed, and a third audio source of a third mode, wherein the third mode is configured such that the user recognizes that the audio generated using the third mode is generated from a third point in a space where the audio moves over time in the first direction along the predetermined path at a third speed, and the first point, the second point, and the third point are different points in the space.

[0018] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers, and the one or more programs detect a first user input for activating a discovery mode, and in response to detecting the first user input for activating the discovery mode, use the two or more speakers to output a first audio source in a first mode, where the first mode is configured such that the user perceives that the audio generated using the first mode is generated from a first point in a space where the audio moves in a first direction along a predetermined path over time at a first speed, a second audio source in a second mode, where the second mode is configured such that the user perceives that the audio generated using the second mode is generated from a second point in a space where the audio moves in a first direction along a predetermined path over time at a second speed, and a third audio source in a third mode, where the third mode is configured such that the user recognizes that the audio generated using the third mode is generated from a third point in a space where the audio moves in a first direction along a predetermined path over time at a third speed, and includes instructions to simultaneously generate audio using the first audio source, the second audio source, and the third audio source, and the first point, the second point, and the third point are different points in the space.

[0019] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The electronic device is operably connected to two or more speakers. The one or more programs detect a first user input for activating a discovery mode, and in response to detecting the first user input for activating the discovery mode, use the two or more speakers to output a first audio source in a first mode, a second audio source in a second mode, and a third audio source in a third mode. The first mode is configured such that the user perceives that the audio generated using the first mode is generated from a first point in a space where the audio moves in a first direction along a predetermined path over time at a first speed. The second mode is configured such that the user perceives that the audio generated using the second mode is generated from a second point in a space where the audio moves in a first direction along a predetermined path over time at a second speed. The third mode is configured such that the user recognizes that the audio generated using the third mode is generated from a third point in a space where the audio moves in a first direction along a predetermined path over time at a third speed. The first point, the second point, and the third point are different points in the space.

[0020] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface to which the electronic device is operably connected to two or more speakers, means for detecting a first user input for activating a discovery mode, and in response to detecting the first user input for activating the discovery mode, using the two or more speakers, a first audio source in a first mode, wherein the first mode is configured such that the audio generated using the first mode is perceived by the user as being generated from a first point in a space that moves in a first direction along a predetermined path over time at a first speed, a second audio source in a second mode, wherein the second mode is configured such that the audio generated using the second mode is perceived by the user as being generated from a second point in a space that moves in a first direction along a predetermined path over time at a second speed, a third audio source in a third mode, wherein the third mode is configured such that the audio generated using the third mode is perceived by the user as being generated from a third point in a space that moves in a first direction along a predetermined path over time at a third speed, and means for simultaneously generating audio using, wherein the first point, the second point, and the third point are different points in the space.

[0021] According to some embodiments, a method is described that is executed in an electronic device having a display and a touch sensing surface, the electronic device being operatively connected to two or more speakers. The method includes displaying a user movable affordance at a first position on the display; operating the electronic device in a first state of ambient sound transparency while the user movable affordance is displayed at the first position; using two or more speakers to generate audio using an audio source in a first mode; detecting user input using the touch sensing surface; and in response to detecting the user input, operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency and transitioning the generation of audio using the audio source from the first mode to a second mode different from the first mode according to a set of one or more conditions being satisfied, the set of one or more conditions including a first condition that is satisfied when the user input is a touch-and-drag operation on the user movable affordance; and maintaining the electronic device in the first state of ambient sound transparency and maintaining the generation of audio using the audio source in the first mode according to the set of one or more conditions not being satisfied.

[0022] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device comprising a display and a touch-sensing surface, the electronic device being operatively connected to two or more speakers, and the one or more programs display a user-movable affordance at a first position on the display and, while the user-movable affordance is displayed at the first position, operate the electronic device in a first state of ambient sound transparency, use the two or more speakers to generate audio using an audio source in a first mode, detect user input using the touch-sensing surface, and, in response to detecting the user input, operate the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency and transition the generation of audio using the audio source from the first mode to a second mode different from the first mode according to a set of one or more conditions including a first condition that is satisfied when the user input is a touch-and-drag operation on the user-movable affordance, and maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode according to the set of one or more conditions not being satisfied, comprising instructions.

[0023] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers, and the one or more programs cause the electronic device to display a user-movable affordance at a first position on the display, operate the electronic device in a first state of ambient sound transparency while the user-movable affordance is displayed at the first position, use the two or more speakers to generate audio using an audio source in a first mode, detect user input using the touch-sensing surface, and in response to detecting the user input, operate the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency, transition the generation of audio using the audio source from the first mode to a second mode different from the first mode, according to a set of one or more conditions including a first condition that is satisfied when the user input is a touch-and-drag operation on the user-movable affordance, and maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode according to the set of one or more conditions not being satisfied, including instructions.

[0024] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The electronic device is operably connected to two or more speakers. The one or more programs display a user-movable affordance at a first position on the display, and while the user-movable affordance is displayed at the first position, operate the electronic device in a first state of ambient sound transparency, generate audio using an audio source in a first mode using the two or more speakers, detect user input using the touch sensing surface, and according to one or more conditions being satisfied, including a first condition that is satisfied when the user input is a touch-and-drag operation on the user-movable affordance, operate the electronic device in a second state of ambient sound transparency different from the first state, transition the generation of audio using the audio source from the first mode to a second mode different from the first mode, and according to the one or more conditions not being satisfied, maintain the electronic device in the first state of ambient sound transparency and maintain generating audio using the audio source in the first mode.

[0025] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch-sensing surface to which the electronic device is operably connected to two or more speakers, means for displaying a user-movable affordance at a first position on the display, while the user-movable affordance is being displayed at the first position, operating the electronic device in a first state of ambient sound transparency, using the two or more speakers to generate audio using an audio source in a first mode, and using the touch-sensing surface to detect user input, means for, in response to detecting user input, operating the electronic device in a second state of ambient sound transparency different from the first state of ambient sound transparency and transitioning the generation of audio using the audio source from the first mode to a second mode different from the first mode according to a set of one or more conditions including a first condition being satisfied when the user input is a touch-and-drag operation on the user-movable affordance, and means for maintaining the electronic device in the first state of ambient sound transparency and maintaining the generation of audio using the audio source in the first mode according to the set of one or more conditions not being satisfied.

[0026] According to some embodiments, a method is described for being executed in an electronic device comprising a display and a touch sensing surface, the electronic device being operably connected to two or more speakers including a first speaker and a second speaker. The method includes using the two or more speakers to generate audio using an audio source in a first mode, the audio source including a plurality of audio streams including a first audio stream and a second audio stream; detecting a first user input using the touch sensing surface; in response to detecting the first user input, simultaneously transitioning, using the two or more speakers, the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode; transitioning, using the two or more speakers, the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode; displaying, on the display, a first visual representation of the first audio stream of the audio source; and displaying, on the display, a second visual representation of the second audio stream of the audio source, wherein the first visual representation is different from the second visual representation.

[0027] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device having a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers including a first speaker and a second speaker, and the one or more programs are to use the two or more speakers to generate audio using an audio source in a first mode, the audio source including a plurality of audio streams including a first audio stream and a second audio stream, use the touch-sensing surface to detect a first user input, and in response to detecting the first user input, simultaneously use the two or more speakers to transition the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode, use the two or more speakers to transition the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode, display a first visual representation of the first audio stream of the audio source on the display, and display a second visual representation of the second audio stream of the audio source on the display, the first visual representation being different from the second visual representation.

[0028] According to some embodiments, a non-transitory computer-readable storage medium is described. The non-transitory computer-readable storage medium stores one or more programs configured to be executed by one or more processors of an electronic device comprising a display and a touch-sensing surface, the electronic device being operably connected to two or more speakers including a first speaker and a second speaker, and the one or more programs are to use the two or more speakers to generate audio using an audio source in a first mode, the audio source including a plurality of audio streams including a first audio stream and a second audio stream, detect a first user input using the touch-sensing surface, and in response to detecting the first user input, simultaneously use the two or more speakers to transition the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode, use the two or more speakers to transition the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode, display a first visual representation of the first audio stream of the audio source on the display, and display a second visual representation of the second audio stream of the audio source on the display, the first visual representation being different from the second visual representation.

[0029] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The electronic device is operably connected to two or more speakers including a first speaker and a second speaker. The one or more programs are to use the two or more speakers to generate audio using an audio source in a first mode, where the audio source includes a plurality of audio streams including a first audio stream and a second audio stream, detect a first user input using the touch sensing surface, and in response to detecting the first user input, simultaneously use the two or more speakers to transition the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode, use the two or more speakers to transition the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode, display a first visual representation of the first audio stream of the audio source on the display, and display a second visual representation of the second audio stream of the audio source on the display, where the first visual representation is different from the second visual representation.

[0030] According to some embodiments, an electronic device is described. The electronic device includes a display, a touch sensing surface to which the electronic device is operably connected to two or more speakers including a first speaker and a second speaker, and means for generating audio using the two or more speakers in a first mode using an audio source, the audio source including a plurality of audio streams including a first audio stream and a second audio stream, means for detecting a first user input using the touch sensing surface, and in response to detecting the first user input, simultaneously using the two or more speakers to transition the generation of the first audio stream of the audio source from the first mode to a second mode different from the first mode, and using the two or more speakers to transition the generation of the second audio stream of the audio source from the first mode to a third mode different from the first mode and the second mode, and displaying on the display a first visual representation of the first audio stream of the audio source,

[0031] means for displaying on the display a second visual representation of the second audio stream of the audio source, wherein the first visual representation is different from the second visual representation.

[0032] The executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured to be executed by one or more processors. The executable instructions for performing these functions are optionally included in a temporary computer-readable storage medium or other computer program product configured to be executed by one or more processors.

[0033] Thus, a faster and more efficient method and interface for managing spatial audio are provided to the device, thereby increasing the effectiveness, efficiency, and user satisfaction of such a device. Such a method and interface can complement or replace other methods for managing spatial audio.

Brief Description of the Drawings

[0034] To better understand the various embodiments described, the following "Modes for Carrying Out the Invention" should be referred to in conjunction with the following drawings, and like reference numerals refer to corresponding parts throughout the following figures.

[0035]

Figure 1A

[0036]

Figure 1B

[0037]

Figure 2

[0038]

Figure 3

[0039]

Figure 4A

[0040]

Figure 4B

[0041]

Figure 5A

[0042]

Figure 5B

[0043]

Figure 5C

Figure 5D

[0044]

Figure 5E

Figure 5F

Figure 5G

Figure 5H

[0045]

Figure 6A

Figure 6B

Figure 6C

Figure 6D

Figure 6E

Figure 6F

Figure 6G

Figure 6H

Figure 6I

Figure 6J

Figure 6K

Figure 6L

Figure 6M

Figure 6N

[0046]

Figure 7A

Figure 7B

Figure 7C

[0047]

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 8E

Figure 8F

Figure 8G

Figure 8H

Figure 8I

Figure 8J

Figure 8K

[0048]

Figure 9A

Figure 9B

Figure 9C

[0049]

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 10E

Figure 10F

Figure 10G

Figure 10H

Figure 10I

Figure 10J

Figure 10K

[0050]

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 11E

Figure 11F

Figure 11G

[0051]

Figure 12A

Figure 12B

[0052]

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 13F

[0053]

Figure 13G

Figure 13H

Figure 13I

Figure 13J

Figure 13K

Figure 13L

Figure 13M

[0054]

Figure 14A

Figure 14B

[0055]

Figure 15

[0056] The following description sets forth illustrative methods, parameters, and the like. However, it should be recognized that such description is not intended as a limitation on the scope of the present disclosure, but rather as an illustration of exemplary embodiments.

[0057] There is a need for an efficient method and interface for managing spatial audio in an electronic device. For example, spatial audio can provide a user with context awareness of the state of the electronic device. Such techniques can reduce the cognitive burden on the user of the electronic device, thereby increasing productivity. Further, such techniques can reduce processor and battery power that would otherwise be wasted on redundant user input.

[0058] Hereinafter, FIGS. 1A-1B, 2, 3, 4A-4B, and 5A-5H provide an illustration of an exemplary device for performing a technique for managing event notifications.

[0059] FIGS. 6A-6N illustrate exemplary techniques for transitioning between visual elements, according to some embodiments. FIGS. 7A-7C are flow diagrams illustrating methods for transitioning between visual elements using an electronic device, according to some embodiments. The user interfaces of FIGS. 6A-6N are used to illustrate processes described hereinafter, including the processes of FIGS. 7A-7C.

[0060] FIGS. 8A-8K illustrate exemplary techniques for previewing audio, according to some embodiments. FIGS. 9A-9C are flow diagrams illustrating methods for previewing audio using an electronic device, according to some embodiments. The user interfaces of FIGS. 8A-8K are used to illustrate processes described hereinafter, including the processes of FIGS. 9A-9C.

[0061] FIGS. 10A-10K illustrate exemplary techniques for discovering music, according to some embodiments. FIGS. 11A-11G illustrate exemplary techniques for discovering music, according to some embodiments. FIGS. 12A-12B are flow diagrams illustrating methods for discovering music using an electronic device, according to some embodiments. The user interfaces of FIGS. 10A-10K and FIGS. 11A-11G are used to illustrate processes described hereinafter, including the processes of FIGS. 12A-12B.

[0062] Figures 13A - 13F illustrate exemplary techniques for managing headphone transparency according to some embodiments. Figures 14A - 14B are flow diagrams showing methods for managing headphone transparency using an electronic device according to some embodiments. The user interfaces of Figures 13A - 13F are used to illustrate processes described later, including the process of Figures 14A - 14B.

[0063] Figures 13G - 13M illustrate exemplary techniques for manipulating multiple audio streams of an audio source according to some embodiments. Figure 15 is a flow diagram showing a method for manipulating multiple audio streams of an audio source using an electronic device according to some embodiments. The user interfaces of Figures 13G - 13M are used to illustrate processes described later, including the process of Figure 15.

[0064] In the following description, terms such as "first", "second", etc. are used to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the various embodiments described, the first touch can be called the second touch, and similarly, the second touch can be called the first touch. Both the first touch and the second touch are touches, but they are not the same touch.

[0065] The terms used in the description of the various embodiments described herein are for the purpose of describing particular embodiments only and are not intended to be limiting. In the description of the various embodiments described and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural as well, unless the context clearly dictates otherwise. Also, as used herein, the term “and / or” refers to and includes any and all possible combinations of one or more of the associated listed items. It should be understood that the terms “includes,” “including,” “comprises,” and / or “comprising,” when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0066] The term "if" is optionally interpreted to mean "when" or "upon", or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" are optionally interpreted to mean "upon determining" or "in response to determining", or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]", depending on the context.

[0067] Embodiments of an electronic device, a user interface for such a device, and related processes for using such a device are described. In some embodiments, the device is a portable communication device such as a cellular phone that also includes other functions such as PDA functionality and / or music player functionality. Exemplary embodiments of portable multifunctional devices include, but are not limited to, the iPhone®, iPod Touch®, and iPad® devices from Apple Inc. of Cupertino, California. Optionally, other portable electronic devices such as laptop or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad) are also used. Also, in some embodiments, it should be understood that the device is not a portable communication device but a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0068] In the following discussion, an electronic device including a display and a touch-sensitive surface is described. However, it should be understood that the electronic device optionally includes one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick.

[0069] The device typically supports various applications such as one or more of a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a phone application, a video conferencing application, an email application, an instant messaging application, a training support application, a photo management application, a digital camera application, a digital video camera application, a web browsing application, a digital music player application, and / or a digital video player application.

[0070] Various applications executed on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface, as well as the corresponding information displayed on the device, are optionally adjusted and / or changed for each application and / or within each respective application. Thus, the common physical architecture of the device (such as a touch-sensitive surface) optionally supports various applications with a user interface that is intuitive and transparent to the user.

[0071] Attention is now directed to an embodiment of a portable device having a touch-sensitive display. FIG. 1A is a block diagram showing a portable multifunctional device 100 having a touch-sensitive display system 112 according to some embodiments. The touch-sensitive display 112 may be referred to herein as a "touch screen" for convenience and may be known or referred to as a "touch-sensitive display system." Device 100 includes a memory 102 (optionally including one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, an RF circuit 108, an audio circuit 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more light sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 (such as a touch-sensitive surface such as the touch-sensitive display system 112 of device 100) for detecting the intensity of contact on device 100. Device 100 optionally includes one or more haptic output generators 167 for generating haptic output on device 100 (such as on a touch-sensitive surface such as the touch-sensitive display system 112 of device 100 or the touch pad 355 of device 300). These components optionally communicate via one or more communication buses or signal lines 103.

[0072] As used herein and in the claims, the term "intensity" of a contact on a touch sensing surface refers to the force or pressure (force per unit area) of the contact (e.g., finger contact) on the touch sensing surface, or a proxy for the force or pressure of the contact on the touch sensing surface. The intensity of the contact has a range of values that includes at least four distinct values, and more typically, hundreds (e.g., at least 256) of distinct values. The intensity of the contact is optionally determined (or measured) using a variety of techniques and a variety of sensors or combinations of sensors. For example, one or more force sensors under or adjacent to the touch sensing surface are optionally used to measure the force at various points on the touch sensing surface. In some implementations, force measurements from multiple force sensors are combined (e.g., weighted averaged) to determine the estimated force of the contact. Similarly, a pressure-sensitive tip of a stylus is optionally used to determine the pressure of the stylus on the touch sensing surface. Alternatively, the size and / or change in size of the contact area detected on the touch sensing surface, the capacitance and / or change in capacitance of the touch sensing surface proximate to the contact, and / or the resistance and / or change in resistance of the touch sensing surface proximate to the contact are optionally used as an alternative to the force or pressure of the contact on the touch sensing surface. In some implementations, an alternative measurement of the force or pressure of the contact is used directly to determine whether it exceeds an intensity threshold (e.g., the intensity threshold is described in units corresponding to the alternative measurement). In some implementations, a proxy measurement of the contact force or pressure is converted to an estimated value of the force or pressure, and the estimated value of the force or pressure is used to determine whether it exceeds an intensity threshold (e.g., the intensity threshold is a pressure threshold measured in units of pressure). By using the intensity of the contact as an attribute of user input, in some situations, it enables user access to additional device functionality that would otherwise not be accessible to the user on a device with a reduced-size having a limited area, for the display of affordances (e.g., on a touch sensing display), and / or the reception of user input (e.g., via a touch sensing display, a touch sensing surface, or a physical / mechanical control such as a knob or button).

[0073] As used in this specification and the claims, the term "haptic output" refers to a physical displacement of the device relative to its previous position, a physical displacement of a component of the device (e.g., a touch-sensing surface) relative to another component of the device (e.g., the housing), or a displacement of a component relative to the center of mass of the device, that is to be detected by the user's sense of touch. For example, in a situation where the device or a component of the device is in contact with a touch-sensitive surface of the user (e.g., the finger, palm, or other part of the user's hand), the haptic output generated by the physical displacement will be interpreted by the user as a tactile sensation corresponding to a perceived change in the physical characteristics of the device or component thereof. For example, the movement of a touch-sensing surface (e.g., a touch-sensing display or a trackpad) may optionally be interpreted by the user as a "down click" or "up click" of a physical actuator button. In some cases, the user may feel a tactile sensation such as a "down click" or "up click" even when there is no movement of the physical actuator button associated with the touch-sensing surface that has been physically pushed (e.g., displaced) by the user's action. As another example, the movement of a touch-sensing surface may optionally be interpreted or perceived by the user as the "roughness" of the touch-sensing surface, even if there is no change in the smoothness of the touch-sensing surface. Such interpretation of touch by the user depends on the user's individual sensory perception, but there are many sensory perceptions of touch that are common to a majority of users. Thus, when a haptic output is described as corresponding to a particular sensory perception of the user (e.g., "up click", "down click", "roughness"), unless otherwise stated, the generated haptic output corresponds to a physical displacement of the device or a component of the device that produces the described sensory perception of a typical (or average) user.

[0074] Device 100 is merely an example of a portable multifunctional device. It should be understood that device 100 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of those components. The various components shown in FIG. 1A are implemented in a combination of hardware, software, or both hardware and software, including one or more signal processing circuits and / or application-specific integrated circuits.

[0075] Memory 102 optionally includes high-speed random access memory and also optionally includes non-volatile memory such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0076] Peripheral interface 118 can be used to couple input and output peripheral devices of the device to CPU 120 and memory 102. One or more processors 120 operate or execute various software programs and / or instruction sets stored in memory 102 to perform various functions for device 100 and process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip such as chip 104. In some other embodiments, they are optionally implemented on separate chips.

[0077] The RF (radio frequency) circuit 108 transmits and receives RF signals, also called electromagnetic signals. The RF circuit 108 converts electrical signals into electromagnetic signals or vice versa and communicates with a communication network and other communication devices via electromagnetic signals. The RF circuit 108 optionally includes well-known circuits for performing these functions, such as, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc. The RF circuit 108 optionally communicates wirelessly with networks such as the Internet, also called the World Wide Web (WWW), an intranet, and / or wireless networks such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN), as well as with other devices. The RF circuit 108 optionally includes well-known circuits for detecting a near field communication (NFC) field, such as by a short-range communication radio. Wireless communication optionally includes, but is not limited to only, Global System for Mobile Communications (GSM) for mobile communication, Enhanced Data GSM Environment (EDGE), high-speed downlink packet access (HSDPA), high-speed uplink packet access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPDA), Long Termevolution, LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Bluetooth Low Energy (BTLE), Wireless Fidelity (Wi-Fi) (e.g., IEEE802.11a, IEEE802.11b, IEEE802.11g, IEEE802.11n, and / or IEEE802.11ac), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email (e.g., Internet Message Access Protocol (IMAP) and / or Post Office Protocol (POP)), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol for Instant Messaging and Presence Leveraging Extensions (SIMPLE), Instant Messaging and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other suitable communication protocol including communication protocols not yet developed as of the filing date of this specification, using any one of a plurality of communication standards, protocols, and technologies.

[0078] The audio circuit 110, speaker 111, and microphone 113 provide an audio interface between the user and the device 100. The audio circuit 110 receives audio data from the peripheral device interface 118, converts this audio data into an electrical signal, and transmits this electrical signal to the speaker 111. The speaker 111 converts the electrical signal into human audible sound waves. Also, the audio circuit 110 receives the electrical signal converted from sound waves by the microphone 113. The audio circuit 110 converts the electrical signal into audio data and transmits this audio data to the peripheral device interface 118 for processing. The audio data is optionally obtained from and / or transmitted to the memory 102 and / or the RF circuit 108 by the peripheral device interface 118. In some embodiments, the audio circuit 110 also includes a headset jack (e.g., 212 of FIG. 2). The headset jack provides an interface between the audio circuit 110 and a detachable audio input / output peripheral device such as an output-only headset or a headset with both output (e.g., mono or stereo headphones) and input (e.g., microphone).

[0079] The I / O subsystem 106 couples input / output peripheral devices on the device 100, such as the touch screen 112 and other input control devices 116, to the peripheral device interface 118. The I / O subsystem 106 optionally includes a display controller 156, a light sensor controller 158, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from and transmit electrical signals to the other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, the input controller 160 is optionally coupled to (or not coupled to any of) a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. One or more buttons (e.g., 208 in FIG. 2) optionally include up / down buttons for volume control of the speaker 111 and / or the microphone 113. One or more buttons optionally include push buttons (e.g., 206 in FIG. 2).

[0080] As described in U.S. Patent Application No. 11 / 322,549, filed December 23, 2005, "Unlocking a Device by Performing Gestures on an Unlock Image," and U.S. Patent No. 7,657,849, which are hereby incorporated by reference in their entirety, a quick press of a push button optionally unlocks the touch screen 112 or, optionally, initiates a process of unlocking the device using gestures on the touch screen. A longer press of a push button (e.g., 206) optionally turns the power to the device 100 on or off. The functionality of one or more of the buttons is optionally customizable by the user. The touch screen 112 is used to implement virtual or soft buttons and one or more soft keyboards.

[0081] The touch-sensitive display 112 provides an input interface and an output interface between the device and the user. The display controller 156 receives electrical signals from the touch screen 112 and / or transmits electrical signals to the touch screen 112. The touch screen 112 displays a visual output to the user. This visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0082] The touch screen 112 has a touch sensing surface, sensor, or set of sensors that accepts input from a user based on tactile and / or haptic contact. The touch screen 112 and the display controller 156 detect contact (and any movement or interruption of the contact) on the touch screen 112 (along with any associated modules and / or instruction sets in the memory 102) and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the touch screen 112. In an exemplary embodiment, the point of contact between the touch screen 112 and the user corresponds to the user's finger.

[0083] The touch screen 112 optionally uses LCD (liquid crystal display) technology, LPD (light emitting polymer display) technology, or LED (light emitting diode) technology, although in other embodiments other display technologies are also used. The touch screen 112 and the display controller 156 optionally use any of a plurality of touch sensing technologies, now known or later developed, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch screen 112, to detect contact and any movement or interruption thereof. In an exemplary embodiment, projected mutual capacitance sensing technology, such as that found in the iPhone® and iPod Touch® from Apple Inc. of Cupertino, California, is used.

[0084] The touch-sensing display in some embodiments of touch screen 112 is optionally similar to a multi-touch sensing touch pad described in U.S. Patent Nos. 6,323,846 (Westerman et al.), 6,570,557 (Westerman et al.), and / or 6,677,932 (Westerman), and / or U.S. Patent Publication No. 2002 / 0015024 (A1), each of which is hereby incorporated by reference in its entirety. However, touch screen 112 displays visual output from device 100, whereas the touch-sensing touch pad does not provide visual output.

[0085] The touch-sensing display in some embodiments of touch screen 112 is described in the following applications. (1) U.S. Patent Application No. 11 / 381,313, filed on May 2, 2006, "Multipoint Touch Surface Controller"; (2) U.S. Patent Application No. 10 / 840,862, filed on May 6, 2004, "Multipoint Touchscreen"; (3) U.S. Patent Application No. 10 / 903,964, filed on July 30, 2004, "Gestures For Touch Sensitive Input Devices"; (4) U.S. Patent Application No. 11 / 048,264, filed on January 31, 2005, "Gestures For Touch Sensitive Input Devices"; (5) U.S. Patent Application No. 11 / 038,590, filed on January 18, 2005, "Mode-Based Graphical User Interfaces For Touch Sensitive Input Devices"; (6) U.S. Patent Application No. 11 / 228,758, filed on September 16, 2005, "Virtual Input Device Placement On A Touch Screen User Interface"; (7) U.S. Patent Application No. 11 / 228,700, filed on September 16, 2005, "Operation Of A Computer With A Touch Screen Interface"; (8) U.S. Patent Application No. 11 / 228,737, filed on September 16, 2005, "Activating Virtual Keys Of A Touch-Screen Virtual Keyboard"; and (9) U.S. Patent Application No. 11 / 367,749, filed on March 3, 2006, "Multi-Functional Hand-Held Device". All of these applications are hereby incorporated by reference in their entirety into this specification.

[0086] The touch screen 112 optionally has a video resolution exceeding 100 dpi. In some embodiments, the touch screen has a video resolution of approximately 160 dpi. The user optionally touches the touch screen 112 using any suitable object or appendage such as a stylus, finger, etc. In some embodiments, the user interface is designed to operate primarily using finger-based contact and gestures, although this may not be as accurate as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the device converts rough input by a finger into an accurate pointer / cursor position or command for performing the action desired by the user.

[0087] In some embodiments, in addition to the touch screen, the device 100 optionally includes a touch pad (not shown) for activating or deactivating certain functions. In some embodiments, the touch pad, unlike the touch screen, is a touch-sensitive area of the device that does not display a visual output. The touch pad is optionally a touch-sensitive surface separate from the touch screen 112 or an extension of the touch-sensitive surface formed by the touch screen.

[0088] The device 100 also includes a power system 162 that supplies power to various components. The power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharge system, a power outage detection circuit, a power converter or inverter, a power status indicator (e.g., light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power within a portable device.

[0089] Device 100 also optionally includes one or more optical sensors 164. FIG. 1A shows an optical sensor coupled to an optical sensor controller 158 within I / O subsystem 106. The optical sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. The optical sensor 164 receives light from the environment projected through one or more lenses and converts that light into data representing an image. The optical sensor 164, in cooperation with imaging module 143 (also referred to as a camera module), optionally captures a still image or a video. In some embodiments, the optical sensor is located on the back surface of device 100 opposite touch screen display 112 on the front of the device, and thus the touch screen display can be used as a viewfinder for acquiring still images and / or videos. In some embodiments, the optical sensor is disposed on the front of the device such that an image of the user is optionally obtained for a video conference while the user is viewing other video conference participants on the touch screen display. In some embodiments, the position of the optical sensor 164 can be changed by the user (e.g., by rotating the lens and sensor within the device housing), and thus a single optical sensor 164 can be used for both video conferencing and for acquiring still images and / or videos together with the touch screen display.

[0090] Device 100 also optionally includes one or more contact intensity sensors 165. FIG. 1A shows a contact intensity sensor coupled to an intensity sensor controller 159 within I / O subsystem 106. The contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electrokinetic force sensors, piezoelectric force sensors, optical force sensors, capacitive touch sensing surfaces, or other intensity sensors (e.g., sensors used to measure the force (or pressure) of contact on a touch sensing surface). The contact intensity sensor 165 receives contact intensity information (e.g., pressure information, or a proxy for pressure information) from the environment. In some embodiments, at least one contact intensity sensor is juxtaposed with, or proximate to, a touch sensing surface (e.g., touch sensing display system 112). In some embodiments, at least one contact intensity sensor is disposed on the back of device 100, opposite a touch screen display 112 disposed on the front of device 100.

[0091] Device 100 also optionally includes one or more proximity sensors 166. FIG. 1A shows a proximity sensor 166 coupled to the peripheral device interface 118. Alternatively, the proximity sensor 166 is optionally coupled to an input controller 160 within the I / O subsystem 106. The proximity sensor 166 functions as described, for example, in U.S. Patent Application Nos. 11 / 241,839, "Proximity Detector In Handheld Device", 11 / 240,788, "Proximity Detector In Handheld Device", 11 / 620,702, "Using Ambient Light Sensor To Augment Proximity Sensor Output", 11 / 586,862, "Automated Response To And Sensing Of User Activity In Portable Devices", and 11 / 638,251, "Methods And Systems For Automatic Configuration Of Peripherals", which are hereby incorporated by reference in their entirety. In some embodiments, when a multifunctional device is placed near the user's ear (e.g., when the user is making a phone call), the proximity sensor turns off and disables the touch screen 112.

[0092] Device 100 also optionally includes one or more haptic output generators 167. FIG. 1A shows a haptic output generator coupled to a haptic feedback controller 161 within I / O subsystem 106. The haptic output generator 167 optionally includes one or more electroacoustic devices such as speakers or other audio components, and / or electromechanical devices that convert energy into linear movement, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output generating components (e.g., components that convert an electrical signal into a haptic output) on the device. The contact intensity sensor 165 receives haptic feedback generation instructions from the haptic feedback module 133 and generates a haptic output on device 100 that can be sensed by a user of device 100. In some embodiments, at least one haptic output generator is juxtaposed with or proximate to a touch sensing surface (e.g., touch sensing display system 112) and optionally generates a haptic output by moving the touch sensing surface in a vertical direction (e.g., in / out of the surface of device 100) or a horizontal direction (e.g., back and forth within the same plane as the surface of device 100). In some embodiments, at least one haptic output generator sensor is disposed on the back of device 100, which is opposite the touch screen display 112 disposed on the front of device 100.

[0093] Additionally, device 100 optionally includes one or more accelerometers 168. FIG. 1A shows an accelerometer 168 coupled to the peripheral device interface 118. Alternatively, the accelerometer 168 is optionally coupled to the input controller 160 within the I / O subsystem 106. The accelerometer 168 functions optionally as described in both U.S. Patent Publication No. 20050190059, "Acceleration-based Theft Detection System for Portable Electronic Devices", and U.S. Patent Publication No. 20060017692, "Methods And Apparatuses For Operating A Portable Device Based On An Accelerometer", which are hereby incorporated by reference in their entireties. In some embodiments, information is displayed on the touch screen display in a portrait or landscape orientation based on an analysis of data received from one or more accelerometers. Device 100 optionally includes, in addition to the accelerometer(s) 168, a magnetometer (not shown), and a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information regarding the position and orientation of device 100 (e.g., portrait or landscape orientation).

[0094] In some embodiments, the software components stored in memory 102 include an operating system 126, a communication module (or instruction set) 128, a touch / motion module (or instruction set) 130, a graphics module (or instruction set) 132, a text input module (or instruction set) 134, a Global Positioning System (GPS) module (or instruction set) 135, and an application (or instruction set) 136. Further, in some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) stores a device / global internal state 157 as shown in FIGS. 1A and 3. The device / global internal state 157 includes an active application state indicating which application is active if there is a currently active application, a display state indicating which application, view, or other information occupies various regions of the touch screen display 112, a sensor state including information obtained from various sensors and input control devices 116 of the device, and one or more of position information regarding the position and / or orientation of the device.

[0095] The operating system 126 (e.g., an embedded operating system such as Darwin, RTXC, LINUX, UNIX, OS X, iOS, WINDOWS, or VxWorks) includes various software components and / or drivers that control and manage general system tasks (e.g., memory management, storage device control, power management, etc.) and facilitate communication between various hardware components and software components.

[0096] The communication module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by the RF circuit 108 and / or the external port 124. The external port 124 (e.g., Universal Serial Bus (USB), FIREWIRE, etc.) is adapted to couple to other devices either directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.). In some embodiments, the external port is a multi-pin (e.g., 30-pin) connector that is the same as or similar to and / or is adapted to the 30-pin connector used on iPod (registered trademark) devices (a trademark of Apple Inc.).

[0097] The contact / motion module 130 optionally detects contact with the touch screen 112 and other touch-sensing devices (e.g., a touch pad or a physical click wheel) (in cooperation with the display controller 156). The contact / motion module 130 includes various software components for performing various operations related to the detection of contact, such as determining whether contact has occurred (e.g., detecting a finger-down event), determining the intensity of the contact (e.g., the force or pressure of the contact, or an alternative to the force or pressure of the contact), determining whether there is movement of the contact, tracking movement across the touch-sensing surface (e.g., detecting one or more events of dragging a finger), and determining whether the contact has stopped (e.g., detecting a finger-up event or an interruption of the contact). The contact / motion module 130 receives contact data from the touch-sensing surface. Determining the movement of the contact point, represented by a series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. These operations are optionally applied to a single contact (e.g., the contact of one finger) or multiple simultaneous contacts (e.g., "multi-touch" / contact of multiple fingers). In some embodiments, the contact / motion module 130 and the display controller 156 detect contact on the touch pad.

[0098] In some embodiments, the contact / motion module 130 uses a set of one or more intensity thresholds to determine whether an action has been performed by a user (e.g., to determine whether a user has "clicked" on an icon). In some embodiments, at least one subset of the intensity thresholds is determined according to software parameters (e.g., the intensity thresholds can be adjusted without changing the physical hardware of the device 100, rather than being determined by the activation threshold of a particular physical actuator). For example, the mouse "click" threshold for a trackpad or touch screen display can be set to any of a wide range of default thresholds without changing the trackpad or touch screen display hardware. Additionally, in some implementations, the user of the device is provided with software settings to adjust one or more of the set of intensity thresholds (e.g., by adjusting individual intensity thresholds and / or by adjusting multiple intensity thresholds at once according to system-level click "intensity" parameters).

[0099] The contact / motion module 130 optionally detects gesture inputs by the user. Different gestures on the touch-sensing surface have different contact patterns (e.g., the detected movement, timing, and / or intensity of the contact is different). Thus, gestures are optionally detected by detecting a particular contact pattern. For example, detecting a finger tap gesture includes detecting a finger down event followed by detecting a finger up (lift off) event at the same position (or substantially the same position) as the finger down event (e.g., the position of an icon). As another example, detecting a finger swipe gesture on the touch-sensing surface includes detecting a finger down event followed by detecting one or more finger drag events and then followed by detecting a finger up (lift off) event.

[0100] The graphic module 132 includes various known software components for rendering and displaying graphics on the touch screen 112 or other display, including components that vary the visual impact of the displayed graphics (e.g., brightness, transparency, saturation, contrast, or other visual characteristics). As used herein, the term "graphic" includes, but is not limited to, any object that can be displayed to a user, including text, web pages, icons (such as user interface objects including soft keys), digital images, videos, animations, and the like.

[0101] In some embodiments, the graphic module 132 stores data representing the graphics that will be used. Each graphic is optionally assigned a corresponding code. The graphic module 132 receives from an application or the like, as needed, one or more codes that specify the graphic to be displayed, along with coordinate data and other graphic characteristic data, and then generates the image data for the screen to be output to the display controller 156.

[0102] The haptic feedback module 133 includes various software components for generating the instructions used by the haptic output generator 167, which generates haptic output at one or more locations on the device 100 in response to interaction of the user with the device 100.

[0103] The text input module 134 is optionally a component of the graphic module 132 and provides a soft keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application that requires text input).

[0104] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for location-based dialing, to the camera 143 as photo / video metadata, and to applications that provide location-based services such as weather widgets, local business phone directories, and map / navigation widgets).

[0105] The application 136 optionally includes the following modules (or sets of instructions) or subsets or supersets thereof. ● Contact module 137 (also sometimes called an address book or contact list), ● Phone module 138, ● Video conferencing module 139, ● Email client module 140, ● Instant messaging (IM) module 141, ● Training support module 142, ● Camera module 143 for still images and / or videos, ● Image management module 144, ● Video player module, ● Music player module, ● Browser module 147, ● Calendar module 148, ● Optionally, a widget module 149 that includes one or more of weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, dictionary widget 149-5, and other widgets obtained by the user, as well as user-created widget 149-6, ● Widget creator module 150 for creating user-created widget 149-6, ● Search module 151, ● Video and music player module 152 that integrates the video player module and the music player module ● Memo module 153, ● Map module 154, and / or, ● Online video module 155.

[0106] Examples of other applications 136 optionally stored in the memory 102 include other word processing applications, other image editing applications, drawing applications, presentation applications, JAVA-compatible applications, encryption, digital rights management, voice recognition, and voice replication.

[0107] The contact module 137 is used in cooperation with the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134 to optionally manage an address book or a contact list (e.g., stored in the internal application state 192 of the contact module 137 in the memory 102 or the memory 370). The management by the contact module 137 includes Adding a name to the address book, deleting a name (singular or plural) from the address book, associating a phone number (singular or plural), an email address (singular or plural), a physical address (singular or plural), or other information with a name, associating an image with a name, classifying and sorting names, providing a phone number or an email address to initiate and / or facilitate communication by the phone 138, the video conferencing module 139, the email 140, or the IM 141, etc.

[0108] The telephone module 138 is used in cooperation with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134 to optionally input a character sequence corresponding to a telephone number, access one or more telephone numbers in the contact module 137, modify the input telephone number, dial each telephone number, execute a call, and perform connection disconnection and call hold at the end of the call. As described above, the wireless communication optionally uses any of a plurality of communication standards, protocols, and technologies.

[0109] The video conferencing module 139 includes executable instructions for starting, executing, and ending a video conference between the user and one or more other participants according to the user's instructions in cooperation with the RF circuit 108, the audio circuit 110, the speaker 111, the microphone 113, the touch screen 112, the display controller 156, the optical sensor 164, the optical sensor controller 158, the contact / motion module 130, the graphic module 132, the text input module 134, the contact module 137, and the telephone module 138.

[0110] The email client module 140 includes executable instructions for creating, sending, receiving, and managing emails according to the user's instructions in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134. In cooperation with the image management module 144, the email client module 140 makes it very easy to create and send emails with still or moving images captured by the camera module 143.

[0111] The instant messaging module 141, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134, includes executable instructions for inputting a character sequence corresponding to an instant message, modifying a previously input character, (e.g., using the Short Message Service (SMS) or Multimedia Message Service (MMS) protocol for instant messages based on telephone communication, or XMPP, SIMPLE, or IMPS for instant messages based on the Internet) transmitting each instant message, receiving an instant message, and viewing the received instant message. In some embodiments, the instant messages transmitted and / or received optionally include graphics, photos, audio files, video files, and / or other attachment files as supported by MMS and / or Enhanced Messaging Service (EMS). As used herein, "instant messaging" refers to both telephone communication-based messages (e.g., messages transmitted using SMS or MMS) and Internet-based messages (e.g., messages transmitted using XMPP, SIMPLE, or IMPS).

[0112] The training support module 142 cooperates with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, GPS module 135, map module 154, and music player module to create training (e.g., having time, distance, and / or calorie burn goals), communicate with a training sensor (sports device), receive training sensor data, calibrate sensors used to monitor the training, select and play music for the training, and includes executable instructions for displaying, storing, and transmitting training data.

[0113] The camera module 143 cooperates with the touch screen 112, display controller 156, light sensor 164, light sensor controller 158, contact / motion module 130, graphic module 132, and image management module 144 to include executable instructions for capturing still images or videos (including video streams) and storing them in the memory 102, modifying the characteristics of still images or videos, or deleting still images or videos from the memory 102.

[0114] The image management module 144 cooperates with the touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, and camera module 143 to include executable instructions for arranging, modifying (e.g., editing), or other operations, labeling, deleting, presenting (e.g., in a digital slide show or album), and storing still images and / or videos.

[0115] The browser module 147 includes executable instructions for browsing the Internet according to user instructions, including searching for, linking to, receiving, and displaying a web page or a portion thereof, as well as attached files and other files linked to the web page, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, and the text input module 134.

[0116] The calendar module 148 includes executable instructions for creating, displaying, modifying, and storing a calendar and data associated with the calendar (e.g., calendar items, to-do lists, etc.) according to user instructions, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the contact / motion module 130, the graphic module 132, the text input module 134, the email client module 140, and the browser module 147.

[0117] The widget module 149 cooperates with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, and browser module 147, and optionally, mini-applications (e.g., weather widget 149-1, stock price widget 149-2, calculator widget 149-3, alarm clock widget 149-4, and dictionary widget 149-5) that are downloaded and used by the user, or mini-applications created by the user (e.g., user-created widget 149-6). In some embodiments, the widget includes an HTML (Hypertext Markup Language) file, a CSS (Cascading Style Sheets) file, and a JavaScript file. In some embodiments, the widget includes an XML (Extensible Markup Language) file and a JavaScript file (e.g., Yahoo! widget).

[0118] The widget creator module 150 cooperates with the RF circuit 108, touch screen 112, display controller 156, contact / motion module 130, graphic module 132, text input module 134, and browser module 147, and is used by the user to optionally create a widget (e.g., make a user-specified portion of a web page into a widget).

[0119] The search module 151 cooperates with the touch screen 112, display controller 156, contact / motion module 130, graphic module 132, and text input module 134, and includes executable instructions for searching for characters, music, sound, images, videos, and / or other files in the memory 102 that match one or more search criteria (e.g., one or more user-specified search terms) according to the user's instructions.

[0120] The video and music player module 152, in cooperation with the touch screen 112, the display controller 156, the touch / motion module 130, the graphic module 132, the audio circuit 110, the speaker 111, the RF circuit 108, and the browser module 147, includes executable instructions that enable a user to download and play recorded music and other sound files stored in one or more file formats such as MP3 or AAC files, and executable instructions for displaying, presenting, or otherwise playing videos (e.g., on the touch screen 112 or on an external display connected via the external port 124). In some embodiments, the device 100 optionally includes the functionality of an MP3 player such as an iPod (a trademark of Apple Inc.).

[0121] The memo module 153, in cooperation with the touch screen 112, the display controller 156, the touch / motion module 130, the graphic module 132, and the text input module 134, includes executable instructions for creating and managing memos, to-do lists, etc. according to the user's instructions.

[0122] The map module 154, in cooperation with the RF circuit 108, the touch screen 112, the display controller 156, the touch / motion module 130, the graphic module 132, the text input module 134, the GPS module 135, and the browser module 147, is optionally used to receive, display, modify, and store maps and map-related data (e.g., driving routes, data on stores and other points of interest near a particular location or its vicinity, and other location-based data) according to the user's instructions.

[0123] The online video module 155 cooperates with the touch screen 112, the display controller 156, the touch / motion module 130, the graphic module 132, the audio circuit 110, the speaker 111, the RF circuit 108, the text input module 134, the email client module 140, and the browser module 147 to enable a user to access a particular online video, browse a particular online video, receive it (e.g., by streaming and / or downloading), play it (e.g., on the touch screen or on an external display connected via the external port 124), send an email having a link to a particular online video, and perform other management of online videos in one or more file formats such as H.264. In some embodiments, instead of the email client module 140, the instant messaging module 141 is used to send a link to a particular online video. For additional explanation of the online video application, see U.S. Provisional Patent Application No. 60 / 936,562, filed Jun. 20, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", and U.S. Patent Application No. 11 / 968,067, filed Dec. 31, 2007, "Portable Multifunction Device, Method, and Graphical User Interface for Playing Online Videos", the entire contents of which are hereby incorporated by reference.

[0124] The modules and applications identified above each correspond to a set of executable instructions that perform one or more of the functions described above and the methods described in this application (e.g., the computer-implemented methods and other information processing methods described herein). These modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus, in various embodiments, various subsets of these modules may optionally be combined or otherwise reconfigured. For example, a video player module may optionally be combined with a music player module to form a single module (e.g., the video and music player module 152 of FIG. 1A). In some embodiments, the memory 102 optionally stores a subset of the modules and data structures identified above. Further, the memory 102 optionally stores additional modules and data structures not described above.

[0125] In some embodiments, the device 100 is a device in which the operation of a set of default functions in the device is performed only via a touch screen and / or a touch pad. By using the touch screen and / or the touch pad as the main input control device for the device 100 to operate, optionally, the number of physical input control devices (push buttons, dials, etc.) on the device 100 is reduced.

[0126] The set of default functions performed only through the touch screen and / or the touch pad optionally includes navigation between user interfaces. In some embodiments, when the touch pad is touched by the user, the device 100 is navigated from any user interface displayed on the device 100 to the main menu, the home menu, or the root menu. In such embodiments, the "menu button" is implemented using the touch pad. In some other embodiments, the menu button is a physical push button or other physical input control device rather than a touch pad.

[0127] FIG. 1B is a block diagram showing exemplary components for event processing according to some embodiments. In some embodiments, memory 102 (FIG. 1A) or 370 (FIG. 3) includes an event sorter 170 (e.g., within operating system 126) and respective applications 136-1 (e.g., any of the aforementioned applications 137-151, 155, 380-390).

[0128] The event sorter 170 receives event information and determines an application 136-1 to which the event information is to be delivered and an application view 191 of the application 136-1. The event sorter 170 includes an event monitor 171 and an event dispatcher module 174. In some embodiments, the application 136-1 includes an application internal state 192 that represents the current application view displayed on the touch-sensitive display 112 when the application is active or running. In some embodiments, the device / global internal state 157 is used by the event sorter 170 to determine which application(s) is / are currently active, and the application internal state 192 is used by the event sorter 170 to determine the application view 191 to which the event information is to be delivered.

[0129] In some embodiments, the application internal state 192 includes additional information such as resume information to be used if the application 136-1 resumes execution, user interface state information indicating or ready to display information being displayed by the application 136-1, a state queue that enables the user to return to a previous state or view of the application 136-1, and one or more of a redo / undo queue of previous actions performed by the user.

[0130] The event monitor 171 receives event information from the peripheral device interface 118. The event information includes information regarding sub-events (e.g., a user touch as part of a multi-touch gesture on the touch-sensitive display 112). The peripheral device interface 118 transmits information received from the I / O subsystem 106, or sensors such as the proximity sensor 166, the accelerometer(s) 168, and / or the microphone 113 (via the audio circuit 110). The information that the peripheral device interface 118 receives from the I / O subsystem 106 includes information from the touch-sensitive display 112 or a touch-sensitive surface.

[0131] In some embodiments, the event monitor 171 transmits requests to the peripheral device interface 118 at predetermined intervals. In response, the peripheral device interface 118 transmits event information. In other embodiments, the peripheral device interface 118 transmits event information only when there is an important event (e.g., receipt of an input that exceeds a predetermined noise threshold and / or exceeds a predetermined duration).

[0132] In some embodiments, the event sorter 170 also includes a hit view determination module 172 and / or an active event recognition unit determination module 173.

[0133] The hit view determination module 172 provides a software procedure for determining where in one or more views a sub-event occurs when the touch-sensitive display 112 is displaying two or more views. A view is composed of control devices and other elements that a user can view on the display.

[0134] Another aspect of the user interface associated with an application is a set of views, sometimes referred to herein as application views or user interface windows, in which information is displayed and touch-based gestures occur. The application view (for each application) where a touch is detected optionally corresponds to a program level within the program hierarchy or view hierarchy of the application. For example, the lowest level view where a touch is detected is optionally referred to as the hit view, and the set of events recognized as appropriate input is optionally determined based at least in part on the hit view of the initial touch that starts the touch-based gesture.

[0135] The hit view determination module 172 receives information related to sub-events of a touch-based gesture. When an application has a plurality of hierarchically structured views, the hit view determination module 172 identifies the hit view as the lowest level view within the hierarchy where the sub-event is to be processed. In most situations, the hit view is the lowest level view where the start sub-event (e.g., the first sub-event in a sub-event sequence that forms an event or potential event) occurs. Once the hit view is identified by the hit view determination module 172, the hit view typically receives all sub-events related to the same touch or input source as the touch or input source identified as the hit view.

[0136] The active event recognition unit determination module 173 determines which view(s) within the view hierarchy should receive a particular sequence of sub-events. In some embodiments, the active event recognition unit determination module 173 determines that only the hit view should receive a particular sequence of sub-events. In other embodiments, the active event recognition unit determination module 173 determines that all views including the physical location of the sub-event are views that are actively involved, and thus determines that all views that are actively involved should receive a particular sequence of sub-events. In other embodiments, even if a touch sub-event is completely limited to an area associated with one particular view, the upper-level views within the hierarchy will still continue to be views that are actively involved.

[0137] The event dispatcher module 174 dispatches event information to an event recognition unit (e.g., event recognition unit 180). In embodiments including the active event recognition unit determination module 173, the event dispatcher module 174 dispatches event information to the event recognition unit determined by the active event recognition unit determination module 173. In some embodiments, the event dispatcher module 174 stores the event information obtained by each event receiver 182 in an event queue.

[0138] In some embodiments, the operating system 126 includes an event sorter 170. Alternatively, the application 136-1 includes an event sorter 170. In still other embodiments, the event sorter 170 is a stand-alone module or a part of another module stored in the memory 102 such as the touch / motion module 130.

[0139] In some embodiments, application 136-1 includes a plurality of event processing units 190 and one or more application views 191, each including instructions for processing touch events that occur within respective views of the application's user interface. Each application view 191 of application 136-1 includes one or more event recognition units 180. Typically, each application view 191 includes a plurality of event recognition units 180. In other embodiments, one or more of the event recognition units 180 are part of a separate module, such as a user interface kit (not shown) or a higher-level object from which application 136-1 inherits methods and other characteristics. In some embodiments, each event processing unit 190 includes one or more of event data 179 received from data update unit 176, object update unit 177, GUI update unit 178, and / or event sorter 170. The event processing unit 190 optionally utilizes or calls the data update unit 176, object update unit 177, or GUI update unit 178 to update the internal state 192 of the application. Alternatively, one or more of the application views 191 include one or more respective event processing units 190. Also, in some embodiments, one or more of the data update unit 176, object update unit 177, and GUI update unit 178 are included in respective application views 191.

[0140] Each event recognition unit 180 receives event information (e.g., event data 179) from event sorter 170 and identifies an event from the event information. The event recognition unit 180 includes an event receiving unit 182 and an event comparing unit 184. In some embodiments, the event recognition unit 180 also includes at least a subset of metadata 183 and event distribution instructions 188 (optionally including sub-event distribution instructions).

[0141] The event receiving unit 182 receives event information from the event sorter 170. The event information includes sub-events, for example, information about a touch or a movement of a touch. Depending on the sub-event, the event information also includes additional information such as the position of the sub-event. When the sub-event is related to the movement of a touch, the event information also optionally includes the speed and direction of the sub-event. In some embodiments, the event includes a rotation of the device from one orientation to another (for example, from portrait to landscape, or vice versa), and the event information includes corresponding information about the current orientation of the device (also referred to as the posture of the device).

[0142] The event comparison unit 184 compares the event information with the definition of a defined event or sub-event, and based on the comparison, determines an event or sub-event, or determines or updates the state of an event or sub-event. In some embodiments, the event comparison unit 184 includes an event definition 186. The event definition 186 includes definitions of events (for example, a sequence of predefined sub-events) such as event 1 (187-1) and event 2 (187-2). In some embodiments, the sub-events within an event (187) include, for example, a touch start, a touch end, a touch movement, a touch cancellation, and multiple touches. In one example, the definition of event 1 (187-1) is a double-tap on a displayed object. The double-tap includes, for example, a first touch (touch start) on the displayed object for a predetermined stage, a first lift-off (touch end) for the predetermined stage, a second touch (touch start) on the displayed object for the predetermined stage, and a second lift-off (touch end) for the predetermined stage. In another example, the definition of event 2 (187-2) is a drag on a displayed object. The drag includes, for example, a touch (or contact) on the displayed object for a predetermined stage, a movement of the touch across the touch-sensitive display 112, and a lift-off of the touch (touch end). In some embodiments, the event also includes information about one or more associated event processing units 190.

[0143] In some embodiments, event definition 187 includes the definition of events for each user interface object. In some embodiments, event comparison unit 184 performs a hit test to determine which user interface object is associated with the sub - event. For example, within an application view where three user interface objects are displayed on touch - sensitive display 112, when a touch is detected on touch - sensitive display 112, event comparison unit 184 performs a hit test to determine which of the three user interface objects is associated with the touch (sub - event). If each of the displayed objects is associated with a respective event processing unit 190, the event comparison unit determines which event processing unit 190 should be activated using the result of the hit test. For example, event comparison unit 184 selects the event processing unit associated with the sub - event and object that triggered the hit test.

[0144] In some embodiments, the definition of each event 187 also includes a delay action that delays the transmission of event information until it is determined whether the sequence of sub - events corresponds to the event type of the event recognition unit.

[0145] If each event recognition unit 180 determines that a series of sub - events does not match any of the events of event definition 186, each event recognition unit 180 enters a state of event impossible, event failure, or event end, and then ignores the next sub - event of the touch - based gesture. In this situation, if there are other event recognition units that remain active for the hit view, that event recognition unit continues to track and process the sub - events of the ongoing touch - based gesture.

[0146] In some embodiments, each event recognition unit 180 includes metadata 183 having configurable properties, flags, and / or lists indicating how the event distribution system should actively participate in the event recognition unit that should execute sub-event distribution. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists indicating how the event recognition units interact with each other or how they can interact with each other. In some embodiments, the metadata 183 includes configurable properties, flags, and / or lists indicating whether sub-events are distributed at various levels in the view hierarchy or program hierarchy.

[0147] In some embodiments, each event recognition unit 180 activates the event processing unit 190 associated with the event when one or more specific sub-events of the event are recognized. In some embodiments, each event recognition unit 180 distributes event information associated with the event to the event processing unit 190. Activating the event processing unit 190 is separate from sending (and deferring sending) sub-events to each hit view. In some embodiments, the event recognition unit 180 sets a flag associated with the recognized event, and the event processing unit 190 associated with the flag catches the flag and executes a predefined process.

[0148] In some embodiments, the event distribution command 188 includes a sub-event distribution command that distributes event information about sub-events without activating the event processing unit. Instead, the sub-event distribution command distributes the event information to the event processing unit associated with a series of sub-events or to the view actively involved. The event processing unit associated with a series of sub-events or the view actively involved receives the event information and executes a predetermined process.

[0149] In some embodiments, the data update unit 176 creates and updates data used in the application 136-1. For example, the data update unit 176 updates the phone numbers used in the contact module 137 or stores video files used in the video player module. In some embodiments, the object update unit 177 creates and updates objects used in the application 136-1. For example, the object update unit 177 creates a new user interface object or updates the position of a user interface object. The GUI update unit 178 updates the GUI. For example, the GUI update unit 178 prepares display information and sends the display information to the graphic module 132 for display on the touch-sensitive display.

[0150] In some embodiments, the event processing unit(s) 190 includes or has access to the data update unit 176, the object update unit 177, and the GUI update unit 178. In some embodiments, the data update unit 176, the object update unit 177, and the GUI update unit 178 are included in a single module of their respective application 136-1 or application view 191. In other embodiments, they are included in two or more software modules.

[0151] The foregoing description regarding event processing of a user's touch on the touch-sensitive display also applies to other forms of user input for operating the multifunctional device 100 using an input device, but it should be understood that not all of them are initiated on the touch screen. For example, movement of a mouse and pressing of a mouse button, movement of contact such as tapping, dragging, and scrolling on a touch pad, pen stylus input, movement of the device, verbal commands, detected eye movement, biometric input, and / or any combination thereof, optionally in association with a single or multiple presses or holding of a keyboard, are used as input corresponding to sub-events that define events to be optionally recognized.

[0152] FIG. 2 shows a portable multifunctional device 100 having a touch screen 112, according to some embodiments. The touch screen optionally displays one or more graphics within a user interface (UI) 200. In this embodiment, as well as in other embodiments described below, the user can select one or more of those graphics by performing gestures on the graphics using, for example, one or more fingers 202 (not drawn to scale in the figure) or one or more styluses 203 (not drawn to scale in the figure). In some embodiments, the selection of one or more graphics is performed when the user interrupts contact with the one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (from left to right, from right to left, upward and / or downward), and / or rolling (from right to left, from left to right, upward and / or downward) of a finger in contact with the device 100. In some implementations or situations, an accidental contact with a graphic does not select the graphic. For example, if the gesture corresponding to the selection is a tap, a swipe gesture across an application icon does not optionally select the corresponding application.

[0153] The device 100 also optionally includes one or more physical buttons, such as a "home" button or a menu button 204. As described above, the menu button 204 is optionally used to navigate to any application 136 within a set of applications optionally running on the device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a GUI displayed on the touch screen 112.

[0154] In some embodiments, device 100 includes a touch screen 112, a menu button 204, a push button 206 for turning the device on / off and locking the device, volume adjustment buttons 208, a subscriber identity module (SIM) card slot 210, a headset jack 212, and a docking / charging external port 124. The push button 206 is optionally used to turn the device on / off by pressing the button and holding it in the depressed state for a predetermined period, to lock the device by pressing the button and releasing it before a predetermined time has elapsed, and / or to unlock the device or initiate an unlock process. In an alternative embodiment, device 100 also accepts verbal input via a microphone 113 to activate or deactivate some functions. Device 100 optionally also includes one or more contact intensity sensors 165 for detecting the intensity of contacts on the touch screen 112 and / or one or more haptic output generators 167 for generating haptic output to the user of device 100.

[0155] FIG. 3 is a block diagram of an exemplary multifunctional device having a display and a touch sensing surface, in accordance with some embodiments. Device 300 need not be portable. In some embodiments, device 300 is a laptop computer, desktop computer, tablet computer, multimedia player device, navigation device, educational device (such as a child's learning toy), gaming system, or a control device (e.g., a home or industrial controller). Device 300 typically includes one or more processing units (CPUs) 310, one or more networks or other communication interfaces 360, memory 370, and one or more communication buses 320 interconnecting these components. Communication bus 320 optionally includes circuitry (sometimes called a chipset) that interconnects and controls communications between system components. Device 300 includes an input / output (I / O) interface 330 that includes a display 340, which is typically a touch screen display. I / O interface 330 also optionally includes a keyboard and / or mouse (or other pointing device) 350, as well as a touch pad 355, a haptic output generator 357 (e.g., similar to haptic output generator 167 described above with reference to FIG. 1A) that generates haptic outputs on device 300, and sensors 359 (e.g., light, acceleration, proximity, touch sensing, and / or contact intensity sensors similar to contact intensity sensor 165 described above with reference to FIG. 1A). Memory 370 includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices, and optionally includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. Memory 370 optionally includes one or more storage devices remotely located from the CPU(s) 310.In some embodiments, memory 370 stores programs, modules, and data structures similar to, or a subset of, the programs, modules, and data structures stored in memory 102 of portable multifunctional device 100 (FIG. 1A). Further, memory 370 optionally stores additional programs, modules, and data structures not present in memory 102 of portable multifunctional device 100. For example, memory 370 of device 300 optionally stores a drawing module 380, a presentation module 382, a word processing module 384, a website creation module 386, a disk authoring module 388, and / or a spreadsheet module 390, whereas memory 102 of portable multifunctional device 100 (FIG. 1A) optionally does not store these modules.

[0156] Each of the elements identified above in FIG. 3 is optionally stored in one or more of the memory devices described above. Each of the modules identified above corresponds to a set of instructions for performing the functions described above. The modules or programs (e.g., sets of instructions) identified above need not be implemented as separate software programs, procedures, or modules, and thus, in various embodiments, various subsets of these modules are optionally combined or otherwise reconfigured. In some embodiments, memory 370 optionally stores a subset of the modules and data structures identified above. Further, memory 370 optionally stores additional modules and data structures not described above.

[0157] Next, optionally direct attention to an embodiment of a user interface, for example, implemented on portable multifunctional device 100.

[0158] Figure 4A shows an exemplary user interface of a menu of applications on a portable multifunctional device 100 according to some embodiments. A similar user interface is optionally implemented on device 300. In some embodiments, the user interface 400 includes the following elements, or a subset or superset thereof. ● Signal strength indicator(s) 402 for wireless communication(s) such as cellular signal and Wi-Fi signal, ● Time 404, ● Bluetooth indicator 405, ● Battery status indicator 406, ● A tray 408 having icons of frequently used applications such as ○ An icon 416 of the phone module 138 labeled "Phone", optionally including an indicator 414 of the number of missed calls or voice mail messages, ○ An icon 418 of the email client module 140 labeled "Mail", optionally including an indicator 410 of the number of unread emails, ○ An icon 420 of the browser module 147 labeled "Browser", and ○ An icon 422 for the video and music player module 152, also referred to as the iPod (trademark of Apple Inc.) module 152, labeled "iPod", and ● Icons of other applications such as ○ An icon 424 of the IM module 141 labeled "Message", ○ An icon 426 of the calendar module 148 labeled "Calendar", ○ An icon 428 of the image management module 144 labeled "Photos", ○ An icon 430 of the camera module 143 labeled "Camera", ○ An icon 432 of the online video module 155 labeled "Online Video", ○ The icon 434 of the stock price widget 149-2, labeled "Stock price", ○ The icon 436 of the map module 154, labeled "Map", ○ The icon 438 of the weather widget 149-1, labeled "Weather", ○ The icon 440 of the alarm clock widget 149-4, labeled "Clock", ○ The icon 442 of the training support module 142, labeled "Training support", ○ The icon 444 of the memo module 153, labeled "Memo", and ○ The icon 446 of the settings application or module, labeled "Settings", which provides access to the settings of the device 100 and its various applications 136.

[0159] Note that the icon labels shown in FIG. 4A are merely illustrative. For example, other labels such as "Music" or "Music player" can be optionally used for the icon 422 of the video and music player module 152 for various application icons. In some embodiments, the label for each application icon includes the name of the application corresponding to that application icon. In some embodiments, the label for a particular application icon is different from the name of the application corresponding to that particular application icon.

[0160] FIG. 4B shows an exemplary user interface on a device (e.g., device 300 of FIG. 3) having a touch sensing surface 451 (e.g., the tablet or touch pad 355 of FIG. 3) separate from the display 450 (e.g., touch screen display 112). The device 300 also optionally includes one or more contact intensity sensors (e.g., one or more of the sensors 359) for detecting the intensity of contact on the touch sensing surface 451, and / or one or more haptic output generators 357 for generating haptic output to the user of the device 300.

[0161] Some of the following examples are given with reference to inputs on touch screen display 112 (where the touch sensing surface and the display are combined), but in some embodiments, the device detects inputs on a touch sensing surface separate from the display, as shown in FIG. 4B. In some embodiments, the touch sensing surface (e.g., 451 in FIG. 4B) has a primary axis (e.g., 452 in FIG. 4B) corresponding to the primary axis (e.g., 453 in FIG. 4B) on the display (e.g., 450). According to these embodiments, the device detects contact (e.g., 460 and 462 in FIG. 4B) with the touch sensing surface 451 at positions corresponding to each position on the display (e.g., in FIG. 4B, 460 corresponds to 468 and 462 corresponds to 470). In this way, user inputs (e.g., contacts 460 and 462 and their movements) detected by the device on the touch sensing surface (e.g., 451 in FIG. 4B) are used by the device to operate the user interface on the display (e.g., 450 in FIG. 4B) of the multifunctional device when the touch sensing surface is separate from the display. It should be understood that similar methods are optionally used for other user interfaces described herein.

[0162] In addition, while the following examples are given mainly with reference to finger inputs (e.g., finger contact, finger tap gesture, finger swipe gesture), it should be understood that in some embodiments, one or more of the finger inputs may be replaced by inputs from another input device (e.g., mouse-based input or stylus input). For example, a swipe gesture may optionally be a mouse click (e.g., instead of a contact), followed by a mouse click that involves movement of the cursor along the path of the swipe (e.g., instead of movement of a contact). As another example, a tap gesture may optionally be replaced by a mouse click while the cursor is positioned over the location of the tap gesture (e.g., instead of detecting a contact and subsequently stopping detection of the contact). Similarly, it should be understood that when multiple user inputs are detected simultaneously, multiple computer mice may optionally be used simultaneously, or mouse and finger contacts may optionally be used simultaneously.

[0163] FIG. 5A shows an exemplary personal electronic device 500. The device 500 includes a body 502. In some embodiments, the device 500 can include some or all of the features described with respect to devices 100 and 300 (e.g., FIGS. 1A - 4B). In some embodiments, the device 500 has a touch-sensitive display screen 504, hereinafter referred to as touch screen 504. Alternatively, or in addition to the touch screen 504, the device 500 has a display and a touch-sensitive surface. Similar to devices 100 and 300, in some embodiments, the touch screen 504 (or touch-sensitive surface) optionally includes one or more intensity sensors that detect the intensity of an applied contact (e.g., a touch). One or more intensity sensors of the touch screen 504 (or touch-sensitive surface) can provide output data representative of the intensity of the touch. The user interface of the device 500 can respond to the touch(es) based on its intensity, which means that touches of different intensities can call different user interface operations on the device 500.

[0164] Exemplary techniques for detecting and processing touch intensity are described, for example, in International Patent Application No. PCT / US2013 / 040061, filed May 8, 2013, published as International Patent No. WO / 2013 / 169849, "Device, Method, and Graphical User Interface for Displaying User Interface Objects Corresponding to an Application," and International Patent Application No. PCT / US2013 / 069483, filed Nov. 11, 2013, published as International Patent No. WO / 2014 / 105276, "Device, Method, and Graphical User Interface for Transitioning Between Touch Input to Display Output Relationships," each of which is hereby incorporated by reference in its entirety.

[0165] In some embodiments, device 500 has one or more input mechanisms 506 and 508. Input mechanisms 506 and 508, if included, can be physical. Examples of physical input mechanisms include push buttons and rotatable mechanisms. In some embodiments, device 500 has one or more attachment mechanisms. Such attachment mechanisms, if included, can enable device 500 to be attached, for example, to hats, glasses, earrings, necklaces, shirts, jackets, bracelets, watch bands, chains, pants, belts, shoes, wallets, backpacks, and the like. These attachment mechanisms enable a user to wear device 500.

[0166] FIG. 5B shows an exemplary personal electronic device 500. In some embodiments, device 500 can include some or all of the components described with respect to FIGS. 1A, 1B, and 3. Device 500 has a bus 512 that operably couples an I / O section 514 to one or more computer processors 516 and a memory 518. The I / O section 514 can be connected to a display 504, which can have a touch sensing component 522 and optionally an intensity sensor 524 (e.g., a contact intensity sensor). Additionally, the I / O section 514 can be connected to a communication unit 530 that receives application and operating system data using Wi-Fi, Bluetooth, near field communication (NFC), cellular, and / or other wireless communication techniques. The device 500 can include an input mechanism 506 and / or 508. The input mechanism 506 can optionally be, for example, a rotatable input device or a depressible and rotatable input device. In some examples, the input mechanism 508 can optionally be a button.

[0167] In some examples, the input mechanism 508 can optionally be a microphone. The personal electronic device 500 can optionally include various sensors such as a GPS sensor 532, an accelerometer 534, a direction sensor 540 (e.g., a compass), a gyroscope 536, a motion sensor 538, and / or combinations thereof, all of which can be operably connected to the I / O section 514.

[0168] The memory 518 of the personal electronic device 500 may include one or more non-transitory computer-readable storage media for storing computer-executable instructions, which, when executed by one or more computer processors 516, can cause, for example, the computer processor to execute the techniques described below, including process 700 (Figs. 7A-7C), process 900 (Figs. 9A-9C), process 1200 (Figs. 12A-12B), process 1400 (Figs. 14A-14B), and process 1500 (Fig. 15). A computer-readable storage media can be any media that can tangibly contain or store computer-executable instructions used by or related to an instruction execution system, apparatus, or device. In some embodiments, the storage media is a transitory computer-readable storage media. In some embodiments, the storage media is a non-transitory computer-readable storage media. Non-transitory computer-readable storage media can include, but are not limited to, magnetic, optical, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, CDs, DVDs, or optical disks based on Blu-ray technology, as well as persistent solid-state memories such as flash, solid-state drives. The personal electronic device 500 is not limited to the components and configurations of Fig. 5B and can include other or additional components in multiple configurations.

[0169] As used herein, the term "affordance" optionally refers to user interaction graphical user interface objects displayed on the display screens of devices 100, 300, and / or 500 (Figs. 1A, 3, and 5A-5B). For example, images (e.g., icons), buttons, and text (e.g., hyperlinks) each optionally constitute an affordance.

[0170] As used herein, the term "focus selector" refers to an input element that indicates the current part of the user interface with which the user is interacting. In some implementations that include a cursor or other position marker, the cursor acts as the "focus selector," and thus while the cursor is positioned over a particular user interface element (e.g., a button, window, slider, or other user interface element), when an input (e.g., a press input) is detected on the touch-sensitive surface (e.g., the touchpad 355 of FIG. 3 or the touch-sensitive surface 451 of FIG. 4B), the particular user interface element is adjusted according to the detected input. In some implementations that include a touch screen display (e.g., the touch-sensitive display system 112 of FIG. 1A or the touch screen 112 of FIG. 4A) that enables direct interaction with user interface elements on the touch screen display, the detected contact on the touch screen acts as the "focus selector," and thus when an input (e.g., a press input by contact) is detected at the position of a particular user interface element (e.g., a button, window, slider, or other user interface element) on the touch screen display, the particular user interface element is adjusted according to the detected input. In some implementations, the focus is moved from one area of the user interface to another area of the user interface without moving the corresponding cursor or contact on the touch screen display (e.g., by using the tab key or arrow keys to move the focus from one button to another), and in these implementations, the focus selector moves in accordance with the movement of the focus between different areas of the user interface. Regardless of the specific form the focus selector takes, the focus selector is generally a user interface element (or a contact on the touch screen display) that is controlled by the user to convey information about the user's intended interaction with the user interface (e.g., by indicating to the device the user interface element through which the user intends to interact).For example, while a press input is detected on a touch sensing surface (e.g., a touch pad or a touch screen), the position of a focus selector (e.g., a cursor, a contact, or a selection box) over the corresponding button indicates that the user intends to activate that corresponding button (rather than other user interface elements shown on the display of the device).

[0171] As used in this specification and the claims, the term "characteristic strength" of a contact refers to the characteristics of that contact based on one or more strengths of the contact. In some embodiments, the characteristic strength is based on a plurality of strength samples. The characteristic strength is optionally based on a set of strength samples collected during a predetermined time (e.g., 0.05, 0.1, 0.2, 0.5, 1, 2, 5, 10 seconds) associated with a predetermined number of strength samples, i.e., a predetermined event (e.g., after detecting the contact, before detecting the lift-off of the contact, before or after detecting the start of movement of the contact, before detecting the end of the contact, before or after detecting an increase in the strength of the contact, and / or before or after detecting a decrease in the strength of the contact). The characteristic strength of a contact is optionally based on one or more of the maximum value of the strength of the contact, the mean value of the strength of the contact, the average value of the strength of the contact, the top 10 percentile value of the strength of the contact, the median value of the strength of the contact, the 90th percentile value of the strength of the contact, etc. In some embodiments, the duration of the contact is used when determining the characteristic strength (e.g., when the characteristic strength is the average of the strength of the contact over time). In some embodiments, the characteristic strength is compared to a set of one or more strength thresholds to determine whether an action has been performed by a user. For example, the set of one or more strength thresholds optionally includes a first strength threshold and a second strength threshold. In this example, a contact having a characteristic strength that does not exceed the first threshold results in a first action, a contact having a characteristic strength that exceeds the first strength threshold but does not exceed the second strength threshold results in a second action, and a contact having a characteristic strength that exceeds the second threshold results in a third action. In some embodiments, the comparison between the characteristic strength and one or more thresholds is not used to determine whether to perform a first action or a second action, but rather is used to determine whether to perform one or more actions (e.g., whether to perform each action or to defer performing each action).

[0172] Figure 5C shows detecting a plurality of contacts 552A - 552E on a touch - sensing display screen 504 by a plurality of intensity sensors 524A - 524D. Figure 5C additionally includes an intensity diagram showing the current intensity measurements of the intensity sensors 524A - 524D in intensity units. In this example, the intensity measurements of intensity sensors 524A and 524D are each 9 intensity units, and the intensity measurements of intensity sensors 524B and 524C are each 7 intensity units. In some implementations, the aggregated intensity is the sum of the intensity measurements of the plurality of intensity sensors 524A - 524D, which is 32 intensity units in this example. In some embodiments, to each contact, a respective intensity that is a portion of the aggregated intensity is assigned. Figure 5D shows assigning the aggregated intensity to contacts 552A - 552E based on the distance from the center of force 554. In this example, to each of contacts 552A, 552B, and 552E, a contact intensity of 8 intensity units of the aggregated intensity is assigned, and to each of contacts 552C and 552D, a contact intensity of 4 intensity units of the aggregated intensity is assigned. More generally, in some implementations, for each contact j, a respective intensity Ij that is a portion of the total intensity A is assigned according to a predetermined mathematical function Ij = A·(Dj / ΣDi), where Dj is the distance from the center of force to each respective contact j, and ΣDi is the sum of the distances from the center of force to all respective contacts (e.g., from i = 1 to the last). The operations described with reference to Figures 5C - 5D can be performed using an electronic device similar or identical to device 100, 300, or 500. In some embodiments, the characteristic intensity of a contact is based on one or more intensities of the contact. In some embodiments, the intensity sensors are used to determine a single characteristic intensity (e.g., the single characteristic intensity of a single contact). Note that the intensity diagram is included in Figures 5C - 5D to assist the reader, rather than being part of the display user interface.

[0173] In some embodiments, for the purpose of determining characteristic intensity, a portion of a gesture is specified. For example, the touch sensing surface optionally receives continuous swipe contacts that transition from a start position and reach an end position, where the intensity of the contact increases at that position. In this example, the characteristic intensity of the contact at the end position is optionally based on only a portion of the continuous swipe contact (e.g., only the portion of the swipe contact at the end position), rather than the entire swipe contact. In some embodiments, optionally, a smoothing algorithm is applied to the intensity of the swipe contact before determining the characteristic intensity of the contact. For example, the smoothing algorithm optionally includes one or more of a non-weighted moving average smoothing algorithm, a triangular smoothing algorithm, a median filter smoothing algorithm, and / or an exponential smoothing algorithm. In some situations, these smoothing algorithms eliminate narrow spikes or drops in the width of the swipe contact intensity for the purpose of determining characteristic intensity.

[0174] The intensity of a contact on the touch sensing surface is optionally characterized relative to one or more intensity thresholds, such as a contact detection intensity threshold, a light press intensity threshold, a deep press intensity threshold, and / or one or more other intensity thresholds. In some embodiments, the light press intensity threshold typically corresponds to the intensity at which the device performs an operation associated with clicking a button or trackpad of a physical mouse. In some embodiments, the deep press intensity threshold typically corresponds to the intensity at which the device performs an operation different from the operation associated with clicking a button or trackpad of a physical mouse. In some embodiments, when a contact having a characteristic intensity below the light press intensity threshold (e.g., and above a nominal contact detection intensity threshold below which the contact is no longer detected) is detected, the device moves the focus selector without performing an operation associated with the light press intensity threshold or the deep press intensity threshold as the contact on the touch sensing surface moves. Generally, unless otherwise specified, these intensity thresholds are consistent among various sets of user interface values.

[0175] An increase in the characteristic strength of a contact from a strength below a light press intensity threshold to a strength between the light press intensity threshold and a deep press intensity threshold may be referred to as an input of "light press". An increase in the characteristic strength of a contact from a strength below a deep press intensity threshold to a strength above the deep press intensity threshold may be referred to as an input of "deep press". An increase in the characteristic strength of a contact from a strength below a contact detection intensity threshold to a strength between the contact detection intensity threshold and the light press intensity threshold may be referred to as a detection of a contact on the touch surface. A decrease in the characteristic strength of a contact from a strength above the contact detection intensity threshold to a strength below the contact detection intensity threshold may be referred to as a detection of a lift-off of the contact from the touch surface. In some embodiments, the contact detection intensity threshold is zero. In some embodiments, the contact detection intensity threshold is greater than zero.

[0176] In some embodiments described herein, one or more operations are performed in response to detecting a gesture that includes each press input, or in response to detecting each press input performed by each contact (or contacts), and each press input is detected based at least in part on detecting an increase in the strength of a contact (or contacts) above a press input intensity threshold. In some embodiments, each operation is performed in response to detecting an increase in the strength of each contact above the press input intensity threshold (e.g., the "downstroke" of each press input). In some embodiments, the press input includes an increase in the strength of each contact above the press input intensity threshold and a subsequent decrease in the strength of the contact below the press input intensity threshold, and each operation is performed in response to detecting a subsequent decrease in the strength of each contact below the press input threshold (e.g., the "upstroke" of each press input).

[0177] Figures 5E-5H show an increase in intensity from below the light press intensity threshold of Figure 5E (e.g., "IT" L ") to the deep press intensity threshold of Figure 5H (e.g., "IT" DIt shows the detection of a gesture including a press - down input corresponding to an increase in the intensity of contact 562 to an intensity exceeding “」). The gesture executed by contact 562 is detected on the touch - sensing surface 560, and on the display user interface 570 including application icons 572A - 572D displayed within a predetermined region 574, a cursor 576 is displayed over the application icon 572B corresponding to App 2. In some embodiments, the gesture is detected on the touch - sensing display 504. The intensity sensor detects the intensity of the contact on the touch - sensing surface 560. The device determines that the intensity of contact 562 has reached a peak exceeding a deep - press intensity threshold (e.g., “IT D ”). Contact 562 is maintained on the touch - sensing surface 560. In response to the detection of the gesture, in accordance with contact 562 having an intensity exceeding a deep - press intensity threshold (e.g., “IT D ”) during the gesture, as shown in FIGS. 5F - 5H, reduced - scale representations 578A - 578C (e.g., thumbnails) of the most recently opened documents for App 2 are displayed. In some embodiments, this intensity compared to one or more intensity thresholds is the characteristic intensity of the contact. Note that the intensity diagram for contact 562 is included in FIGS. 5E - 5H to assist the reader, not as part of the display user interface.

[0178] In some embodiments, the display of representations 578A - 578C includes an animation. For example, as shown in FIG. 5F, representation 578A is first displayed proximate to application icon 572B. As the animation progresses, as shown in FIG. 5G, representation 578A moves upward and representation 578B is displayed proximate to application icon 572B. Then, as shown in FIG. 5H, representation 578A moves upward, representation 578B moves upward toward representation 578A, and representation 578C is displayed proximate to application icon 572B. Representations 578A - 578C form an array over icon 572B. In some embodiments, the animation progresses in accordance with the intensity of contact 562, as shown in FIGS. 5F - 5G, and the intensity of contact 562 exceeds a deep - press intensity threshold (e.g., “ITD As it increases towards ") ", expressions 578A to 578C appear and move upward. In some embodiments, the intensity based on the progress of the animation is the characteristic intensity of the contact. The operations described with reference to FIGS. 5E to 5H can be executed using an electronic device similar or identical to devices 100, 300, or 500.

[0179] In some embodiments, the device employs intensity hysteresis to avoid spurious inputs sometimes referred to as "jitter", and the device defines or selects a hysteresis intensity threshold having a predefined relationship with the press input intensity threshold (e.g., the hysteresis intensity threshold is X intensity units lower than the press input intensity threshold, or the hysteresis intensity threshold is 75%, 90%, or some reasonable percentage of the press input intensity threshold). Thus, in some embodiments, a press input includes an increase in the intensity of each contact that exceeds the press input intensity threshold, and a subsequent decrease in the intensity of the contact that falls below the hysteresis intensity threshold corresponding to the press input intensity threshold, and each operation is executed in response to detecting a subsequent decrease in the intensity of each contact that falls below the hysteresis intensity threshold (e.g., the "upstroke" of each press input). Similarly, in some embodiments, a press input is detected only when the device detects an increase in the intensity of a contact from an intensity below the hysteresis intensity threshold to an intensity above the press input intensity threshold, and optionally, a subsequent decrease in the intensity of the contact to an intensity below the hysteresis intensity, and each operation is executed in response to detecting a press input (e.g., an increase in the intensity of the contact or a decrease in the intensity of the contact, depending on the situation).

[0180] For ease of explanation, the description of an operation performed in response to a press input associated with a press input strength threshold, or a gesture including the press input, is optionally triggered in response to detecting any one of an increase in the intensity of contact exceeding the press input strength threshold, an increase in the intensity of contact from an intensity below the hysteresis strength threshold to an intensity exceeding the press input strength threshold, a decrease in the intensity of contact below the press input strength threshold, and / or a decrease in the intensity of contact below the hysteresis strength threshold corresponding to the press input strength threshold. Further, in an example where an operation is described to be performed in response to detecting a decrease in the intensity of contact below the press input strength threshold, the operation is optionally performed in response to detecting a decrease in the intensity of contact corresponding to and below a hysteresis strength threshold lower than the press input strength threshold.

[0181] As used herein, an “installed application” refers to a software application that has been downloaded onto an electronic device (e.g., device 100, 300, and / or 500) and is ready to be launched (e.g., opened) on the device. In some embodiments, a downloaded application becomes an installed application by an installation program that extracts a program portion from the downloaded package and integrates the extracted portion with the operating system of the computer system.

[0182] As used herein, the terms “open application” or “running application” refer to a software application that has retained state information (e.g., as part of device / global internal state 157 and / or application internal state 192). An open or running application is optionally any one of the following types of applications. ● An active application currently displayed on the display screen of the device being used by the application, ● Background applications (or background processes) for which one or more processes for an application are being processed by one or more processors, although not currently being displayed, and ● An application in an interrupted or suspended state that is stored in memory (both volatile and non-volatile, respectively) and has state information that can be used to resume execution of the application, although not currently running.

[0183] As used herein, the term "closed application" refers to a software application that does not have retained state information (e.g., state information for a closed application is not stored in the device's memory). Thus, closing an application includes stopping and / or removing the application process for the application and removing the state information for the application from the device's memory. Generally, opening a second application while a first application is open does not close the first application. When the second application is being displayed and the first application has its display terminated, the first application becomes a background application.

[0184] Audio spatial management includes techniques for modifying the characteristics of sound (e.g., by applying filters) so that a listener perceives the sound as being emitted from a specific location within a space (e.g., a three-dimensional (3D) space). Such techniques can be achieved using speakers such as headphones, earbuds, or loudspeakers. In some embodiments, for example, when a listener is using headphones, binaural simulation is used to reproduce binaural cues that give the listener the illusion that sound is arriving from a specific location within the space. For example, the listener perceives that the sound source is arriving from the left of the listener. In another embodiment, the listener perceives that the sound source is passing from left to right in front of the listener. This effect can be enhanced by using head tracking, even when the listener's head is moving or rotating, creating the illusion that the position of the sound source remains stationary within the space. In some embodiments, for example, when a listener is using a loudspeaker, similar effects can be achieved using cross-talk cancellation, giving the listener the illusion that sound is arriving from a specific location within the space.

[0185] The Head-Related Transfer Function (HRTF) characterizes how a human ear receives sound from various points in space. The HRTF can be based on one or more of the direction, elevation, and distance of the sound. By using the HRTF, a device (e.g., device 100) can apply different functions to audio to reproduce the directivity pattern of a human ear. In some embodiments, a pair of HRTFs for two ears are used to synthesize binaural audio that is perceived to be arriving from a specific point in space relative to a listener, e.g., above, below, in front of, behind, to the left or right of the user, or combinations thereof. A personalized HRTF provides better results for the listener for whom the HRTF is personalized as compared to a general HRTF. In some embodiments, the HRTF is applied to a listener using a listening device such as headphones, earphones, and earbuds.

[0186] In another embodiment, when a device (e.g., device 100) uses two or more loudspeakers to generate sound, the sound from each loudspeaker can be heard through each of the listener's closest ears, but also through the opposite ear, resulting in crosstalk. Effectively managing the cancellation of this unintended crosstalk helps to modify the sound so that the listener perceives that the sound is being emitted from a specific position in space.

[0187] The device (e.g., device 100, 300, 500) can also simultaneously modify the characteristics of multiple audio sources (e.g., by applying different filters to each source) to give the listener the illusion that the sound from different audio sources is arriving from different positions in space. Such techniques can be achieved using headphones or loudspeakers.

[0188] In some embodiments, modifying a stereo sound source such that a listener perceives that the sound is being emitted from a particular location in space (e.g., 3D space) includes generating a monaural sound from the stereo sound. For example, the stereo sound includes a left audio channel and a right audio channel. The left audio channel includes, for example, a first device without including a second device. The second audio channel includes, for example, a second device without including the first device.

[0189] When the stereo sound source is placed in space, the device optionally combines the left audio channel and the right audio channel to form a composite channel audio, and then applies an interaural time difference to the composite channel audio. Further, the device optionally (or alternatively) applies an HRTF and / or cross-cancellation to the composite channel audio before generating the composite channel audio with different speakers.

[0190] When the stereo sound source is not placed in space, the device optionally does not combine the left audio channel and the right audio channel and does not apply any of the interaural time difference, HRTF, or cross-cancellation. Instead, the device generates a stereo sound by using the left loudspeaker of the device to generate the left audio channel and using the right loudspeaker of the device to generate the right channel. As a result, the device generates sound in stereo and the listener perceives the audio in stereo.

[0191] Many of the techniques described below modify sound using various processes such that a listener perceives the sound as coming from a particular location in space.

[0192] Next, attention is directed to embodiments of a user interface ("UI") and related processes implemented on an electronic device such as the portable multifunctional device 100, the device 300, or the device 500.

[0193] Figures 6A - 6N illustrate exemplary techniques for transitioning between visual elements, according to some embodiments. The techniques of these figures are used to illustrate the processes described below, including the processes of Figures 7A - 7C.

[0194] Figures 6A - 6G show a user 606 sitting in front of a device 600 (e.g., a laptop computer) having a display 600a, left and right loudspeakers, and a touch - sensing surface 600b (e.g., a touchpad). Throughout Figures 6A - 6G and Figures 6K - 6N, additional enlarged views of the touch - sensing surface 600b of the device 600 are shown on the right side of the user, providing the reader with a better understanding of the techniques described with respect to exemplary user inputs. Similarly, the overhead view 650 is a visual representation of the spatial configuration of the audio generated by the device 600 and is shown throughout Figures 6A - 6G, providing the reader with a better understanding of the techniques with respect to the position where the sound is arriving (e.g., as a result of the device 600 placing audio in space) as perceived by the user 606. The overhead view 650 is not part of the device's user interface. Similarly, visual elements displayed outside the display device, as represented by the dotted outline, are not part of the displayed user interface but are illustrated to provide the reader with a better understanding of the techniques. Similarly, visual elements displayed outside the device's display, as represented by the dotted outline, are not part of the displayed user interface but are illustrated to provide the reader with a better understanding of the techniques. Throughout Figures 6A - 6G, the device 600 is running (1) a web browser 604a that includes the playback of a basketball game with game audio and game video, (2) a music player 604b that includes the playback of music, and (3) a video player 604c that includes the playback of show audio and show video. The audio element 654a corresponds to the audio supplied from the web browser 604a, the audio element 654b corresponds to the audio supplied from the music player 604b, and the audio element 654c corresponds to the audio supplied from the video player 604c. Throughout Figures 6A - 6G, the device 600 simultaneously generates the audio supplied from each of the web browser 604a, music player 604b, and video player 604c.

[0195] In FIG. 6A, device 600 displays web browser 604a on display 600a. While web browser 604a is being displayed, device 600 uses the left and right loudspeakers to generate game audio supplied from web browser 604a. In FIG. 6A, while web browser 604a is being displayed, device 600 does not spatially position the game audio of web browser 604a (e.g., device 600 does not apply any interaural time differences, HRTF, or cross-cancellation). As a result, user 606 perceives the audio element 654a to be in front of user 606 stereophonically, as represented in FIG. 6A by the position of the audio element 654a relative to user 606 within the overhead view 650. For example, the left audio channel of the game audio includes a sports commentator and no noise from the audience, while the right audio channel includes noise from the audience and no sports commentator. While device 600a is displaying web browser 604a, device 600 generates stereo audio by using the game audio to generate the left audio channel using the left loudspeaker of device 600a and the right channel using the right loudspeaker of device 600a. As a result, user 606 perceives the audio element 654a to be in front of user 606 stereophonically, as represented in FIG. 6A by the position of the audio element 654a relative to user 606 within the overhead view 650. In some embodiments, if web browser 604a is part of the currently accessed desktop, the device generates the game audio for web browser 604a in the same manner even if web browser 604a is not actively displayed (e.g., when different visual elements such as a word processing application are displayed on top of web browser 604a, thus blocking the display of web browser 604a on display 600a). Optionally, the audio of various applications follows a curved path 650a when device 600 relocates the audio in space. In some embodiments, device 600a positions various sound sources equidistant in space from adjacent sound sources.In some embodiments, device 600a moves various sound sources along only two axes (e.g., left - right and front - back, but not up - down).

[0196] In FIG. 6A, device 600 does not display music player 604b. While music player 604b is not displayed, device 600 uses the left and right loudspeakers to generate music supplied from music player 604b. In FIG. 6A, device 600 positions the music supplied from music player 604b within the space (e.g., device 600 applies an inter - aural time difference, HRTF, and / or cross - cancellation to the music). Device 600 positions the music such that the user perceives that the music is arriving from a position within the space to the right of display 600a (and user 606), as indicated by audio element 654b. Thereby, the user can recognize that music player 604b is being executed to generate audio even when music player 604b is not displayed. Further, the positioning of the music within the space helps the user recognize how to access the display of music player 604c, as described in FIGS. 6B - 6G.

[0197] In FIG. 6A, device 600 is not displaying video player 604c. While video player 604c is not being displayed, device 600 uses the left and right loudspeakers to generate the audio supplied from video player 604c. In FIG. 6A, device 600 positions the show audio supplied from video player 604b in space (e.g., device 600 applies interaural time differences, HRTF, and / or cross-cancellation to the music). Device 600 positions the show audio such that the user perceives that the show audio is coming from a position in space that is further to the right of display 600a (and user 606) than the music from music player 604d. This allows the user to recognize that video player 604c is being executed to generate audio even when video player 604c is not being displayed. Further, the positioning of the show audio in space helps the user recognize how to access the display of video player 604c.

[0198] In addition, the device 600 optionally applies a low-pass (or high-pass, or band-pass) filter to the audio corresponding to applications that are not on the display (e.g., not part of the currently accessed desktop), thereby attenuating (e.g., removing) audio that exceeds a particular frequency threshold before the audio is generated by the loudspeaker. As a result, the user perceives such audio as background noise as compared to audio to which the low-pass filter is not applied. Thereby, the user can more easily make the audio from a particular application, such as an application not currently being displayed, inaudible. In some embodiments, the same low-pass filter is applied to all audio corresponding to applications that are not displayed. In some embodiments, different low-pass filters are applied to respective audio based on how far away from the display the corresponding application is to be perceived by the user. In some embodiments, the device 600 optionally attenuates the audio corresponding to applications that are not on the display (e.g., across all frequencies of the audio).

[0199] In FIG. 6A, for example, the device 600 does not apply a low-pass filter to the game audio supplied from the web browser 604a, the device 600 applies a first low-pass filter having a first cut-off frequency to the music supplied from the music player 604b, and the device 600 applies a second low-pass filter having a second cut-off frequency (lower than the first cut-off frequency) to the show audio supplied from the video player 604c. Thus, in FIG. 6A, the device 600 simultaneously generates the audio from each of the web browser 604a, the music player 604b, and the video player 604c.

[0200] In FIGS. 6B - 6C, device 600 receives a left - swipe user input 610a at touch - sensing surface 600b. In response to receiving the left - swipe user input 610a, as shown in FIGS. 6B - 6C, device 600 transitions the display of web browser 604a away from display 600a by sliding web browser 604a to the left, and transitions the display of music player 604b onto display 600a by sliding music player 604b to the left. Further, in response to receiving the left - swipe user input 610a, as shown in the overhead view 650 of FIGS. 6B - 6C, device 600 changes the position in the space where the user perceives the audio from the corresponding application.

[0201] In FIG. 6D, device 600 is not displaying web browser 604a. While web browser 604a is not being displayed, device 600 uses the left and right loudspeakers to position the audio supplied from web browser 604a in the space (e.g., device 600 applies inter - aural time difference, HRTF, and / or cross - cancellation to the audio), such that the user perceives that the game audio is coming from a position in the space to the left of display 600a (and user 606). Thereby, even when web browser 604a is not being displayed, the user can recognize that web browser 604a is being executed and generating audio. Further, the positioning of the game audio in the space helps the user recognize how to access the display of web browser 604a (e.g., using a right - swipe user input).

[0202] In FIG. 6D, device 600 displays music player 604b on display 600a. While music player 604b is being displayed, device 600 uses the left and right loudspeakers to generate the audio supplied from music player 604b. In FIG. 6D, while music player 604b is being displayed, device 600 does not place the music supplied from music player 604b in space (e.g., device 600 does not apply any interaural time differences, HRTFs, or cross-cancellation). As a result, user 606 perceives that the audio element 654c is in front of user 606 stereophonically, as represented in the overhead view 650 of FIG. 6D by the position of the audio element 654c relative to user 606.

[0203] In FIG. 6D, device 600 is not displaying video player 604c. While video player 604c is not being displayed, device 600 uses the left and right loudspeakers to place the audio supplied from video player 604c in space (e.g., device 600 applies interaural time differences, HRTFs, and / or cross-cancellation to the audio), such that the user perceives that the audio supplied from video player 604c is coming from a position not as far to the right as was previously perceived by the user in FIG. 6A, within the space to the right of display 600a (and user 606). Thereby, the user can recognize that video player 604c is being executed to generate audio even when video player 604c is not being displayed. Further, the placement of the show audio in space helps the user recognize how to access the display of video player 604c (e.g., using a left swipe user input).

[0204] In FIG. 6D, for example, device 600 does not apply a low-pass filter to the music supplied from music player 604b. Device 600 applies a first low-pass filter having a first cut-off frequency to the game audio supplied from web browser 604a, and device 600 applies a first low-pass filter having a first cut-off frequency to the show audio supplied from video player 604c. Thus, in FIG. 6D, device 600 simultaneously generates audio from each of web browser 604a, music player 604b, and video player 604c.

[0205] In FIGS. 6E - 6F, device 600 receives a right-swipe user input 610b at touch-sensing surface 600b. In response to receiving the right-swipe user input 610b, as shown in FIGS. 6E - 6F, device 600 causes the display of web browser 604a to transition on display 600a by sliding web browser 604a to the right, and causes the display of music player 604b to transition away from display 600a by sliding music player 604b to the right. Further, in response to receiving the right-swipe user input 610b, as shown in the overhead view 650 of FIGS. 6E - 6F, device 600 changes the position within the space where the user perceives the audio from the corresponding application. In this embodiment, the device modifies the audio so that the user perceives the audio as described in FIG. 6A.

[0206] Figure 6G shows an embodiment corresponding to Figure 6A. In Figure 6G, the user is listening to audio generated by device 600a using headphones. As a result, instead of perceiving that the game audio of web browser 604a is in front of the user, the user perceives that the game audio is being generated within the user's head. Optionally, the audio of various applications follows a straight path 650b when device 600 relocates the audio in space. In some embodiments, device 600a places various sound sources equidistant in space from an adjacent sound source. In some embodiments, device 600a moves various sound sources in only two axes (e.g., left - right and front - back but not up - down).

[0207] Figures 6H - 6J show a device 660 (e.g., a mobile phone) having a display 660a (e.g., a touch screen), a touch - sensing surface 660b (e.g., a part of the touch screen), and connected (wirelessly or wired) to headphones. In this embodiment, user 606 listens to device 660 using headphones.

[0208] The overhead view 670 is a visual representation of the spatial configuration of the audio being generated by device 660, shown throughout Figures 6H - 6J, and provides a better understanding of the technique with respect to the positions where the sound is arriving and being perceived by user 606 (e.g., as a result of device 660 placing the audio in space). The overhead view 670 is not part of the user interface of device 660. Similarly, visual elements displayed outside the display device are not part of the displayed user interface but are illustrated to provide a better understanding of the technique, as represented by a dotted - line contour. Throughout Figures 6H - 6J, device 660 is running a music player that includes the playback of music having corresponding album art.

[0209] The audio element 674a corresponds to the audio of track 1 supplied from the music player, the audio element 674b corresponds to the audio of track 2 supplied from the music player, the audio element 674c corresponds to the audio of track 3 supplied from the music player, and the audio element 674d corresponds to the audio of track 4 supplied from the music player.

[0210] In FIG. 6H, the device 660 does not display the album art 664a of track 1 on the display 660a. While the album art 664a of track 1 is not being displayed, the device 660 generates the audio of track 1 by placing the audio of track 1 in space using the left and right speakers of the headphones (e.g., the device 660 applies the interaural time difference, HRTF, and / or cross cancellation), whereby the user perceives that the audio is coming from a first point in space (e.g., to the left of the user, to the left of the device) as shown in FIG. 6H by the position of the audio element 670a relative to the user 606 in the overhead view 670. In this embodiment, the device 660 additionally modifies the audio of track 1 by attenuating the audio and / or applying a low-pass (or high-pass, or band-pass) filter to the audio. In some embodiments, the device 660 does not generate the audio of track 1 if the corresponding album art is not on the display.

[0211] In FIG. 6H, device 660 displays album art 664b of track 2 on display 660a. While album art 664a of track 2 is being displayed, device 600 generates the audio of track 2 by placing the audio of track 2 in space using the left and right headphones (e.g., device 600 applies an interaural time difference, HRTF, or cross-cancellation), whereby the user perceives that the audio is arriving from a second point in space (e.g., different from a first point in space, at a position corresponding to device 660, to the right of the first point in space in front of the user) as indicated in FIG. 6H by the position of audio element 670b relative to user 606 within the overhead view 670. In this embodiment, device 660 does not modify the audio of track 2 by attenuating the audio or applying a low-pass (or high-pass, or band-pass) filter to the audio.

[0212] In FIG. 6H, device 660 does not display album art 664c of track 3 on display 660a. Device 660 also does not generate the audio of track 3 using the left or right speaker of the headphones.

[0213] In FIGS. 6I - 6J, device 660 receives a left - swipe user input 666 at touch - sensing surface 660a. In response to receiving the left - swipe user input 666, as shown in FIGS. 6I - 6J, device 660 transitions the display of album art 664b away from display 660a by sliding album art 664b to the left, and transitions the display of album art 664c onto display 660a by sliding album art 664c to the left. Further, in response to receiving the left - swipe user input 666, device 660 starts generating the audio of track 3 (simultaneously with track 2) and changes the position within the space where the user perceives audio from tracks 2 and 3, as represented in the overhead view 670 of FIGS. 6I - 6J. In some embodiments, generating the audio of track 3 in response to receiving the left - swipe user input 666 includes skipping a predetermined amount of time of the audio (e.g., the first 0.5 seconds of track 2). This provides the user with the sense that track 3 has been previously played, even if device 660 had not previously generated the audio of track 3.

[0214] In FIG. 6J, device 660 stops generating the audio of track 1. Device 660 generates the audio of track 2 by placing the audio of track 2 in space using the left and right speakers of the headphones (e.g., device 660 applies the interaural time difference, HRTF, and / or cross-cancellation), whereby the user perceives that the audio is coming from a first point in space (e.g., to the left of the user, to the left of the device) as shown in FIG. 6J by the position of the audio element 670b with respect to the user 606 in the overhead view 670. In this embodiment, device 660 additionally modifies the audio of track 2 by attenuating the audio and / or applying a low-pass (or high-pass, or band-pass) filter to the audio. In some embodiments, device 660 fades out the audio of track 2 (by stopping generating the audio) as the corresponding album art moves away from the display.

[0215] In FIG. 6J, device 660 displays the album art 664c of track 3 on the display 660a. While displaying the album art 664b of track 3, device 600 generates the audio of track 3 by placing the audio of track 3 in space using the left and right headphones (e.g., device 600 applies the interaural time difference, HRTF, or cross-cancellation), whereby the user perceives that the audio is coming from a second point in space (e.g., at a position corresponding to device 660, to the right of the first point in space in front of the user, different from the first point in space) as shown in FIG. 6H by the position of the audio element 670c with respect to the user 606 in the overhead view 670. In this embodiment, device 660 does not modify the audio of track 3 by attenuating the audio or applying a low-pass (or high-pass, or band-pass) filter to the audio.

[0216] As a result, user 606 perceives the music passing in front of the user as the user swipes through various album arts.

[0217] Figures 6K - 6N show user 606 sitting in front of a device 600 (e.g., a laptop computer) having a display 600a, left and right loudspeakers, and a touch - sensing surface 600b (e.g., a touchpad). Throughout Figures 6K - 6N, additional enlarged views of the touch - sensing surface 600b of device 600 are shown to the right of the user, providing the reader with a better understanding of the techniques described, particularly with respect to user input (or the lack thereof). Similarly, the overhead view 680 is a visual representation of the spatial configuration of the audio being generated by device 600 and is shown throughout Figures 6K - 6N to provide the reader with a better understanding of the techniques, particularly with respect to the position at which the sound is perceived by user 606 (e.g., as a result of device 600 placing the audio in the space). The overhead view 680 is not part of the user interface of device 600. Similarly, visual elements displayed outside of the display device are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the techniques, as represented by the dotted outline.

[0218] In Figure 6K, music player 604b is playing music. While music player 604a is being displayed, device 600 uses the left and right speakers to generate the audio supplied from music player 604a. For example, device 600 does not place the audio supplied from music player 604a in the space (e.g., device 660 does not apply any of the inter - aural time difference, HRTF, or cross - cancellation). As a result, user 606 perceives stereo music to be in front of user 606 as represented in Figure 6K by the position of audio element 680a relative to user 606 within overhead view 680. Throughout Figures 6K - 6N, audio element 680a corresponds to the audio provided by music player 604b.

[0219] In FIG. 6L, device 600 receives a notification (e.g., a message notification received via a network connection). In response to receiving the notification (and without receiving user input), device 600 transitions to generating music supplied from music player 604a by placing the music in space using the left and right speakers (e.g., device 600 applies an interaural time difference, HRTF, or cross cancellation), whereby the user perceives that the audio is coming from a point in the space to the left of device 600 as shown in FIGS. 6L-6M by the position of audio element 680a relative to user 606 in the overhead view 680. Device 600 maintains a display of music player 604a on display 600a. In some embodiments, device 600 also displays the notification in response to receiving the notification. Throughout FIGS. 6K-6N, audio element 680b corresponds to the audio provided by the notification.

[0220] In further response to receiving the notification (and without receiving user input), since the device 600 does not generate the audio of the notification, the device 600 transitions to generating the audio of the notification by placing the audio of the notification in the space using the left and right speakers (e.g., the device 600 applies the interaural time difference, HRTF, or cross-cancellation), whereby the user perceives that the audio of the notification is coming from a point in the space to the right of the device 600 as shown in FIG. 6L by the position of the audio element 680b with respect to the user 606 in the overhead view 680. In some embodiments, the device 600 emphasizes the audio of the notification by ducking the music supplied from the music player 604a. For example, the device 600 attenuates the music supplied from the music player 604a while generating the audio of the notification. Next, the device 600 transitions to generating the audio of the notification without placing the audio in the space (e.g., the device 600 does not apply any of the interaural time difference, HRTF, or cross-cancellation) as shown in FIG. 6M by the position of the audio element 680b with respect to the user 606 in the overhead view 680.

[0221] Subsequently, as shown in FIG. 6N, the device 600 generates the audio supplied from the music player 604a without placing the audio supplied from the music player 604a in the space using the left and right speakers (e.g., the device 600 does not apply any of the interaural time difference, HRTF, or cross-cancellation).

[0222] In some embodiments, devices 600 and 660 include a digital assistant that generates audio feedback, such as by uttering the results to return the results of a query. In some embodiments, devices 600 and 660 generate audio for the digital assistant by placing the audio for the digital assistant at a location in space (e.g., above the user's right shoulder), whereby the user perceives that the digital assistant remains stationary in space even when other audio moves within the space. In some embodiments, devices 600 and 660 emphasize the audio of the digital assistant by ducking one or more (or all) other audio.

[0223] Figures 7A - 7C are flow diagrams showing a method for transitioning between visual elements using an electronic device, according to some embodiments. Method 700 is executed on a device having a display (e.g., 100, 300, 500, 600, 660), and the electronic device is operably connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earphones, left and right earbuds). Some operations of method 700 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0224] As described below, method 700 provides an intuitive way to transition between visual elements. This method reduces the user's cognitive burden for transitioning between visual elements, thereby creating a more efficient human - machine interface. In the case of a battery - operated computing device, enabling the user to transition between visual elements more quickly and efficiently saves power and extends the time between battery charges.

[0225] The electronic device displays (702) a first visual element (e.g., 604a of a first application, a video playback window, album art) at a first position on the display.

[0226] The electronic device accesses (704) a first audio (e.g., 654a from a first source audio) corresponding to a first visual element (audio of a video within a playback window, audio of a song from an album corresponding to album art, audio generated by or from a first application).

[0227] According to some embodiments, the electronic device accesses (706) a second audio (e.g., 654b from a second source audio) corresponding to a second visual element (e.g., 604b).

[0228] While displaying (708) a first visual element (e.g., 604a) at a first position (e.g., a position on a display that is substantially horizontally centered on the display), the electronic device generates audio in a first mode (e.g., without modifying the first audio shown in FIG. 6A, without applying any of interaural time difference, HRTF, or cross cancellation to 654a) using the first audio (e.g., 654a) at two or more speakers.

[0229] According to some embodiments, the first mode is configured such that audio generated using the first mode is perceived by the user as being generated from a first direction (and optionally a position) corresponding to (e.g., aligned with) the display (712).

[0230] While displaying a first visual element (e.g., 604a) at a first position on a display (e.g., a position on the display that is substantially horizontally centered on the display) (708), the electronic device generates audio (e.g., a discrete audio output or a combined audio output including a second audio-based component) in two or more speakers using a second audio (e.g., 654b) in a third mode different from a first mode and a second mode (e.g., by applying an interaural time difference, HRTF, and / or cross cancellation to 654b shown in FIG. 6A). The third mode is configured such that the user perceives that the audio generated in the third mode is generated from a direction (and optionally a position) away from the display (e.g., an unaligned right side) (e.g., modifying the source audio before generating the audio such that the user perceives that the audio is generated from a direction to the right of the display or the user). In some embodiments, the first mode places the audio in front of the user or the display, the second mode places the audio to the left of the user or the display, and the third mode places the audio to the right of the user or the display.

[0231] While displaying a first visual element (e.g., 604a) at a first position on a display (e.g., a position on the display that is substantially horizontally centered on the display) (708), the electronic device refrains from displaying (716) on the display a second visual element (e.g., 604b in FIG. 6A, a second video playback window of a second application, a second album art) corresponding to a second audio (e.g., 654b from a second source audio) (e.g., audio of a video within a playback window, audio of a song from an album corresponding to album art, audio generated by a second application, or audio received from a second application).

[0232] While displaying a first visual element (e.g., 604a) at a first position on the display (e.g., a position on the display that is substantially horizontally centered on the display) (708), the electronic device receives a first user input (e.g., 610a, a swipe input on the touch sensing surface 600b) (718).

[0233] In response to receiving the first user input (e.g., 610a) (720), the electronic device transitions the display of the first visual element (e.g., 604a of FIG. 6B) from the first position on the display to a first visual element that is not displayed on the display (e.g., 604a of FIG. 6C) (e.g., by sliding) (e.g., by sliding off the edge (e.g., the left edge) of the display).

[0234] Further, in response to receiving a first user input (e.g., 610a) (720), while not displaying a first visual element (e.g., 604a in FIG. 6C) on the display, the electronic device uses a first audio (e.g., 654a) in a second mode different from the first mode to generate audio at two or more speakers (724). The second mode is configured such that the audio generated in the second mode is perceived by the user as being generated from a direction (and optionally a position) away from the display (e.g., the unaligned left side) (e.g., for 654a shown in FIG. 6C, the interaural time difference, HRTF, and / or cross-cancellation is applied to modify the source audio before generating the audio such that the user perceives the audio as being generated from a direction away from the display in a direction corresponding to the position where the first visual element was last displayed (e.g., to the left of the display or the user)). By generating audio related to content having changing characteristics, the user can visualize where the content is relative to the user without the need for the display of the content. Thereby, the user can quickly and easily recognize the input required to access the content (e.g., to cause the display of the content). Generating audio having changing characteristics also provides the user with context feedback regarding the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user-device interface more efficient, which in turn reduces power usage and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0235] In some embodiments, the volume of the audio generated using the first mode is greater than the volume of the audio generated using the second mode or the third mode. In some embodiments, the volume of the audio generated using the second mode is the same as the volume of the audio generated using the first mode.

[0236] In some embodiments, a frequency filter such as a low-pass filter, a high-pass filter, or a band-pass filter is not applied to the audio generated in the first mode (e.g., when the audio is perceived to be in the center, such as in front of the user). In some embodiments, a first frequency filter such as a low-pass filter, a high-pass filter, or a band-pass filter is applied to the audio generated using the first mode. In some embodiments, a second frequency filter such as a low-pass filter, a high-pass filter, or a band-pass filter is applied to the audio generated using the second mode. In some embodiments, a third frequency filter such as a low-pass filter, a high-pass filter, or a band-pass filter is applied to the audio generated using the third mode. In some embodiments, the second frequency filter is the same as the third frequency filter. In some embodiments, the frequency filter is applied to the audio generated using the second and third modes, but not to the audio generated using the first mode.

[0237] According to some embodiments, while displaying a first visual element (e.g., 604a in FIG. 6A, a video playback window of a first application, album art) at a first position on a display, an electronic device aligns with displaying a second visual element (e.g., 604b in FIG. 6A, a second video playback window of a second application, second album art) corresponding to a second audio (e.g., 654b from a second source audio) (e.g., audio of a video within a playback window, audio of a song from an album corresponding to album art, audio generated by a second application, or audio received from a second application) on the display. Not displaying content saves display space and enables the device to provide other visual feedback to the user on the display. Providing improved visual feedback to the user improves the operability of the device, (e.g., by assisting the user to provide appropriate input when operating / interacting with the device and reducing user errors), makes the user-device interface more efficient, and in addition, reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0238] According to some embodiments, in further response to receiving a first user input (e.g., 610a) (720), since the second visual element (e.g., 604b) is not displayed on the display, the electronic device transitions (726) the display of the second visual element to a fourth position on the display (e.g., 604b in FIG. 6C at the same position as the first position) (e.g., by sliding) (e.g., by sliding the second visual element onto the display from an edge (such as the right edge) of the display).

[0239] According to some embodiments, in further response to receiving a first user input (e.g., 610a) (720), while generating audio using a first audio (e.g., 654a in FIG. 6C) in a second mode, the electronic device generates audio at two or more speakers using a second audio (e.g., 654b in FIG. 6C) in a first mode (e.g., without modifying the second audio). In other examples, the second audio is generated in a mode different from the first mode.

[0240] According to some embodiments, the second mode is configured such that audio generated using the second mode is perceived by the user as being generated from a second direction (e.g., different from the first direction). According to some embodiments, the third mode is configured such that audio generated using the third mode is perceived by the user as being generated from a third direction different from the second direction (and optionally different from the first direction). Thus, optionally, the perceived positions of the sources of audio generated using different modes are different.

[0241] According to some embodiments, following the display of a first visual element (e.g., 604a in FIG. 6A) at a first position on the display (e.g., a position on the display that is substantially horizontally centered on the display), before the first visual element disappears from the display (e.g., 604a in FIG. 6C), the electronic device displays the first visual element (e.g., 604a in FIG. 6B) at a second position on the display (e.g., a position on the display that is to the left of the first position on the display, a position on the display that is not substantially horizontally centered on the display, a position on the display adjacent to an edge (e.g., the left edge) of the display). According to some embodiments, following the display of a first visual element (e.g., 604a in FIG. 6A) at a first position on the display (e.g., a position on the display that is substantially horizontally centered on the display), before the first visual element disappears from the display (e.g., 604a in FIG. 6C), while the first visual element (e.g., 604a in FIG. 6B) is being displayed at the second position on the display, the electronic device generates audio using a first audio (e.g., 654a in FIG. 6B) in a fourth mode that is different from the first mode, the second mode, and the third mode in two or more speakers. (e.g., modify the source audio before generating the audio such that the audio is perceived by the user as being generated from a direction different from the position where the audio was perceived when the audio was generated in the first mode, such as to the left of or near the display).

[0242] According to some embodiments, the first mode does not include modifying the interaural time difference of the audio. In some examples, the audio generated in the first mode (e.g., 654a in FIG. 6A) is not modified using any of HRTF, the interaural time difference of the audio, or cross-cancellation. In some examples, while the first visual element is being displayed, the electronic device uses an HRTF configured such that the first audio is perceived to be in front of the user, such as within a range of 10 degrees to the left and 10 degrees to the right of the direction the user is facing, to play the first audio. In some examples, while the first visual element is being displayed on the display, the electronic device uses an HRTF that modifies the interaural time difference of the first audio by less than a first predetermined amount (e.g., a minimal change) to generate the first audio.

[0243] According to some embodiments, the second mode includes modifying the interaural time difference of the audio. In some examples, the second mode includes a first degree of modification of the interaural time difference of the audio, and the second mode includes a second degree of modification of the interaural time difference of the audio that is greater than the first degree. In some examples, the third mode includes modifying the interaural time difference of the audio. In some examples, the audio generated in the second mode (e.g., 654b in FIG. 6C) is modified using one or more of HRTF, the interaural time difference of the audio, and cross cancellation. In some examples, the audio generated in the third mode (e.g., 654b in FIG. 6A) is modified using one or more of HRTF, the interaural time difference of the audio, and cross cancellation. By generating audio related to content having varying characteristics, the user can visualize where the content is relative to the user without the need for a display of the content. Thereby, the user can quickly and easily recognize the input required to access the content (e.g., to cause a display of the content). Generating audio having varying characteristics also provides the user with context feedback regarding the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user device interface more efficient, which in turn reduces power usage and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0244] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio of the audio (e.g., the right channel) and a second channel audio of the audio (e.g., the left channel) to form a composite channel audio, updating the second channel audio to include the composite channel audio with a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio with a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount).

[0245] In some examples, modifying the interaural time difference of the audio includes introducing a time delay into a first channel audio (e.g., the right channel) without introducing a time delay into a second channel audio different from the first channel audio (e.g., the left channel).

[0246] According to some embodiments, in response to beginning to receive (e.g., beginning to detect) a first user input (e.g., a swipe input), the electronic device transitions to generating audio using a second audio (e.g., audio of a video within a playback window, audio of a song from an album corresponding to album art, audio generated by or received from a second application) corresponding to a second visual element (e.g., a second video playback window of a second application, second album art) from two or more speakers, rather than generating audio using a second audio (e.g., from a second source audio) corresponding to the second visual element. For example, prior to receiving the first input, audio is not generated by the speakers using the second audio. When the start of the first input is detected (or after a portion of the first input is detected, or after the first input is detected), the device generates audio using the second audio corresponding to the second visual element in two or more speakers. In some examples, the second audio is an audio file (e.g., a song), and transitioning to generating audio using the second audio includes pausing generating audio using a first predetermined portion (e.g., the first 0.1 seconds) of the second audio. For example, this provides the effect that the audio was being played before it was audible to the user.

[0247] According to some embodiments, generating audio using the second mode (and optionally, the third mode) includes one or more of attenuating the audio, applying a high-pass filter to the audio, applying a low-pass filter to the audio, and changing the volume balance between two or more speakers. Optionally, generating audio is to use one or more of attenuating the audio, applying a high-pass filter to the audio, applying a low-pass filter to the audio, and changing the volume balance between two or more speakers. By generating audio related to content having changing characteristics, a user can visualize where the content is relative to the user without the need to view a display of the content. Thereby, the user can quickly and easily recognize the input required to access the content (e.g., to cause a display of the content). Generating audio having changing characteristics also provides context feedback to the user regarding the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user device interface more efficient, which in turn reduces power usage and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0248] In some embodiments, generating audio using the second mode (or the third and fourth modes) includes applying crosstalk cancellation techniques to the audio such that the audio is configured to be perceived by the user as coming from a particular direction. In some embodiments, generating audio using the third mode includes applying crosstalk cancellation techniques to the audio such that the audio is configured to be perceived as coming from a particular direction. In some embodiments, generating audio using the first mode does not include applying crosstalk cancellation techniques to the audio.

[0249] Note that the details of the processes described above with respect to method 700 (e.g., FIGS. 7A - 7C) are also applicable in a similar manner to the methods described later. For example, methods 900, 1200, 1400, and 1500 optionally include one or more of the various method characteristics described above with respect to method 700. For example, the same or similar techniques are used to place audio in space. In another embodiment, the same audio source can be used with various techniques. In yet another embodiment, the audio currently being played in each of the various methods can be manipulated using the techniques described in other methods. For the sake of brevity, these details are not repeated below.

[0250] FIGS. 8A - 8K illustrate exemplary techniques for previewing audio according to some embodiments. The techniques of these figures are used to illustrate the processes described later, including the processes of FIGS. 9A - 9C.

[0251] Figures 8A - 8K illustrate a device 800 (e.g., a mobile phone) having a display, a touch - sensing surface, and left and right speakers (e.g., headphones). The overhead view 850 is a visual representation of the spatial configuration of the audio being generated by the device 800, shown throughout Figures 8A - 8K, and provides the reader with a better understanding of the technique, particularly with respect to the position where the sound is arriving as perceived by the user 856 (e.g., as a result of the device 800 positioning the audio within the space). The overhead view 850 is not part of the user interface of the device 800. In some embodiments, with the techniques described below, the user can more easily and efficiently preview and select music for playback.

[0252] In Figure 8A, the device 800 displays a music player 802 for playing music. The music player 802 includes a plurality of affordances 804 each corresponding to a different song that can be played using the device 800. As shown in the overhead view 850, the device 800 is not playing any audio.

[0253] In Figure 8B, the device 800 detects a tap - and - hold input 810a on the affordance 804a for track 3. As shown in the overhead view 850, the device 800 is not playing any audio.

[0254] In Figure 8C, in response to detecting the input 810a on the affordance 804a and while continuing to detect the input 810a on the affordance 804a, the device 800 updates the plurality of affordances 804 to distinguish the affordance 804a for track 3. In the embodiment of Figure 8C, the device 800 blurs the plurality of affordances 804 other than the selected affordance 804a for track 3. This indicates to the user that the song corresponding to the affordance 804a is being previewed.

[0255] For example, the preview is limited to a predetermined audio playback period (e.g., shorter than the duration of a song). After the predetermined audio playback period has reached during the playback of a song, the device stops playing that song. In some embodiments, after stopping the playback of a song, the device proceeds to provide a preview of a different song. In some embodiments, after stopping the playback of a song, the device proceeds to provide another preview of the same song (e.g., loop the same portion of the song).

[0256] As shown in the top-down view 850 of FIG. 8C, in response to detecting the input 810a on the affordance 804a and while continuing to detect the input 810a on the affordance 804a, the device 800 uses two or more speakers to generate a preview of the audio of track 3 by placing the audio of track 3 within the space (e.g., the device 800 applies the interaural time difference, HRTF, and / or cross-cancellation to the music). The device 800 places the audio of track 3 such that the user perceives the music as coming from a position within the space to the left of the user 856 (or the device 800), as indicated by the audio element 850a of the top-down view 850 of FIG. 8C. In FIGS. 8C-8D, the device 800 continues to detect the input 810a on the affordance 804a and updates the position within the space of the audio of track 3 that the user perceives such that the user perceives the music as moving towards the user.

[0257] In FIG. 8D, the device 800 continues to generate the audio of track 3 but stops placing the audio within the space, such that as a result, the user perceives the audio as being within the user's head, as indicated by the audio element 850a of the top-down view 850 of FIG. 8D. For example, in FIG. 8D, the device 800 generates the audio of track 3 using the left and right speakers without placing the audio (e.g., the device 800 does not apply any of the interaural time difference, HRTF, or cross-cancellation).

[0258] Device 800 continues to play a preview of track 3 for the user until (1) a predetermined audio playback duration is reached, (2) the device detects a movement of input 810a on the touch sensing surface to affordances corresponding to different songs, or (3) the device detects a lift-off of input 810a.

[0259] In Figure 8E, device 800 detects the movement of input 810a from affordance 804a to affordance 804b (without detecting the lift-off of input 810a). It should be noted that device 800 does not scroll (or move in another way) the plurality of affordances 804 in response to detecting the movement of input 810a. In response to detecting the input at affordance 804b, device 800 updates the visual aspect and the spatial arrangement of the audio of the device. Device 800 blurs the affordance 804a of track 3 and stops blurring the affordance 804b of track 4. Device 800 also transitions to generating the audio of track 3 by arranging the audio in space using the left and right speakers (e.g., device 800 applies the interaural time difference, HRTF, and / or cross-cancellation), whereby the user perceives that the audio is moving away from the user's head and moving away to the user's right as indicated by the audio element 850a in the overhead view 850 of Figures 8E - 8F. Device 800 also optionally starts attenuating the audio of track 3 and then stops generating the audio of track 3 as indicated by the audio element 850a which is no longer shown in the overhead view 850 of Figure 8G. Further, in response to detecting the input at affordance 804b, device 800 generates a preview of the audio of track 4 by arranging the audio in space (e.g., device 800 applies the interaural time difference, HRTF, and / or cross-cancellation), whereby the user perceives that the audio of track 4 is coming from the user's left and moving into the user's head as indicated by the audio element 850b in the overhead view 850 of Figures 8E - 8G. As described above, the preview is limited to a predetermined audio playback period (e.g., shorter than the duration of the song). After reaching the predetermined audio playback period, the device stops playing the song.

[0260] In FIGS. 8F - 8G, device 800 generates the audio of track 4 using the left and right speakers without placing the audio (e.g., device 800 does not apply any interaural time difference, HRTF, or cross - cancellation). For example, as a result of not placing the audio in space, a user of device 800 wearing headphones perceives the audio to be inside the user's head, as indicated by the audio element 850b in the overhead view 850 of FIGS. 8F - 8G.

[0261] In FIG. 8H, device 800 detects the lift - off of input 810a. In response to detecting the lift - off of input 810a, device 800 transitions to generating the audio of track 4 by placing the audio in space using the left and right speakers (e.g., device 800 applies an interaural time difference, HRTF, and / or cross - cancellation), whereby the user perceives the audio to be moving away from the user's head and to the right of the user, as indicated by the audio element 850b in the overhead view 850 of FIG. 8H. In response to detecting the lift - off of input 810a, device 800 stops the blurring of a plurality of affordances, as shown in FIG. 8H. Further, in response to detecting the lift - off of input 810a, device 800 optionally starts attenuating the audio of track 4 and then stops generating the audio of track 4, as indicated by the audio element 850b that is no longer shown in the overhead view 850 of FIG. 8I.

[0262] In FIG. 8J, device 800 detects a tap input 810b on affordance 804c of track 6. In FIG. 8K, as shown in the overhead view 850, device 800 begins to generate the audio of track 6 without placing the audio of track 6 in space, as indicated by audio element 850c of the overhead view 850 of FIG. 8J, using two or more speakers. As a result, the user perceives the audio as being within the user's head. For example, in FIG. 8J, device 800 generates the stereo audio of track 6 using the left and right speakers without placing the audio (e.g., device 800 does not apply any interaural time difference, HRTF, or cross cancellation). In response to detecting the tap input 810b, device 800 does not blur any of the plurality of affordances 804. Device 800 also optionally updates affordance 804c to include media control and / or additional information regarding the track. Track 6 continues to play until it reaches the end of the track without being limited to a preview of a predetermined audio playback duration.

[0263] In some embodiments, device 800 includes a digital assistant that generates audio feedback, such as returning the result of a query by uttering the result. In some embodiments, device 800 generates the audio of the digital assistant by placing the audio of the digital assistant at a location in space (e.g., above the user's right shoulder), such that the user perceives the digital assistant as remaining stationary in space even when other audio moves within the space. In some embodiments, device 800 emphasizes the audio of the digital assistant by ducking one or more (or all) other audio.

[0264] Figures 9A - 9C are flowcharts showing a method for previewing audio using an electronic device according to some embodiments. Method 900 is executed on a device (e.g., 100, 300, 500, 800) having a display and a touch - sensing surface. The electronic device is operably connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earphones, left and right earbuds). Some operations of method 900 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0265] As described below, method 900 provides an intuitive method for previewing audio. This method reduces the user's cognitive burden for previewing audio, thereby creating a more efficient human - machine interface. In the case of a battery - operated computing device, enabling the user to preview audio faster and more efficiently saves power and extends the time between battery charges.

[0266] The electronic device displays (902) a list (e.g., in a column) of a plurality of media elements (e.g., 804) on the display. Each media element of the plurality of media elements (e.g., 804a - 804c) (or at least two media elements) corresponds to a respective media file (e.g., an audio file, a song, a video file). In some examples, all of the respective media files are different from each other.

[0267] The electronic device uses the touch - sensing surface to detect (904) a user contact (e.g., a touch - and - hold user input of 810a, a touch input that persists for longer than a predetermined period exceeding 0 seconds) at a position corresponding to a first media element (e.g., 804a).

[0268] In response to detecting user contact (e.g., 810a in FIGS. 8B - 8C) at a position corresponding to the first media element (e.g., 804a) (906), and in accordance with the user contact (e.g., 810a) including a touch - and - hold input (as determined by the electronic device), the electronic device uses two or more speakers to generate audio (e.g., 850a) using a first audio file corresponding to the first media element without exceeding a predetermined audio playback period. For example, the audio is played for a maximum of 5 seconds. For example, the device generates a preview of the audio. In some embodiments, the predetermined audio playback period is shorter than the duration of the audio file.

[0269] Furthermore, in response to detecting user contact (e.g., 810a in FIGS. 8B - 8C) at a position corresponding to the first media element (e.g., 804a) (906), and in accordance with the user contact (e.g., 810a) including a touch - and - hold input (as determined by the electronic device), and while the user contact (e.g., 810a) remains at the position corresponding to the first media element (e.g., 804a) (without a lift - off event) (910), and in accordance with not exceeding a predetermined audio playback period, the electronic device continues to generate audio (e.g., 850a) using the first audio file using two or more speakers (912).

[0270] Further, in response to detecting user contact (e.g., 810a in FIGS. 8B - 8C) at a position corresponding to the first media element (e.g., 804a) (906), and in accordance with the user contact (e.g., 810a) including a touch - and - hold input (as determined by the electronic device), and while the user contact (e.g., 810a) remains at the position corresponding to the first media element (e.g., 804a) (without a lift - off event) (910), and in accordance with exceeding a predetermined audio playback period, the electronic device stops generating audio (e.g., 850a) using the first audio file (and optionally, without starting to generate audio using a different audio file) using two or more speakers (914).

[0271] The electronic device uses a touch - sensing surface to detect movement of user contact (e.g., 810a in FIGS. 8D - 8E, moving from the position corresponding to the first media element (e.g., 804a) to the position corresponding to the second media element (e.g., 804b), away from the top of the display towards the bottom of the display) (916).

[0272] In response to detecting user contact (e.g., 810a in FIG. 8F) at a position corresponding to the second media element (918), and in accordance with the user contact (e.g., 810a) including a touch - and - hold input, the electronic device generates audio (e.g., 850b) using two or more speakers, without exceeding a predetermined audio playback period, using a second audio file (different from the first audio file) corresponding to the second media element (e.g., 804b) (920). For example, the audio plays for a maximum of 5 seconds (less than the total duration of the second audio file).

[0273] Further, in response to detecting a user contact (e.g., 810a in FIG. 8F) at a position corresponding to the second media element (918), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input, and while the user contact remains (without a lift-off event) at the position corresponding to the second media element (e.g., 804b) (922), and in accordance with not exceeding a predetermined audio playback period, the electronic device continues to generate audio (e.g., 850b) using the second audio file using two or more speakers (924).

[0274] Further, in response to detecting a user contact (e.g., 810a in FIG. 8F) at a position corresponding to the second media element (918), and in accordance with the user contact (e.g., 810a) including a touch-and-hold input (as determined by the electronic device), and while the user contact remains at the position corresponding to the second media element (e.g., 804b) (without a lift-off event), and in accordance with exceeding a predetermined audio playback period, the electronic device stops generating audio (e.g., 850b) using the second audio file using two or more speakers (and, optionally, without starting to generate audio using a different audio file) (926).

[0275] The electronic device uses the touch-sensitive surface to detect the lift-off of a user contact (e.g., 810a in FIGS. 8G and 8H) (928).

[0276] In response to detecting the lift-off of the user contact (930), the electronic device stops generating audio (e.g., 850a and 850b in FIG. 8I) using the first audio file or the second audio file using two or more speakers (and, optionally, stops generating any audio using audio files corresponding to a plurality of media elements) (932). Accordingly, the playback of any song being previewed stops.

[0277] According to some embodiments, further in accordance with (906) user contact including touch-and-hold input (e.g., as determined by an electronic device), generating audio using a first audio file without exceeding a predetermined audio playback period includes transitioning the audio among a plurality of modes including a first mode, a second mode different from the first mode (e.g., 850a of FIG. 8D, 850b of FIG. 8F), and a third mode different from the first mode and the second mode. Generating audio having varying characteristics also provides the user with context feedback regarding the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user-device interface more efficient, which in turn reduces power usage and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0278] According to some embodiments, the first mode is configured such that audio generated using the first mode is perceived by the user as being generated from a first direction (e.g., 850a of FIG. 8C, 850b of FIG. 8E) (and optionally location). In some examples, the first direction is a direction originating from a location to the left of the user (e.g., not in front of the user). In some examples, the electronic device modifies the source audio before generating the audio such that the audio is perceived by the user as being generated from a direction originating from a location to the left of the display or the user.

[0279] According to some embodiments, the second mode is configured such that audio generated using the second mode is perceived by the user as being within the user's head (e.g., stereo not a point source). For example, the second mode does not include applying any of HRTF, interaural time difference of the audio, or cross cancellation.

[0280] According to some embodiments, the third mode is configured such that audio generated using the third mode is perceived by the user as being generated from a third direction (e.g., 850a in FIG. 8F, 850b in FIG. 8H) different from the first direction. In some examples, the third direction is a direction originating from a position on the right of the user (e.g., not in front of the user). In some examples, the electronic device modifies the source audio before generating the audio such that the user perceives the audio as being generated from a direction originating from a position on the right of the display or the user.

[0281] According to some embodiments, the first mode includes a first modification of the interaural time difference of the audio, and the third mode includes a second modification of the interaural time difference of the audio different from the first modification. In some examples, the first mode includes a first degree of modification of the interaural time difference of the audio, and the third mode includes a second degree of modification of the interaural time difference of the audio greater than the first degree. In some examples, the first mode includes a modification of the interaural time difference of the audio such that the audio is perceived by the user as being generated from a direction originating from a position on the left of the display or the user. In some examples, the third mode includes a modification of the interaural time difference of the audio such that the audio is perceived by the user as being generated from a direction originating from a position on the right of the display or the user. By modifying the interaural time difference of the audio, the user can perceive the audio as coming from a specific direction. This direction information of the audio provides additional feedback to the user about the arrangement of different contents. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate inputs), makes the user-device interface more efficient, which in turn reduces the power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0282] In some embodiments, the audio generated in the first mode (e.g., 850a in FIG. 8C) is modified using one or more of HRTF, the interaural time difference of the audio, and cross-cancellation. In some examples, the audio generated in the second mode (e.g., 850a in FIG. 8D) is not modified using any of HRTF, the interaural time difference of the audio, or cross-cancellation. In some examples, the audio generated in the third mode (e.g., 850a in FIG. 8F) is modified using one or more of HRTF, the interaural time difference of the audio, and cross-cancellation.

[0283] According to some embodiments, the second mode does not include modification of the interaural time difference of the audio. In some examples, the audio generated in the second mode is not modified using HRTF. In some examples, the audio generated in the second mode (e.g., 850a in FIG. 8D) is not modified using any of HRTF, the interaural time difference of the audio, or cross-cancellation.

[0284] According to some embodiments, modifying the interaural time difference of audio includes combining a first channel audio of the audio (e.g., right channel) and a second channel audio of the audio (e.g., left channel) to form a composite channel audio, updating the second channel audio to include the composite channel audio with a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio with a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of audio includes introducing a time delay into a first channel audio (e.g., right channel) without introducing a time delay into a second channel audio different from the first channel audio (e.g., left channel). Modifying the interaural time difference of audio enables a user to perceive the audio as coming from a specific direction This direction information of the audio provides additional feedback to the user about the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user device interface more efficient, which in turn reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0285] According to some embodiments, in further response to detecting user contact (e.g., 810a) at a location corresponding to a first media element (e.g., 804a), the first visual characteristic (e.g., focus / blur level) of the displayed media elements within a plurality of media elements other than the first media element (e.g., 804b in FIG. 8C) is changed. In response to detecting user contact (e.g., 810a) at a location corresponding to a second media element (e.g., 804ba), the change in the first visual characteristic of the second media element (e.g., 804b in FIG. 8E) is reverted, and the first visual characteristic of the first media element (e.g., 804a in FIG. 8E) is changed. In some examples, the electronic device fades out media elements that are not activated. In some examples, the electronic device adds a blur effect to media elements that are not activated. In some examples, the electronic device changes the color of media elements that are not activated. By visually differentiating the content being played from the content not being played, feedback regarding the state of the device is provided to the user. By providing improved visual feedback to the user, the operability of the device is improved, (e.g., assisting the user to provide appropriate input when operating / interacting with the device and reducing user errors), making the user device interface more efficient, and in addition, reducing power usage and improving the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0286] According to some embodiments, in response to detecting user contact at a position corresponding to a first media element, the electronic device changes a second visual characteristic of the first media element without changing a second visual characteristic of the displayed media elements within a plurality of media elements other than the first media element. In response to detecting user contact at a position corresponding to a second media element, the electronic device reverts the change of the second visual characteristic of the first media element and changes the second visual characteristic of the second media element. In some examples, the electronic device emphasizes the activated media element. In some examples, the electronic device brightens the activated media element. In some examples, the electronic device changes the color of the activated media element. By visually differentiating the content being played from the content not being played, feedback regarding the state of the device is provided to the user. By providing improved visual feedback to the user, the operability of the device is enhanced, (e.g., by assisting the user to provide appropriate input when operating / interacting with the device and reducing user errors), making the user device interface more efficient, and in addition, reducing power usage and improving the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0287] According to some embodiments, in response to detecting user contact at a position corresponding to a first media element and in accordance with the user contact including a tap input (such as determined by the electronic device), the electronic device generates audio using a first audio file corresponding to the first media element without automatically stopping generating audio using the first audio after a predetermined audio playback period and without automatically stopping generating audio using the first audio file when detecting a lift-off of the user contact. For example, the audio plays for longer than a 5-second preview time.

[0288] According to some embodiments, generating audio using a first audio file in accordance with user contact including a tap input (such as as determined by an electronic device) includes generating audio using the first audio file in a second mode without transitioning the audio between a first mode and a third mode of a plurality of modes.

[0289] Note that the details of the processes described above with respect to method 900 (e.g., FIGS. 9A - 9C) are also applicable in a similar manner to the methods described below and above. For example, methods 700, 1200, 1400, and 1500 optionally include one or more of the various method characteristics described above with respect to method 900. For example, the same or similar techniques are used to place audio in space. In another example, the same audio source can be used with various techniques. In yet another example, the audio currently being played in each of the various methods can be manipulated using the techniques described in other methods. For the sake of brevity, these details are not repeated below.

[0290] FIGS. 10A - 10K illustrate exemplary techniques for discovering music according to some embodiments. The techniques of these figures are used to illustrate the processes described below, including the processes of FIGS. 12A - 12B.

[0291] Figures 10A - 10K illustrate a device 1000 (e.g., a mobile phone) having a display, a touch - sensing surface, and left and right speakers (e.g., headphones). The overhead view 1050 is a visual representation of the spatial configuration of the audio generated by the device 1000, shown throughout Figures 10A - 10K, and provides a better understanding of the technique with respect to the position where the user 1056 perceives the sound to be arriving (e.g., as a result of the device 1000 positioning the audio within the space). The overhead view 1050 is not part of the user interface of the device 1000. Similarly, visual elements displayed outside the display device, as represented by the dashed - line contour, are not part of the displayed user interface but are illustrated to provide a better understanding of the technique to the reader. In some embodiments, with the techniques described below, a user can more easily and efficiently discover new songs from a song repository.

[0292] In Figure 10A, the device 1000 displays a music player 1002 that includes an affordance 1002a for discovering new audio content. In Figure 10A, the device 1000 is not generating any audio, as indicated by the lack of audio elements in the overhead view 1050 of Figure 10A. While the device 1000 is not generating any audio, the device 1000 detects a tap input 1010a on the affordance 1002a. In response to detecting the tap input 1010a on the affordance 1002a, the device 1000 enters a discovery mode.

[0293] In Figure 10B, the device 1000 replaces the display of the affordance 1002a with the display of album art 1004a. The album art 1004a corresponds to a first song (e.g., the album art is from the album to which the song belongs), and the album art 1004b corresponds to a second song different from the first song. The audio element 1054a corresponds to the first song, and the audio element 1054b corresponds to the second song.

[0294] In FIG. 10B, device 1000 generates the audio of the first song and the second song by arranging the audio of the first song and the second song along path 1050a in the space using the left speaker and the right speaker (e.g., device 1000 applies interaural time difference, HRTF, and / or cross cancellation). For example, path 1050a is a fixed curved path along which the device arranges the audio of various songs while in discovery mode. Device 100 generates the audio of the first song and the second song and updates the arrangement of the audio so that the user perceives the audio of the song drifting from left to right in front of the user along path 1050a as shown by audio element 1054a and audio element 1054b of the overhead view 1050 in FIGS. 10B - 10G. In some embodiments, the direction in which the song moves along path 1050a is not based on the input provided by the user (e.g., not based on tap input 1010a).

[0295] In FIG. 10C, as the first song and the second song progress along path 1050a, device 1000 begins to generate the audio of the third song by arranging and transitioning the audio in the space using the left speaker and the right speaker (e.g., device 1000 applies interaural time difference, HRTF, and / or cross cancellation), whereby the user perceives the audio drifting from left to right in front of the user along path 1050a as shown by audio element 1054c of the overhead view 1050 in FIGS. 10C - 10G.

[0296] In FIGS. 10D - 10F, device 1000 begins to generate additional song audio (while continuously generating the audio of the first and second songs) by using the left and right speakers to place and transition the respective song audio within the space (e.g., device 1000 applies interaural time differences, HRTFs, and / or cross - cancellation), such that the user perceives the audio as drifting from left to right in front of the user along path 1050a, as shown by audio elements 1054a - 1054e in the overhead view 1050 of FIGS. 10C - 10G. In some embodiments, device 1000 places the songs at equal distances along path 1050a (e.g., perceived as being 2 to 3 meters away from the user). In some embodiments, the distance between the songs changes (e.g., increases) as the songs approach a point on path 1050a (e.g., a point in front of the user), and the distance between the songs changes (e.g., decreases) as the songs move away from a point on path 1050a. In some embodiments, each song moves at the same speed along path 1050a. In some embodiments, device 1000 stops generating songs that reach a particular point along path 1050a (e.g., they move beyond a threshold distance from the user). In some embodiments, device 1000 attenuates the songs based on their position along path 1050a (e.g., the songs are attenuated as they move farther from the user).

[0297] As shown in FIGS. 10B - 10I, album arts 1004a - 1004e correspond to various songs. Device 1000 displays the respective album art for the songs placed in front of the user. For example, device 1000 displays the album art corresponding to the songs placed along a particular subset of path 1050a. The displayed album art moves in the same direction on the display as the various audio moves along path 1050a (e.g., from left to right).

[0298] The first song (visualized as audio element 1054a within the overview view 1050) corresponds to the album art 1004a (e.g., the album art is that of the album to which the song belongs). The second song (visualized as audio element 1054b within the overview view 1050) corresponds to the album art 1004b. The third song (visualized as audio element 1054c within the overview view 1050) corresponds to the album art 1004c. The fourth song (visualized as audio element 1054d within the overview view 1050) corresponds to the album art 1004d. The fifth song (visualized as audio element 1054e within the overview view 1050) corresponds to the album art 1004e.

[0299] In FIGS. 10G - 10H, device 1000 detects a left - swipe input 1010b. In response to detecting the left - swipe input 1010b, device 1000 moves the album art across the display and changes the direction corresponding to the left - swipe input 1010b, and moves various audio along path 1050a and changes the direction corresponding to the left - swipe input 1010b. In FIG. 10G, when the user places their finger on the touch - sensitive surface, album arts 1004d and 1004c stop moving on the display, and the audio corresponding to audio elements 1054a - 1054e stops moving along path 1050a. In FIG. 10H, when device 1000 detects a left - swipe input 1010b from right to left, device 100 updates the display so that album arts 1004d and 1004c moving from right to left on the display and the audio corresponding to audio elements 1054a - 1054e move from right to left along path 1050a. In some embodiments, the speed at which album arts 1004a - 1004e move across the display and the speed at which the audio corresponding to audio elements 1054a - 1054e move along path 1050a are based on one or more characteristics of the left - swipe input 1010b (e.g., length, speed, characteristic strength). In some embodiments, a faster swipe results in faster movement of the album art on the display and the audio along path 1050a. In some embodiments, a longer swipe results in faster movement of the album art on the display and the audio along path 1050a.

[0300] In FIG. 10I, after device 1000 stops detecting the left - swipe input 1010b, device 1000 continues to display album arts 1004a - 1004e moving across the display from left to right, and the audio corresponding to audio elements 1054a - 1054e continues to move along path 1050a from left to right. Thus, the left - swipe input 1010b changes the direction in which the user perceives the audio to move along path 1050a and changes the direction in which device 1000 moves the corresponding album art on the display.

[0301] In FIG. 10J, device 1000 detects a tap input 1010c on album art 1004b. In response to detecting the tap input 1010c on album art 1004b, as shown in FIGS. 10J-10K, device 1000 transitions to generating audio of a second song corresponding to album art 1004b without placing the audio using the left and right speakers (e.g., device 1000 does not apply any of interaural time difference, HRTF, or cross cancellation). For example, as a result of not placing audio in space, a user of device 1000 wearing headphones perceives that the audio of the second song is within the user's head, as indicated by audio element 1054b of the overhead view 1050 of FIG. 10K. Further, in response to detecting the tap input 1010c on album art 1004b, as shown in FIGS. 10J-10K, device 1000 moves the audio corresponding to audio elements 1054c and 1054a in the opposite direction away from the user before stopping generating the audio corresponding to audio elements 1054c and 1054a. Thus, the user perceives that the audio of the unselected song is floating.

[0302] In some embodiments, device 1000 includes a digital assistant that generates audio feedback, such as by uttering a result to return the result of a query. In some embodiments, device 1000 generates the audio of the digital assistant by placing the audio of the digital assistant at a location in space (e.g., above the user's right shoulder), such that the user perceives that the digital assistant remains stationary in space even when other audio moves within the space. In some embodiments, device 1000 emphasizes the audio of the digital assistant by ducking one or more (or all) other audio.

[0303] Figures 11A - 11G illustrate exemplary techniques for discovering music according to some embodiments. Figures 11A - 11G illustrate an exemplary user interface for display by a device (e.g., a laptop) having a display, a touch - sensing surface (e.g., 1100), and left and right speakers (e.g., headphones). The touch - sensing surface 1100 of the device is illustrated to provide the reader with a better understanding of the techniques described particularly with respect to exemplary user input. The touch - sensing surface 1100 is not part of the displayed user interface of the device. The overhead view 1150 illustrated throughout Figures 11A - 11G is displayed by the device. The overhead view 1150 is also a visual representation of the spatial configuration of the audio being generated by the device, and provides the reader with a better understanding of the technique, particularly with respect to the position where the user 1106 perceives the sound as arriving (e.g., as a result of the device placing the audio in space). The arrows of the audio elements 1150a - 1150g indicate the direction and speed in which the audio elements 1150a - 1150g are moving (including the visual display and as perceived by the user through the speakers), and provide the reader with a better understanding of the technique. In some embodiments, the displayed user interface does not include the arrows of the audio elements 1150a - 1150g. Similarly, the visual elements displayed outside the display of the device are not part of the displayed user interface and are illustrated to provide the reader with a better understanding of the technique. In some embodiments, by the techniques described below, a user can more easily and efficiently discover new songs from a song repository.

[0304] Figures 11A - 11D show a device that displays a user representation 1106 in a space. The device also illustrates audio elements 1150a - 1150f corresponding to songs 1 - 6 respectively in the same space. The device uses left and right speakers to simultaneously generate audio for each of songs 1 - 6 by placing and transitioning individual songs in the space (e.g., the device applies interaural time differences, HRTFs, and / or cross - cancellation), such that the user perceives the audio of the songs to drift past the user as shown by the audio elements 1150a - 1150f in the overhead views 1150 of Figures 11A - 11D. In some embodiments, the audio elements are displayed to be equidistant from each other. In some embodiments, the generated audio is arranged such that the user perceives the song sources to be equidistant from each other. In Figure 6A, for example, the user perceives song 1 (corresponding to audio element 1150a) and song 2 (corresponding to audio element 1150b) to be substantially in front of the user, song 6 (corresponding to audio element 1150f) to be substantially to the right of the user, and song 4 (corresponding to audio element 1150d) to be substantially to the left of the user. As the device continues to generate the songs, the user perceives the audio sources to be passing by. For example, in Figure 11D, the user perceives song 6 to be substantially behind the user. Thereby, the user can listen to various songs simultaneously.

[0305] In FIGS. 11D - 11E, the device receives a swipe input 1110 at the touch - sensing surface 1100. In response to receiving the swipe input 1110, as shown in FIGS. 11E - 11F, the device updates, on the display, the direction and / or speed of the displayed visual elements 1150a - 1150h (e.g., according to the direction and / or speed of the swipe input 1110). Further, in response to receiving the swipe input 1110, the device transitions the audio generated for individual songs in the space (e.g., the device applies inter - aural time differences, HRTF, and / or cross - cancellation), such that the user perceives that the audio of the song is drifting past the user at the updated direction and / or speed (e.g., according to the direction and / or speed of the swipe input 1110).

[0306] In some embodiments, the device includes a digital assistant that generates audio feedback, such as by uttering the results to return the results of a query. In some embodiments, the device generates the audio for the digital assistant by placing the audio for the digital assistant at a location in the space (e.g., above the user's right shoulder), such that the user perceives that the digital assistant remains stationary in the space even when other audio moves within the space. In some embodiments, the device emphasizes the audio of the digital assistant by ducking one or more (or all) of the other audio.

[0307] In some embodiments, the device detects a tap input on a displayed audio element (e.g., at a corresponding location on a touch-sensitive surface). In response to detecting a tap input on the audio element, the device transitions to generating the audio of each respective song corresponding to the selected audio element without placing the audio in space using the left and right speakers (e.g., device 1000 does not apply any interaural time differences, HRTF, or cross-cancellation). For example, as a result of not placing the audio in space, a user wearing headphones perceives the audio of the selected song as being within the user's head. Further, in response to detecting a tap input on the displayed audio element, the device stops generating the audio of the remaining audio elements. Thus, the user is enabled to select individual songs to listen to.

[0308] Figures 12A - 12B are flow diagrams showing a method for discovering music using an electronic device, according to some embodiments. Method 1200 is executed on a device (e.g., 100, 300, 500, 1000) having a display and a touch-sensitive surface. The electronic device is operatively connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earphones, left and right earbuds). Some operations of method 1200 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0309] As described below, method 1200 provides an intuitive method for discovering music. This method reduces the cognitive burden on the user for discovering music, thereby creating a more efficient human-machine interface. In the case of a battery-operated computing device, enabling the user to discover music faster and more efficiently saves power and extends the time between battery charges.

[0310] The electronic device detects (1202) a first user input (e.g., 1010a) to activate the discovery mode (e.g., by tapping a "discover" affordance to provide a voice input to the digital assistant such as "sample some songs").

[0311] In response to detecting (1204) a first user input (e.g., 1010a) to activate the discovery mode, the electronic device uses two or more speakers to simultaneously generate (1204) audio (e.g., 1054a, 1054b, 1054c) from a first audio source (e.g., an audio file, a music file, a media file) in a first mode (1206), a second audio source (e.g., an audio file, a music file, a media file) in a second mode (1210), and a third audio source (e.g., an audio file, a music file, a media file) in a third mode (1214).

[0312] The first mode (1206) is configured such that the user perceives that the audio generated using the first mode is generated from a first point in a space where the audio moves in a first direction (e.g., from left to right) along a predetermined path (e.g., 1050a) at a first speed over time.

[0313] According to some embodiments, the first audio source (e.g., an audio file, a music file, a media file) corresponds to (1208) a first visual element (e.g., the audio of a video within a playback window of 1004a, the audio of a song from an album corresponding to album art, audio generated by or received from a first application).

[0314] The second mode (1210) is configured such that the user perceives that the audio generated using the second mode is generated from a second point in a space where the audio moves over time in a first direction (e.g., from left to right) along a predefined path (e.g., 1050a) at a second speed.

[0315] According to some embodiments, a second audio source (e.g., an audio file, a music file, a media file) corresponds to a second visual element (e.g., the audio of a video within a playback window of 1004b, the audio of a song from an album corresponding to album art, audio generated by or received from a first application) (1212).

[0316] The third mode (1214) is configured such that the user perceives that the audio generated using the third mode is generated from a third point in a space where the audio moves in a first direction (e.g., from left to right) along a predefined path (e.g., 1050a) at a third speed.

[0317] According to some embodiments, a third audio source (e.g., an audio file, a music file, a media file) corresponds to a third visual element (e.g., the audio of a video within a playback window of 1004c, the audio of a song from an album corresponding to album art, audio generated by or received from a first application) (1216).

[0318] The first point (e.g., the position of 1054a in FIG. 10C), the second point (e.g., the position of 1054b in FIG. 10C), and the third point (e.g., the position of 1054c in FIG. 10C) are different points in space (e.g., different points in space perceived by the user) (1218). Generating audio with varying characteristics also provides the user with context feedback regarding the arrangement of different content, enabling the user to quickly and easily recognize what input is required to access the content (e.g., to cause the display of the content). For example, within that space, the user can recognize that a particular audio can access a particular input based on the location where the user can perceive the incoming audio. Providing improved audio feedback to the user enhances the operability of the device and makes the user-device interface more efficient (e.g., by assisting the user in providing the appropriate input), which in turn reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0319] According to some embodiments, while generating audio using a first audio source, a second audio source, and a third audio source, the electronic device displays, on a display, the simultaneous movement of two or more of a first visual element (e.g., 1004a), a second visual element (e.g., 1004b), and a third visual element (e.g., 1004c) at a fourth speed (1220). In some examples, the device displays the simultaneous movement of all of a first visual element (e.g., 1004a), a second visual element (e.g., 1004b), and a third visual element (e.g., 1004c) at a fourth speed.

[0320] According to some embodiments, the movement of two or more of a first visual element (e.g., 1004a), a second visual element (e.g., 1004b), and a third visual element (e.g., 1004c) is in a first direction (e.g., from left to right).

[0321] According to some embodiments, a given path (e.g., 1050a) varies along a first dimension (e.g., the x-dimension, left / right dimension). The given path varies along a second dimension (e.g., the z-dimension, near / far dimension) that is different from the first dimension. The given path does not vary along a third dimension (e.g., the y-dimension, up / down dimension, height) that is different from the first and second dimensions. Thus, in some examples, the device generates audio along the path such that the user perceives the audio as moving from left to right and / or from right to left, and also from far to near or from near to far, but does not perceive the audio as moving up and down or down and up.

[0322] According to some embodiments, using two or more speakers to simultaneously generate audio (1204) further includes using a fourth audio source (e.g., an audio file, a music file, a media file) in a fourth mode. The fourth mode is configured such that the user perceives that the audio (e.g., 1054d) generated using the fourth mode is generated from a fourth point (e.g., the position of 1054a) in a space that moves along a first direction (e.g., the direction from left to right) along a given path (e.g., 1050a) over time, and the fourth point in the space is farther from the user than the first point, the second point, and the third point. In some examples, visual elements of audio that are perceived as being at a distance (e.g., farther than a predetermined distance) are not displayed on the display (e.g., 1004a in FIG. 10E is not displayed on the display). While the display is displaying the simultaneous movement of two or more of the first visual element (e.g., 1004a), the second visual element (e.g., 1004b), and the third visual element (e.g., 1004c), the electronic device refrains from displaying a fourth visual element (e.g., 1004d) corresponding to the fourth audio source on the display.

[0323] According to some embodiments, the first speed, the second speed, and the third speed are the same speed. According to some embodiments, the first speed, the second speed, and the third speed are different speeds.

[0324] According to some embodiments, while using two or more speakers to simultaneously generate audio (e.g., as shown at 1050 in FIG. 10F), the electronic device detects a second user input (e.g., 1010b) in a second direction different from the first direction (e.g., a direction from left to right). In response to detecting the second user input (e.g., 1010b) in the second direction, the electronic device uses a mode configured such that the user perceives that the audio source is moving over time in the second direction (e.g., a direction from right to left as shown at 1050 in FIG. 10H) along a predefined path (e.g., 1050a) to update the generation of the audio of the first audio source, the second audio source, and the third audio source. Further, in response to detecting the second user input (e.g., 1010b) in the second direction, the electronic device updates the display of the movement of one or more (e.g., two or more, all) of the first visual element (e.g., 1004a), the second visual element (e.g., 1004a), and the third visual element (e.g., 1004a) on the display such that the movement is in the second direction (e.g., a direction from right to left as shown in FIGS. 10H - 10I).

[0325] According to some embodiments, while using two or more speakers to simultaneously generate audio, the electronic device detects a third user input (e.g., a first direction, a left-to-right direction). In response to detecting the third user input, the electronic device configures the generation of the audio of the first audio source, the second audio source, and the third audio source to move along a predetermined path in a first direction (e.g., a right-to-left direction) over time at a fifth speed that is faster than a first speed, as perceived by the user. Further, in response to detecting the third user input, the electronic device updates, on a display, the display of the simultaneous movement of the first visual element, the second visual element, and the third visual element to move in a first direction (e.g., left-to-right) at a sixth speed that is faster than a fourth speed.

[0326] According to some embodiments, the electronic device detects a selection input (e.g., 1010c, a tap input, a tap-and-hold input) at a position corresponding to a second visual element (e.g., 1004b). In response to detecting the selection input (e.g., 1010c), the electronic device uses two or more speakers to generate the audio of a second audio file in a fifth mode (rather than in a second mode). The audio generated using the fifth mode is not perceived by the user as being generated from a point within a space that moves over time.

[0327] According to some embodiments, the fifth mode does not include modifying the interaural time difference of the audio. In some examples, the audio generated in the fifth mode (e.g., 1054b in FIG. 10K) is not modified using any of HRTF, the interaural time difference of the audio, or cross cancellation.

[0328] According to some embodiments, the first mode, the second mode, and the third mode include modifying the interaural time difference of the audio. In some examples, the audio generated in the first mode is modified using one or more of HRTF, the interaural time difference of the audio, and cross-cancellation. In some examples, the audio generated in the second mode is modified using one or more of HRTF, the interaural time difference of the audio, and cross-cancellation. In some examples, the audio generated in the third mode is modified using one or more of HRTF, the interaural time difference of the audio, and cross-cancellation. By modifying the interaural time difference of the audio, the user is enabled to perceive the audio as coming from a specific direction. This direction information of the audio provides additional feedback to the user about the placement of different content. Providing improved audio feedback to the user enhances the operability of the device, makes the user-device interface more efficient (e.g., by assisting the user in providing appropriate input), which in turn reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0329] According to some embodiments, modifying the interaural time difference of the audio includes combining the first channel audio of the audio (e.g., the right channel) and the second channel audio of the audio (e.g., the left channel) to form a composite channel audio, updating the second channel audio to include the composite channel audio with a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio with a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay to the first channel audio (e.g., the right channel) without introducing a time delay to a second channel audio different from the first audio channel (e.g., the left channel).

[0330] Note that the details of the processes described above with respect to method 1200 (e.g., FIGS. 12A-12B) should also be noted as being applicable in a similar manner to the methods described below and above. For example, methods 700, 900, 1400, and 1500 optionally include one or more of the various method characteristics described above with respect to method 1200. For example, the same or similar techniques are used to place audio in space. In another example, the same audio source can be used with various techniques. In yet another example, the audio currently being played in each of the various methods can be manipulated using the techniques described in other methods. For the sake of brevity, these details are not repeated below.

[0331] FIGS. 13A-13F show exemplary techniques for managing headphone transparency according to some embodiments. The techniques of these figures are used to illustrate the processes described below, including the processes of FIGS. 14A-14B.

[0332] FIGS. 13G-13M show exemplary techniques for manipulating multiple audio streams of an audio source according to some embodiments. The techniques of these figures are used to illustrate the processes described below, including the process of FIG. 15.

[0333] Figures 13A - 13M illustrate a device 1300 (e.g., a mobile phone) having a touch - sensing surface connected to a display and left and right speakers (e.g., headphones 1358). In some embodiments, the left and right speakers are operable to operate at a noise - cancellation level (e.g., the noise from outside the headphones is suppressed, so the noise heard by the user is less than that, and thus in a state of low noise permeability) and at a full - permeability level (e.g., the noise from outside the headphones is completely passed to the user or passes through the headphones as much as possible, so the user can hear those noises, and thus in a state of high noise permeability), individually. In FIGS. 13A - 13C, the device 1300 operates the left speaker 1358a and the right speaker 1358b at the noise - cancellation level, as indicated by the speakers being shaded. However, for example, the device 1300 can operate the left speaker 1358a at the noise - cancellation level while operating the right speaker 1358b at the full - permeability level.

[0334] The top - down view 1350 is a visual representation of the spatial configuration of the audio generated by the device 1300 and is shown throughout FIGS. 13A - 13M to provide the reader with a better understanding of the technique, particularly with respect to the position where the user 1356 perceives the sound to be arriving (e.g., as a result of the device 1300 placing audio in space). Time 1350 is not part of the user interface of the device 1300. In some embodiments, the techniques described below enable the user to more easily and efficiently hear audio sources that are not generated by the device the user is currently listening to. For example, with this technique, while the user is listening to music using headphones, the user can more clearly hear the person speaking to the user. In some embodiments, the techniques described below enable the user to easily manipulate various audio streams.

[0335] In FIG. 13A, device 1300 displays music player 1304 including album art 1304a and optionally axis 1304b. The album art corresponds to the song being played by device 1300. In some embodiments, the song includes a plurality of audio streams. In this embodiment, the song includes five audio streams, and each audio stream corresponds to a particular device.

[0336] In FIG. 13A, device 1300 generates the audio of the song (including all five audio streams) without placing the audio in space using left speaker 1358a and right speaker 1358b (e.g., device 1300 does not apply any interaural time differences, HRTF, or cross cancellation). For example, this causes the user to perceive the audio as being within the user 1356's head, as shown by audio element 1354 in the overhead view 850 of FIGS. 8F-8G. Audio element 1354 corresponds to the song and includes five audio streams. Device 1300 is configured to control the music based on user input received at affordance 1304c, such as pausing, playing, fast forwarding, and rewinding the music.

[0337] In some embodiments, user 1356 sees the person on their right speaking and provides a drag input 1310a. In FIGS. 13B-13D, device 1300 detects drag input 1310a that displaces album art 1304a. For example, device 1300 updates the display of the position of album art 1304a to correspond to the movement of drag input 1310a.

[0338] In FIGS. 13B and 13C, the displacement of the album art 1304a does not exceed a predetermined distance, and the device 1300 uses the left speaker 1358a and the right speaker 1358b to continue generating the audio of the song (including all five audio streams) without placing the audio in the space (e.g., the device 1300 does not apply any of the interaural time difference, HRTF, or cross cancellation), while maintaining the operation of the left speaker 1358a and the right speaker 1358b at the noise cancellation level.

[0339] In FIG. 13D, the device 1300 determines that the displacement of the album art 1304a exceeds a predetermined distance (e.g., the user has moved the album art 1304a far enough away). In response to the determination that the displacement of the album art 1304a exceeds a predetermined distance, the device 1300 transitions to generating the audio of the song (including all five audio streams) by placing the audio in the space using the left speaker 1358a and the right speaker 1358b (e.g., the device 1300 applies the interaural time difference, HRTF, and / or cross cancellation), and as shown in the top - down view 1350 of FIG. 13D, the user perceives that the audio is at a position in the space away from the person on the user's right (e.g., in front and to the left). In some embodiments, the audio is moved to a position in the space based on the direction and / or distance of the drag input 1310a. Further, in response to the determination that the displacement of the album art 1304a exceeds a predetermined distance, the device 1300 maintains the operation of the left speaker 1358a at the noise cancellation level, the left speaker 1358a is filled, and the right speaker 1358b is not filled, thereby transitioning the operation of the right speaker 1358b to the full transparency level as shown in the top - down view 1350 of FIG. 13D.

[0340] While device 1300 maintains the audio placement within the space (e.g., device 1300 applies interaural time difference, HRTF, and / or cross cancellation), operates the left speaker 1358a at the noise cancellation level, and operates the right speaker 1358b at the full transparency level, device 1300 detects the lift-off of the drag input 1310a in FIG. 13E.

[0341] In response to detecting the lift-off of the drag input 1310a, as shown in FIG. 13F, device 1300 returns the display of the album art 1304a to its original pre-drag position and transitions to generating the audio of the song without placing the audio in the space (including all five audio streams) using the left 1358a and right 1358b speakers (e.g., device 1300 does not apply any of the interaural time difference, HRTF, and / or cross cancellation), whereby the user perceives the audio as being within the user's head. Further, in response to detecting the lift-off of the drag input 1310a, device 1300 maintains the operation of the left speaker 1358a at the noise cancellation level and transitions the operation of the right speaker 1358b to the noise cancellation level as shown in the overhead view 1350 of FIG. 13D by the left speaker 1358a and the right speaker 1358b being filled.

[0342] In some embodiments, in response to detecting the lift-off of the drag input 1310a, device 1300 maintains the placement of the audio within the space and maintains the operation of the right speaker 1358b at the full transparency level.

[0343] In FIG. 13G, device 1300 continues to generate the audio of the song without placing the audio in the space (including all five audio streams) using the left speaker 1358a and the right speaker 1358b (e.g., device 1300 does not apply any of the interaural time difference, HRTF, or cross cancellation). In FIG. 13G, device 1300 detects the input 1310b.

[0344] In some embodiments, device 1300 determines whether the characteristic intensity of input 1310b exceeds a strength threshold. In response to the device 1300 determining that the characteristic intensity of input 1310b does not exceed the strength threshold, the device 1300 continues to generate the audio of the song (including all five audio streams) without placing the audio in the space using the left speaker 1358a and the right speaker 1358b.

[0345] In response to the device 1300 determining that the characteristic intensity of input 1310b exceeds the strength threshold, the device 1300 transitions to generating the audio of various audio streams (including all five audio streams) for the song by placing various audio in the space using the left speaker 1358a and the right speaker 1358b, such that the user perceives that the various audio streams are departing from the user's head in different directions, as indicated by the audio elements 1354 divided into audio elements 1354a - 1354e in FIGS. 13G - 13I. Further, in response to the device 1300 determining that the characteristic intensity of input 1310b exceeds the strength threshold, the device 1300 updates the display of the music player 1304 to show the animation of the stream affordances 1306a - 1306e that are displayed and spread apart, as shown in FIGS. 13H - 13I. For example, each of the stream affordances 1306a - 1306e corresponds to a respective audio stream of the song. For example, stream affordance 1306a corresponds to a first audio stream that includes the singer (but does not include a guitar, keyboard, drums, etc.). Similarly, audio element 1354a corresponds to the first audio stream. In another example, stream affordance 1306d corresponds to a second audio stream that includes a guitar (but does not include the singer, keyboard, drums, etc.). Similarly, audio element 1354d corresponds to the second audio stream.

[0346] As shown in FIGS. 13H - 13I, the device 1300 displays stream affordances 1306a - 1306e at positions on the display corresponding to positions within the space where the audio stream corresponding to the device 1300 is placed (e.g., relative to each other, relative to points on the display).

[0347] In FIGS. 13J - 13K, the device 1300 detects a drag gesture 1310c on the stream affordance 1306d. In response to detecting the drag gesture 1310c, the device 1300 transitions to generating audio for a second audio stream (corresponding to the audio element 1354e) using the left speaker 1358a and the right speaker 1358b without placing the audio in the space, such that the user perceives the audio for the second audio stream as being within the user's head, and continues to generate various other audio streams for the song at various positions in the space (e.g., corresponding to the audio elements 1354a - 1354d) as shown by the audio elements 1354a - 1354e in FIG. 13K. Thus, the device moves individual song audio streams inside and outside the user's head, and to different positions in the space, based on detecting a drag input on the corresponding stream affordances 1306a - 1306d. In some embodiments, the device places various audio streams at positions in the space based on the detected drag input (e.g., direction of movement, lift - off placement).

[0348] In FIG. 13L, device 1300 detects a tap input 1310d on stream affordance 1306a. In response to detecting the tap input 1310d on stream affordance 1306a, device 1300 stops generating audio for a second audio stream and thus invalidates the audio stream. Further, in response to detecting the tap input 1310d on stream affordance 1306a, device 1300 updates one or more visual characteristics (e.g., shading, size, movement) of stream affordance 1306a to indicate that the corresponding audio stream is invalidated. In some embodiments, device 1300 receives an input (e.g., a drag gesture on stream affordance 1306a) that moves the invalidated audio stream to a new position in the space without generating the invalidated audio stream. In response to detecting an additional tap input on the stream affordance corresponding to the invalidated audio stream, device 1300 enables the audio stream and begins generating the audio of the enabled audio stream at a new position in the space.

[0349] In some embodiments, device 1300 includes a digital assistant that generates audio feedback, such as by uttering the results to return the results of a query. In some embodiments, device 1300 generates the audio of the digital assistant by placing the audio for the digital assistant at a position in the space (e.g., above the user's right shoulder), whereby the user perceives that the digital assistant remains stationary in the space even when other audio moves within the space. In some embodiments, device 1300 emphasizes the audio of the digital assistant by ducking one or more (or all) other audio.

[0350] Figures 14A-14B are flow diagrams illustrating a method for managing headphone transparency using an electronic device, according to some embodiments. Method 1400 is executed on a device having a display and a touch-sensing surface (e.g., 100, 300, 500, 1300). The electronic device is operably connected to two or more speakers (e.g., left and right speakers, left and right headphones, left and right earphones, left and right earbuds). Some operations of method 1400 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0351] As described below, method 1400 provides an intuitive method for managing headphone transparency. This method reduces the user's cognitive burden for managing headphone transparency, thereby creating a more efficient human-machine interface. In the case of a battery-operated computing device, it saves power and extends the time between battery charges by enabling the user to manage headphone transparency faster and more efficiently.

[0352] The electronic device displays a user-movable affordance (e.g., 1304a, album art) at a first position on the display (1402).

[0353] While the user-movable affordance (e.g., 1304a) is displayed at the first position (e.g., 1304a as shown in FIG. 13A) (1404), the electronic device operates the electronic device in a first state of ambient sound transparency (e.g., a state where transparency is ineffective for both the first speaker and the second speaker of the two or more speakers, and external noise is suppressed with respect to the first speaker and the second speaker of the two or more speakers, 1358a and 1358b as shown in FIG. 13A) (1408).

[0354] While a user movable affordance (e.g., 1304a) is being displayed at a first position (e.g., 1304a as shown in FIG. 13A) (1404), the electronic device uses two or more speakers to generate audio (e.g., 1354) using an audio source (e.g., an audio file, a music file, a media file) in a first mode (e.g., 1354 in FIG. 13A, HRTF, cross-cancellation, or no interaural time difference applied, stereo mode) (1410).

[0355] While a user movable affordance (e.g., 1304a) is being displayed at a first position (e.g., 1304a as shown in FIG. 13A) (1404), the electronic device uses a touch sensing surface to detect user input (e.g., 1310a, a drag gesture) (1412).

[0356] A set of one or more conditions includes a first condition that is satisfied when the user input (e.g., 1310a) is a touch-and-drag operation on the user movable affordance (e.g., 1304a) (1416). According to some embodiments, the set of one or more conditions further includes a second condition that is satisfied when the user input causes a displacement of the user movable affordance by at least a predetermined amount from the first position (1418). For example, a small movement of the movable affordance (e.g., 1304a) does not cause a change, and the movable affordance (e.g., 1304a) must move a minimum distance while the mode is being changed from the first mode to the second mode and before the state is changed from the first state to the second state. This helps to avoid inadvertent mode and state changes.

[0357] In response to detecting a user input (e.g., 1310a) (1414), and in accordance with one or more sets of conditions being satisfied (1416), the electronic device operates the electronic device in a second ambient sound transmissivity state that is different from the first ambient sound transmissivity state (e.g., a state in which transparency is enabled for a first speaker of two or more speakers and disabled for a second speaker, two or more speakers, 1358a and 1358b as shown in FIG. 13D, where external noise is suppressed for the first speaker and not suppressed for the second speaker) (1420). By changing the ambient sound transmissivity state, the user can better hear sounds from the user's environment, specifically sounds from a specific direction within the environment.

[0358] Furthermore, in response to detecting a user input (e.g., 1310a) (1414), and in accordance with one or more sets of conditions being satisfied (1416), the electronic device transitions the generation of audio using an audio source (e.g., 1354 shown in FIG. 13D) from a first mode to a second mode that is different from the first mode (e.g., a monomode applying one or more of HRTF, cross-cancellation, and interaural time difference) (1422). By changing the mode in which the audio is generated, the user can better hear sounds from the user's environment, specifically sounds from a specific direction within the environment (e.g., a specific direction away from the direction in which the generation of the audio has been moved in space).

[0359] Furthermore, in response to detecting a user input (e.g., 1310a) (1414), and in accordance with one or more sets of conditions not being satisfied (1424), the electronic device maintains the electronic device in a first ambient sound transmissivity state (e.g., a state in which transmissivity is disabled for both a first speaker and a second speaker) (1426).

[0360] Further, in response to detecting a user input (e.g., 1310a) (1414), and in accordance with one or more sets of conditions not being satisfied (1424), the electronic device maintains generating audio (e.g., 1354 as shown in FIG. 13C) using an audio source in a first mode.

[0361] In some embodiments, a speaker (e.g., headphones) is operable to operate at a full noise cancellation level (e.g., noise from outside the headphones is suppressed and the headphones are suppressed as much as possible so that the user cannot hear that noise, and thus a state of low noise permeability), and a full transparency level (noise from outside the headphones is completely passed to the user or passed as much as possible by the headphones so that the user can hear that noise, and thus a state of high noise permeability).

[0362] Further, in response to detecting a user input (e.g., 1310a) (1414), the electronic device updates (1430) the display of a user - movable affordance (e.g., 1304a) on the display from a first position on the display (e.g., the position of 1304a in FIG. 13A) to a second position on the display (e.g., the position of 1304a in FIG. 13D) according to the movement on the touch - sensing surface of the user input (e.g., 1310a). In some embodiments, the display position of the user - movable affordance (e.g., 1304a) is based on the user input, such as by updating the display of the user - movable affordance (e.g., 1304a) so as to correspond to the position of contact of the user input. Thus, the farther the user input (e.g., 1310a) moves, the farther the user - movable affordance (e.g., 1304a) moves on the display.

[0363] According to some embodiments, after (e.g., during) one or more sets of conditions are satisfied, the electronic device detects the end of a user input (e.g., by detecting a lift-off of contact of the user input on the touch-sensing surface, such as 1310a). In response to detecting the end of the user input, the electronic device updates the display of a user-movable affordance (e.g., 1304a in FIG. 13F) to a first position on the display. Further, in response to detecting the end of the user input, the electronic device transitions to operate the electronic device in a first state of ambient sound transparency (e.g., as shown by 1358a and 1358b in FIG. 13F). Further, in response to detecting the end of the user input, the electronic device transitions the generation of audio from a second mode to a first mode using an audio source (e.g., as shown by a change in position of 1354 in FIGS. 13E - 13F).

[0364] According to some embodiments, after (e.g., during) one or more sets of conditions are satisfied, the electronic device detects the end of a user input (e.g., by detecting a lift-off of contact of the user input on the touch-sensing surface). In response to detecting the end of the user input, the electronic device maintains the display of a user-movable affordance at a second position on the display. Further, in response to detecting the end of the user input, the electronic device maintains the operation of the electronic device in a second state of ambient sound transparency. Further, in response to detecting the end of the user input, the electronic device maintains the generation of audio using an audio source in a second mode.

[0365] According to some embodiments, the user input (e.g., 1310a) includes a direction of movement. According to some embodiments, the second state of ambient sound transmissibility (as shown by 1358a and 1358b in FIG. 13D) is based on the direction of movement of the user input (e.g., 1310a). For example, the electronic device selects a particular speaker (e.g., 1358b) and changes the transmissibility state based on the direction of movement of the user input (e.g., 1310a). According to some embodiments, the second mode is based on the direction of movement of the user input. For example, in the second mode, the electronic device generates audio such that the audio is perceived to be generated from a point in space corresponding to a second position of the displayed user movable affordance, and the separating direction is based on the direction of movement of the first input (and / or based on the position of the user movable affordance 1304a).

[0366] According to some embodiments, the second mode is configured such that the user perceives that the audio generated using the second mode is generated from a point in space corresponding to a second position of the displayed user movable affordance.

[0367] According to some embodiments, the first mode does not include a modification of the interaural time difference of the audio. In some examples, the audio generated in the first mode is not modified using any of HRTF, cross cancellation, or interaural time difference.

[0368] According to some embodiments, the second mode includes modifying the interaural time difference of the audio. In some examples, the second mode includes modifying the interaural time difference of the audio such that the audio is perceived by the user as being generated from a direction corresponding to a position of a user movable affordance such as a display or the user's left or right. In some examples, the audio generated in the second mode is modified using one or more of HRTF, cross cancellation, and interaural time difference. Modifying the interaural time difference of the audio enables the user to perceive the audio as coming from a particular direction. This directional information of the audio provides additional feedback to the user about the placement of different content. Providing improved audio feedback to the user enhances the operability of the device, makes the user device interface more efficient (e.g., by assisting the user in providing appropriate input), which in turn reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0369] According to some embodiments, modifying the interaural time difference of the audio includes combining a first channel audio of the audio (e.g., right channel) and a second channel audio of the audio (e.g., left channel) to form a composite channel audio, updating the second channel audio to include the composite channel audio with a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio with a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of the audio includes introducing a time delay to a first channel audio (e.g., right channel) without introducing a time delay to a second channel audio (e.g., left channel) different from the first audio channel.

[0370] Note that the details of the processes described above with respect to method 1400 (e.g., FIGS. 14A-14B) are also applicable in a similar manner to the methods described below and above. For example, methods 700, 900, 1200, and 1500 optionally include one or more of the various method characteristics described above with respect to method 1400. For example, the same or similar techniques are used to place audio in space. In another example, the same audio source can be used with various techniques. In yet another example, the currently playing audio in each of the various methods can be manipulated using the techniques described in other methods. For the sake of brevity, these details are not repeated below.

[0371] FIG. 15 is a flow diagram illustrating a method for operating multiple audio streams of an audio source using an electronic device according to some embodiments. Method 1500 is executed on a device having a display and a touch-sensing surface (e.g., 100, 300, 500, 1300). The electronic device is operably connected to two or more speakers, including a first (e.g., left, 1358a) speaker and a second (e.g., right, 1358b) speaker. For example, the two or more speakers are left and right speakers, left and right headphones, left and right earphones, or left and right earbuds. Some operations of method 1500 are optionally combined, the order of some operations is optionally changed, and some operations are optionally omitted.

[0372] As described below, method 1500 provides an intuitive way to operate multiple audio streams of an audio source. This method reduces the user's cognitive burden for operating multiple audio streams of an audio source, thereby creating a more efficient human-machine interface. In the case of a battery-operated computing device, by enabling the user to more quickly and efficiently operate multiple audio streams of an audio source, power is conserved and the time between battery charges is extended.

[0373] The electronic device uses two or more speakers to generate (1502) audio (e.g., 1354 in FIG. 13F) using an audio source (e.g., an audio file, a music file, a media file) in a first mode (e.g., HRTF, cross-cancellation, or stereo mode without applying any of the interaural time differences). The audio source includes a plurality (e.g., five) of audio streams, including a first audio stream and a second audio stream. In some embodiments, each audio stream of the audio source is a stereo audio stream. In some embodiments, each audio stream is limited to a single respective device. In some embodiments, each audio stream is limited to the voice of a single respective singer. Thus, each audio stream of the audio source is generated in the first mode.

[0374] The electronic device uses a touch sensing surface to detect (1504) a first user input (e.g., 1310b, a tap on an affordance, an input having a characteristic intensity that exceeds an intensity threshold).

[0375] In response to detecting the first user input (e.g., 1310b) (1506), the electronic device uses two or more speakers to transition (1508) the generation of the first audio stream of the audio source (e.g., 1354a in FIG. 13H) from the first mode to a second mode different from the first mode.

[0376] Furthermore, in response to detecting the first user input (e.g., 1310b) (1506), the electronic device uses two or more speakers to transition (1510) the generation of the second audio stream of the audio source (e.g., 1354b in FIG. 13H) from the first mode to a third mode different from the first mode and the second mode.

[0377] By placing various audio streams at different positions within a space, a user can better distinguish between different audio streams. Thereby, the user can more quickly and efficiently turn off (and on) specific portions of the audio (e.g., a specific audio stream) that the user desires to exclude from the listening experience. Providing improved audio feedback to the user improves the operability of the device, makes the user-device interface more efficient (e.g., by assisting the user to provide appropriate input when operating / interacting with the device and reducing user errors), and in addition, reduces power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0378] Further, in response to detecting a first user input (e.g., 1310b) (1506), the electronic device displays (1512) a first visual representation (e.g., 1306a) of a first audio stream of an audio source on a display.

[0379] Further, in response to detecting a first user input (e.g., 1310b) (1506), the electronic device displays (1514) a second visual representation (e.g., 1306b) of a second audio stream of the audio source on the display, and the first visual representation (e.g., 1306a) is different from the second visual representation (e.g., 1306b).

[0380] According to some embodiments, blocks 1508 - 1514 occur simultaneously.

[0381] According to some embodiments, the first mode does not include correction of the interaural time difference of the audio. In some examples, the audio generated in the first mode is not corrected using any of HRTF, cross-cancellation, or the interaural time difference.

[0382] According to some embodiments, the second mode includes modifying the interaural time difference of the audio. In some examples, the second mode includes modifying the interaural time difference of the audio such that the audio is perceived by the user as being generated from a direction corresponding to a position corresponding to the position of the corresponding visual representation. In some examples, the second mode includes applying one or more of HRTF, cross cancellation, and interaural time difference. In some examples, the third mode includes modifying the interaural time difference of the audio. In some examples, the third mode includes applying one or more of HRTF, cross cancellation, and interaural time difference. Modifying the interaural time difference of the audio enables the user to perceive the audio as coming from a particular direction. This direction information of the audio provides additional feedback to the user about the placement of different content. Providing improved audio feedback to the user enhances the operability of the device (e.g., by assisting the user in providing appropriate input), makes the user-device interface more efficient, which in turn reduces the power consumption and improves the battery life of the device by enabling the user to use the device more quickly and efficiently.

[0383] According to some embodiments, modifying the interaural time difference of audio includes combining the first channel audio of the audio stream (e.g., the right channel) and the second channel audio of the audio stream (e.g., the left channel) to form a composite channel audio, updating the second channel audio to include the composite channel audio with a first delay (e.g., 0 ms delay, less than the delay amount), and updating the first channel audio to include the composite channel audio with a second delay different from the first delay (e.g., 100 ms delay, more than the delay amount). In some examples, modifying the interaural time difference of audio includes introducing a time delay to the first channel audio (e.g., the right channel) without introducing a time delay to the second channel audio (e.g., the left channel) different from the first audio channel.

[0384] According to some embodiments, an electronic device includes displaying a first visual representation at a first position and sliding the first visual representation in a first direction toward a second position (e.g., 1306a in FIG. 13I) on a display, to display a first visual representation (e.g., 1306a) of a first audio stream of an audio source on the display. According to some embodiments, an electronic device includes displaying a second visual representation at a first position and sliding the second visual representation in a second direction different from the first direction toward a third position (e.g., 1306b in FIG. 13I) different from the second position, to display a second visual representation (e.g., 1306b) of a second audio stream of an audio source on the display.

[0385] According to some embodiments, the electronic device uses a touch-sensing surface to detect a second user input (e.g., 1310c), and the second input (e.g., 1310c) starts at a position corresponding to a second position and ends at a position corresponding to a first position. In response to detecting the second user input, the electronic device slides the first visual representation on the display from the second position to the first position while maintaining a second visual representation at a third position on the display. Further, in response to detecting the second user input, the electronic device uses two or more speakers to transition the generation of audio using a first audio stream from a second mode to a first mode while maintaining the generation of audio using a second audio stream in a third mode.

[0386] According to some embodiments, detecting a first user input (e.g., 1310a) includes accessing a characteristic intensity of the first user input and determining that the characteristic intensity of the first user input exceeds an intensity threshold.

[0387] According to some embodiments, while the electronic device uses two or more speakers to generate audio using a first audio stream of an audio source and a second audio stream of the audio source, the electronic device uses a touch-sensing surface to detect a third user input (e.g., 1310d, tap input) at a position on the touch-sensing surface corresponding to a first visual representation (e.g., 1306a). In response to detecting the third user input (e.g., 1310d), the electronic device stops generating audio (e.g., 1354a) using the first audio stream of the audio source using two or more speakers. Further, in response to detecting the third user input (e.g., 1310d), the electronic device maintains the generation of audio using the second audio stream (e.g., 1354b) of the audio source using two or more speakers. Thus, the electronic device detects a tap on the device affordance and deactivates the audio stream corresponding to that device.

[0388] According to some embodiments, while the electronic device does not generate audio using a first audio stream of an audio source using two or more speakers and generates audio using a second audio stream of the audio source using two or more speakers, the electronic device uses a touch-sensing surface to detect a fourth user input (e.g., tap input) at a position on the touch-sensing surface corresponding to a first visual representation. In response to detecting the fourth user input, th...

Claims

1. An electronic device comprising a display and a touch sensing surface, the electronic device being operably connected to two or more speakers including a first speaker and a second speaker, generating audio using the two or more speakers in a first mode using a single audio source, the single audio source including a plurality of audio streams including a first audio stream and a second audio stream, detecting a first user input using the touch sensing surface, in response to detecting the first user input, using the two or more speakers to transition the generation of the first audio stream of the single audio source from the first mode to a second mode different from the first mode, using the two or more speakers to transition the generation of the second audio stream of the single audio source from the first mode to a third mode different from the first mode and the second mode, displaying, via the display, a first visual representation of the first audio stream of the single audio source at a first position on the display, the first position corresponding to a first position in a space corresponding to the first audio stream of the single audio source, displaying, via the display, a second visual representation of the second audio stream of the single audio source at a second position on the display different from the first position on the display, the second position corresponding to a second position in a space corresponding to the second audio stream of the single audio source, the first visual representation being different from the second visual representation, simultaneously performing, A method, comprising.

2. The method according to claim 1, wherein the first mode does not include correction of the interaural time difference of the audio.

3. The method according to claim 1, wherein the second mode includes correction of the interaural time difference of the audio.

4. The correction of the interaural time difference of the audio is, combining the first channel audio of the audio stream and the second channel audio of the audio stream to form a composite channel audio, Updating the second channel audio to include the composite channel audio with the first delay; Updating the first channel audio to include the composite channel audio with a second delay different from the first delay; The method according to claim 3, comprising:

5. Displaying the first visual representation of the first audio stream of the single audio source via the display includes displaying the first visual representation at a third position and sliding the first visual representation in a first direction toward the first position of the display; Displaying the second visual representation of the second audio stream of the single audio source via the display includes displaying the second visual representation at the third position and sliding the second visual representation in a second direction different from the first direction toward a second position different from the first position; The method according to claim 1.

6. Detecting a second user input using the touch-sensitive surface, the second user input starting at a position corresponding to the first position and ending at a position corresponding to the third position; In response to detecting the second user input; Sliding the first visual representation from the first position to the third position of the display while maintaining the second visual representation at the second position of the display; Transitioning the generation of audio using the first audio stream from the second mode to the first mode using the two or more speakers while maintaining the generation of audio using the second audio stream in the third mode; The method according to claim 5, further comprising:

7. Detecting the first user input includes: Accessing a characteristic intensity of the first user input; Determining that the characteristic intensity of the first user input exceeds an intensity threshold; The method according to claim 1, comprising:

8. While the electronic device is generating audio using the two or more speakers with the first audio stream of the single audio source and the second audio stream of the single audio source; Detecting a third user input at a position on the touch-sensing surface corresponding to the first visual representation using the touch-sensing surface; In response to detecting the third user input, Stopping generating audio using the first audio stream of the single audio source using the two or more speakers; Maintaining generating audio using the second audio stream of the single audio source using the two or more speakers; The method according to claim 1, further comprising.

9. While the electronic device is not generating audio using the first audio stream of the single audio source using the two or more speakers and is generating audio using the second audio stream of the single audio source using the two or more speakers, Detecting a fourth user input at a position on the touch-sensing surface corresponding to the first visual representation using the touch-sensing surface; In response to detecting the fourth user input, Generating audio using the first audio stream of the single audio source using the two or more speakers; Maintaining generating audio using the second audio stream of the single audio source using the two or more speakers; The method according to claim 8, further comprising.

10. Detecting a volume control input; In response to detecting the volume control input, modifying the volume of the audio generated using the two or more speakers for each audio stream in the plurality of audio streams; The method according to claim 1, further comprising.

11. A computer program for causing a computer to execute the method according to any one of claims 1 to 10.

12. A memory storing the computer program according to claim 11, An electronic device comprising one or more processors capable of executing the computer program stored in the memory.

13. An electronic device comprising means for executing the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Anode of xxray tube

    JP1977046790A

  • Multiplexed arithmetic processing synchronous controller

    JP1985065369A

  • Multi-task management system

    JP1998055260A

  • Display and display method

    JP2010074258A

  • JPP4171675B