User Interface for Image Capture

The described method addresses inefficiencies in image data capture by using server-based or local processing to generate personalized audio experiences, improving efficiency and reducing power consumption while enhancing user satisfaction.

JP2026500532APending Publication Date: 2026-01-07DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025536468
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-21
Filing Date
2023-12-21
Publication Date
2026-01-07

AI Technical Summary

Technical Problem

Existing techniques for capturing image data on electronic devices are cumbersome, complex, inefficient, and error-prone, leading to unsuitable data for personalized applications like PHRTF generation, and consume excessive power.

Method used

Electronic devices capture images, transmit them to a network server, and receive PHRTF data to generate personalized audio experiences, or process images locally to determine pose data and apply PHRTF, with user interfaces that abort the capture process based on sufficiency criteria.

Benefits of technology

The method reduces cognitive load, improves efficiency, conserves power, and enhances user satisfaction by providing faster and more efficient image data capture for device personalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026500532000001_ABST
    Figure 2026500532000001_ABST
Patent Text Reader

Abstract

[0001] The present disclosure generally relates to user interfaces and techniques for capturing image data. In some embodiments, a method includes: displaying a user interface on a display, the user interface including a preview portion; displaying images in the preview portion corresponding to image data captured by a camera; performing a capture process while displaying the images in the preview portion, capturing a series of images corresponding to the preview images in the preview portion using the camera; determining pose data associated with the series of images; outputting a first set of instructional prompts in accordance with the pose data meeting or exceeding a first threshold; and outputting a second set of instructional prompts in accordance with the pose data meeting or exceeding a second threshold; and aborting the capture process in response to determining that a set of sufficiency criteria has been met.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to computer user interfaces, and more particularly to user interfaces and techniques for capturing image data for device personalization. [Background technology]

[0002] Image data can be captured by a camera of an electronic device to enable device or experience personalization for a user of the electronic device. Information about the image capture process can be presented to the user in a user interface (UI) displayed on the screen of the electronic device. Existing techniques for image data capture using electronic devices are generally cumbersome, complex, and inefficient. For example, some existing techniques use complex and time-consuming UIs that may involve multiple key presses or keystrokes. Furthermore, some existing techniques are error-prone, resulting in a collection of image data that is not suitable for implementing complex image-driven personalization, such as generating a personalized head-related transfer function (PHRTF) for audio playback. Existing techniques also take longer than necessary to complete image capture, consuming more device power than necessary, which is particularly important for battery-powered electronic devices. Summary of the Invention

[0003] Thus, the disclosed embodiments provide electronic devices with faster, more efficient methods and interfaces for capturing image data. In some embodiments, the electronic device (e.g., a desktop or notebook computer, a smartphone, or a tablet computer) captures images, transmits the images to a network server computer, and then receives PHRTF data from the server computer and applies the PHRTF data to the audio to generate a personalized audio experience (e.g., a spatial or immersive audio experience) for the user. Types of audio signals that can be processed using the PHRTF data include, but are not limited to, channel-based audio, object-based audio (e.g., 5.1.2, 7.1.2, 5.1.4, 7.1.4, 9.1.2), etc. In other embodiments, the electronic device captures images, selects specific images from the images, calculates the PHRTF, and applies the PHRTF to the audio to generate a personalized audio experience for the user.

[0004] According to some embodiments, a method is described for operation on an electronic device including a display device, the method including: displaying a user interface on the display including a preview portion displaying images corresponding to image data captured by a camera; performing a capture process while displaying the images in the preview portion; and aborting the capture process in response to determining that a set of sufficiency criteria is met, the capture process including capturing, with the camera, a series of images corresponding to preview images in the preview portion, determining pose data associated with each image in the series of images, outputting a first set of instructional prompts in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts in accordance with the pose data satisfying or exceeding a second threshold.

[0005] According to some embodiments, a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device including a display device and a camera is described, the one or more programs including instructions for displaying on the display a user interface including a preview portion displaying images corresponding to image data captured by the camera, instructions for performing a capture process when the images are displayed in the preview portion, and instructions for aborting the capture process in response to determining that a set of sufficiency criteria is met, the capture process including capturing a series of images with the camera, determining pose data associated with each image in the series of images, outputting a first set of instructional prompts in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts in accordance with the pose data satisfying or exceeding a second threshold.

[0006] According to some embodiments, a transient computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device including a display device and a camera is described, the one or more programs including instructions for displaying on the display a user interface including a preview portion displaying images corresponding to image data captured by the camera, instructions for performing a capture process while displaying the images in the preview portion, and instructions for aborting the capture process in response to determining that a set of sufficiency criteria is met, the capture process including capturing a series of images with the camera, determining pose data associated with each image in the series of images, outputting a first set of instructional prompts in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts in accordance with the pose data satisfying or exceeding a second threshold.

[0007] According to some embodiments, an electronic device is described. The electronic device has a display device, a camera, one or more processors, and a memory storing one or more programs configured to be executed by the one or more processors. The one or more programs include instructions for displaying a user interface on the display including a preview portion displaying images corresponding to image data captured by the camera, instructions for performing a capture process when the images are displayed in the preview portion, and instructions for aborting the capture process in response to determining that a set of sufficiency criteria is met, the capture process including capturing a series of images with the camera, determining pose data associated with each image in the series of images, outputting a first set of instructional prompts in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts in accordance with the pose data satisfying or exceeding a second threshold.

[0008] According to some embodiments, an electronic device is described that includes a display device, a camera, means for displaying on the display a user interface including a preview portion that displays images corresponding to image data captured by the camera, means for performing a capture process when the images are displayed in the preview portion, and means for aborting the capture process in response to determining that a set of sufficiency criteria is met, the capture process including capturing a series of images with the camera, determining pose data associated with each image in the series of images, outputting a first set of instructional prompts in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts in accordance with the pose data satisfying or exceeding a second threshold.

[0009] According to some embodiments, a method includes presenting a user interface on a display of the electronic device including a preview portion displaying images captured by a camera of the electronic device; presenting a first instruction to a user in the user interface to rotate the user's head in a first direction; capturing a first set of images of a first ear of the user in the preview portion; presenting a second instruction to the user in the user interface to rotate the user's head in a second direction opposite the first direction; capturing a second set of images of a second ear of the user in the preview portion; generating a final set of images from the first and second sets of images by selecting a subset of images of the first and second sets of images that have a minimum angular distance of the user's head rotation captured in the images; and generating data corresponding to a set of personalized head-related transfer functions (PHRTFs) for the user based on the final set of images.

[0010] According to some embodiments, the final set of images has a maximum angular extent of the user's head rotation of less than 110 degrees.

[0011] According to some embodiments, the method further comprises scaling at least some of the first and second sets of images based on the one or more front view images.

[0012] According to some embodiments, a method includes, on an electronic device having a display, displaying on the display a user interface including a graphical object of a first size, a slider affordance including an indicator located at a first position of the slider affordance, and a first text description corresponding to the first position of the indicator of the slider affordance; detecting a user input while displaying the graphical object; and in response to detecting the user input, updating a position of the indicator at the slider affordance from the first position to a second position, displaying the graphical object at a second size smaller or larger than the first size, and displaying the second text description corresponding to the second position of the indicator of the slider affordance in place of the first text description.

[0013] According to some embodiments, the method further comprises generating PHRTF data corresponding to a set of personalized head-related transfer functions (PHRTFs) based on data corresponding to the display size and the first or second text description of the graphical object and the generative model.

[0014] According to some embodiments, generating PHRTF data corresponding to a set of functions associated with the PHRTF occurs in response to detecting user input on a creation profile affordance presented on a touch-sensitive display.

[0015] According to some embodiments, generating the PHRTF data includes applying demographic data associated with the user to a generative model.

[0016] According to some embodiments, generating the PHRTF data excludes the use of image data corresponding to the user.

[0017] According to some embodiments, the first size of the graphical object is smaller than the second size of the graphical object.

[0018] According to some embodiments, the first size of the graphical object is larger than the second size of the graphical object.

[0019] Executable instructions for implementing these functions are optionally contained in a non-transitory computer-readable storage medium, computing device, or other computer program product configured for execution by one or more processors. Executable instructions for implementing these functions are optionally contained in a transitory computer-readable storage medium, computing device, or other computer program product configured for execution by one or more processors.

[0020] The disclosed embodiments provide at least one or more of the following advantages: The disclosed method and user interface optionally complement or replace other methods for capturing image data through a user interface; The disclosed method and interface reduce the cognitive load on the user and provide a more efficient human-machine interface; For battery-powered computing devices, the disclosed method and user interface conserve power and extend the time between battery charges; Thus, electronic devices are provided with a faster, more efficient method and user interface for capturing image data, thereby improving the effectiveness, efficiency, and user satisfaction of such electronic devices, and further reducing power consumption by such electronic devices.

[0021] For a better understanding of the various embodiments described, reference should be made to the following detailed description taken in conjunction with the following drawings, in which like reference numerals refer to corresponding parts throughout the drawings, in which: [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a block diagram illustrating a portable multifunction device with a touch-sensitive display in accordance with some embodiments. [Figure 2]1 illustrates a portable multifunction device with a touch screen according to some embodiments. [Figure 3A] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3B] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3C] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3D] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3E] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3F] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3G] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3H] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3I] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3J] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3K] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3L]1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3M] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3N] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3O] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3P] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3Q] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3R] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3S] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3T] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3U] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 3V] 1 is a screenshot of a user interface for capturing an image using an electronic device according to some embodiments. [Figure 4A] FIG. 1 is a flow diagram depicting a process for capturing image data using an electronic device according to some embodiments. [Figure 4B] FIG. 1 is a flow diagram depicting a process for capturing image data using an electronic device according to some embodiments. [Figure 5A] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5B] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5C] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5D] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5E] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5F] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5G] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5H] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5I] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5J] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 5K] 1 is a screenshot of a first alternative user interface for capturing image data using an electronic device according to some embodiments. [Figure 6A] 10 is a screenshot of a second alternative user interface for generating a PHRTF without capturing an image, according to some embodiments. [Figure 6B] 10 is a screenshot of a second alternative user interface for generating a PHRTF without capturing an image, according to some embodiments. [Figure 6C] 10 is a screenshot of a second alternative user interface for generating a PHRTF without capturing an image, according to some embodiments. [Figure 6D] 10 is a screenshot of a second alternative user interface for generating a PHRTF without capturing an image, according to some embodiments. [Figure 6E] 10 is a screenshot of a second alternative user interface for generating a PHRTF without capturing an image, according to some embodiments. [Figure 7] FIG. 10 is a flow diagram depicting a process for generating a PHRTF without capturing an image according to some embodiments. [Figure 8] FIG. 1 is a flow diagram depicting a process for capturing image data using an electronic device according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0023] The following description sets forth example methods, parameters, etc. However, it should be recognized that such description is not intended as a limitation on the scope of the present disclosure, but is instead provided as a description of exemplary embodiments.

[0024] There is a need for electronic devices that provide efficient methods and interfaces for capturing image data. For example, there is a need for electronic devices that provide a user with information about the ongoing image capture process in an easily understandable and convenient manner. In another example, there is a need for electronic devices that provide effective feedback to a user of the electronic device while capturing data used to achieve device personalization, such as personalized audio playback. Such techniques can reduce the cognitive load placed on a user operating an electronic device, thereby increasing productivity. Furthermore, such techniques can reduce processor and battery power that would otherwise be wasted on redundant user input.

[0025] Below, Figures 1, 2, 3A-3V, and 4A-4B describe an exemplary device for performing embodiments of image data capture. Figures 3A-3V depict exemplary user interfaces for capturing image data. Figures 4A and 4B are flow diagrams depicting methods of capturing image data using an electronic device, according to some embodiments. The user interfaces of Figures 3A-3V are used to illustrate processes described below, including the processes of Figures 4A and 4B. Figures 5A-5K are used to describe alternative methods and user interfaces for capturing image data. Figures 6A-6E are used to describe alternative methods of generating a PHRTF without image capture.

[0026] In the following description, terms such as "first" and "second" are used to describe various elements, but these elements should not be limited by such terms. These terms are used only to distinguish between elements. For example, a first touch may be referred to as a second touch, and similarly, a second touch may be referred to as a first touch, without departing from the scope of various embodiments described. Although a first touch and a second touch are both touches, they are not the same touch.

[0027] The terminology used in the description of the various embodiments described herein is intended to describe particular embodiments only and is not intended to be limiting. As used in the description of the various embodiments described and the appended claims, the singular forms (a, an, the) are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, as used herein, the term "and / or" is understood to refer to and encompass any and all possible combinations of one or more of the associated listed terms. Furthermore, the terms "includes," "including," and / or "comprises," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0028] The term "if" is, depending on the context, interpreted to mean "when" or "when" or "in response to determining" or "in response to detecting," as appropriate. Similarly, the phrases "when determined" or "when (a stated condition or event) is detected" are, depending on the context, interpreted to mean "upon determining" or "in response to determining" or "upon detection of (a stated condition or event)" or "in response to detecting (a stated condition or event)," as appropriate.

[0029] Embodiments of electronic devices, user interfaces for such devices, and associated processes for using such devices are described. In some embodiments, the device is a portable communication device, such as a mobile phone, that also includes other functions, such as PDA functionality and / or music playback functionality. Other portable electronic devices, such as a laptop computer or tablet computer, with a touch-sensitive surface (e.g., a touchscreen display and / or a touchpad) are optionally used. Also, in some embodiments, the device is not a portable communication device, but rather a device such as a desktop computer, laptop computer, tablet computer, multimedia player device, gaming system, or augmented reality device / system (AR / VR). In some embodiments, the electronic device includes a touch-sensitive surface. In some embodiments, the electronic device does not include a touch-sensitive surface, but rather includes a display device (e.g., a display) and one or more separate control device displays (e.g., a mouse, keyboard, stylus, etc.) for interaction with the device, including graphics displayed on the display device.

[0030] In the discussion that follows, electronic devices are described that include a display and a touch-sensitive surface, although it should be understood that the electronic device optionally includes one or more other physical user interface devices, such as a physical keyboard, a mouse, and / or a joystick.

[0031] The device typically supports a variety of applications, such as one or more of a gaming application, a phone application, a video conferencing application, an email application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital video camera application, a web searching application, a digital music player application, and / or a digital video player application. In addition to the applications listed above, the device may support other applications that include media (e.g., audio) playback capabilities.

[0032] Various applications running on the device optionally use at least one common physical user interface device, such as a touch-sensitive surface. One or more features of the touch-sensitive surface and corresponding information displayed on the device are optionally tailored and / or modified for each application and / or within each application. In this way, the device's common physical architecture (e.g., touch-sensitive surface) supports various applications with user interfaces that are intuitive and transparent to the user.

[0033] Attention is now directed to embodiments of portable devices with touch-sensitive displays. FIG. 1 is a block diagram illustrating a portable multifunction device 100 with a touch-sensitive display surface 112, according to some embodiments. Touch-sensitive surface 112 may conveniently be referred to as a "touch screen" and is sometimes known or referred to as a "touch-sensitive display system." Device 100 includes memory 102 (optionally including one or more computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral interface 118, RF circuitry 108, audio circuitry 110, a speaker 111, a microphone 113, an input / output (I / O) subsystem 106, other input control devices 116, and an external port 124. Device 100 optionally includes one or more light sensors 164. Device 100 optionally includes one or more contact intensity sensors 165 that detect the intensity of a contact on device 100 (e.g., on a touch-sensitive surface such as touch-sensitive display system 112 of device 100). Device 100 optionally includes one or more haptic output generators 167 that generate haptic output on device 100 (e.g., on a touch-sensitive surface such as touch-sensitive display surface 112 or a touchpad (e.g., a touch-sensitive surface separate from a display device) of device 100). These components optionally communicate via one or more communication buses or signals 103.

[0034] Of course, device 100 is merely one example of a portable multifunction device, and device 100 optionally has more or fewer components than shown, optionally combines two or more components, or optionally has a different configuration or arrangement of components. The various components shown in Figure 1 may be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing circuits and / or application specific integrated circuits.

[0035] Memory 102 optionally includes high-speed random access memory, and optionally also includes non-volatile memory, such as one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory controller 122 optionally controls access to memory 102 by other components of device 100.

[0036] Peripheral interface 118 may be used to couple input and output peripherals of the device to CPU 120 and memory 102. One or more processors 120 run or execute various software programs and / or sets of instructions stored in memory 102 to implement various functions of device 100 or to process data. In some embodiments, peripheral interface 118, CPU 120, and memory controller 122 are optionally implemented on a single chip, such as chip 104. In some other embodiments, they are optionally implemented as separate chips.

[0037] The radio frequency (RF) circuitry 108 receives and transmits RF signals, also known as electromagnetic signals. The RF circuitry 108 converts electrical signals to and from electromagnetic signals and communicates with communication networks and other communication devices via electromagnetic signals. The RF circuitry 108 optionally includes well-known circuitry, including, but not limited to, an antenna system, an RF transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a CODEC chipset, a subscriber identity module (SIM) card, memory, etc., to accomplish these functions. The RF circuitry 108 optionally communicates wirelessly with networks, such as the Internet, also known as the World Wide Web (WWW), an intranet, and / or wireless networks, such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and other devices.

[0038] Audio circuitry 110, speaker 111, and microphone 113 provide an audio interface between a user and device 100. Speaker 111 converts electrical signals into sound waves that humans can hear. Audio circuitry 110 also receives electrical signals converted from sound waves by microphone 113.

[0039] The I / O subsystem 106 couples input / output peripherals of the device 100, such as the touchscreen 112 and other input control devices 116, to a peripheral interface 118. The I / O subsystem 106 optionally includes a display controller 156, a light sensor controller 158, a depth camera controller 169, an intensity sensor controller 159, a haptic feedback controller 161, and one or more input controllers 160 for other input or control devices. The one or more input controllers 160 receive electrical signals from or send electrical signals to the other input control devices 116. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, the input controllers 160 are optionally coupled to any (or none) of a keyboard, an infrared port, a USB port, and a pointer device such as a mouse. The one or more buttons (e.g., 208 in FIG. 2) optionally include up / down buttons for volume control of the speaker 111 and / or microphone 113. The one or more buttons optionally include a push button (e.g., 206 in FIG. 2).

[0040] Touch-sensitive surface 112 provides an input and output interface between the device and a user. Display controller 156 receives electrical signals from and / or sends electrical signals to touch surface 112. Touch surface 112 displays visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (collectively referred to as "graphics"). In some embodiments, some or all of the visual output optionally corresponds to user interface objects.

[0041] Touch surface 112 has a touch-sensitive surface, sensor, or set of sensors that accepts input from a user based on haptic and / or tactile contact. Touch surface 112 and display controller 156 (along with any associated modules and / or sets of instructions in memory 102) detect contacts (and any movement or cessation of contact) on touch screen 112 and translate the detected contacts into interactions with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on touch surface 112. In an exemplary embodiment, the point of contact between touch screen 112 and the user corresponds to the user's finger.

[0042] In some embodiments, in addition to the touchscreen, device 100 optionally includes a touchpad for enabling or disabling certain functions. In some embodiments, the touchpad is a touch-sensitive area of ​​the device that, unlike the touchscreen, does not display visual output. The touchpad is optionally a touch-sensitive surface that is separate from touchscreen 112 or an extension of the touch-sensitive surface formed by the touchscreen.

[0043] Device 100 also includes a power system 162 for powering the various components. Power system 162 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharging system, power fault detection circuitry, power converters or inverters, power status indicators (e.g., light emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in a portable device.

[0044] Device 100 also optionally includes one or more light sensors 164. FIG. 1 shows a light sensor coupled to light sensor controller 158 in I / O subsystem 106. Light sensor 164 optionally includes a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) phototransistor. Light sensor 164 receives light from the environment, projected through one or more lenses, and converts the light into data representing an image. In conjunction with imaging module 143 (also referred to as a camera module), light sensor 164 optionally captures still images or video. In some embodiments, the light sensor is located on the back of device 100, opposite touchscreen surface 112 on the front of the device, thereby allowing the touchscreen display to be used as a viewfinder for capturing still and / or video images. In some embodiments, the light sensor is located on the front of the device, thereby optionally allowing an image of a user to be captured for a video conference while the user is viewing other videoconference participants on the touchscreen display.

[0045] Device 100 also optionally includes one or more depth camera sensors 175. FIG. 1 shows a depth camera sensor coupled to depth camera controller 169 in I / O subsystem 106. Depth camera sensor 175 receives data from the environment and generates a three-dimensional model of an object (e.g., a face) in the scene from the viewpoint (e.g., the depth camera sensor). In some embodiments, depth camera sensor 175, together with imaging module 143 (also referred to as a camera module), is optionally used to determine a depth map of different portions of an image captured by imaging module 143. In some embodiments, the depth camera sensor is placed in front of device 100.

[0046] Device 100 also optionally includes one or more haptic output generators 167. Figure 1 shows haptic output generators coupled to haptic feedback controller 161 in I / O subsystem 106. Haptic output generator 167 optionally includes one or more electroacoustic devices, such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other haptic output-generating components (e.g., components that convert electrical signals into haptic output at the device).

[0047] Device 100 also optionally includes one or more accelerometers 168. Figure 1 shows accelerometer 168 coupled to peripheral interface 118. Alternatively, accelerometer 168 is optionally coupled to input controller 160 in I / O subsystem 106. Device 100 optionally includes, in addition to accelerometer 168, a magnetometer or GPS (or GLONASS or other global navigation satellite system (GNSS)) receiver that obtains information regarding the location and orientation (portrait or landscape) of device 100.

[0048] In some embodiments, the software components stored in memory 102 include an operating system 126, a communications module (or set of instructions) 128, a touch / motion module (or set of instructions) 130, a graphics module (or set of instructions) 132, a text input module (or set of instructions) 134, a global positioning system (GPS) module (or set of instructions) 135, and an application (or set of instructions) 136.

[0049] Operating system 126 (e.g., Android®, Tizen®, Darwin, RTXC, LINUX®, UNIX®, OSX®, iOS®, WINDOWS®, or an embedded operating system such as VxWorks®) includes various software components and / or drivers for controlling and managing common system tasks (e.g., memory management, storage device control, power management, etc.) and facilitating communication between various hardware and software modules.

[0050] Communications module 128 facilitates communication with other devices via one or more external ports 124 and also includes various software components for processing data received by RF circuitry 108 and / or external ports 124. External ports 124 (e.g., Universal Serial Bus (USB), FIREWIRE®, etc.) are adapted to couple to other devices directly or indirectly via a network (e.g., the Internet, wireless LAN, etc.).

[0051] Contact / action module 130 optionally (in conjunction with display controller 156) detects contact with touch surface 112 and other touch-sensitive devices (eg, a touchpad or physical click wheel).

[0052] Graphics module 132 includes various known software components for rendering and displaying graphics on touchscreen 112 or other display, including components that modify the visual impact (e.g., brightness, transparency, saturation, contrast, or other visual characteristics) of the displayed graphics. As used herein, the term "graphics" includes any object that can be displayed to a user, including, without limitation, text, web pages, icons (e.g., user interface objects including softkeys), digital images, video, animation, etc.

[0053] The haptic feedback module 133 includes various software components for generating instructions used by the haptic output generator 167 to generate haptic outputs at one or more locations on the device 100 in response to user interaction with the device 100.

[0054] Text input module 134 is optionally a component of graphics module 132 and provides a software keyboard for entering text in various applications (e.g., contacts 137, email 140, IM 141, browser 147, and any other application requiring text input).

[0055] The GPS module 135 determines the location of the device and provides this information for use in various applications (e.g., to the phone 138 for use in location-based calling, to the camera 143 as picture / video metadata, and to applications that provide location-based services such as weather widgets, local yellow pages widgets, and maps / navigation widgets).

[0056] The application 136 optionally includes the following modules (or sets of instructions), or a subset or superset thereof: a contacts module 137 (also called an address book or contact list); Telephone module 138, Videoconferencing module 139, an email client module 140; an instant messaging (IM) module 141; workout support module 142, a camera module 143 for still and / or video images; an image management module 144; Video player module, Music player module, Browser module 147, a video and music player module 152 that combines a video player module and a music player module; Map module 154, and / or Online Video Module 155.

[0057] Each of the above modules and applications corresponds to a set of executable instructions that perform one or more of the functions described above and methods described herein (e.g., the computer-implemented methods and other information processing information described herein). These modules (e.g., sets of instructions) need not be implemented as separate software programs, procedures, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. For example, a video player module may optionally be combined with a music player module into a single module (e.g., video and music player module 152 of FIG. 1). In some embodiments, memory 102 optionally stores a subset of the modules and data structures described above. Additionally, memory 102 optionally stores additional modules and data structures not described above.

[0058] In some embodiments, device 100 is a device in which operation of a predefined set of functions on the device is performed exclusively by a touch surface and / or touchpad. By using a touch surface (e.g., a touchscreen) and / or a touchpad as the primary input control device for operation of device 100, the number of physical input control devices (e.g., push buttons, dials, etc.) on the device is optionally reduced.

[0059] The set of predefined functions performed exclusively by the touchscreen and / or touchpad optionally includes navigation between user interfaces. In some embodiments, the touchpad, when touched by a user, navigates device 100 to a main, home, or root menu from any user interface displayed on device 100. In such embodiments, a "menu button" is implemented using the touchpad. In some other embodiments, the menu button is not a touchpad, but rather a physical push button or other physical input control device.

[0060] FIG. 2 depicts portable multifunction device 100 having touch surface 112, according to some embodiments. In this exemplary embodiment, touch surface 112 is a touch screen that optionally displays one or more graphics within a UI. In this embodiment, as well as others described below, a user can select one or more of the graphics by making a gesture on the graphics, for example, with one or more fingers (not shown) or one or more styluses (not shown). In some embodiments, selection of one or more graphics occurs when the user breaks contact with one or more graphics. In some embodiments, the gesture optionally includes one or more taps, one or more swipes (left to right, right to left, upward, and / or downward), and / or rotating a finger interacting with device 100 (right to left, left to right, upward, and / or downward). In some implementations or situations, inadvertent contact with a graphic does not cause that graphic to be selected. For example, a swipe gesture sweeping over an application icon optionally does not select the corresponding application if the gesture corresponding to selection is a tap.

[0061] Device 100 optionally includes one or more physical buttons, such as a "home" or menu button, that are optionally used to navigate any application 136 in a set of applications optionally running on device 100. Alternatively, in some embodiments, the menu button is implemented as a soft key within a graphical UI (GUI) displayed on touch surface 112. In some embodiments, device 100 includes touch surface 112, a menu button, a back button, and an overview button.

[0062] As used herein, the term "affordance" refers to a user-interactive GUI object optionally displayed on the display screen of device 100 (FIG. 1). For example, images (e.g., icons), buttons, and text (hyperlinks) each optionally constitute an affordance.

[0063] Attention is now directed to an embodiment of a UI and associated processes implemented on an electronic device such as portable multifunction device 100 shown in FIG.

[0064] 3A-3V depict exemplary user interfaces for capturing images, according to some embodiments. The illustrated user interfaces are used to explain the processes described below, including the processes described with reference to FIGS. 4A and 4B.

[0065] 3A , device 100 includes speaker 111, display 112 (e.g., a display device), microphone 113, one or more light sensors 164 (e.g., a camera, a front-facing camera), one or more output generators 167, and one or more depth sensors 175. In some embodiments, device 100 includes light sensor 164 but does not include depth sensor 175. In some embodiments, device 100 is a mobile phone, such as a smartphone. In some embodiments, device 100 includes one or more features of device 100.

[0066] As depicted in Figure 3A, capture user interface 302 is displayed on display 112 and includes a preview portion 304 that includes a target area (304-2) visually distinct from a non-target area (304-1), and an instruction prompt 306. Preview portion 304 continuously displays images (e.g., live video) as they are captured by camera 164. As depicted in Figure 3A, preview portion 304 displays an image of device user 303 who is positioned in front of camera 164 but off-center (e.g., relative to the camera's boresight and / or device centerline), such that user 303's head does not fall within target area 304-2.

[0067] 3B , in response to the text instructions (and / or spoken audio prompts) presented on display 112, user 303 repositions his or her head relative to camera 164 so that his or her head is within target area 304-2, as indicated by an image of user 303 being more centered relative to target area 304-2 and by a change in the visual appearance (e.g., color shading) of target area 304-2. In response to detecting this change in user 303's head position, device 100 displays a capture button 308 in user interface 302 to initiate the image capture process and ceases displaying instruction prompt 306 in user interface 302. In some embodiments, one or more other criteria are verified prior to displaying capture button 308 and / or ceasing display of instruction prompt 306.

[0068] Figure 3C depicts device 100 receiving user input 310-1 (e.g., a tap) on capture button 308. In response to receiving user input 310-1, device 100 displays user interface 302 as depicted in Figure 3D, which includes countdown animation graphic 312. After countdown animation graphic 312 completes playing, device 100 displays user interface 302 as depicted in Figure 3E.

[0069] As shown in Figure 3E, device 100 begins capturing a series of images (e.g., real-time images captured by a camera) corresponding to the image of device user 303 displayed in preview portion 304, and displays instructional prompt 314 to guide user 303 in repositioning himself / herself relative to the camera. As shown in Figure 3E, instructional prompt 314 includes both a text component and a symbolic component (e.g., a directional arrow). In some embodiments, device 100 outputs an audible prompt over speaker 111 corresponding to instructional prompt 314.

[0070] 3F represents user interface 302 after user 303 has rotated their head relative to the boresight of camera 164 (or more generally, relative to device 100) beyond a predetermined rotational angle (e.g., x degree yaw in a first rotational direction (e.g., x=50 degrees)). In user interface 302, instructional prompt 314 has been replaced by instructional prompt 316 ("Stop"), and the image is updated in preview portion 304 to reflect the change in user 303's head position.

[0071] 3G depicts device 100 after displaying instruction prompt 316 when user 303 begins to turn to face the camera. As depicted in FIG. 3G, instruction prompt 316 has been replaced by example instruction prompt 318 ("Turn around slowly and look at your phone").

[0072] 3H depicts user interface 302 after user 303 has rotated their head back to a centered position (e.g., user 303 is looking at the camera) relative to the boresight of camera 164 (or more generally, relative to device 100). As depicted in FIG. 3H, in user interface 302, instruction prompt 318 has been replaced by prompt 320 (a "check" mark) that indicates to user 303 that their portion of the capture process has been successfully completed.

[0073] 31 depicts device 100 after displaying instruction prompt 320 if the user begins to rotate away from the camera in another direction. As depicted in FIG. 31, instruction prompt 320 has been replaced by exemplary instruction prompt 322 that provides further positional guidance to user 303.

[0074] 3J represents user interface 302 after user 303 has rotated his or her head relative to the boresight of camera 164 (or more generally, relative to device 100) beyond a predetermined rotational angle (e.g., x degree yaw in a second rotational direction (e.g., x=50 degrees)). In user interface 302, instructional prompt 322 has been replaced by instructional prompt 324 ("Stop"), and the image is updated in preview portion 304 to reflect the change in user 303's head position.

[0075] 3K depicts device 100 after displaying instruction prompt 324 when user 303 begins to turn to face the camera. As depicted in FIG. 3K, instruction prompt 324 has been replaced by example instruction prompt 326 ("Turn around slowly and look at your phone").

[0076] 3L depicts user interface 302 after user 303 has rotated their head relative to the boresight of camera 164 (or more generally, relative to device 100) and returned to a centered position (e.g., user 303 is looking at the camera). As depicted in FIG. 3L, in user interface 302, instruction prompt 326 has been replaced by prompt 320 that indicates to user 303 that a portion of the capture process has been successfully completed, and a continue button 328 is displayed.

[0077] 3M depicts device 100 receiving user input 310-2 (e.g., touch input) on continue button 328. In response to receiving user input 310-2, device 100 displays demographic input user interface 330 as depicted in FIG. 3N, which includes birth gender affordances 332-1, 332-2, age affordance 334, and skip affordance 336.

[0078] 3O depicts device 100 receiving user input 310-3 (e.g., touch input) for birth gender affordance 332-1 (“Male”). In response to receiving user input 310-3, device 100 updates demographic input user interface 330 as depicted in FIG. 3P. As depicted in FIG. 3P, birth gender affordance 332-1 is highlighted and continue button 338 is displayed.

[0079] Figure 3Q depicts device 100 receiving user input 310-4 (e.g., touch input) for age affordance 334. In response to receiving user input 310-4, device 100 updates demographic input user interface 330 as depicted in Figure 3R, which includes a virtual keyboard 340 for entry of numerical data. Figure 3S depicts device 100 receiving user input 310-5 (e.g., tap) for completion affordance 341, this time displaying user interface 330 including numerical age data displayed in age affordance 334.

[0080] 3T depicts device 100 receiving user input 310-6 (e.g., touch input) at continue affordance 338. In response to receiving user input 310-6, device 100 displays PHRTF processing interface 342 as depicted in FIG. 3U while generating personalized PHRTF data based on the series of images captured throughout the guided image capture process and the data entered via demographic input user interface 330.

[0081] FIG. 3V depicts device 100 displaying PHRTF processing interface 342 after generation of personalized PHRTF data is complete, including a done button 344.

[0082] 4A and 4B are flow diagrams illustrating a process for capturing image data using an electronic device, according to some embodiments. Process 400 is performed on an electronic device (e.g., 100) that includes a display (e.g., 112) and a camera (e.g., 164). In some embodiments, the electronic device also includes a set of sensors (e.g., motion sensors such as a gyroscope, accelerometer, etc.). Some operations in process 400 may be arbitrarily combined, the order of some operations may be arbitrarily changed, and some operations may be arbitrarily omitted.

[0083] As described below, process 400 provides an intuitive method for capturing image data, particularly image data relevant to a user of an electronic device suitable for personalizing audio output from the device. The method reduces the cognitive load on a user attempting to capture image data suitable for image-based device personalization, thereby producing a more efficient human-machine interface. For battery-powered computing devices, allowing users to monitor sound exposure levels more quickly and efficiently can more efficiently conserve power and extend the time between battery charges.

[0084] In step 402, the electronic device displays a user interface on a display that includes a preview portion that displays images (individual images or frames of video) that correspond to image data captured by the camera.

[0085] At step 408, the electronic device performs a capture process that includes, at step 408-1, using a camera, capturing a series of images corresponding to a preview image in the preview portion and determining pose data associated with the series of images, and at step 408-2, outputting a first set of instructional prompts (e.g., a prompt instructing the device user to stop rotating their head away from the device (e.g., 316) and / or a prompt instructing the device user to rotate their head towards the device (e.g., 318)) in accordance with the pose data satisfying or exceeding a first threshold, and outputting a second set of instructional prompts (e.g., a prompt instructing the device user to stop rotating their head away from the device (e.g., 324) and / or a prompt instructing the device user to rotate their head towards the device (e.g., 326)) in accordance with the pose data satisfying or exceeding a second threshold.

[0086] In some embodiments, the series of images includes the head and / or torso of a user of the device. In some embodiments, the current pose data includes angular values ​​(e.g., degrees or radians) associated with the orientation (e.g., estimated pitch, yaw, or roll) of the user's head displayed in each image. In some embodiments, the current pose data includes velocity and / or acceleration values ​​associated with the angular values.

[0087] In some embodiments, the threshold is a fixed angle value of rotation in a first direction (e.g., x degrees counterclockwise (x=50 degrees, 45 degrees, etc.)). In some embodiments, the threshold is a fixed angle value of rotation in a first direction (e.g., x degrees clockwise (x=50 degrees, 45 degrees, etc.)). In some embodiments, the first threshold and the second threshold have the same magnitude.

[0088] In some embodiments, the first thresholds have opposite magnitudes (eg, −50 degrees and +50 degrees).

[0089] At step 410, the electronic device, in response to determining that the set of sufficiency criteria has been met, aborts the capture process (eg, causes the device to cease simultaneously capturing and determining relevant pose data).

[0090] At step 412, the device generates data corresponding to the set of PHRTFs based on a subset of the series of images associated with pose data satisfying or exceeding a first threshold, a subset of the series of images associated with pose data satisfying or exceeding a second threshold, and a generative model (e.g., a machine learning model (e.g., a deep learning neural network) trained to output PHRTF data from the input data), where the subsets of the series of images are non-overlapping subsets of the series of images. In some embodiments, the input data may be derived from the subset of the series of images associated with pose data satisfying or exceeding the first threshold and the subset of the series of images associated with pose data satisfying or exceeding a second threshold. In some embodiments, the input data may be derived from user input received at the electronic device (see, e.g., FIGS. 3N-3T and corresponding description). In some embodiments, the generative model includes a convolutional neural network.

[0091] In steps 414-418, respectively, the device receives original audio data, processes the original data based on data corresponding to a pair of PHRTFs in the set of PHRTFs for the original audio to generate personalized audio data, and outputs the personalized audio data (e.g., via a speaker of the electronic device or a speaker physically or wirelessly coupled to the electronic device).

[0092] In some embodiments, generating the data corresponding to the set of PHRTFs is further based on a subset(s) of images in the series of images corresponding to a frontal pose (e.g., images in which the user's head is facing towards the camera; images captured while the countdown animated graphic 312 is displayed; images captured after capturing the subset of images in the series associated with pose data satisfying or exceeding a first threshold and before capturing the subset of images in the series associated with pose data satisfying or exceeding a second threshold), where the subsets of images in the series corresponding to a frontal pose are non-overlapping subsets of the series of images.

[0093] In some embodiments, the subset of images in the series of images corresponding to a frontal pose, the subset of images in the series of images associated with pose data satisfying or exceeding a first threshold, or the subset of images in the series of images associated with pose data satisfying or exceeding a second threshold are selected based on at least one of image metadata (e.g., ISO or shutter speed) for each image in the series of images, image pose data (e.g., yaw) for each image in the series of images, and device position data (e.g., motion data at the time of capture) associated with each image in the series of images.

[0094] In some embodiments, the device obtains demographic data associated with the device user and generates data corresponding to the set of PHRTFs further based on the demographic data. In some embodiments, generating the data corresponding to the set of PHRTFs occurs at the device. In some embodiments, generating the data corresponding to the set of PHRTFs includes sending the series of images or a subset of the series of images to a server device different from the device and receiving data corresponding to the set of PHRTFs from the server device, where the subset of the series of images is a non-overlapping subset of the series of images.

[0095] In some embodiments, after halting the capture process, the device displays a second user interface (e.g., 330) including one or more affordances related to at least one of the user's age and the user's birth sex (or gender), and receives demographic data via user input in one or more of the affordances related to at least one of the user's age and the user's birth sex (or gender). Birth sex refers to the sex (male or female) assigned to an infant, most often based on anatomy and other biological characteristics during infancy (e.g., reproductive organs and functions derived from chromosomal complement (typically XX for females and XY for males)). It is also sometimes referred to as birth sex, natal sex, biological sex or sex, sex assigned at birth, or gender assigned at birth.

[0096] In some embodiments, the current pose data represents a yaw angle value, and the first and second thresholds are yaw angle values ​​related to a pose of the user's head. In some embodiments, the pose data includes a yaw velocity value and / or a yaw acceleration value.

[0097] In some embodiments, capturing the series of images using a camera is performed at a frame rate that varies based on at least one of image metadata (e.g., ISO or shutter speed) for each image in the series, pose data (e.g., yaw) for each image in the series, and device position data (e.g., movement data at the time of capture) associated with each image in the series.

[0098] In some embodiments, the device outputs one or more prompts (e.g., auditory and / or visual commands to slow the user's head) in response to determining that pose data associated with angular velocity and / or angular acceleration exceeds a threshold.

[0099] In some embodiments, the device outputs the first or second set of instructional prompts after a period of time pursuant to a determination that current pose data is unavailable (e.g., the device is unable to determine a pose for one or more images / frames).

[0100] In some embodiments, the duration is calculated based on characteristic landmarks determined in each image in the sequence of images and at least one of a pose velocity value or a pose acceleration value associated with each image. In some embodiments, outputting the first set of instructional prompts is performed in accordance with a determination that one or more velocity and / or acceleration values ​​are below a respective velocity and / or acceleration threshold. In some embodiments, outputting the second set of instructional prompts is performed in accordance with a determination that one or more velocity and / or acceleration values ​​are below a respective velocity and / or acceleration threshold.

[0101] In some embodiments, after satisfying or exceeding the first or second threshold, a third affordance (e.g., a confirmation graphic, check mark, or continue graphic; 328) is displayed according to the current pose data satisfying or exceeding a third threshold between the first and second thresholds (e.g., a reference pose angle corresponding to a pose facing the camera; see, e.g., 303 in FIG. 3L).

[0102] In some embodiments, the device repeats the capture process following a determination that the first or second thresholds have not been met or exceeded.

[0103] In some embodiments, after the first and second thresholds are met or exceeded, a third affordance is displayed (e.g., a checkmark or continue graphic, 328) pursuant to a determination that the number of images in the series of images representing the user's left ear, right ear, and frontal views exceeds a predetermined threshold.

[0104] In some embodiments, satisfying the sufficiency criteria includes determining that the series of images includes a set of one or more images associated with pose data that meets or exceeds a first threshold and a set of one or more images associated with pose data that meets or exceeds a second threshold. In some embodiments, satisfying the set of sufficiency criteria includes determining that the current pose data met or exceeded at least the first threshold and / or the second threshold during the capture process.

[0105] In some embodiments, when the device is displaying a first preview portion displaying an image captured by the camera, the device does not display a first affordance (e.g., 308), but calculates first data associated with a first set of capture criteria, and in accordance with a determination that the first set of capture criteria have been met based on the first data, displays the first affordance in a first user interface, outputs a third set of instruction prompts, and determines reference pose data (e.g., a reference pose) associated with the image initially captured by the camera before performing the capture process.

[0106] In some embodiments, the device, pursuant to a determination based on the first data that the first set of capture criteria are not met, does not display the first affordance and causes the display of visual feedback (e.g., 306) or the output of auditory feedback associated with at least one criterion in the first set of capture criteria.

[0107] In some embodiments, the first set of capture criteria includes a device position or device orientation, which in some embodiments relates to the device being positioned vertically (e.g., the camera is oriented approximately perpendicular to gravity), the device being within a specified distance from the user's face (e.g., within a range estimated based on image data from the device's camera or depth sensor (e.g., Lidar), such as between 15 cm and 60 cm), or the device's camera being positioned to squint the user's face in the center of the image (e.g., the user's face is centered relative to the camera position).

[0108] In some embodiments, the first set of capture criteria includes at least one of: a face is detected (e.g., the device detects a face in image data captured by the camera); adequate lighting is detected (e.g., image metadata indicates that the ISO is below the iso threshold); or no ear or eye occlusion is detected (e.g., the device detects that eyewear or hair is covering the ears in image data captured by the camera).

[0109] In some embodiments, an augmented reality framework on the electronic device is used for head pose estimation and camera pose estimation, and the subset of image frames provided by the framework is selected such that the angular distance between frames on each side of the user's head is minimal (e.g., 0.6 degrees).

[0110] In some embodiments, image frames are selected for each side of the user's head within a specified angular distance range (e.g., 20-50 degrees in yaw). The selected images are input to a machine learning model (e.g., a deep learning neural network) to detect the user's ears. In some embodiments, the network (e.g., a convolutional neural network) is trained to detect human ears, for example, using images of ears under different conditions (e.g., different angles, listening situations, etc.). In some embodiments, synthetic or augmented images of ears are included in the training data. Ear landmark detection is then performed on frames in which ears are detected by the machine learning model.

[0111] In some embodiments, the subsets of captured images (e.g., the subset of images in the series associated with pose data satisfying or exceeding a first threshold, the subset of images in the series associated with pose data satisfying or exceeding a second threshold) are selected so that the angular distance of head rotation between image frames is minimal (e.g., 0.6 degrees). Constraining the angular distance between images has several technical effects, including ensuring that the image capture process is efficient by limiting capture to only data useful for image-based PHRTF generation. For example, image-based PHRTF generation techniques that utilize multi-angle adjustment techniques (e.g., to project 2D data into 3D coordinates) can suffer from reduced accuracy when the angular distance between pose angles of the 2D input image data is insufficient.

[0112] In some embodiments, a subset of the captured images (e.g., a subset of the series of images associated with pose data meeting or exceeding a first threshold, a subset of the series of images associated with pose data meeting or exceeding a second threshold) is selected to constrain the images by an overall angular distance range (e.g., a yaw angle range). Constraining the angular distance range has several technical effects, including ensuring that the image capture process is efficient by limiting capture to only reliable data. For example, pose angles provided by a face recognition algorithm may become unreliable beyond a certain maximum angle of head rotation.

[0113] In some embodiments, image data from one ear (e.g., the left ear) is used to derive image data to be used in place of data from the other ear (e.g., the right ear) under the assumption that the ears are symmetrical.

[0114] In some embodiments, if there are an insufficient number of suitable images, demographic head information (e.g., average head diameter / radius for males or females) may be used to calculate a default PHRTF.

[0115] 5A-5K are screenshots of alternative user interfaces for capturing image data using an electronic device according to some embodiments.

[0116] As shown in FIG. 5A , capture user interface 502 includes a preview portion 504 that includes a target area (504-2) that is visually distinct from a non-target area (504-1), and an instructional prompt 506. In the example shown, instructional prompt 506 states, "Keep your head in the center of the frame." Preview portion 504 displays images as they are captured by camera 164. As shown in FIG. 5A , preview portion 504 displays an image of user 503 positioned in front of light sensor 164 and depth sensor 175, in accordance with instructional prompt 506, so that user 503's head is centered within target area 504-2.

[0117] As shown in FIG. 5B , in response to determining that the user's head is centered in target area 504-2, instructional prompt 506 is replaced by exemplary feedback prompt 507 (in this example, "Great!"). This exemplary feedback prompt 507 informs the user that their head is centered in the frame. Also shown in FIG. 5B are directional indicators 508-1 through 508-4 located above, below, and on either side of target area 504-2. Directional indicators 508-1 and 508-2 indicate head pitch direction (tilting the head up or down), and directional indicators 508-3 and 508-4 indicate head yaw direction (turning the head right or left). In the example shown, the directional indicators are arrows. In other embodiments, the directional indicators can be any other suitable graphic capable of indicating left and right (e.g., yaw angle direction).

[0118] 5C , instructional prompt 506 now includes the exemplary text, "Tilt your head as instructed." Additionally, directional indicator 508-3 (e.g., an arrow) changes to a different color than directional indicators 508-1, 508-2, and 508-4 (also arrows in this example) to indicate to user 503 the requested tilt direction. For example, directional indicator 508-3 can change from green to red, while directional indicators 508-1, 508-2, and 508-4 remain green. In some embodiments, different colors or graphics may be used for the directional indicators, and / or different directional indicator styles (e.g., flashing arrows) may be used, which may be static, animated, or both.

[0119] As shown in FIG. 5D, user 503 follows instruction prompt 506, which results in another instruction prompt 506 being displayed, including the exemplary instruction "Stay Still." An image of user 503's head is captured by light sensor 164 and depth sensor 175. In some embodiments, a countdown animated graphic is displayed to indicate that the image is being captured. After the countdown animated graphic has completed playing, device 100 displays user interface 502, as shown in FIG. 5E.

[0120] 5E, instructional prompt 506 now displays the exemplary instruction, "Turn your head slowly and look over your right shoulder." A directional arrow 509 is also displayed to indicate the direction in which the user should turn their head. Also, in some embodiments, the right boundary (right half) of target portion 504-2 is highlighted (as shown) or otherwise animated to indicate the direction in which user 503 should turn their head.

[0121] 5F, user 503 follows the prompts and an image of user 503's right ear is captured by light sensor 164 and depth sensor 175. In some embodiments, a countdown animated graphic is displayed to indicate that the image is being captured. After the image is captured, instruction prompt 506 includes the exemplary text, "Turn around and look at your phone."

[0122] As shown in Figure 5G, user 503 follows the prompt and exemplary feedback prompt 507 is displayed, stating "Great!" and directional indicators 508-1 through 508-4 are again displayed. As shown in Figure 5H, second instructional prompt 506 includes exemplary text that reads, "Stay Still."

[0123] 5I and 5J are similar to Figures 5E and 5F, but are mirror images and are directed to capturing an image of the left ear of user 503. As shown in Figure 5K, feedback prompt 507 includes exemplary text "Capture complete!", which indicates to the user that the image capture process is complete.

[0124] In some embodiments, instruction prompts 506 and feedback prompts 507 are accompanied by or replaced by spoken audio for instruction prompts 506 and haptic / tactile feedback for feedback prompts 507 .

[0125] 6A-6E are screenshots of another alternative user interface 600 for generating a PHRTF without capturing an image, according to some embodiments. The screenshots shown in user interface 600 depict user input of head size data using a slider affordance and can be used in place of the process described with reference to FIGS. 3A-3L.

[0126] Referring to FIG. 6A, slider affordance 601-1 updates silhouette 602-1 (e.g., increases silhouette size) and updates text description 603-1 (e.g., "extra small," "approximately 6 3 / 4 (XS) hat size") when slider affordance 601-1 is repositioned by the user.

[0127] In addition to the user-selected head size data, the user inputs their demographic data (e.g., selection of an affordance corresponding to the user's birth sex or gender), which, along with the head size data, is used to generate the PHRTF without using captured image data. When affordance 604-1 (e.g., a virtual button labeled "Create Profile") is selected by the user, generation of the PHRTF based on the head size data and demographic data begins. Figures 6B-6E illustrate the same process described with reference to Figure 6A for small, medium, large, and extra-large head sizes, respectively.

[0128] 7 is a flow diagram illustrating a process for generating a PHRTF according to some embodiments. Process 700 is performed on an electronic device (e.g., 100) that includes a display (e.g., 112). In some embodiments, the electronic device also includes a set of sensors (e.g., motion sensors such as gyroscopes, accelerometers, etc.). Some operations in process 700 may be arbitrarily combined, the order of some operations may be arbitrarily changed, and some operations may be arbitrarily omitted.

[0129] As described below, process 700 provides an intuitive way to generate a PHRTF without capturing an image. The process reduces the cognitive load on a user attempting to capture image data suitable for image-based device personalization, thereby creating a more efficient human-machine interface. For battery-powered computing devices, allowing a user to more quickly and efficiently capture image data for personalizing audio output from the device can more efficiently conserve power and extend the time between battery charges.

[0130] Process 700 may be performed on an electronic device with a display. Process 700 includes displaying (701) a user interface on the display including a graphical object of a first size, a slider affordance including an indicator located at a first position of the slider affordance, and a first text description corresponding to the first position of the indicator of the slider affordance. While displaying the graphical object, process 700 continues by detecting (702) a user input, and in response to detecting the user input, includes updating (703) the position of the indicator at the slider affordance from the first position to a second position, displaying (704) the graphical object at a second size smaller or larger than the first size, and displaying (705) a second text description corresponding to the second position of the indicator of the slider affordance in place of the first text description.

[0131] In some embodiments, the process 700 further includes generating PHRTF data corresponding to a set of personalized head-related transfer functions (PHRTFs) based on the data corresponding to the display size and the first or second text description of the graphical object and the generative model.

[0132] In some embodiments, generating PHRTF data corresponding to a set of functions associated with the PHRTF occurs in response to detecting user input in a creation profile affordance (e.g., 604) presented on a touch-sensitive display.

[0133] In some embodiments, generating the PHRTF data includes applying demographic data associated with the user (e.g., the user's age and the user's birth gender; see Figures 3N-3T) to a generative model.

[0134] In some embodiments, generating the PHRTF data excludes the use of image data corresponding to the user.

[0135] In some embodiments, the first size of the graphical objects (see, eg, 602-1) is smaller than the second size of the graphical objects (see, eg, 602-2, 602-3, 602-4, etc.).

[0136] In some embodiments, the first size of the graphical object is larger than the second size of the graphical object.

[0137] In some embodiments, detecting user input includes, but is not limited to, detecting slide and / or drag input (e.g., detecting the start of input at a first location and the end of input at a second location). For example, finger down, stylus down, mouse click and hold at a location corresponding to the indicator in the first location can be followed by finger up, stylus up, and mouse click release at a location corresponding to the indicator in the second location.

[0138] 8 is a flow diagram illustrating another process 800 for capturing image data using an electronic device, according to some embodiments. Process 800 is performed by portable multifunction device 100 described with reference to FIG.

[0139] As described below, process 800 provides an intuitive method for capturing image data, particularly image data relevant to a user of an electronic device, suitable for personalizing audio output from the device. The process reduces the cognitive load on a user attempting to capture image data for image-based device personalization, thereby producing a more efficient human-machine interface. For battery-powered computing devices, enabling a user to more quickly and efficiently capture image data for personalizing audio output from the device can more efficiently conserve power and extend the time between battery charges. The process also results in computationally efficient selection of a subset of captured image data suitable for downstream device personalization processes (e.g., PHRTF generation).

[0140] Process 800 includes presenting (801) a user interface on a display of an electronic device. The user interface includes a preview portion having a target area displaying images captured by the device's camera. Process 800 continues by presenting (802) a first instruction to the user in the user interface to rotate the user's head in a first direction and capturing (803) a first set of images including the user's first ear within the target area. Process 800 continues by presenting (804) a second instruction to the user in the user interface to rotate the user's head in a second direction opposite the first direction and capturing (805) a second set of images including the user's second ear within the target area. Process 800 continues by generating (806) a final set of images from the first and second sets of images. A subset of the final set of images is selected to minimize the angular head rotation distance between image frames and to constrain the images by a total angular distance range (e.g., a maximum yaw angle range of less than 110 degrees). Process 800 continues by generating 807 data corresponding to a set of PHRTFs for the user based on the final set of images.

[0141] In some embodiments, capturing a first set of images of a first ear of the user within the target area and capturing a second set of images of a second ear of the user within the target area includes determining a respective angular head rotation for each image. In some embodiments, the final set of images is selected such that the overall maximum angular distance of head rotation between the first image in the final set of images and the last image in the final set of images is less than 110 degrees (e.g., 90 degrees). That is, selecting the final set of images includes a process based on the determined respective angular head rotation data.

[0142] In some embodiments, the device scales at least a portion of the first and second sets of images based on one or more front view images of the user before generating 807 data corresponding to a set of PHRTFs for the user based on the final set of images. The term “scaling” refers to ratio-based adjustment (e.g., resizing) of pose coordinates or reference coordinates (e.g., 3D world coordinates, anatomical coordinates, photogrammetric coordinates, landmark coordinates) corresponding to the image data. In some embodiments, the device captures a front view image of the user before capturing the first set of images of the user's first ear. In some embodiments, the device captures a front view image of the user after determining the alignment of the user with the camera and during the display of a countdown animation (e.g., during one or more of the steps depicted in FIGS. 3D , 5D , 5H , etc.). In some embodiments, the device captures a front view image of the user after capturing the first set of images of the user's first ear. In some embodiments, the device may capture a front view image during a transition phase between capturing the first set of images of the user's first ear and the second set of images of the user's second ear (e.g., during one or more of the steps depicted in FIGS. 3H, 5G, 5H, etc.). In some embodiments, the device captures a front view image of the user after capturing the first set of images of the user's first ear and after capturing the second set of images of the user's second ear (e.g., during one or more of the steps depicted in FIGS. 3L, 5K, etc.).

[0143] The foregoing description has been described with reference to specific embodiments for purposes of explanation. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the technology and their practical application so that those skilled in the art can best utilize the technology and various embodiments with various modifications as suited to the particular use contemplated.

[0144] Although the disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art, and such changes and modifications are to be understood as being included within the scope of the disclosure and examples as defined by the claims.

Claims

1. An electronic device having a display and a camera, displaying a user interface on the display including a preview portion displaying an image corresponding to image data captured by the camera; performing a capture process while displaying the image in the preview portion; halting the capture process in response to determining that a set of sufficiency criteria has been met; and and The capture process comprises: capturing a series of images using the camera corresponding to preview images of the preview portion; determining pose data associated with the series of images; outputting a first set of instructional prompts in accordance with the pause data meeting or exceeding a first threshold; outputting a second set of instruction prompts in accordance with the pause data meeting or exceeding a second threshold; and A method comprising:

2. generating data corresponding to a set of personalized head-related transfer functions (PHRTFs) based on a subset of the series of images associated with the pose data satisfying or exceeding the first threshold, a subset of the series of images associated with the pose data satisfying or exceeding the second threshold, and a generative model. The method of claim 1.

3. receiving the original audio data; processing the original audio data based on data corresponding to a pair of PHRTFs in the set of PHRTFs for the original audio data to generate personalized audio data; outputting the personalized audio data; The method of claim 2 further comprising:

4. generating the data corresponding to the set of PHRTFs is further based on scaling information derived from a subset of images in the series of images corresponding to a frontal pose of a user. The method of claim 2.

5. The subset of images in the series of images corresponding to the frontal pose, the subset of images in the series associated with pose data satisfying or exceeding the first threshold, or the subset of images in the series associated with pose data satisfying or exceeding the second threshold, image metadata for each image in the series of images; pose data for each image in the sequence of images; and device location data associated with each image in the sequence of images based on at least one of The method of claim 4.

6. obtaining demographic data associated with the user; generating the data corresponding to the set of PHRTFs is further based on the demographic data. The method of claim 2.

7. generating the data corresponding to the set of PHRTFs is performed on the electronic device; The method of claim 2.

8. Generating the data corresponding to the set of PHRTFs includes: sending the series of images or a subset of the series of images to a server device different from the electronic device; receiving the data corresponding to the set of PHRTFs from the server device; Including, The method of claim 2.

9. After terminating the capture process, displaying a second user interface including one or more affordances associated with at least one of the user's age or the user's birth gender; receiving demographic data via user input in one or more of the affordances related to the at least one of the user's age or the user's birth sex; The method of claim 1 further comprising:

10. the pose data represents a yaw angle value; the first threshold and the second threshold are yaw angle values ​​associated with a user's head; The method of claim 1.

11. the pose data includes at least one of a yaw velocity value or a yaw acceleration value; The method of claim 1.

12. Capturing the series of images with the camera includes: image metadata for each image in the series of images; current pose data for each image in the sequence of images; or device location data associated with each image in the sequence of images at a frame rate that varies based on at least one of: The method of claim 1.

13. outputting one or more prompts in response to determining that the pose data related to angular velocity or acceleration exceeds the first threshold or the second threshold. The method of claim 1.

14. and outputting the first set of instruction prompts or the second set of instruction prompts after a period of time in accordance with determining that the pause data is unavailable. The method of claim 1.

15. the duration is calculated based on characteristic landmarks determined in each image in the sequence of images and at least one of pose velocity data or pose acceleration data associated with each image; 15. The method of claim 14.

16. outputting the first set of instructional prompts occurs in accordance with a determination that one or more velocity or acceleration values ​​are below a respective velocity or acceleration threshold; The method of claim 1.

17. outputting the second set of instructional prompts occurs in accordance with a determination that one or more velocity or acceleration values ​​are below a respective velocity or acceleration threshold. The method of claim 1.

18. and displaying a third affordance according to the pose data satisfying or exceeding a third threshold between the first and second thresholds after the first threshold or the second threshold has been satisfied or exceeded. The method of claim 1.

19. repeating the capture process in response to a determination that the first threshold or the second threshold has not been met or exceeded. The method of claim 1.

20. and displaying a third affordance in accordance with a determination that the set of images includes a number of images representing the user's left ear, right ear, and frontal views that exceeds a predetermined threshold after the first threshold and the second threshold are met or exceeded. The method of claim 1.

21. Satisfying the set of sufficiency criteria means the series of images a set of one or more images associated with the pose data meeting or exceeding the first threshold; and a set of one or more images associated with said pose data meeting or exceeding said second threshold; determining that the The method of claim 1.

22. Satisfying the set of sufficiency criteria means determining whether the pose data meets or exceeds at least the first threshold or the second threshold during the capture process; The method of claim 1.

23. When the preview portion displaying the image captured by the camera is displayed, a first affordance is not displayed, calculating first data associated with a first set of capture criteria; upon determining, based on the first data, that the first set of capture criteria has been met; Displaying the first affordance in the user interface; outputting a third set of instruction prompts; and determining reference pose data associated with an image captured by the camera prior to performing the capture process; The method of claim 1 further comprising:

24. in response to a determination based on the first data that the first set of capture criteria is not satisfied; not displaying the first affordance; causing a display of visual feedback or an output of auditory feedback associated with at least one criterion in the first set of capture criteria; 24. The method of claim 23, further comprising:

25. the first set of capture criteria includes a device position or a device orientation; 24. The method of claim 23.

26. the first set of capture criteria includes at least one of: a face is detected; adequate lighting is detected; or no ear or eye occlusion is detected; 24. The method of claim 23.

27. determining that at least one image in the series of images does not meet the sufficiency criteria; replacing said at least one image in said series of images with another image in said series of images; The method of claim 1 further comprising:

28. determining whether a specified number of images in the set of images meet the sufficiency criteria; calculating a default PHRTF for the user using demographic head information in accordance with the specified number of images in the series of images not meeting the sufficiency criteria; and The method of claim 1 further comprising:

29. a subset of the series of images is selected such that the angular distance of head rotation between image frames is minimized; The method of claim 1.

30. the subset of images in the sequence of images is constrained by a total angular range of head rotation of less than 110 degrees; The method of claim 1.

31. scaling a subset of images in the series of images capturing the ears of the user based on one or more front view images of the user in the series of images. The method of claim 1.

32. presenting a user interface on a display of an electronic device, the user interface including a preview portion displaying an image captured by a camera of the electronic device; presenting a first command to the user at the user interface to rotate the user's head in a first direction; capturing a first set of images including a first ear of the user in the preview portion; presenting a second command to the user at the user interface to rotate the user's head in a second direction opposite the first direction; capturing a second set of images in the preview portion that includes a second ear of the user; generating a final set of images from the first and second sets of images by selecting a subset of the first and second sets of images that has a minimum angular distance in head rotation between pairs of adjacent images in the final set of images; generating data corresponding to a set of personalized head-related transfer functions (PHRTFs) for the user based on the final set of images; A method having the following.

33. the final set of images is selected such that the maximum overall angular distance of the user's head rotation between a first image in the final set of images and a last image in the final set of images is less than 110 degrees.

33. The method of claim 32.

34. scaling at least a portion of the first set of images and the second set of images based on one or more front view images of the user.

33. The method of claim 32.

35. An electronic device having a display, displaying on the display a user interface including a slider affordance including a graphical object of a first size, an indicator located at a first position of the slider affordance, and a first text description corresponding to the first position of the indicator of the slider affordance; When displaying said graphical object, Detect user input, in response to detecting the user input; updating a position of the indicator from the first position to a second position with the slider affordance; displaying the graphical object at a second size that is smaller or larger than the first size; displaying, in place of the first text description, a second text description corresponding to the second position of the indicator of the slider affordance; and A method having the following.

36. generating personalized head-related transfer function (PHRTF) data corresponding to a set of PHRTF-related transfer functions based on a generative model and data corresponding to a display size of the graphical object and the first text description or the second text description.

36. The method of claim 35.

37. generating the PHRTF data corresponding to a set of transfer functions associated with the PHRTFs in response to detecting user input at a creation profile affordance presented on the display; 37. The method of claim 36.

38. generating the PHRTF data includes applying demographic data associated with the user to the generative model; 37. The method of claim 36.

39. generating the PHRTF data excludes the use of image data corresponding to the user; 37. The method of claim 36.

40. the first size of the graphical object is smaller than the second size of the graphical object; 39. The method of claim 38.

41. the first size of the graphical object is greater than the second size of the graphical object; 39. The method of claim 38.

42. A computer program comprising instructions which, when executed by a computing device, cause the computing device to carry out a method according to any one of claims 1 to 41.

43. 42. A non-transitory computer readable storage medium storing instructions that, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 41.

44. 1. A computing device, comprising: The display and at least one processor; a memory storing instructions; The instructions, when executed by the at least one processor, cause the computing device to perform the method of any one of claims 1 to 41. computing device.