Real-time feedback to improve image capture for faces
Real-time sensory feedback systems help blind and low-vision users capture better self-portraits by guiding them to align their face and camera angles and adjust lighting, addressing the challenges of poor framing and angles in existing technologies.
Patent Information
- Application Number
- PCT/US2024/041908
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-12
- Publication Date
- 2026-02-19
AI Technical Summary
Blind and low-vision users face difficulties in capturing high-quality self-portrait photographs due to challenges in positioning themselves and objects within the camera's field-of-view, leading to poor framing and angles.
Providing real-time sensory feedback, such as haptic, audio, and visual cues, to guide users in aligning their face and camera angles, and adjusting lighting conditions for optimal image capture.
Enhances the quality of self-portrait photographs by ensuring accurate framing and alignment, even for users with visual impairments, by using sensory feedback to adjust positioning and lighting.
Smart Images

Figure US2024041908_19022026_PF_FP_ABST
Abstract
Description
REAL-TIME FEEDBACK TO IMPROVE IMAGE CAPTURE FOR FACESBACKGROUND
[0001] Front-facing cameras on computing devices allow a user of the computing device to capture self-portrait photographs, also referred to as “selfies.” When the user takes a selfie, they may look at a display of the computing device to position themselves, and any other desired objects for the photograph, respective to the field-of-view of the camera (e.g., a viewing frame). Positioning a camera for a selfie can be particularly challenging for blind and low- vision users (collectively “low-vision users” herein). If it is difficult for the low-vision user to see the display screen, the quality of the selfie may be poor with undesirable framing and angles.SUMMARY
[0002] This document describes systems and techniques directed to providing real-time feedback to improve self-portrait photographs (aka “selfies”) as well as other image capture for a camera user (e.g., a low-vision user). A selfie is captured by a device or component capable of capturing an image (e.g., camera, front-facing camera) of a computing device (e.g., mobile device, smartphone, tablet computer) and may include one or more objects (e.g., a user’s face). The disclosed systems and techniques may track a user’s face and provide sensory feedback (e.g., haptic, audio, or visual feedback) to guide the user to position the computing device and / or the user so that the user becomes positioned in a center of frame of the camera. Additional sensory feedback may then be provided to guide the user to align an angle of the object (e.g., the user’s face) with an angle of the camera. By aligning the two angles, a more desirable image may be captured. Additional examples involve providing guidance for lighting. Further examples involve facilitating image capture when multiple objects are present, by considering saliency.
[0003] In one aspect, a method includes receiving an input from a user to initiate a camera of a computing device to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angledifference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
[0004] In another aspect, a computing device is configured to receive an input from a user to initiate a camera to capture an image of an environment; detect an object of interest in the environment; provide sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determine an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, provide additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiate capture of an image of the object.
[0005] In a further aspect, a non-transitory computer readable medium is provided comprising program instructions executable by one or more processors to perform functions comprising: receiving an input from a user to initiate a camera to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
[0006] In another aspect, a computer program is provided comprising program instructions executable to perform functions comprising: receiving an input from a user to initiate a camera to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
[0007] In another aspect, means are provided for performing the method described above.
[0008] This Summary is provided to introduce simplified concepts for providing realtime feedback to improve self-portrait photographs and other types of image capture for camera users, which is further described below in the Detailed Description and is illustrated in the Drawings. This Summary is intended neither to identify essential features of the claimed subject matter nor for use in determining the scope of the claimed subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The details of one or more aspects of providing real-time feedback to improve self-portrait photographs and other types of image capture for camera users are described in this document with reference to the following drawings, where the use of same numbers in different instances may indicate similar features or components:Figure 1 is an example computing device, in accordance with examples described herein;Figure 2 is a simplified block diagram showing some of the components of an example computing system, in accordance with examples described herein;Figure 3 is a schematic illustration of an example computing device, in accordance with examples described herein;Figure 4 illustrates different angle positions for selfies, in accordance with examples described herein;Figure 5 illustrates a sequence of face positions and angles for a selfie, in accordance with examples described herein;Figure 6 illustrates components configured to provide angle and lighting guidance, in accordance with examples described herein;Figure 7 illustrates components configured to handle images with multiple objects, in accordance with examples described herein; andFigure 8 is a block diagram of an example method, in accordance with examples described herein.DETAILED DESCRIPTION
[0010] Blind and low vision users have difficulty taking good photos, including selfies. In particular, these users may find it difficult to center an object of interest for image capture. Sensory (haptic, audio, and / or visual) feedback may be provided to guide a user to center the object of interest within a viewfinder of the camera. However, if the object (particularly a face for a selfie) is at an angle that significantly differs from an angle of the camera, a capturedimage may be suboptimal. Additional sensory feedback may be provided to facilitate guiding the user to better align the object angle with the camera angle before capturing an image. To identify a relevant object of interest for which to provide sensory feedback, a saliency model may be applied to remove background objects from consideration and to compare saliency of objects sharing a particular object label. Lighting guidance may also be provided to warn a user of a low light condition that may produce a suboptimal image.
[0011] Face angle guidance may be provided once a detected face is positioned within a viewfinder of a camera. In particular, an angle difference between the face and the camera may be determined. The angle difference may be set to zero degrees when a detected face is oriented directly towards the camera. The angle difference may be determined using a face angle which is based on detection of facial landmarks (e.g., the eyes, nose, and mouth). The angle difference may be compared to a threshold angle, and when the angle difference is greater than the threshold angle, sensory feedback may be provided to direct a user to reorient the face and / or reorient the camera to better align the face angle with the camera angle. When the angle difference is reduced below the threshold angle, autocapture of an image may be initiated (e.g., to occur within a fixed amount of time such as three seconds).
[0012] In some examples, separate face angles for pan and tilt may be considered. In such examples, a user may be directed to first tilt their head to reduce a tilt angle difference below a first threshold angle and subsequently directed to rotate their head to reduce a pan angle difference below a second threshold angle. The first and second threshold angles may be the same or different. In further examples, different sequences of guidance may be provided rather than directing the user to first center the face, then adjust the tilt angle, and then adjust the pan angle.
[0013] In further examples, a second threshold angle may be used to abort an image capture process. For example, the first threshold angle may be set to 25 degrees. When the face angle difference exceeds 25 degrees, the user may be directed to reduce the angle difference below the first threshold angle. Separately, a second threshold angle may be set to a greater amount than the first threshold angle (e.g., the second threshold angle may be set to 27.5 degrees). Subsequently, if the face angle exceeds the second threshold angle, an auto capture process may be aborted. By using separate thresholds to initiate and abort auto capture, state jumping back and forth between capture and abort states may be avoided when the face angle is near the first threshold angle.
[0014] In further examples, a saliency model may be used which determines the saliency of different portions of the environment shown in the viewfinder of a camera. The saliency model may facilitate identification of objects of interest before providing guidance to facilitate image capture. In some examples, the saliency model may be used to identify a background area and a foreground area. Objects in the background area may be disregarded when identifying relevant objects of interest. Considering saliency in this manner may prevent locking into a face in the background and attempting to provide guidance for that face.
[0015] In general, objects of different types may have different labels which may be used to determine a particular object or objects of interest to focus on for purposes of providing guidance. For instance, a face may be prioritized over a pet, which in turn may be prioritized over an inanimate object. In situations where multiple objects with the same label are present in the viewfinder, a saliency model may be used to identify a most relevant object (e.g., a particular face) to focus on for purposes of providing sensory feedback.
[0016] In further examples, lighting guidance may be provided to help direct a user to increase lighting when needed before capturing an image such as a selfie. In some examples, the user may be directed to move the camera until the lighting condition improves. In some examples, the camera may be configured with a night sight mode for capturing images in low light conditions. Night sight mode may require longer exposure times which may make it difficult to capture a good selfie image. Accordingly, in some examples, night sight mode may be disabled when providing sensory guidance to facilitate image capture. Night sight mode may be configured to trigger by comparing a lighting level of the environment to a threshold level. The same threshold level may be used to trigger sensory feedback to direct the user to improve the lighting condition before image capture.
[0017] Figure 1 illustrates an example computing device 100. In examples described herein, computing device 100 may be an image capturing device and / or a video capturing device. Computing device 100 is shown in the form factor of a mobile phone. However, computing device 100 may be alternatively implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. Computing device 100 may include various elements, such as body 102, display 106, and buttons 108 and 110. Computing device 100 may further include one or more cameras, such as front-facing camera 104 and at least one rear-facing camera 112. In examples with multiple rear-facing cameras such as illustrated in Figure 1, each of the rear-facing cameras may have a different field of view. For example, the rear facing cameras may include a wide angle camera, a main camera,and a telephoto camera. The wide angle camera may capture a larger portion of the environment compared to the main camera and the telephoto camera, and the telephoto camera may capture more detailed images of a smaller portion of the environment compared to the main camera and the wide angle camera.
[0018] Front-facing camera 104 may be positioned on a side of body 102 typically facing a user while in operation (e.g., on the same side as display 106). Rear-facing camera 112 may be positioned on a side of body 102 opposite front-facing camera 104. Referring to the cameras as front and rear facing is arbitrary, and computing device 100 may include multiple cameras positioned on various sides of body 102.
[0019] Display 106 could represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some examples, display 106 may display a digital representation of the current image being captured by front-facing camera 104 and / or rear-facing camera 112, an image that could be captured by one or more of these cameras, an image that was recently captured by one or more of these cameras, and / or a modified version of one or more of these images. Thus, display 106 may serve as a viewfinder for the cameras. Display 106 may also support touchscreen functions that may be able to adjust the settings and / or configuration of one or more aspects of computing device 100.
[0020] Front-facing camera 104 may include an image sensor and associated optical elements such as lenses. Front-facing camera 104 may offer zoom capabilities or could have a fixed focal length. In other examples, interchangeable lenses could be used with front-facing camera 104. Front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. Front-facing camera 104 also could be configured to capture still images, video images, or both. Further, front-facing camera 104 could represent, for example, a monoscopic, stereoscopic, or multiscopic camera. Rear-facing camera 112 may be similarly or differently arranged. Additionally, one or more of front-facing camera 104 and / or rear-facing camera 112 may be an array of one or more cameras.
[0021] One or more of front-facing camera 104 and / or rear-facing camera 112 may include or be associated with an illumination component that provides a light field to illuminate a target object. For instance, an illumination component could provide flash or constant illumination of the target object. An illumination component could also be configured to provide a light field that includes one or more of structured light, polarized light, and light withspecific spectral content. Other types of light fields known and used to recover three- dimensional (3D) models from an object are possible within the context of the examples herein.
[0022] Computing device 100 may also include an ambient light sensor that may continuously or from time to time determine the ambient brightness of a scene that cameras 104 and / or 112 can capture. In some implementations, the ambient light sensor can be used to adjust the display brightness of display 106. Additionally, the ambient light sensor may be used to determine an exposure length of one or more of cameras 104 or 112, or to help in this determination.
[0023] Computing device 100 could be configured to use display 106 and front-facing camera 104 and / or rear-facing camera 112 to capture images of a target object. The captured images could be a plurality of still images or a video stream. The image capture could be triggered by activating button 108, pressing a softkey on display 106, or by some other mechanism. Depending upon the implementation, the images could be captured automatically at a specific time interval, for example, upon pressing button 108, upon appropriate lighting conditions of the target object, upon moving computing device 100 a predetermined distance, or according to a predetermined capture schedule.
[0024] Figure 2 is a simplified block diagram showing some of the components of an example computing system 200, such as an image capturing device and / or a video capturing device. By way of example and without limitation, computing system 200 may be a cellular mobile telephone (e.g., a smartphone), a computer (such as a desktop, notebook, tablet, server, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a gaming console, a robotic device, a vehicle, or some other type of device. Computing system 200 may represent, for example, aspects of computing device 100.
[0025] As shown in Figure 2, computing system 200 may include communication interface 202, user interface 204, processor 206, data storage 208, and camera components 224, all of which may be communicatively linked together by a system bus, network, or other connection mechanism 210. Computing system 200 may be equipped with at least some image capture and / or image processing capabilities. It should be understood that computing system 200 may represent a physical image processing system, a particular physical hardware platform on which an image sensing and / or processing application operates in software, or other combinations of hardware and software that are configured to carry out image capture and / or processing functions.
[0026] Communication interface 202 may allow computing system 200 to communicate, using analog or digital modulation, with other devices, access networks, and / or transport networks. Thus, communication interface 202 may facilitate circuit-switched and / or packet-switched communication, such as plain old telephone service (POTS) communication and / or Internet protocol (IP) or other packetized communication. For instance, communication interface 202 may include a chipset and antenna arranged for wireless communication with a radio access network or an access point. Also, communication interface 202 may take the form of or include a wireline interface, such as an Ethernet, Universal Serial Bus (USB), or High- Definition Multimedia Interface (HDMI) port, among other possibilities. Communication interface 202 may also take the form of or include a wireless interface, such as a Wi-Fi, BLUETOOTH®, global positioning system (GPS), or wide-area wireless interface (e.g., WiMAX or 3 GPP Long-Term Evolution (LTE)), among other possibilities. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used over communication interface 202. Furthermore, communication interface 202 may comprise multiple physical communication interfaces (e.g., a Wi-Fi interface, a BLUETOOTH® interface, and a wide-area wireless interface).
[0027] User interface 204 may function to allow computing system 200 to interact with a human or non-human user, such as to receive input from a user and to provide output to the user. Thus, user interface 204 may include input components such as a keypad, keyboard, touch-sensitive panel, computer mouse, trackball, joystick, microphone, and so on. User interface 204 may also include one or more output components such as a display screen, which, for example, may be combined with a touch-sensitive panel. The display screen may be based on CRT, LCD, LED, and / or OLED technologies, or other technologies now known or later developed. User interface 204 may also be configured to generate audible output(s), via a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface 204 may also be configured to receive and / or capture audible utterance(s), noise(s), and / or signal(s) by way of a microphone and / or other similar devices.
[0028] In some examples, user interface 204 may include a display that serves as a viewfinder for still camera and / or video camera functions supported by computing system 200. Additionally, user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of a camera function and the capturing of images. It may be possible that some or all of these buttons, switches, knobs, and / or dials are implemented by way of a touch-sensitive panel.
[0029] Processor 206 may comprise one or more general purpose processors - e.g., microprocessors - and / or one or more special purpose processors - e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, or application-specific integrated circuits (ASICs). In some instances, special purpose processors may be capable of image processing, image alignment, and merging images, among other possibilities. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated in whole or in part with processor 206. Data storage 208 may include removable and / or non-removable components.
[0030] Processor 206 may be capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in data storage 208 to carry out the various functions described herein. Therefore, data storage 208 may include a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by computing system 200, cause computing system 200 to carry out any of the methods, processes, or operations disclosed in this specification and / or the accompanying drawings. The execution of program instructions 218 by processor 206 may result in processor 206 using data 212.
[0031] By way of example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device driver(s), and / or other modules) and one or more application programs 220 (e.g., camera functions, address book, email, web browsing, social networking, audio-to-text functions, text translation functions, and / or gaming applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be accessible primarily to operating system 222, and application data 214 may be accessible primarily to one or more of application programs 220. Application data 214 may be arranged in a file system that is visible to or hidden from a user of computing system 200.
[0032] Application programs 220 may communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs may facilitate, for instance, application programs 220 reading and / or writing application data 214, transmitting or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on.
[0033] In some cases, application programs 220 may be referred to as “apps” for short. Additionally, application programs 220 may be downloadable to computing system 200through one or more online application stores or application markets. However, application programs can also be installed on computing system 200 in other ways, such as via a web browser or through a physical interface (e.g., a USB port) on computing system 200.
[0034] Camera components 224 may include, but are not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or an image sensor), lens, shutter button, infrared projectors, and / or visible-light projectors. Camera components 224 may include components configured for capturing of images in the visible-light spectrum (e.g., electromagnetic radiation having a wavelength of 380 - 700 nanometers) and / or components configured for capturing of images in the infrared light spectrum (e.g., electromagnetic radiation having a wavelength of 701 nanometers - 1 millimeter), among other possibilities. Camera components 224 may be controlled at least in part by software executed by processor 206.
[0035] In further examples, one or more remote cameras 230 may be controlled by computing system 200. For instance, computing system 200 may transmit control signals to the one or more remote cameras 230 through a wireless or wired connection. Such signals may be transmitted as part of an ambient computing environment. In such examples, inputs received at the computing system 200 (for instance, physical movements of a wearable device) may be mapped to movements or other functions of the one or more remote cameras 230. Images captured by the one or more remote cameras 230 may be transmitted to the computing system 200 for further processing. Such images may be treated as images captured by cameras physically located on the computing system 200.
[0036] Figure 3 illustrates an example computing device 300 in which the described systems and techniques can be implemented. As illustrated in Figure 3, the computing device 300 is a smartphone. In some implementations, the computing device can be a variety of other electronic devices (e.g., digital cameras, tablet computers, computers, smartwatches, and so forth). The computing device 300 may include a processor, an input / output device (e.g., a display, a speaker, a haptic mechanism (e.g., DC motor that creates vibrations)), and a camera module. In aspects, the camera module is a device or component capable of capturing an image (e.g., a front-facing camera). The camera module may include a lens and an image sensor.
[0037] The computing device 300 may also include a computer-readable media (CRM) that stores device data (e.g., user data, multimedia data, applications, an operating system). The device data may include instructions of an image capture application that, responsive to execution by the processor, cause the processor to perform operations described in thisdocument to utilize the image sensor, image processing capabilities, and feedback capabilities of the computing device to provide real-time feedback that improves selfies for a camera user. The image capture application may include a detection controller, a haptics controller, a camera sound player, and a viewfinder overlay.
[0038] The entities of Figure 3 may be further divided, combined, used along with other sensors (e.g., accelerometers or gyros) or components, and so on. In this way, different implementations of computing devices, with different configurations, can be used to implement the providing of real-time feedback to improve selfies for a camera user. The example computing device 300 of Figure 3 illustrates but some of many possible devices capable of employing the described techniques. Although described with reference to a frontfacing camera, any of the described aspects may also be applied to assist users with capturing self-portraits or other types of photos using a back-facing camera, a remote camera (e.g., wirelessly linked), or the like.
[0039] The techniques include methods of providing real-time feedback to improve self-portrait photographs for camera users. The described methods may operate separately or together in whole or in part. The methods may be performed by a computing device (e.g., computing system 300) in accordance with one or more aspects of providing real-time feedback to improve selfies. The methods are described as operations performed but are not necessarily limited to the order or combinations shown for performing the operations. Further, one or more of the operations may be repeated, combined, reorganized, reordered, or linked to provide a wide array of additional and / or alternate methods. In portions of the following discussion, reference may be made to the example computing device 300 of Figure 3 or to entities or processes as detailed in other Figures, reference to which is made for example only. The techniques are not limited to performance by one entity or multiple entities operating on one device. A method may use elements of any of the Figures.
[0040] In example methods, assume that a user initiates a front-facing camera of a smartphone (computing device) to take a selfie. The image capture application is initiated, and the detection controller may detect a face or a portion of a face (collectively a “face portion”) in the viewing area of the front-facing camera (e.g., in a center portion of the camera frame). The detection controller may calculate a distance from the face to a center portion of the camera frame and provide it to the haptics controller. The haptics controller may segment the camera’s viewfinder into zones. As a user’s face transitions between zones, the haptic controller may generate (cause) haptic feedback (i.e., vibrations) and audio feedback to change states. Forexample, the vibration of the phone and / or an audio pattern produced can become more rapid, intense, and high-pitched as the face approaches a zone at the center of the camera frame that is desirable for selfies.
[0041] In some examples, the detection controller may calculate an angle of the face relative to an angle of the camera, and provide this information to the haptics controller, the camera sound player, and / or the viewfinder overlay. In additional examples, the detection controller may calculate lighting information, and provide this information to the haptics controller, the camera sound player, and / or the viewfinder overlay.
[0042] Haptics may indicate a direction (e.g., up, down) to move the smartphone and / or a distance (e.g., near, far). The haptics controller may compose haptic feedback for the haptic mechanism in combination with audio feedback. Audio feedback may be communicated to the user via the camera sound player and the speaker. Audio feedback may facilitate improvement of angle and / or lighting conditions for image capture. A viewfinder overlay initiates communication of visual information via a user interface (UI) to a user. The UI may overlay a display and include a visual indicator. For example, a visual indicator that flashes, a checkmark, and / or a high-contrast outline. The user may receive the haptic, audio, and / or visual feedback and continue to adjust the position or orientation of the smartphone. The image capture application and its modules may repeat the process (e.g., every 100-200 milliseconds) to track the face, calculate updated distance, angle, and / or lighting information, and communicate feedback.
[0043] Also disclosed are systems that may include an apparatus comprising a processor and a computer-readable storage media (CRM) having stored thereon instructions that, responsive to execution by the processor, cause the processor to perform a disclosed method.
[0044] Figure 4 illustrates different angle positions for selfies. As described herein, sensory feedback may guide a user to move the user’s face to the center of an image and to avoid cropping. But angle is another important factor to capture a good selfie. Most users do not like a double-chin photo. In examples described herein, the relevant angle information does not only come from the face, but also from the device. Figure 4 illustrates different face and camera angles in images 402, 404, 406, 408, and 410.
[0045] As illustrated in Figure 4, image 402 illustrates camera positioning for a selfie where the angle of the camera and the angle of the user’s face are aligned. In situations where the angle of the user’s face and the angle of the camera differ significantly, suboptimal selfiesmay result. Image 404 illustrates an example situation where the angles are not aligned. In some examples described herein, sensory feedback may be provided to cause the user to adjust their face angle to align with the camera angle, as illustrated in image 402. In alternative examples, sensory feedback may be provided to cause the user to adjust the camera angle to align with the face angle, resulting in positioning as illustrated in image 406. The positioning illustrated in image 402 and image 406 may both provide adequate alignment for good selfie capture. In further examples, sensory feedback may be provided both to cause the user to adjust the camera angle and the face angle.
[0046] Image 408 also illustrates camera positioning where the angle of the camera and the angle of the user’s face are not aligned. In some examples described herein, sensory feedback may be provided to cause the user to adjust their face angle to align with the camera, as illustrated in image 410. In alternative examples, sensory feedback may be provided to cause the user to adjust the camera angle to align with the face angle, resulting in positioning as illustrated in image 402. The positioning illustrated in image 402 and image 410 may both provide adequate alignment for good selfie capture.
[0047] Figure 5 illustrates a sequence of face positions and angles for a selfie. Image 502 illustrates an initial positioning and angle of a user’s face for a selfie. In this example, face positioning may first be adjusted to center the user’s face before addressing angle. Auditory feedback may be provided to indicate that the user’s face has been detected, but is cropped and not centered. The feedback may guide the user to adjust positioning of the camera or the user’s face to center the user’s face in the image. In additional examples, other types of sensory feedback such as haptic and / or visual feedback may be provided as well or instead.
[0048] Image 504 illustrates positioning and angle of the user’s face after centering. In this example, the angle of the user’s face may be determined to be misaligned with the camera angle in tilt angle (e.g., by more than a first threshold angle). Accordingly, the tilt angle may be addressed first by instructing the user to tilt their face to align with the camera (in this case, by tilting downward). Image 506 illustrates positioning and angle when the user’s face is determined to be aligned with the camera angle in the tilt direction (e.g., to within the first threshold angle).
[0049] It may then be determined that the pan angle of the user’s face is misaligned with the camera (e.g., by more than a second threshold angle). Accordingly, the pan angle may be addressed by instructing the user to rotate their face to align with the camera (in this case, by rotating to the right). Image 508 illustrates positioning and angle when the user’s face isdetermined to be aligned with the camera angle in the tilt and pan directions (e.g., to within the first threshold angle and the second threshold angle, respectively). In some examples, autocapture may then be triggered. In alternative examples, sensory feedback may be provided to the user to indicate that positioning and angle are both well aligned for good selfie image capture.
[0050] Figure 6 illustrates components configured to provide angle and lighting guidance. More specifically, an object detection controller 602 may be configured to detect and identify objects within a viewfinder. In some examples, a face may be identified for which positioning and angle guidance may be provided. In addition to identifying relevant objects, pan and tilt angles may be determined and used to determine proper alignment with the camera. In the case of detected faces, facial landmarks including the eyes, nose, mouths, and / or ears may be used to determine relevant pan and tilt angle information.
[0051] Device angle module 604 may provide pan and tilt angle information for the camera. As illustrated here, an existing class RotationVectorSensor may read data and handle register and unregister sensor behaviors. Accordingly, RotationVectorSensor may be used directly to provide relevant pan and tilt information for the camera.
[0052] Lighting module 606 may provide lighting information for the environment surrounding the camera. It may be desirable to prevent users from taking blurry photos due to low light conditions. Accordingly, a lighting level associated with the environment may be determined and compared with a threshold level to determine whether sensory feedback to adjust camera location is needed. In some examples, the threshold level may be aligned with a threshold level that is used to cause the camera to switch to a night sight mode (e.g., auto- NightSight).
[0053] Face angle module 608 may provide pan and tilt angle information for objects identified as faces for purposes of selfie image capture. The pan and tilt angles may separately be compared to camera pan and tilt angles to ensure both the pan angle and the tilt angle of the face are within respective thresholds of the camera pan and tilt angles.
[0054] Sensory feedback 612 may then be provided in the form of visual effects, sound effects, haptic effects, and / or voice effects in order to cause the user to adjust obj ect positioning, angle, and / or lighting as needed.
[0055] With respect to face angle guidance, an initial angle threshold may be set (e.g., to 25 degrees) for both tilt and pan angles. Furthermore, a tolerance angle may be set (e..g, to 2.5 degrees) to avoid the state jump back and forth. That means, for example, if the currentface angle is 30 degrees relative to the camera angle (in pan or tilt), the user may be asked to rotate their face to reduce the angle to less than the 25 degrees. Then this threshold angle may be extended (e.g., to 27.5 degrees) to prevent jumping at the boundary value.
[0056] Figure 7 illustrates components configured to handle images with multiple objects. More specifically, objection detection controller 702 may detect and identify relevant objects, including faces, in the viewfinder. Storage 704 may provide information to facilitate object identification, as well as determining saliency measures for individual objects within the viewfinder.
[0057] Objection detection interface 706 may track objects, and filter and prioritize objects for purposes of providing relevant sensory feedback for user guidance. More specifically, faces which appear in the background may be removed from consideration. In some complex environments, background faces may appear and disappear, and will impact auto-capture workflow if not removed from consideration. Regarding object prioritization, in order to have better experience to trigger auto-capture for single objects, for those non-face objects, the list of objects may be filtered to only leave the first one non-face object (e.g., with highest saliency). In some examples, objects may be sorted both by saliency and label (e.g., object type). In such examples, saliency may be used to select between objects with a given label.
[0058] Figure 8 is a block diagram of an example method 800. Method 800 may be executed by one or more computing systems (e.g., computing system 200 of Figure 2) and / or one or more processors (e.g., processor 206 of Figure 2). Method 800 may be carried out on a computing system, such as computing device 100 of Figure 1, or a computing device, such as computing device 300 of Figure 3.
[0059] At block 802, method 800 may involve receiving an input from a user to initiate a camera of the computing device to capture an image of an environment. Such input may involve activating a viewfinder to preview an image that may be captured by a front-facing or rear-facing camera associated with a computing device.
[0060] At block 804, method 800 may involve detecting one or more objects of interest in the environment. In some examples, the one or more objects of interest may be faces. In further examples, different types of objects of interest may be detected (e.g., pets, documents). In some examples, objects of interest may be identified from a predefined list. One or more object recognition algorithms (e.g., machine learning algorithms) may be used to identify types of objects for purposes of detecting objects of interest in the environment.
[0061] At block 806, method 800 may involve providing sensory feedback to guide the user to position the computing device so that the one or more objects of interest become positioned within a viewfinder of the camera. The feedback may be haptic, auditory, visual, and / or other types of sensory feedback.
[0062] At block 808, method 800 may involve, when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera. In some examples, both a pan angle difference and a tilt angle difference may separately be determined.
[0063] At block 810, method 800 may involve, based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference. In some examples, the sensory feedback may be provided to cause the user to adjust the object (e.g., face) angle only. In some examples, the sensory feedback may be provided to cause the user to adjust the camera angle only. In some examples, the sensory feedback may be provided to cause the user to adjust both the object angle and the camera angle. In some examples, separate feedback may be provided to guide the user to adjust pan angle difference and / or tilt angle difference.
[0064] At block 812, method 800 may involve when the angle difference is less than the threshold angle, initiating capture of an image of the object. In some examples, capture may be initiated automatically after a set period of time (e.g., 3 seconds). In other examples, a user may be notified through sensory feedback that image capture is now available after angle alignment has been achieved.
[0065] Some examples of method 800 may involve, when the angle difference becomes greater than a second threshold angle, aborting capture of the image, where the second threshold angle is greater than the threshold angle. For instance, the first threshold angle may be 25 degrees and the second threshold angle may be 27.5 degrees.
[0066] In some examples of method 800, additional sensory feedback is provided to cause the user to sequentially adjust a tilt and then a pan of the object. In further examples, because a face’s rotation will affect the face rectangle position which may cause the face to become cropped, a second round of moving phone guidance may be provided.
[0067] Some examples of method 800 may further involve based on a lighting level associated with the environment being less than a threshold level, providing sensory feedback indicative of a low light condition. In such examples, the threshold level may be the same threshold level used to cause the camera to switch to a night sight mode.
[0068] Some examples of method 800 may further involve detecting the object of interest in the environment based on application of a saliency model (e.g., a machine learning model). In such examples, detecting the object of interest in the environment may comprise determining an object label associated with the object and comparing a saliency of the object to one or more other objects sharing the object label. Such examples may further involve determining a foreground and a background based on the saliency model, wherein the object of interest is conditioned on the object being in the foreground.
[0069] As used herein, a phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, c-c-c, or any other ordering of a, b, and c).
[0070] Although implementations of systems and techniques related to providing realtime feedback to improve self-portrait photographs for camera users have been described in language specific to features and / or methods, it is to be understood that the subject of the appended claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of devices and systems related to providing real-time feedback to improve image capture for camera users.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: receiving an input from a user to initiate a camera of a computing device to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
2. The method of claim 1, wherein the additional sensory feedback guides the user to adjust an angle of the object.
3. The method of claim 1, wherein the additional sensory feedback guides the user to adjust an angle of the camera.
4. The method of claim 1, wherein the additional sensory feedback guides the user to adjust a first angle of the object and a second angle of the camera.
5. The method of claim 1, wherein the sensory feedback comprises at least one of haptic, audio, or visual feedback.
6. The method of claim 1, wherein the capture of the image of the object is performed a fixed amount of time after determining that the angle difference is less than the threshold angle.
7. The method of claim 1, wherein the angle difference comprises at least one of a pan angle difference or a tilt angle difference.
8. The method of claim 1, further comprising: when the angle difference becomes greater than a second threshold angle, aborting capture of the image, wherein the second threshold angle is greater than the threshold angle.
9. The method of claim 1, wherein the object of interest is a face and the image is a selfie.
10. The method of claim 9, wherein determining the angle difference is based on detecting a plurality of facial landmarks on the face.
11. The method of claim 1, wherein the additional sensory feedback is provided to cause the user to sequentially adjust a tilt and then a pan of the object.
12. The method of claim 1, further comprising: based on a lighting level associated with the environment being less than a threshold level, providing sensory feedback indicative of a low light condition.
13. The method of claim 12, wherein the threshold level is used to cause the camera to switch to a night sight mode.
14. The method of claim 1, wherein detecting the object of interest in the environment is based on application of a saliency model.
15. The method of claim 14, wherein detecting the object of interest in the environment comprises: determining an object label associated with the object; and comparing a saliency of the object to one or more other objects sharing the object label.
16. The method of claim 14, further comprising: determining a foreground and a background based on the saliency model, wherein the object of interest is conditioned on the object being in the foreground.
17. A computing device configured to: receive an input from a user to initiate a camera to capture an image of an environment; detect an object of interest in the environment; provide sensory feedback to guide the user to position the computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determine an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, provide additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiate capture of an image of the object.
18. The computing device of claim 17, wherein the camera is a front-facing camera on the computing device, and wherein the image is a selfie.
19. A non-transitory computer readable medium comprising program instructions executable by one or more processors to perform functions comprising: receiving an input from a user to initiate a camera to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position a computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
20. A computer program comprising program instructions executable to perform functions comprising:receiving an input from a user to initiate a camera to capture an image of an environment; detecting an object of interest in the environment; providing sensory feedback to guide the user to position a computing device so that the object becomes positioned within a viewfinder of the camera; when the object is positioned within the viewfinder of the camera, determining an angle difference between the object and the camera; based on the angle difference being greater than a threshold angle, providing additional sensory feedback to guide the user to reduce the angle difference; and when the angle difference is less than the threshold angle, initiating capture of an image of the object.
Citation Information
Patent Citations
Identifying User Intent for Auto Selfies
US20190208115A1
Audio assisted enrollment
US20220180667A1
Systems and methods for providing displayed feedback when using a rear-facing camera
US20230053026A1
Real-time feedback to improve image capture
WO2024076631A1