Automatic Camera Guidance and Setting Adjustment

The image capture device recognizes objects and uses machine learning models to generate camera settings adjustment guidance, which solves the problem that photographers do not understand camera settings in different scenarios, and realizes visual optimization of images.

CN115668967BActive Publication Date: 2025-07-25QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180035537.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-15
Filing Date
2021-05-19
Publication Date
2025-07-25
Estimated Expiration
2041-05-19

AI Technical Summary

Technical Problem

Photographers do not understand how to apply camera settings to optimize image composition in different scenes, resulting in poor image effects.

Method used

The object is identified by the image capture device and the visual differences are identified using machine learning models, and attribute changes guiding the image capture device are generated to optimize the image composition.

Benefits of technology

Automatically adjust camera settings to improve the visual appeal and quality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115668967B_ABST
    Figure CN115668967B_ABST
Patent Text Reader

Abstract

An image capture and processing device captures an image. Based on the image and / or one or more additional images, the image capture and processing device generates and outputs guidance for optimizing image composition, image capture settings, and / or image processing settings. The guidance can be generated based on determination of the direction the object in the image faces, based on sensor measurements indicating that a horizontal line may be tilted, another image of the same scene captured using a wide-angle lens, another image of the same object, another image of a different object, and / or the output of a machine learning model trained using an image set. The image capture and processing device can automatically apply certain aspects of the generated guidance, such as image capture settings and / or image processing settings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to image capture and image processing. More specifically, the present application relates to systems and methods for automatically guiding image capture and automatically adjusting settings to visually optimize image composition and / or apply a specific style. Background Art

[0002] In photography, certain rules or guidelines of image composition can help a photographer frame an object in an image to make the image more visually appealing. However, many photographers are not familiar with many different rules and guidelines of image composition, do not know how to best apply these rules and guidelines to different types of photos, or when to ignore certain rules and guidelines.

[0003] A camera can apply various image capture and image processing settings to change the appearance of an image. Some camera settings are determined and applied before or during the capture of a photo, such as ISO, exposure time, aperture size, f / stop, shutter speed, focal length, and gain. Other camera settings can configure the post-processing of a photo, such as changes in contrast, brightness, saturation, sharpness, levels, curves, or color. Different camera settings can emphasize different aspects of an image. However, a large number of different camera settings can confuse a user. The user may not know which settings are helpful in which scenarios, or may not understand how to adjust certain camera settings to be helpful in those scenarios. Summary of the Invention

[0004] Systems and techniques for generating and outputting guidance for image capture are described herein. An image capture device captures a first image. Based on the first image, the image capture device identifies changes to the attributes of the image capture device. These changes cause a visual difference between the first image and a second image that will be captured by the image capture device after the capture of the first image. The image capture device can identify the changes based on the settings of the attributes in cases where other images besides the first image are captured. For example, the other images can be images depicting the same object as the object depicted in the first image, or an object similar to the object depicted in the first image. In some examples, these changes can be based on a machine learning model trained on these other images. The image capture device generates and outputs guidance to indicate the changes that result in the visual difference when the image capture device captures the second image. These attributes can include the positioning of the image capture device (to affect image composition), image capture settings, and / or image processing settings.

[0005] In one example, a device for guiding image capture is provided. The device includes one or more connectors coupled to one or more image sensors, where the one or more connectors receive image data from the one or more image sensors. The device includes one or more storage units that store instructions and one or more processors that execute the instructions. Execution of the instructions by the one or more processors causes the one or more processors to perform a method for guiding image capture. The method includes receiving a first image of a scene captured by an image sensor. The method includes identifying an object depicted in the first image. The method includes inputting the first image into a machine learning model that is trained using a plurality of training images having the identified object. The method includes using the machine learning model to identify one or more changes to one or more attributes associated with the image capture that cause a visual difference between the first image and a second image that will be captured by the image sensor after the first image is captured. The method includes outputting, before the image sensor captures the second image, guidance indicating the one or more changes that produce the visual difference.

[0006] In another example, a method for guiding image capture is provided. The method includes receiving a first image of a scene captured by an image sensor of an image capture device. The method includes identifying an object depicted in the first image. The method includes inputting the first image into a machine learning model that is trained using a plurality of training images having the identified object. The method includes using the machine learning model to identify one or more changes to one or more attributes of the image capture device that cause a visual difference between the first image and a second image that will be captured by the image sensor after the first image is captured. The method includes outputting, before the image sensor captures the second image, guidance indicating the one or more changes that produce the visual difference.

[0007] In another example, a non-transitory computer-readable storage medium having a program thereon is provided. The program is executable by a processor to perform a method for guiding image capture. The method includes receiving a first image of a scene captured by an image sensor of an image capture device. The method includes identifying an object depicted in the first image. The method includes inputting the first image into a machine learning model that is trained using a plurality of training images having the identified object. The method includes using the machine learning model to identify one or more changes to one or more attributes of the image capture device that cause a visual difference between the first image and a second image that will be captured by the image sensor after the first image is captured. The method includes outputting, before the image sensor captures the second image, guidance indicating the one or more changes that produce the visual difference.

[0008] In another example, an apparatus for guiding image capture is provided. The apparatus includes means for receiving a first image of a scene captured by an image sensor of an image capture device. The apparatus includes means for identifying an object depicted in the first image. The apparatus includes means for inputting the first image into a machine learning model that is trained using a plurality of training images with the identified object. The apparatus includes means for using the machine learning model to identify one or more changes to one or more attributes of the image capture device, the one or more changes causing a visual difference between the first image and a second image that will be captured by the image sensor after the first image. The apparatus includes means for outputting, before the image sensor captures the second image, a guidance indicating the one or more changes that result in the visual difference.

[0009] In some aspects, identifying the object depicted in the first image includes performing at least one of feature detection, object detection, face detection, feature recognition, object recognition, face recognition, and generation of a saliency map. In some aspects, the above method, apparatus, and computer-readable medium further include: after outputting the guidance, receiving a second image from the image sensor; and outputting the second image, where outputting the second image includes at least one of displaying the second image using a display and transmitting the second image using a transmitter.

[0010] In some aspects, identifying one or more changes to one or more attributes of the image capture device includes identifying a movement of the image capture device from a first position to a second position, where outputting the guidance includes outputting an indication for moving the image capture device from the first position to the second position. In some aspects, the machine learning model is used to identify the second position. In some aspects, the indication includes at least one of a visual indicator, an auditory indicator, and a vibration indicator. In some aspects, the indication identifies at least one of a translation direction of the movement, a translation distance of the movement, a rotation direction of the movement, and a rotation angle of the movement.

[0011] In some aspects, the indicator identifies a translation direction from the first position to the second position. In some aspects, the indicator identifies a translation distance from the first position to the second position. In some aspects, the indicator identifies a rotation direction from the first position to the second position. In some aspects, the indicator identifies a rotation angle from the first position to the second position. In some aspects, the indicator includes one or more position coordinates of the second position. In some aspects, the visual difference between the first image and the second image levels a horizontal line in the second image, where the horizontal line is not horizontal as depicted in the first image.

[0012] In some aspects, the above methods, apparatuses, and computer-readable media further include: receiving attitude sensor measurement data from one or more attitude sensors; and determining the attitude of the apparatus based on the attitude sensor measurement data, wherein identifying the movement of the apparatus from a first position to a second position is based on the attitude of the apparatus, and wherein the attitude of the apparatus includes at least one of the position of the apparatus and the orientation of the apparatus. In some aspects, the one or more attitude sensors include at least one of an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit, a global navigation satellite system (GNSS) receiver, and an altimeter.

[0013] In some aspects, the above methods, apparatuses, and computer-readable media further include: determining the position of an object in a first image; and determining the direction the object is facing in the first image, wherein identifying the movement of the image capture device from a first position to a second position is based on the position of the object in the first image and the direction the object is facing in the first image. In some aspects, determining the direction the object is facing in the first image is based on the relative positioning of two features of the object. In some aspects, determining the direction the object is facing in the first image is based on the positioning of multiple features of the object relative to each other within the first image. In some aspects, the object is a person, and the multiple features of the object include at least one of a person's ears, a person's cheeks, a person's eyes, a person's eyebrows, a person's nose, a person's mouth, a person's chin, and a person's appendages.

[0014] In some aspects, determining the direction the object is facing in the first image is based on the direction of movement of the object as it moves between the first image and a third image captured by an image sensor. In some aspects, the above methods, apparatuses, and computer-readable media further include: receiving a third image captured by the image sensor, the third image depicting the object; and determining the direction of movement of the object based on the position of the object in the first image and the position of the object in the third image, wherein determining the direction the object is facing in the first image is based on the direction of movement of the object. In some aspects, the visual difference between the first image and the second image includes an increase in the negative space adjacent to the object in the direction the object is facing.

[0015] In some aspects, the above methods, apparatuses, and computer-readable media further include: receiving a third image of a scene captured by a second image sensor, wherein the first image of the scene and the third image of the scene are captured within a time window, wherein the second image sensor has a wider field of view than the image sensor, and wherein the guidance is based on the depiction in the third image of a portion of the scene that was not depicted in the first image.

[0016] In some aspects, the guidance instructs the image capture device to remain stationary between the capture of the first image and the capture of the second image.

[0017] In some aspects, the plurality of training images include training images depicting at least one of an object and a second object sharing one or more similarities with the object, wherein the one or more changes to the one or more attributes indicated by the guidance are based on one or more settings for the one or more attributes used to capture the training images. In some aspects, the one or more similarities shared between the second object and the object include: one or more saliency values associated with the second object being within a predetermined range of one or more saliency values associated with the object. In some aspects, the visual difference between the first image and the second image includes: the second image being more similar to the training image than the first image is to the training image. In some aspects,

[0018] In some aspects, the one or more changes to the one or more attributes associated with image capture include: applying image capture settings before the image sensor captures the second image, where the image capture settings correspond to at least one of zoom, focus, exposure time, aperture size, ISO, depth of field, analog gain, and aperture scale. In some aspects, the output guidance includes an output indicator that identifies the one or more changes to the one or more attributes associated with image capture corresponding to the application of the image capture settings. In some aspects, the output guidance includes automatically applying the one or more changes to the one or more attributes associated with image capture corresponding to the application of the image capture settings.

[0019] In some aspects, the above methods, apparatuses, and computer-readable media further include: receiving a second image captured by the image sensor, wherein the one or more changes to the one or more attributes associated with image capture include applying image processing settings to the second image, where the image processing settings correspond to at least one of brightness, contrast, saturation, gamma, level, histogram, color adjustment, blur, sharpness, level, curve, filtering, and cropping. In some aspects, the output guidance includes an output indicator that identifies the one or more changes to the one or more attributes associated with image capture corresponding to the application of the image processing settings. In some aspects, the output guidance includes automatically applying the one or more changes to the one or more attributes associated with image capture corresponding to the application of the image processing settings.

[0020] In some aspects, the device includes a camera, a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wireless communication device, a wearable device, a head-mounted display (HMD), an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other devices. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device further includes an image sensor. In some aspects, the device further includes one or more connectors coupled to the image sensor, wherein the one or more processors receive a first image from the image sensor through the one or more connectors. In some aspects, the device further includes a display for at least displaying a second image. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data.

[0021] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0022] The foregoing and other features and embodiments will become more apparent with reference to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The illustrative embodiments of the present application are described in detail below with reference to the following drawings;

[0024] Figure 1 is a block diagram showing the architecture of an image capture and processing device;

[0025] Figure 2A is a conceptual diagram showing an object at the center of an image;

[0026] Figure 2B is a conceptual diagram showing an object aligned with two lines representing one-third of an image Figure 2A ;

[0027] Figure 3A is a conceptual diagram showing a left-moving object depicted on the left hand side of an image;

[0028] Figure 3B is a conceptual diagram showing a Figure 2A left-moving object depicted on the right hand side of an image;

[0029] Figure 4 is a conceptual diagram showing three images of a face with markings on certain features that can be used to determine the direction the face is facing;

[0030] Figure 5 is a conceptual diagram showing a user interface of an image capture device, the user interface having a positioning guidance indication for guiding a user to move the image capture device a specific distance in a specific direction;

[0031] Figure 6 is a flowchart showing an operation for guiding image capture based on the direction in which an object in an image faces;

[0032] Figure 7A is a conceptual diagram showing a user interface of an image capture device, the user interface having a positioning guidance indication for guiding a user to tilt the image capture device counterclockwise to make the horizontal line in the image horizontal;

[0033] Figure 7B is a conceptual diagram showing a user interface of an image capture device, the user interface having a positioning guidance indication for guiding a user to tilt the image capture device clockwise to make the horizontal line in the image horizontal;

[0034] Figure 8 is a flowchart showing an operation for guiding image capture based on sensor measurement data from one or more positioning sensors of an image capture device;

[0035] Figure 9 is a conceptual diagram showing a view visible to a first image sensor with a normal lens superimposed on a view visible to a second image sensor with a wide - angle lens;

[0036] Figure 10 is a flowchart showing an operation for guiding image capture using a first image sensor with a first lens based on image data from a second image sensor with a second lens, where the second lens has a wider angle than the first lens;

[0037] Figure 11 is a conceptual diagram showing a user interface of an image capture device, where an image of a previously captured object is overlaid on an image of an object captured by the image sensor of the image capture device;

[0038] Figure 12 is a flowchart showing an operation for guiding the capture and / or processing of an image of an object based on another image of the same object;

[0039] Figure 13 is a conceptual diagram showing a user interface of an image capture device, where an image of a previously captured object is used to generate a guidance overlaid on an image of a different object captured by the image sensor of the image capture device;

[0040] Figure 14is a flowchart showing operations for guiding the capture and / or processing of an image of an object based on another image of a different object;

[0041] Figure 15 is a conceptual diagram showing a user interface of an image capture device, where a machine learning model trained using an image set is used to generate guidance overlaid on an image of an object captured by an image sensor of the image capture device;

[0042] Figure 16 is a flowchart showing operations for guiding the capture and / or processing of an image of an object based on a machine learning model trained using an image set;

[0043] Figure 17 is a flowchart showing a method for guiding image capture; and

[0044] Figure 18 is a diagram showing an example of a system for implementing certain aspects of the present technology. Detailed Description

[0045] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments may be applied independently, and some of them may be applied in combination, which will be obvious to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, it is obvious that the various embodiments can be practiced without these specific details. The drawings and the description are not intended to be restrictive.

[0046] The following description only provides exemplary embodiments and is not intended to limit the scope, application, or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application set forth in the appended claims.

[0047] An image capture and processing device captures an image. Based on the image and / or one or more additional images, guidance for optimizing the image composition, image capture settings, and / or image processing settings is generated and output. For example, the guidance can be generated based on the determination of the direction the object in the image is facing, sensor measurements indicating that the horizontal line may be tilted, another image of the same scene captured using a wide-angle lens, another image of the same object, another image of a different object, and / or the output of a machine learning model trained using an image set. The image capture and processing device can automatically apply certain aspects of the generated guidance, such as image capture settings and / or image processing settings.

[0048] Figure 1FIG. 0 is a block diagram showing the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing an image of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture still images (or photos) and / or can capture video including a plurality of images (or video frames) in a particular sequence. The lens 115 of the system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the light towards the image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0049] One or more control mechanisms 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. One or more control mechanisms 120 can include multiple mechanisms and components; for example, the control mechanism 120 can include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. One or more control mechanisms 120 can include additional control mechanisms in addition to the illustrated control mechanisms, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0050] The focus control mechanism 125B of the control mechanism 120 can obtain a focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or a servo mechanism to adjust the focus. In some cases, additional lenses can be included in the system 100, such as one or more microlenses on each photodiode of the image sensor 130, where each microlens bends the light received from the lens 115 towards the corresponding photodiode before the light reaches the photodiode. The focus setting can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The control mechanism 120, the image sensor 130, and / or the image processor 150 can be used to determine the focus setting. The focus setting can be referred to as an image capture setting and / or an image processing setting.

[0051] The exposure control mechanism 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A can control the size of the aperture (e.g., aperture size or aperture f-number), the duration for which the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting can be referred to as an image capture setting and / or an image processing setting.

[0052] The zoom control mechanism 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of a lens assembly (lens mounting) including the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo mechanisms to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly can include a focusing lens (which can be the lens 115 in some cases) that first receives light from the scene 110, and then the light passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system can include two positive lenses (e.g., converging lenses, convex lenses) with equal or similar focal lengths (e.g., within a threshold difference), with a negative lens (e.g., diverging lens, concave lens) between them. In some cases, the zoom control mechanism 125C moves one or more lenses in the afocal zoom system, such as the negative lens and one or two positive lenses.

[0053] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a particular pixel in the image produced by the image sensor 130. In some cases, different photodiodes may be covered by different color filters and can thus measure light that matches the color of the color filter covering the photodiode. For example, a Bayer filter includes a red filter, a blue filter, and a green filter, where each pixel of the image is generated based on red light data from at least one photodiode covered by the red filter, blue light data from at least one photodiode covered by the blue filter, and green light data from at least one photodiode covered by the green filter. Other types of color filters may use yellow, magenta, and / or cyan (also known as "emerald") filters in place of the red, blue, and / or green filters, or may use yellow, magenta, and / or cyan (also known as "emerald") filters in addition to the red, blue, and / or green filters. Some image sensors may have no color filter at all and instead use different photodiodes throughout the pixel array (in some cases, vertically stacked). Different photodiodes throughout the pixel array can have different spectral sensitivity curves and thus respond to different wavelengths of light. Monochrome image sensors may also lack a color filter and thus lack color depth.

[0054] In some cases, the image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). The image sensor 130 may also include an analog gain amplifier that amplifies the analog signal output by the photodiode and / or an analog-to-digital converter (ADC) that converts the analog signal output by the photodiode (and / or amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 can be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide semiconductor (CMOS), N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0055] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more main processors (including main processor 152), and / or one or more any other type of processors 1810 discussed with respect to the computing device 1800. The main processor 152 may be a digital signal processor (DSP) and / or other type of processor. In some embodiments, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system on a chip or SoC) that includes the main processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connection components (e.g., Bluetooth TM , Global Positioning System (GPS), etc.), any combination thereof, and / or other components. The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as an inter-integrated circuit 2 (I2C) interface, an inter-integrated circuit 3 (I3C) interface, a serial peripheral interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port, and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0056] The image processor 150 may perform many tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 may store the image frames and / or the processed images in a random access memory (RAM) 140 / 1820, a read-only memory (ROM) 145 / 1825, a cache, a memory cell, another storage device, or some combination thereof.

[0057] Various input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device 1835, any other input device 1845, or some combination thereof. In some cases, captions may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O device 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the system 100 and one or more peripheral devices, through which the system 100 may receive data from and / or send data to one or more peripheral devices. The I / O device 160 may include one or more wireless transceivers that implement a wireless connection between the system 100 and one or more peripheral devices, through which the system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed above, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they may themselves be considered I / O devices 160.

[0058] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled via one or more wireless transceivers. In some embodiments, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0059] As Figure 1 shown, the vertical dashed line divides Figure 1 the image capture and processing system 100 into two parts representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a main processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components shown in the image capture device 105A, such as the ISP 154 and / or the main processor 152, may be included in the image capture device 105A.

[0060] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some embodiments, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.

[0061] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art will understand that the image capture and processing system 100 may include more components than Figure 1 those shown. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.

[0062] Traditional camera systems (e.g., image sensors and ISPs) are tuned with parameters and process images according to the tuned parameters. The ISP is typically tuned using a fixed tuning method during production. Camera systems (e.g., image sensors and ISPs) typically also perform global image adjustment based on predefined conditions (such as light level, color temperature, exposure time, etc.). Typical camera systems are also tuned using coarse-precision heuristic-based tuning (e.g., window-based local tone mapping). As a result, traditional camera systems cannot enhance images based on the content contained in the images.

[0063] Figure 2Ais a conceptual diagram showing object 205 at the center of image 210. In Figure 2A the example image 210, object 205 is a person whose face is horizontally centered in image 210 and vertically oriented at the top third of image 210.

[0064] Four dashed lines pass through image 210, dividing image 210 into nine regions of equal size. The four dashed lines include two vertical lines and two horizontal lines. The two vertical lines are parallel to each other and parallel to the left and right sides of image 210, and perpendicular to the horizontal lines and the top and bottom of image 210. The first vertical dashed line includes one-third of image 210 on its left side and two-thirds of image 210 on its right side. The second vertical dashed line includes one-third of image 210 on its right side and two-thirds of image 210 on its left side. The two horizontal lines are parallel to each other and parallel to the top and bottom of image 210, and perpendicular to the vertical lines and the left and right sides of image 210. The first horizontal dashed line includes one-third of image 210 above it and two-thirds of image 210 below it. The second horizontal dashed line includes one-third of image 210 below it and two-thirds of image 210 above it. These dashed lines and the portions of the image they represent may be referred to as guide lines, rule of thirds lines, grid lines, or some combination thereof.

[0065] The rule of thirds is a rule or guideline for image composition that indicates that aligning an object along one of these lines of thirds, or to the intersection of two of these lines of thirds, is visually more interesting and creates more tension, energy, and interest than simply aligning the object at the center of the image or other locations of the object within the image. In Figure 2A it, object 205—the human face—is between the two vertical lines of thirds and above the top horizontal line of thirds, which is sub-optimal according to the rule of thirds.

[0066] Figure 2B is a conceptual diagram showing Figure 2A object 205 aligned with two lines representing one-third of image 220. Specifically, Figure 2A the dashed lines of thirds in Figure 2B are also shown in Figure 2B and object 205—the human face—is centered at the intersection of the right vertical line of thirds and the top horizontal line of thirds. Based on the rule of thirds, then, Figure 2B image 220 in

[0067] Some images have multiple objects. Since there are four three-point lines and four three-point line intersection points, the image composition in an image with multiple objects can be improved by aligning each object with at least one of these four three-point lines and / or at least one of these four three-point line intersection points. In some cases where an image depicts multiple objects, the image composition can be improved by aligning at least one subset of the objects with at least one of these four three-point lines and / or aligning with at least one of these four three-point line intersection points. For example, the most prominent object can be selected to be aligned with at least one of these four three-point lines and / or aligned with at least one of these four three-point line intersection points.

[0068] Figure 3A is a conceptual diagram showing the leftward-moving object 305 depicted on the left-hand side of the image 310. In Figure 3A the example image 310, the object 305 is a person walking a dog. The object 305 faces left and walks leftward.

[0069] Another rule or guideline regarding image composition indicates that negative space should be left in front of an object facing a specific direction, especially if the object is moving in that direction. The viewer's eye is drawn to where the object is looking and / or where it is moving. Including negative space in front of the object in the image allows the viewer to see more space where the object is looking and / or moving towards, making the viewer more interested in the image. On the other hand, failing to include too much negative space in front of the object in the image causes the viewer's gaze to suddenly terminate when looking in front of the object, but includes more area behind the object, which is not as visually interesting as the area in front of the object.

[0070] In Figure 3A the object 305 faces left and walks leftward, and is located on the left-hand side of the image 310, seemingly aligned with the left vertical three-point line in the image 310. However, despite following the rule of thirds, the image composition of the image 310 is not very good because very little negative space is included in front of the object 305 in the image 310, and a large amount of space is included behind the object 305 in the image 310.

[0071] Figure 3B is a conceptual diagram showing the Figure 2A leftward-moving object 305 depicted on the right-hand side of the image 320. The object 305 still faces left and walks leftward in the image 320, but is now aligned with the right vertical three-point line in the image 320. Therefore, compared with the image 310, the image 320 includes more negative space in front of the object 305, which means that the image 320 has a better image composition than the image 310 in terms of negative space.

[0072] Figure 2A - 3B The rule of thirds demonstrated andFigure 3A - 3B The negative space rules shown are just two of many image composition rules or guidelines that can help a photographer frame the objects in an image in a way that makes the image more visually appealing. Another image composition rule or guideline indicates that image composition is improved when lines (whether straight or curved) in the image are "guiding" lines that lead the viewer's eye to the object of the image. These lines can be roads, railroads, coastlines, rivers, converging buildings, a row of trees, a row of clouds, a row of birds or people or other creatures, a row of cars, a person's limbs, other types of lines, or combinations thereof. Another image composition rule or guideline indicates that image composition is improved when the image shows symmetry with matching or similar elements, whether the symmetry is vertical, horizontal, radial, or otherwise. Symmetry provides balance to the scene and can be achieved, for example, by providing two objects at opposite ends of the image. On the other hand, intentional asymmetry can also improve image composition if it also provides balance to the image. Asymmetrical balance can be achieved via tonal balance (dark versus light), color balance (coarse / bright versus subtle / neutral), size balance (large versus small), texture balance (high texture versus smooth), spatial balance (the direction of the viewer's eye or the movement of the object into space versus movement to the edge of the frame), abstract balance (contrasting two viewpoints such as nature versus industry, old versus new, happy versus sad, etc.), or some combination thereof.

[0073] Another image composition rule or guideline indicates that image composition is improved by including patterns that imply harmony. Patterns may include a row of columns, books on a bookshelf, a row of people, bricks on a brick wall, petals on a flower, ocean waves rolling onto a beach, and other patterns. Another image composition rule or guideline indicates that image composition is improved by filling the image frame (boundary) as much as possible with one or more objects of the image such that the objects of the image are clear. Filling the frame can be achieved when the image capture device 105A is close to the object, when the image capture device 105A uses a zoom lens to magnify the object, or when the image processing device 105B crops the image after capture to remove the space in the image that is not occupied by one or more objects. Filling the frame can improve image composition, especially when the area around the object is cluttered or otherwise distracting. Conversely, when the area around the object is simple (e.g., a blue sky) and not cluttered or distracting, providing negative space around the object can also help draw the viewer's eye to the object and thus improve image composition. As discussed above regarding Figure 3A - 3B when the object faces or moves within the image, including negative space is useful.

[0074] Another image composition rule or guideline indicates that the image composition can be improved by including multiple layers of depth in the scene, where objects (e.g., living beings, targets, or other visual elements of interest that are the focus in the image) are in the foreground of the image, in the background of the image, and in one or more intermediate layers. The related image composition rule or guideline indicates that the image composition can be improved by using depth of field to ensure that the objects in the image are sharp due to the depth of field, while the less important areas are blurred. For example, a shallow (narrow) depth of field can provide an improved image composition for portrait images and make the object clear and sharp in the image, while anything in the background of the object (behind the object and farther from the image capture device 105A) and / or anything in the foreground of the object (in front of the object and closer to the image capture device 105A) appears more blurred than the object. On the other hand, a deep (wide) depth of field can provide an improved image composition for landscape images and generally allows most of the image to be sharp. The image composition can also be improved by reducing distractions in other ways, even if not through the depth of field, such as by blurring the background relative to the object, darkening the background relative to the object, or lightening the background relative to the object.

[0075] Another image composition rule or guideline indicates that the image composition can be improved by framing the object with visual elements that are also included in the image. For example, if the object in the image is visually framed by one or more arches, doorways, openings, bridges, trees, branches, caves, mountains, walls, arms, limbs, or some combination thereof, the image composition is improved. Another image composition rule or guideline indicates that the image composition can be improved by including diagonals and / or triangles in the image, which can provide tension and / or a more natural feel to the image, where the image is typically captured and / or stored as a square or rectangle. Another image composition rule or guideline indicates that an unusual perspective of a familiar object can improve the image composition by making the resulting image more interesting. For example, a portrait of a person or a group of people may be more visually interesting if captured from an aerial view above the person or group of people or from a worm's-eye view below the person or group of people, rather than simply captured straight ahead at the eye level of the person or group of people. Another image composition rule or guideline indicates that the image composition is improved if a moving object moves from the left side of the image to the right side of the image because most viewers read from left to right. The image composition can also be improved if the image includes an odd number of objects or visual elements.

[0076] A large number of different image composition rules and guidelines can make them confusing and difficult to learn for new photographers, and even difficult to master for professional photographers. Therefore, an image capture device 105A and / or an image processing device 105B that provides guidance to users to improve image composition and / or can automatically adjust settings to improve image composition will produce excellent images with better image composition than those without such guidance for the image capture device 105A and / or the image processing device 105B.

[0077] Figure 4 is a conceptual diagram showing three images of a human face with markings on certain features, which can be used to determine the direction the face is facing. Figure 4 The first image 410 of shows a human face facing the right side of the image. Figure 4 The second image 420 of shows a human face facing forward in the image and towards the image capture device 105A that captured the second image 420. Figure 4 The third image 430 of shows a human face facing the left side of the image.

[0078] Several facial features are marked in the three images 410, 420, and 430. Specifically, in the three images 410, 420, and 430, all eyes, ears, cheeks, and noses are identified and marked with white circle indicators and / or labels. Other facial features that can be detected but not marked in Figure 4 the images 410, 420, and 430 of include the mouth, chin, eyebrows, and nostrils. When an image is received from the image sensor, the image processing device 105B can use a feature detection algorithm to detect any combination of these features. The feature detection algorithm can include feature detection, object detection, face detection, landmark detection, edge detection, feature recognition, object recognition, face recognition, landmark recognition, image classification, computer vision, or some combination thereof.

[0079] Once the image processing device 105B receives an image from the image sensor and detects these features, the image processing device 105B can determine the direction the object is facing based on these features. This can be done by comparing the distance between two features on the left side of the object (e.g., the left side of the human face) in the image with the distance between two features on the right side of the object (e.g., the right side of the human face) in the image. For example, the "left distance" between the left cheek of the object and the left ear of the object can be compared with the "right distance" between the right cheek of the object and the right ear of the object.

[0080] If the image processing device 105B receives an image and determines that the left distance and the right distance of the object in the image are equal to each other or within a threshold of each other, then the image processing device 105B determines that the object is facing forward, as in the second image 420. If the image processing device 105B receives an image and determines that the left distance of the object in the image exceeds the right distance by at least a threshold amount, then the image processing device 105B determines that the object is facing right, as in the first image 410. If the image processing device 105B receives an image and determines that the right distance of the object in the image exceeds the left distance by at least a threshold amount, then the image processing device 105B determines that the object is facing left, as in the third image 430. The left distance and the right distance can also be calculated as the distances from the left or right feature to the center feature (e.g., a person's nose, mouth, or chin), respectively. For example, the left distance can be the distance from the tip of the object's nose to the object's left eye, and the right distance is the distance from the tip of the object's nose to the object's right eye. Depending on the most clearly visible feature in the object, different features can be used. For example, if the object has long hair covering the object's ears, then the object's eyes or cheeks can be used instead of the object's ears as the features for calculating the left distance value and the right distance value. On the other hand, if some features are not visible, this can also be used to determine the direction the object is facing. For example, if the image processing device 105B detects the left ear of the object in the image but cannot detect the right ear of the object in the image, then the image processing device 105B can determine that the object is facing right because the right ear of the object is hidden behind the object. Similarly, if the image processing device 105B detects the right ear of the object in the image but cannot detect the left ear of the object in the image, then the image processing device 105B can determine that the object is facing left because the left ear of the object is hidden behind the object.

[0081] Note that, as discussed herein, the left side of the object means the side of the object that is closest to the left side of the image and / or the left side of the object as depicted in the image, and the right side of the object means the side of the object that is closest to the right side of the image and / or the right side of the object as depicted in the image. Thus, in some cases, the left ear or left eye or left cheek of the object discussed herein can be the right ear or right eye or right cheek that the object itself might consider, and vice versa. Therefore, it should be understood that these directions can be reversed so as to discuss the directions from the perspective of the object rather than from the perspective of the captured image.

[0082] Feature detection and / or recognition algorithms may be performed using any suitable feature recognition and / or detection techniques. In some embodiments, the feature detection and / or recognition algorithms applied by the image processing device 105B may include and / or incorporate image detection and / or recognition algorithms, object detection and / or recognition algorithms, face detection and / or recognition algorithms, feature detection and / or recognition algorithms, landmark detection and / or recognition algorithms, edge detection algorithms, boundary tracking functions, or some combination thereof. Feature detection is a technique for detecting (or locating) target features from an image or video frame. The detected features or objects may be represented using a bounding region that identifies the location and / or approximate boundary of the object (e.g., face) in the image or video frame. The bounding region of the detected object may include a bounding box, a bounding circle, a bounding ellipse, a bounding polygon, or a region of any other suitable shape that represents and / or includes the detected object. Object detection and / or recognition may be used to identify the detected object and / or classify the detected object into a category or type of object. For example, feature recognition may identify multiple edges and corners in a scene region. Object detection may detect that the edges and corners detected in that region all belong to a single object. Object detection and / or object recognition and / or face detection may identify that the object is a human face. Object recognition and / or face recognition may further identify the identity of the person corresponding to the face.

[0083] In some embodiments, feature detection and / or recognition algorithms may be performed using any suitable feature recognition and / or detection techniques. In some embodiments, the feature detection and / or recognition algorithms may be based on a machine learning model trained on images of the same type of object and / or feature using a machine learning algorithm, which may extract features of the image and detect and / or classify objects including those features based on the training of the model by the algorithm. For example, the machine learning algorithm may be a neural network (NN), such as a convolutional neural network (CNN), a time delay neural network (TDNN), a deep feedforward neural network (DFFNN), a recurrent neural network (RNN), an autoencoder (AE), a variational AE (VAE), a denoising AE (DAE), a sparse AE (SAE), a Markov chain (MC), a perceptron, or some combination thereof. The machine learning algorithm may be a supervised learning algorithm, a deep learning algorithm, or some combination thereof.

[0084] In some embodiments, computer vision-based feature detection and / or recognition techniques can be used. Different types of computer vision-based object detection algorithms can be used. In an illustrative example, template matching-based techniques can be used to detect one or more hands in an image. Various types of template matching algorithms can be used. An example of a template matching algorithm can perform Haar or Haar-like feature extraction, integral image generation, Adaboost training, and cascade classification. This object detection technique performs detection by applying a sliding window (e.g., having a rectangular, circular, triangular, or other shape) on the image. The integral image can be calculated as an image representation that evaluates the features of a specific region (e.g., rectangular or circular features) from the image. For each current window, the Haar features of the current window can be calculated from the integral image mentioned above, which can be calculated before calculating the Haar features.

[0085] Harr features can be calculated by computing the sum of the image pixels within a specific feature region of the target image (such as those of the integral image). For example, in a face, the area with eyes is usually darker than the area with the nose bridge or cheeks. Haar features can be selected by choosing the best features and / or a learning algorithm (e.g., Adaboost learning algorithm) that trains a classifier using them, and can be used to effectively classify a window as a face (or other target) window or a non-face window using a cascade classifier. The cascade classifier includes a plurality of classifiers combined in cascade, which allows the background region of the image to be quickly discarded while performing more calculations on regions of similar targets. Using the face as an example of a body part of an external viewer, the cascade classifier can classify the current window as a face category or a non-face category. If a classifier classifies the window as a non-face category, the window is discarded. Otherwise, if a classifier classifies the window as a face category, the next classifier in the cascade arrangement will be used to test again. Until all classifiers determine that the current window is a face (or other target), the window will be marked as a candidate for a hand (or other target). After detecting all windows, a non-maximum suppression algorithm can be used to group the windows around each face to generate the final result of one or more detected faces.

[0086] Figure 5 is a conceptual diagram showing a user interface 510 of an image capture device 500, which has a positioning guidance indication for guiding a user to move the image capture device 500 a specific distance in a specific direction. Figure 5 The user interface 510 is an image capture user interface and shows a preview image of an image recently received from the image sensor 130 of the image capture device 500, at least until the user of the image capture device 500 presses the shutter button 560 to capture an image. Figure 5The image capture device 500 may include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the computing device 1800, or some combination thereof.

[0087] The preview image displayed by the user interface 510 is an image with an object 520. The object 520 is a person, and the face of the person is aligned with the center of the preview image. Since the face of the person is aligned with the center of the preview image, moving the image capture device 500 to align the object 520 (the face of the person) with one or more of the third lines of the preview image will improve the image composition based on the rule of thirds. The image processing device 105B of the image capture device 500 identifies the object 520, and identifies, for example, the intersection of the third line closest to the object 520, and generates and outputs a positioning guidance indication that guides the user of the image capture device 500 to move the image capture device 500 to align the object 520 with the intersection of the third line. In Figure 5In the user interface 510, the positioning guidance indication is a visual indicator, which is shown as a small icon 530 representing the image capture device 500, an arrow 550 extending from the icon 530 indicating the direction of moving the image capture device 500, and a target rectangle 540 located at the other end of the arrow 550, where the target rectangle 540 represents the position where the movement of the image capture device 500 should stop. The movement can be a translational movement (different from a rotational movement), and thus the direction can be a translational direction. In some cases, the movement can include rotation, and the direction can include a rotational direction. The arrow in the user interface 510 points to the lower left, indicating that the user should move the image capture device 500 to the lower left. When the user moves the image capture device 500 to the lower left in the direction indicated by the arrow 550, the icon 530 can move along the arrow towards the target rectangle 540 until the icon 530 reaches the target rectangle. In this way, the direction of the arrow 550 and the direction of the target rectangle 540 relative to the icon 530 indicate the direction in which the image processing device 150B of the image capture device 500 guides the user to move the image capture device 500. The length of the arrow 550 and the distance between the target rectangle 540 and the icon 530 show a representation of the distance in which the image processing device 150B of the imaging device 500 guides the user to move the imaging device 500 in this direction. In some cases, the arrow 550 or the target rectangle 540 can be omitted. Alternative interfaces can be used, such as an audio interface that tells the user to move the image capture device in a certain direction and / or move the image capture device a certain distance in that direction. A tactile interface element can also be used. For example, once the image capture device 500 reaches the correct position, the image processing device 150B of the imaging device 500 can actuate one or more motors to vibrate the image capture device 500. Alternatively, the image processing device 150B of the imaging device 500 can actuate one or more motors to vibrate the image capture device 500 until the image capture device 500 reaches the correct position.

[0088] Although the image 510 includes a single object 520—a human face—some images can include more than one object. For example, an image can depict multiple people, pets, documents, and display screens, all of which can be detected by the image capture device 500 and determined by the image capture device 500 as objects depicted within the image. In some cases where the image depicts multiple objects, the image composition can be improved by aligning at least a subset of the objects with at least one of these four trisectors and / or aligning them to at least one of the intersection points of these four trisectors. In some examples, the image capture device 500 can select a subset of the objects depicted in the image as the selected objects. The image capture device 500 can output guidance for guiding the movement of the image capture device 500, such as along Figure 5The depicted arrow 550 is directed to the small icon 530 of the target rectangle 540 such that each selected object is aligned with and / or aligned to at least one of the four trisectors. For example, one or more selected objects may be selected to include one or more of the most prominent objects among all the objects depicted in the image. The image capture device 500 may identify the most prominent objects by generating a saliency map of the image and selecting one or more objects corresponding to the highest saliency regions of the image. The image capture device 500 may identify the most prominent objects by, for example, detecting which object(s) is / are most in the foreground of the image (closest to the image capture device 500 during capture) based on depth sensor information and / or based on the dimensions depicted in the image, and selecting the closest object(s) in the foreground of the image as the selected objects. The image capture device 500 may select the object(s) depicted as the largest in the image as the selected objects. The image capture device 500 may receive one or more user inputs identifying the objects, such as by a user touching, clicking, gesturing, or otherwise selecting one or more objects depicted in the preview image, and may select the selected objects based on the objects identified by the user input. In some cases, the image capture device 500 may select the selected object(s) based on a combination of the above selection techniques.

[0089] Figure 6 is a flowchart showing an operation 600 for guiding image capture based on the direction an object in an image faces. Although operation 600 refers to the image capture device 105A, operation 600 may be performed by various devices, which may include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0090] In operation 605, the device receives an image captured by the image sensor 130 of the image capture device 150A. The term "captured" as used herein may refer to temporary storage (e.g., in a temporary image buffer of the device), long-term storage in a non-transitory computer-readable storage medium, or some combination thereof. In operation 610, the device identifies the objects in the image using, for example, one of the object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein.

[0091] In operation 615, the device determines the location of the objects in the image and the direction the objects in the image face. Determining the direction the objects in the image face may be based on the positioning of multiple features of the objects in the image relative to each other. As referred to Figure 4As discussed, the device can identify two features in the image along the left side and / or at the center of the image and determine the left distance between the two features. The device can identify two features in the image along the right side and / or at the center of the image and determine the right distance between the two features. The device can determine the direction in which the object in the image is facing by comparing the left distance and the right distance. If the left distance is equal to the right distance or within a threshold of the right distance, the object is facing forward. If the left distance exceeds the right distance by at least the threshold, the object is facing right. If the right distance exceeds the left distance by at least the threshold, the object is facing left. If the object is or includes a person, the features can include, for example, ears, cheeks, eyes, eyebrows, nose, mouth, chin, chest, abdomen, back, rear, legs, arms, shoulders, elbows, knees, ankles, hands, feet, another appendage, or a part thereof, or some combination of them.

[0092] In some cases, determining the direction in which the object in the image is facing can be based on receiving an additional image also captured by the image sensor 130, where the device determines the direction of movement of the object based on the image and the second image, and determines that the direction in which the object is facing is the direction of movement of the object. For example, if an additional image is captured after the image, and the object appears to be more to the left in the additional image than in the image, the device determines that the object is moving left and thus facing left. Or, if an additional image is captured before the image, and the object appears to be more to the left in the additional image than in the image, the device determines that the object is moving right and thus facing right.

[0093] At operation 620, the device generates and outputs an indication for positioning the image capture device based on the direction in which the object in the image is facing and the position of the object in the image. The indication can identify the direction in which the device is to be moved so as to improve the framing of the object in the second image to be captured after the image capture device is moved. The indication can include at least one of a visual indicator, an auditory indicator, or a vibration indicator. For example, the visual indicator can look similar to Figure 5 the visual indicators 530 / 540 / 550.

[0094] In some cases, the device can identify that the image capture device has moved in that direction and can receive a second image captured by the image sensor 130 after identifying that the image capture device has moved in that direction, where the image sensor 130 has captured the second image. In some cases, outputting guidance for positioning the image capture device includes: outputting, at the image capture device, an indicator indicating that the image capture device is to remain stationary between the capture of the first image and the capture of the second image.

[0095] Figure 7AFIG. 0 is a conceptual diagram showing a user interface of an image capture device 700, the user interface 700 having a positioning guidance indicator 730 that guides a user to tilt the image capture device counterclockwise to make a horizontal line in the image horizontal. Figure 7A The user interface of FIG. 1 is an image capture user interface and shows a preview image 710 of an image most recently received from an image sensor 130 of the image capture device 700. Figure 7A - 7B The image capture device 700 of FIG. 2 may include an image capture device 105A, an image processing device 105B, an image capture and processing system 100, a computing device 1800, or some combination thereof.

[0096] Figure 7A The preview image 710 shown in the UI of FIG. 3 includes a horizontal line that is not horizontal. In other words, the horizontal line is not level and may, for example, be more than a threshold angle away from horizontal. Figure 7A The UI of FIG. 4 includes a horizontal dashed line 720 for reference so that it can be more clearly seen that the horizontal line in the image 710 is not horizontal. Figure 7A The image capture device 700 of FIG. 5 can detect that the image capture device 700 is tilted, for example, using an accelerometer, gyroscope, magnetometer, or inertial measurement unit (IMU) of the imaging device 700. The image capture device 700 generates a positioning guidance indicator 730 that has an icon representing the image capture device 700, an arrow showing that the image capture device 700 is to be rotated counterclockwise, and a counterclockwise guidance square of the image capture device 700 representing the position that the image capture device 700 is to be in order to make the horizontal line horizontal.

[0097] Figure 7B FIG. 6 is a conceptual diagram showing a user interface of an image capture device 700, the user interface 700 having a positioning guidance indicator760 that guides a user to tilt the image capture device clockwise to make a horizontal line in the image horizontal. Figure 7B The user interface of FIG. 7 is an image capture user interface and shows a preview image 740 of an image most recently received from an image sensor 130 of the image capture device 700. Figure 7B The preview image 740 shown in the UI of FIG. 8 includes a horizontal line that is not horizontal. Figure 7B The UI of FIG. 9 includes a horizontal dashed line 750 for reference so that it can be more clearly seen that the horizontal line in the image 740 is not horizontal. Figure 7BThe image capture device 700 can detect that the image capture device 700 is tilted, for example, using an accelerometer, gyroscope, magnetometer, or IMU of the imaging device 700. The image capture device 700 generates a positioning guidance indication 760 that has an icon representing the image capture device 700, an arrow showing that the image capture device 700 is to be rotated counterclockwise, and a counterclockwise guidance square of the image capture device 700 representing the position that the image capture device 700 is to be in order to make the horizontal line horizontal.

[0098] Figure 8 is a flowchart showing an operation 800 for guiding image capture based on sensor measurement data from one or more positioning sensors of the image capture device 105A. Although the operation 800 refers to the image capture device 105A, the operation 800 can be performed by various devices, which can include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0099] In operation 805, the device receives sensor measurement data from one or more positioning sensors of the image capture device 105A. The one or more positioning sensors can include at least one of an accelerometer, gyroscope, magnetometer, inertial measurement unit, global navigation satellite system (GNSS) receiver, or altimeter.

[0100] In operation 810, the device determines the orientation of the image capture device 105A based on the sensor measurement data. In operation 815, the device generates and outputs an indication for positioning the image capture device 105A based on the orientation of the image capture device 105A. In some cases, the indication identifies the direction in which the image capture device 105A is to be tilted so that the horizontal line in the image to be captured is horizontal after the image capture device 105A is tilted. In this case, tilting refers to the rotation of the image capture device 105A about one or more axes. These axes can include, for example, an axis perpendicular to the front surface and / or the rear surface of the image capture device 105A, i.e., an axis perpendicular to the display surface of the image capture device 105A, an axis perpendicular to the surface of the image sensor of the image capture device 105A, an axis perpendicular to the surface of the lens of the image capture device 105A, or some combination thereof. The indication includes at least one of a visual indicator, an auditory indicator, and a vibration indicator. For example, the visual indicator can include Figure 7A - 7BAny element that indicates 730 and 760. In some cases, the indicator also identifies the angle by which the image capture device 105A is to be tilted in that direction to improve the framing of an object in the image. The tilt can be referred to as a rotational movement. The direction of the tilt can be referred to as the rotational direction. The angle of the tilt can be referred to as the rotational angle. In some cases, the rotational movement can occur in pairs with a translational movement.

[0101] In some cases, the device also identifies that the device has been tilted in that direction, and after identifying that the image capture device 105A has been tilted in that direction, receives an image from the image sensor 130 of the image capture device 105A, where the image sensor 130 has captured the image. In some cases, the output for guiding the positioning of the image capture device 105A includes: outputting an indicator at the image capture device indicating that the image capture device is to remain stationary (e.g., between the capture of the first image and the capture of the second image).

[0102] Figure 9 is a conceptual diagram showing a view visible to a first image sensor with a normal lens superimposed on a view visible to a second image sensor with a wide - angle lens. Figure 9 The image capture device 900 in Figure 9 displays a preview image that includes imaging data captured by a first image sensor of the image capture device 900 with a normal lens and a second image sensor of the image capture device 900 with a wide - angle lens, where the wide - angle lens has a wider angle than the normal lens. The image capture device 900 can include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the computing device 1800, or some combination thereof.

[0103] The entire preview image represents a view 920 visible to the second image sensor with a wide - angle lens. A rectangle with a black outline is shown within the preview image, and the area inside the rectangle with the black outline represents a view 910 visible to the first image sensor with a normal lens. If the image capture device 900 only considers the view 910 visible to the first image sensor with a normal lens, the image capture device 900 may not detect that an object 940 (a face) has been cut off and is not included in the view 910 visible to the first image sensor with a normal lens. However, if the image capture device 900 views the view 920 visible to the second image sensor with a wide - angle lens, the image capture device 900 can detect the object 940, and if the user wishes to capture the object 940, can warn the user of the image capture device 900 to move the image capture device 900.

[0104] Figure 10It is a flowchart showing an operation of guiding image capture using a first image sensor with a first lens based on image data from a second image sensor with a second lens, where the second lens has a wider angle than the first lens. Although image capture device 105A is referred to in operation 1000, operation 1000 can be performed by various devices, which can include image capture device 105A, image processing device 105B, image capture and processing system 100, image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, computing device 1800, or some combination thereof.

[0105] In operation 1005, the device receives a first image of a scene captured by a first image sensor of image capture device 105A, where the first image sensor is associated with the first lens. In operation 1010, the device receives a second image of the scene captured by a second image sensor of the image capture device, where the second image is captured by the second image sensor within a threshold time after the first image sensor captures the first image, and the second image sensor is associated with a second lens having a wider angle than the first lens.

[0106] In operation 1015, the device determines that the image composition of the first image is sub-optimal based on the second image. In operation 1020, the device generates and outputs an indication for positioning the image capture device such that the image composition of a third image to be captured by the first image sensor is better than the image composition of the first image.

[0107] In some cases, the device also receives a third image from the first image sensor, where the third image is captured by the first image sensor after the first image sensor captures the first image. Determining that the image composition of the first image is sub-optimal is based on at least a part of an object being outside the frame of the first image, and that part of the object is included in the third image. For example, object 940 is at least partially outside the frame in view 910 of Figure 9 but once Figure 9 image capture device 900 is moved to the right, for a later image, object 940 will be within the frame in view 910 of Figure 9

[0108] In some cases, determining that the image composition of the first image is sub-optimal includes identifying a horizontal line in the second image and determining that the image capture device is to be tilted to make the horizontal line horizontal. Outputting guidance for positioning the image capture device includes outputting an indicator at the image capture device that identifies the direction in which the image capture device is to be tilted to make the horizontal line in the third image horizontal. The indicator includes at least one of a visual indicator, an auditory indicator, and a vibration indicator. In some cases, the indicator also identifies the angle by which the image capture device is to be tilted in that direction to improve the framing of the object in the second image.

[0109] In some cases, at least one subset of operation 1000 can be performed remotely by one or more web servers of a cloud service that performs image analysis (e.g., steps 1010 and / or 1015), generates and / or outputs the indication and / or guidance of operation 1020, or some combination thereof.

[0110] Figure 11 is a conceptual diagram showing the user interface of an image capture device, where an image of a previously captured object is overlaid on an image of an object captured by the image sensor of the image capture device. Figure 11 The image capture device 1100 displays a preview image 1110 via an image capture interface. The image capture device 1100 identifies an object in the image 1110, which is the Eiffel Tower in the image 1110. The object can be determined using one of object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein. The object can be determined based on a caption received from an input device that receives user input (e.g., "I'm at the Eiffel Tower!"). The object can be determined based on the user's schedule in a calendar or clock or reminder application, e.g., if the schedule has an event corresponding to a tour of the Eiffel Tower at a date and time that matches or is within a threshold time of the date and time of image capture. The object can be determined based on a specific image capture setting selected by the user (e.g., "sports mode", "food mode", "pet mode", "portrait mode", "landscape mode", "group photo mode", "night mode"). The object can be determined by simply prompting the user to provide the object (or select from a set of possible objects determined by the device) before, during, and / or after capturing the image. The object can be determined based on the image capture device 1100 identifying that its position during or within a threshold time of capture is within a threshold distance of the known position of the object (here, the known position of the Eiffel Tower). The image capture device 1100 can determine its position based on signals received by its GNSS / GPS receiver.

[0111] The image capture device 1100 identifies a second image 1120 of the same object (i.e., the Eiffel Tower). The second image 1120 can be a previously captured image. The second image 1120 can be an image that the image capture device 1100 or another device has determined to have good image composition based on various image composition rules and guidelines. The second image 1120 can be an image captured by a well-known photographer. The second image 1120 can be an image that has received a positive rating on a photography rating website. The second image 1120 can be an image that has received positive feedback (e.g., above a specific threshold of "likes" and / or "shares") on a social media website.

[0112] The image capture device 1100 then generates an overlay based on the second image 1120 and overlays the overlay on the preview image (or a preview image captured later), as shown in the composite image 1130. The overlay is shown in the composite image 1130 using a dashed line. The overlay can be combined with the preview image using alpha compositing or transparency. In some cases, the overlay can simply be the image data corresponding to the object in the second image 1120 rather than all of the image data of the image 1120. In some cases, the overlay can simply be the outline of the object in the second image 1120 rather than all of the image data of the object in the image 1120. By showing the overlay in the composite image 1130 to the user of the image capture device 1100, the user of the image capture device 1100 can better understand what is the optimal image composition for an image of the object and how to reposition the image capture device 1100 to achieve the optimal image composition for an image of the object.

[0113] In some cases, the second image 1120 may include metadata that identifies the geographic coordinates from which the second image 1120 was captured (e.g., determined using a GPS / GNNS receiver and / or altimeter of the image capture device that captured the second image 1120) and / or the direction the camera faced during capture (e.g., determined using a GPS / GNNS receiver and / or accelerometer and / or gyroscope of the image capture device that captured the second image 1120). In such cases, the image capture device 1100 may also display or otherwise output an indicator that identifies the coordinates to which the image capture device 1100 should be moved in order to capture an image similar to the second image 1120, and in some cases the direction the image capture device 1100 should face in order to capture an image similar to the second image 1120. These coordinates may include geographic coordinates such as latitude and longitude coordinates. These coordinates may include elevation coordinates instead of or in addition to latitude and longitude coordinates. In some cases, the indicator may include a map that may show the current location of the image capture device 1100 (the “first location” of the image capture device 1100), and the location to which the image capture device 1100 should be moved in order to capture an image similar to the second image 1120 (the “second location” of the image capture device 1100). In some cases, a path from the first location of the image capture device 1100 to the second location of the image capture device 1100 may be shown on the map. The path may be generated based on navigation for walking, driving, public transportation, or some combination thereof. The indicator may also include the direction the image capture device 1100 should face at the second location in order to capture an image similar to the second image 1120. The direction may include a compass direction (e.g., north, east, south, west, or some direction therebetween), which may be shown on the map if a map is used. The direction may include an angle corresponding to yaw, pitch, roll, or some combination thereof.

[0114] In some cases, multiple overlays and / or other location indicators corresponding to multiple objects visible in the preview image 1110 and / or possible objects known to be in the vicinity can be provided. For example, if the current location (first location) of the image capture device 1100 is in Times Square in midtown Manhattan (New York), the image capture device 1100 can output a list of the following: objects visible in the preview image 1110, objects known to be within Times Square, objects known to be within a predetermined radius of Times Square, objects known to be within a predetermined radius of the first location of the image capture device 1100, or some combination thereof. A user of the image capture device 1100 can select one or more of these objects from the list, and can generate and output an overlay (as in the composite image 1130) for that object based on the second image 1120, and can also generate and output any other location indicator discussed herein (e.g., coordinates, map, etc.) for that object. If an object different from the object in the second image 1120 is selected from the list, a different second image 1120 of the selected object can be identified and used.

[0115] Figure 12 is a flowchart showing operations for guiding the capture and / or processing of an image of an object based on another image of the same object. Although the image capture device 105A is referenced in operation 1200, operation 1200 can be performed by various devices, which can include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0116] In operation 1205, the device receives a first image of a scene captured by an image sensor of the image capture device. In operation 1210, the device identifies the object depicted in the first image. In operation 1215, the device identifies a second image that also depicts the object. Operations 1210 and / or 1215 can be performed using one of object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein.

[0117] In some cases, the device receives sensor measurement data from one or more location sensors of the image capture device, including a global navigation satellite system (GNSS) receiver. The device determines the location of the image capture device 105A within a threshold time of the time the first image was captured. Identifying the object depicted in the first image in operation 1210 is based on identifying that the location of the image capture device is within a threshold distance of the location of the object.

[0118] After operation 1215 are operation 1220, operation 1225, operation 1230, or some combination thereof. In operation 1220, the device generates and outputs an indication for positioning the image capture device such that the position of an object in a third image to be captured by the image sensor matches the position of the object in the second image. In some cases, outputting the indicator includes, for example, using alpha compositing to overlay at least a portion of the second image on a preview image displayed by the image capture device (as in Figure 11 the composite image 1130).

[0119] In operation 1225, the device generates and outputs capture setting guidance for adjusting one or more attributes of the image capture device based on one or more image capture settings used to capture the second image. In some cases, outputting the capture setting guidance includes automatically adjusting one or more attributes of the image capture device based on one or more image capture settings before the image sensor captures the third image. Image capture settings can include, for example, zoom, focus, exposure time, aperture size, ISO, depth of field, analog gain, aperture scale, or some combination thereof.

[0120] In operation 1230, the device generates and outputs processing setting guidance for processing a third image to be captured by the image sensor based on one or more image processing settings applied to the second image. In some cases, outputting the processing setting guidance includes automatically applying one or more image processing settings to the third image in response to receiving the third image from the image sensor. Image processing settings can include, for example, brightness, contrast, saturation, gamma, levels, histogram, color levels, color warmth, blur, sharpness, levels, curves, filters, cropping, or some combination thereof. Filters can include high-pass filters, low-pass filters, band-pass filters, band-stop filters, or some combination thereof. Filters can also refer to visual effects applied to an image that automatically adjust one or more of the previously mentioned image processing settings to apply a specific "look" to the image, e.g., a filter specifically applies a "vintage photo" look that mimics a photo captured using a film camera from a certain era, or modifies the image to make it look painted or hand-drawn, or some other visual modification.

[0121] In some cases, at least a subset of operation 1200 can be remotely executed by one or more web servers of a cloud service that performs image analysis (e.g., step 1210), locates the second image (e.g., step 1215), generates and / or outputs indicators and / or guidance (e.g., operations 1220, 1225, and / or 1230), or some combination thereof.

[0122] Figure 13is a conceptual diagram showing a user interface of an image capture device, where an image of a previously captured object is used to generate guidance overlaid on an image of a different object captured by an image sensor of the image capture device. Figure 13 The image capture device 1300 displays a preview image 1310 via an image capture interface. The image capture device 1300 identifies an object in the preview image 1310, which is a standing person in the preview image 1310. The object can be determined using one of object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein. The object can be determined based on the image capture device 1300 identifying that its position during capture or within a threshold time of capture is within a threshold distance of a known position of the object (e.g., if the person shares their position via social media or other means). The image capture device 1300 can determine its position based on signals received by a GNSS / GPS receiver of the image capture device 1300.

[0123] The image capture device 1300 identifies a second image 1320 of different objects (i.e., different people sitting down). In some cases, the different objects depicted in the second image 1320 can be the same type of object as the object depicted in the preview image 1310. For example, the object in the preview image 1310 can be a person, and the different objects in the second image 1320 can be different people, or the same person in a different pose and / or outfit. The object in the preview image 1310 can be an object of a specific object type (e.g., a building, a statue, a monument), and the different objects in the second image 1320 can be different objects of the same object type with similar dimensions.

[0124] The image capture device 1300 can determine that the object in the preview image 1310 shares one or more similarities with the different objects in the second image 1320. These similarities can include similarities in object type as described above. These similarities can include similarities in dimensions. These similarities can include similarities in color or color scheme. These similarities can include similarities in lighting.

[0125] These similarities can include one or more saliency values associated with an object in the preview image 1310 being within a predetermined range of one or more saliency values associated with a different object in the second image. For example, the image capture device 1300 can generate a first saliency map of the preview image 1310 and a second saliency map of the second image 1320. The first saliency map includes saliency values corresponding to each pixel of the preview image 1310 and can also include confidence values corresponding to each saliency value. The second saliency map includes saliency values corresponding to each pixel of the second image 1320 and can also include confidence values corresponding to each saliency value. The image capture device 1300 can locate an object in the preview image 1310 based on the pattern of saliency values in the first saliency map. The image capture device 1300 can locate a different object in the second image 1320 based on the pattern of saliency values in the second saliency map. The image capture device 1300 can determine that an object in the preview image 1310 is similar to a different object in the second image 1320 based on the similarity between the pattern of saliency values in the first saliency map and the pattern of saliency values in the second saliency map.

[0126] The second image 1320 can be a previously captured image. Just as Figure 11 the second image 1120, the second image 1320 can be an image that the image capture device 1300 (or another device) has determined to have good image composition, captured by a well-known photographer, received a positive rating on a photography rating website, received positive feedback (e.g., above a certain threshold of "likes" and / or "shares") on a social media website, or some combination thereof. In some cases, the image capture device 1300 (or another device) can select the second image 1320 from an image set based on one or more similarities between an object in the preview image 1310 and a different object in the second image 1320. In some cases, if the second saliency map includes a pattern of saliency values having corresponding confidence values that are on average below a threshold, the image capture device 1300 can reject the selection of the second image 1320. The threshold can be determined based on the average of the confidence values corresponding to the pattern of saliency values, where the pattern of saliency values corresponds to an object in the first saliency map.

[0127] The image capture device 1300 then generates an overlay based on the second image 1320 and overlays the overlay on the preview image 1310 (or a later captured preview image), as shown in the composite image 1330. The overlay is shown in the composite image 1330 using a dashed line. The overlay can be combined with the preview image using alpha compositing or transparency.

[0128] In some cases, the overlay may include image data corresponding to a second image 1320, an object in the second image 1320, or a contour or other abstract representation of an object in the second image 1320. Alternatively, as shown in the composite image 1330 of Figure 13 , the overlay may include image data corresponding to a preview image 1310, an object in the second image 1320, or a contour or other abstract representation of an object in the second image 1320. Even in cases where the overlay is based on image data corresponding to at least some of the first images in the first image 1310, the positioning, size, and orientation of the overlay may be based on the positioning, size, and orientation of the object in the second image 1320. By displaying the overlay in the composite image 1330 to a user of the image capture device 1300, the user of the image capture device 1300 can better understand what an optimal image composition is for an image of the object and how to reposition the image capture device 1300 to achieve an optimal image composition for an image of the object. For example, in the composite image 1330, the object in the overlay is larger and further to the right in the composite image 1330, indicating that the user should move the image capture device 1300 closer to the object (or zoom in) and move the image capture device 1300 to the left so that the object appears further to the right in the image to be captured by the image capture device 1300. Figure 13 Even in cases where the overlay is based on image data corresponding to at least some of the first images in the first image 1310, the positioning, size, and orientation of the overlay may be based on the positioning, size, and orientation of the object in the second image 1320. By displaying the overlay in the composite image 1330 to a user of the image capture device 1300, the user of the image capture device 1300 can better understand what an optimal image composition is for an image of the object and how to reposition the image capture device 1300 to achieve an optimal image composition for an image of the object. For example, in the composite image 1330, the object in the overlay is larger and further to the right in the composite image 1330, indicating that the user should move the image capture device 1300 closer to the object (or zoom in) and move the image capture device 1300 to the left so that the object appears further to the right in the image to be captured by the image capture device 1300.

[0129] In some cases, the composite image 1330 may further include an indication that identifies a location to which the image capture device 1300 should move from its current location (the "first location" of the image capture device 1300) to a location where the image capture device 1300 can capture an image similar to the overlay, where the overlay is based on the second image 1320. The indication may include geographical coordinates of the first location and / or the second location, which may include latitude coordinates, longitude coordinates, and / or altitude coordinates. In some cases, the indication may include a map that may show the first location of the image capture device 1300, the second location of the image capture device 1300, a path from the first location to the second location, or some combination thereof. The indication may include the direction that the image capture device 1300 should face at the second location in order to capture an image similar to the overlay based on the second image 1320. The direction may include a compass direction (e.g., north, east, south, west, or some direction therebetween), which may be shown as an arrow on the map. The direction may include an angle corresponding to yaw, pitch, roll, or some combination thereof.

[0130] In some cases, multiple overlays and / or other location indicators corresponding to multiple objects visible in the preview image 1310 and / or possible objects known to be nearby can be provided. For example, if the current location (first location) of the image capture device 1300 is in Times Square in midtown Manhattan (New York), the image capture device 1300 can output a list of the following: objects visible in the preview image 1310, objects known to be within Times Square, objects known to be within a predetermined radius of Times Square, objects known to be within a predetermined radius of the first location of the image capture device 1300, or some combination thereof. A user of the image capture device 1300 can select one or more of these objects from the list and can generate and output an overlay (as in the composite image 1330) for that object based on the second image 1320, and can also generate and output any other location indication discussed herein (e.g., coordinates, map, etc.) for that object. In some cases, the second image 1320 can be selected based on the object type of the object selected from the list. For example, if the object selected from the list is a building, the second image 1320 can be selected as a view of the building. If the object selected from the list is a person, the second image 1320 can be selected as a view of the person.

[0131] The image capture device 1300 also provides image capture setting guidance and image processing setting guidance overlaid on the composite image 1330. Specifically, the image capture device 1300 generates and displays a guidance box indicating "turn off flash" to suggest to the user to turn off the flash. This suggestion can be based on the second image 1320 having been captured without flash. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1300 can automatically turn off the flash. The image capture device 1300 also generates and displays a guidance box indicating "extend exposure time" to suggest to the user to extend the exposure time before capture. This suggestion can be based on the second image 1320 having been captured with an exposure time longer than the exposure time that the image capture device 1300 is currently set to. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1300 can automatically extend the exposure time. The image capture device 1300 also generates and displays a guidance box indicating "increase contrast" to suggest to the user to increase the contrast during image processing after image capture. This suggestion can be based on the second image 1320 having been processed to increase the contrast after capture, or based on the second image 1320 having only a higher contrast than the image currently received from the image sensor of the image capture device 1300. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1300 can automatically increase the contrast after capturing the image.

[0132] The advantage of generating guidance (e.g., overlays, image capture settings, and / or image processing settings) based on a second image 1320 having an object different from the object of the preview image 1310 is flexibility. For example, if the second image 1320 is selected from an image set, the image capture device 1300 (or another device) that performs the selection does not need to find an image having exactly the same object as the preview image 1310. Thus, if the preview image 1310 depicts a person as its object, the image capture device 1300 (or another device) that performs the selection only needs to find a second image 1320 having another person or another object similar to the person in the preview image 1310. Similarly, if the preview image 1310 depicts the Eiffel Tower as its object, the image capture device 1300 (or another device) that performs the selection only needs to find a second image 1320 having another building or another object similar to the Eiffel Tower in the preview image 1310. Thus, the image capture device 1300 can generate and output useful guidance even for blurry or unusual objects.

[0133] Figure 14 is a flowchart showing operations for guiding the capture and / or processing of an image of an object based on another image of a different object. Although the image capture device 105A is referenced in operation 1400, operation 1400 can be performed by various devices, which can include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0134] In operation 1405, the device receives a first image of a scene captured by an image sensor of an image capture device. In operation 1410, the device identifies a first object depicted in the first image, e.g., using one of object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein. In operation 1415, the device identifies a second image depicting a second object.

[0135] After operation 1415 are operations 1420, 1425, 1430, or some combination thereof. In operation 1420, the device generates and outputs an indication for positioning the image capture device such that the position of the first object in a third image to be captured by the image sensor matches the position of the second object in the second image. In some cases, the output indicator includes, for example, overlaying at least a portion of the second image on a preview image displayed by the image capture device using alpha compositing (or an edited portion of the first image in a combined image 1330 as Figure 13 of the combined image 1330).

[0136] At operation 1425, the device generates and outputs capture setting guidance for adjusting one or more attributes of the image capture device based on one or more image capture settings used to capture the second image. In some cases, outputting the capture setting guidance includes automatically adjusting one or more attributes of the image capture device based on one or more image capture settings before the image sensor captures the third image.

[0137] At operation 1430, the device generates and outputs processing setting guidance for processing a third image to be captured by the image sensor based on one or more image processing settings applied to the second image. In some cases, outputting the processing setting guidance includes automatically applying one or more image processing settings to the third image in response to receiving the third image from the image sensor.

[0138] In some cases, at least a subset of operation 1400 may be remotely executed by one or more web servers of a cloud service that performs image analysis (e.g., step 1410), locates the second image (e.g., step 1415), generates and / or outputs indicators and / or guidance (e.g., operations 1420, 1425, and / or 1430), or some combination thereof.

[0139] Figure 15 is a conceptual diagram showing a user interface of an image capture device, where a machine learning model trained using an image set is used to generate guidance overlaid on an image of an object captured by an image sensor of the image capture device. Figure 15 The image capture device 1500 displays a preview image 1510 via an image capture interface. The image capture device 1500 identifies an object in the preview image 1510, which is a standing person in the preview image 1510. The object can be determined using one of object detection, feature detection, face detection, or other image detection or recognition techniques discussed herein.

[0140] The image capture device 1500 inputs the preview image 1510 into a machine learning model 1520. The machine learning model 1520 is trained using an image set with the identified object. The machine learning model 1520 outputs one or more insights based on the preview image 1510 and its training. These insights can include, for example, an alternative positioning of the object in the preview image 1510 generated using the machine learning model, image capture settings generated using the machine learning model, image processing settings generated using the machine learning model, or some combination thereof.

[0141] The image set used to train the machine learning model 1520 can be selected based on an image set entirely captured by a specific photographer, painter, or other artist. The user of the image capture device 1500 can select this photographer. For example, the user of the image capture device 1500 can select a machine learning model 1520 trained using photos captured by the photographer Ansel Adams. Thus, the insights generated by the machine learning model 1520 can help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to be more similar to the image composition, image capture settings, and image processing settings used by the photographer Ansel Adams. Similarly, the user of the image capture device 1500 can select a machine learning model 1520 trained using paintings created by the artist Monet. The insights generated by the machine learning model 1520 can thus help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to generate an appearance and style similar to that of the paintings created by Monet.

[0142] The image set used to train the machine learning model 1520 can be selected based on an image set all having a similar scene type and / or object type. For example, if the object is a building, then even if the exact identity of the building is unrecognizable, a machine learning model 1520 trained using an image set all having buildings can be employed. The insights generated by the machine learning model 1520 can thus help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to be suitable for photographing buildings and can, for example, reduce glare from windows. In another example, if the image is of a baby at the Grand Canyon, the image capture device 1500 can prompt the user to specify whether the focus of the image is the baby or the scenery and can select either a machine learning model 1520 trained using baby images or a machine learning model 1520 trained using natural scenery images. Some image capture devices 1500 have features that allow the image capture device 1500 to receive input from the user. The object can be determined based on the specific image capture settings selected by the user. For example, if the user selects "sports mode", the scene is likely to be a sports scene and the object is a player or a game situation. If the user selects "food mode", the scene is likely to be a kitchen or a dining scene and the object is food. If the user selects "pet mode", the scene / object is likely to be a fast-moving pet. If the user selects "portrait mode", the scene / object is likely to be a person holding a specific pose. If the user selects "landscape mode", the scene / object is likely to be a natural or urban landscape. If the user selects "group photo mode", the scene / object is likely to be a group of people. If the user selects "night mode", the scene / object is likely to be the night sky or a dimly lit outdoor scene. A machine learning model 1520 can be selected that is trained using training images having the same kind of object and / or scene in order to obtain suitable insights.

[0143] The image set used to train the machine learning model 1520 can be selected based on an image set all captured within a specific time of day, which can be determined based on the clock of the image capture device 1500 and / or a calendar indicating the time of year. Thus, the insights generated by the machine learning model 1520 can help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to suit the time of day (e.g., sunrise, daytime, sunset, dusk, night) when the user wishes to capture a photo. The image set used to train the machine learning model 1520 can be selected based on an image set all captured indoors or outdoors, such that the insights generated by the machine learning model 1520 can thus help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to suit indoor or outdoor photography. The image set used to train the machine learning model 1520 can be selected based on an image set all captured during a specific type of weather (e.g., sunny, cloudy, rainy, snowy), such that the insights generated by the machine learning model 1520 can thus help adjust the image composition, image capture settings, and image processing settings used by the user of the image capture device 1500 to suit the weather during the photography time and at the location where the image capture device 1500 captures images. The image set used to train the machine learning model 1520 can be selected based on an image set that the image capture device 1500 (or another device) has determined to have good image composition, received a positive rating on a photography rating website, or received positive feedback (e.g., above a specific threshold of "likes" and / or "shares") on a social media website. In some cases, the image set used to train the machine learning model 1520 can be selected based on an image set having some combination of the above characteristics.

[0144] The image capture device 1500 then generates an overlay based on the alternative positioning of the object in the preview image 1510 generated using the machine learning model and overlays the overlay on the preview image (or a later captured preview image), as shown in the composite image 1530. The overlay is shown in the composite image 1530 using a dashed line. The overlay can be combined with the preview image using alpha compositing or translucency. As Figure 15 shown in the composite image 1530, the overlay can include image data corresponding to the preview image 1510, the object in the preview image 1510, or the outline or other abstract representation of the object in the preview image 1510. By displaying the overlay in the composite image 1530 to the user of the image capture device 1500, the user of the image capture device 1500 can better understand what is the optimal image composition for the image of the object and how to reposition the image capture device 1500 to achieve the optimal image composition for the image of the object.

[0145] In some cases, the composite image 1530 may also include an indicator that identifies a location where the image capture device 1500 should move from its current location (the "first location" of the image capture device 1500) to a location where the image capture device 1300 can capture an image similar to the overlay map, where the overlay map is based on the machine learning model 1520. The indicator may include the geographical coordinates of the first location and / or the second location, which may include latitude coordinates, longitude coordinates, and / or altitude coordinates. In some cases, the indicator may include a map that may show the first location of the image capture device 1500, the second location of the image capture device 1500, the path from the first location to the second location, or some combination thereof. The indicator may also include the direction that the image capture device 1500 should face at the second location in order to capture an image similar to the overlay map based on the machine learning model 1520. The direction may include a compass direction (e.g., north, east, south, west, or some direction in between), which may be shown as an arrow on the map. The direction may include an angle corresponding to yaw, pitch, roll, or some combination thereof.

[0146] In some cases, a plurality of overlay maps and / or other location indicators corresponding to a plurality of objects visible in the preview image 1510 and / or possible objects known to be in the vicinity may be provided. For example, if the current location (first location) of the image capture device 1500 is in Times Square in midtown Manhattan (New York), the image capture device 1500 may output a list of the following: objects visible in the preview image 1510, objects known to be within Times Square, objects known to be within a predetermined radius of Times Square, objects known to be within a predetermined radius of the first location of the image capture device 1500, or some combination thereof. A user of the image capture device 1500 may select one or more of these objects from the list, and an overlay map (as in the composite image 1530) may be generated and output for that object based on the machine learning model 1520, and any other location indicators discussed herein (e.g., coordinates, maps, etc.) may also be generated and output for that object. In some cases, the machine learning model 1520 may be selected such that the object types in the training image set used to train the machine learning model 1520 match the object types of the objects selected from the list. For example, if the object selected from the list is a building, the machine learning model 1520 trained on a training image set of buildings may be selected. If the object selected from the list is a person, the machine learning model 1520 trained on a training image set of people may be selected.

[0147] The image capture device 1500 also provides image capture setting guidance and image processing setting guidance overlaid on the combined image 1530. Specifically, the image capture device 1500 generates and displays a guidance box indicating "improve focus on the object" to advise the user to adjust the focus (e.g., controlled by the focus control mechanism 125B) to ensure that the object is in focus, e.g., based on a machine learning model trained using an image set that typically has better focus on its objects. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1500 can automatically improve the focus on the object. The image capture device 1500 also generates and displays a guidance box indicating "increase the aperture size" to advise the user to increase the aperture size before capture, e.g., based on a machine learning model trained using an image set that typically has an aperture size larger than the aperture size currently set for the image capture device 1500. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1500 can automatically increase the aperture size. The image capture device 1500 also generates and displays a guidance box indicating "increase saturation" to advise the user to increase the saturation after image capture during image processing, e.g., based on a machine learning model trained using an image set in which the saturation is increased during processing or simply an image set with higher saturation. In some cases, instead of displaying such a guidance box or in addition to displaying such a guidance box, the image capture device 1500 can automatically increase the contrast after capturing the image.

[0148] A machine learning model 1520 can be trained using an image set with a machine learning algorithm. The machine learning algorithm can be a neural network (NN), such as a convolutional neural network (CNN), a time-delay neural network (TDNN), a deep feedforward neural network (DFFNN), a recurrent neural network (RNN), an autoencoder (AE), a variational AE (VAE), a denoising AE (DAE), a sparse AE (SAE), a Markov chain (MC), a perceptron, or some combination thereof. The machine learning algorithm can be a supervised learning algorithm, a deep learning algorithm, or some combination thereof.

[0149] In some cases, the image capture device 1500 may include two or more cameras (e.g., two image sensors 130 with two corresponding lenses), both cameras pointing at the same scene. In some cases, the image capture device 1500 applies the image capture settings and / or image processing settings generated using the machine learning model 1520 to only one of these cameras, while allowing the other camera to capture an image simultaneously (or within a threshold time of capturing another image) with the previously set image capture settings and / or image processing settings of the image capture device 105A. In some cases, the two images may then be displayed to the user of the image capture device 105A, and the user of the image capture device 105A may choose to retain only one of the two images and delete the other, or may choose to retain both images.

[0150] Figure 16 FIG. 4 is a flow chart showing operations 1600 for guiding the capture and / or processing of an image of an object based on a machine learning model 1520 trained using a training image set. Although operations 1600 refer to the image capture device 105A, operations 1600 may be performed by various devices, which may include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0151] In operation 1605, the device receives a first image of a scene captured by an image sensor of the image capture device. In operation 1610, the device identifies an object depicted in the first image using, for example, one or more of object detection, feature detection, face detection, feature recognition, object recognition, face recognition, saliency mapping, other image detection or recognition techniques discussed herein, or a combination thereof. In operation 1615, the device inputs the first image into the machine learning model 1520, which is trained using multiple images with the identified object. In operation 1620, the device uses the machine learning model 1520 to generate an alternative positioning of the object within the first image, an alternative positioning of the image capture device, one or more image capture settings, one or more image processing settings, or some combination thereof.

[0152] After operation 1620, there is operation 1625, operation 1630, operation 1635, or some combination thereof. In operation 1625, the device generates and outputs an indicator for positioning the image capture device. The indicator can be based on an alternative positioning of the image capture device determined during operation 1620. The indicator can be based on an alternative positioning of an object within the first image determined during operation 1620. For example, the indicator for positioning the image capture device can direct the repositioning of the image capture device such that the position of the object in the second image to be captured by the image sensor matches the alternative positioning generated using the machine learning model 1520. In some cases, outputting the indicator includes, for example, using alpha compositing to overlay an edited portion of the first image on a preview image displayed by the image capture device (as in the composite image 1530 of Figure 15 ). In some cases, outputting the indicator includes displaying a set of coordinates in the world to which the image capture device should be moved, a map highlighting the location corresponding to the set of coordinates, a map highlighting the path to the set of coordinates, a set of directions to the set of coordinates, or some combination thereof. In some cases, outputting the indicator includes displaying one or more arrows indicating the direction in which the image capture device is to be moved, tilted, and / or rotated, such as arrow 550, or the arrows in indicators 730 and 760.

[0153] In operation 1630, the device generates and outputs capture setting guidance for adjusting one or more attributes of the image capture device based on one or more image capture settings generated using the machine learning model 1520. In some cases, outputting the capture setting guidance includes automatically adjusting one or more attributes of the image capture device based on the one or more image capture settings before the image sensor captures the second image.

[0154] In operation 1635, the device generates and outputs processing setting guidance for processing the second image to be captured by the image sensor based on one or more image processing settings generated using the machine learning model 1520. In some cases, outputting the processing setting guidance includes automatically applying one or more image processing settings to the second image in response to receiving the second image from the image sensor.

[0155] In some cases, at least a subset of operation 1600 can be remotely executed by one or more web servers of a cloud service that performs image analysis (e.g., step 1610), trains a machine learning model, inputs the first image into the machine learning model (e.g., step 1615), generates and / or outputs indicators and / or guidance (e.g., operations 1620, 1625, 1630, and / or 1635), or some combination thereof.

[0156] Figure 17is a flowchart showing a method 1700 for guiding image capture. Although the image capture device 105A is referenced in the method 1700, the method 1700 can be performed by various devices, which can include the image capture device 105A, the image processing device 105B, the image capture and processing system 100, the image capture devices 500 / 700 / 900 / 1100 / 1300 / 1500, one or more web servers of a cloud service, the computing device 1800, or some combination thereof.

[0157] The method 1700 includes operation 1705. At operation 1705, the image capture device 105A receives a first image of a scene captured by the image sensor 130 of the image capture device 105A. The image capture device 105A can include the image sensor 130. The image capture device 105A can include one or more connectors coupled to the image sensor 130. The one or more connectors can couple the image sensor 130 to a portion of the image capture device 105A, such as the image processor 150 of the image capture device 105A. The image capture device 105A (or its processor) can receive the first image from the image sensor 130 via the one or more connectors.

[0158] At operation 1710, the image capture device 105A identifies an object depicted in the first image. For example, the image capture device 105A can use object detection, feature detection, face detection, feature recognition, object recognition, face recognition, saliency mapping, one or more other image detection or recognition techniques discussed herein, or a combination thereof to identify the object depicted in the first image.

[0159] At operation 1715, the image capture device 105A inputs the first image into a machine learning model 1520 that is trained using a plurality of training images with the identified object. The machine learning model 1520 can be based on any type of neural network (NN), machine learning algorithm, artificial intelligence algorithm, other algorithms discussed herein, or a combination thereof. For example, the machine learning model 1520 can be based on a convolutional neural network (CNN), a time-delay neural network (TDNN), a deep feedforward neural network (DFFNN), a recurrent neural network (RNN), an autoencoder (AE), a variational AE (VAE), a denoising AE (DAE), a sparse AE (SAE), a Markov chain (MC), a perceptron, or some combination thereof.

[0160] At operation 1720, the image capture device 105A uses the machine learning model 1520 to identify one or more changes to one or more attributes of the image capture device 105A that cause a visual difference between a first image and a second image that will be captured by the image sensor after the first image is captured. The one or more attributes can be one or more attributes associated with image capture. The one or more attributes can include the pose of the image capture device 105A. The pose of the image capture device 105A can refer to the position of the image capture device 105A, the orientation of the image capture device 105A (e.g., pitch, roll, and / or yaw), or both. The one or more attributes can include one or more image capture settings. The one or more attributes can include one or more image processing settings.

[0161] At operation 1725, the image capture device 105A outputs guidance indicating one or more changes that result in a visual difference before the image sensor 130 captures the second image. Outputting the guidance can include outputting a visual indicator, an auditory indicator, a vibration indicator, or a combination thereof. Outputting the guidance can include outputting one or more instructions that guide the user to move the image capture device 105A to achieve a visual difference that includes a change in perspective caused by the movement of the image capture device 105A.

[0162] Outputting the guidance can include outputting one or more instructions that guide the user to apply a specific image capture setting to the image capture device 105A such that the image capture setting is applied during the capture of the second image. Outputting the guidance can include automatically applying the image capture setting to the image capture device 105A such that the image capture setting is applied during the capture of the second image. Outputting the guidance can include outputting one or more instructions that guide the user to apply a specific image processing setting to the second image. Outputting the guidance can include automatically applying the image processing setting to the second image.

[0163] In some cases, the image capture device 105A receives the second image from the image sensor after outputting the guidance. In some cases, the image capture device 105A outputs the second image, such as by displaying the second image using a display (e.g., a display coupled to the image capture device 105A) and / or sending the second image to a receiving device using a transmitter. The receiving device can use a display (e.g., a display coupled to the receiving device) to display the second image.

[0164] Identifying one or more changes to one or more attributes of the image capture device 105A can include identifying a movement of the image capture device 105A from a first position to a second position. Outputting guidance at the image capture device 105A can include outputting an indication for moving the image capture device 105A from the first position to the second position. The second position can be identified using a machine learning model. The indication can include at least one of a visual indicator, an auditory indicator, and a vibration indicator. The indication can include one or more position coordinates of the second position, a map with an overlay marker highlighting the second position on the map, a map highlighting the path to the second position, a set of directions to the second position, or some combination thereof.

[0165] The indicator can identify movement information indicating a movement of the indicating device from a first position to a second position. The movement information of the indicator can identify a translation direction from the first position to the second position and / or a translation distance from the first position to the second position. For example, see Figure 5 the indicators 530, 540, and 550 depicted in. The movement information of the indicator can identify a rotation direction from the first position to the second position and / or a rotation angle from the first position to the second position. For example, see Figure 7A and 7B the indicators 730 and 760 depicted in. The rotation direction can include any rotation about any axis or combination of axes, such as roll, pitch, yaw, or another type of rotation direction. The rotation angle can be expressed in degrees, radians, graphical representation, or some combination thereof, and can indicate how far the image capture device 105A will rotate in the corresponding rotation direction. The movement information of the indicator can identify at least one of a translation direction of the movement, a translation distance of the movement, a rotation direction of the movement, a rotation angle of the movement, or a combination thereof.

[0166] The image capture device 105A can determine the direction the object is facing based on features of the object. More specifically, the device can determine the position of the object in a first image and the direction the object is facing in the first image. The image capture device 105A identifies a movement of the image capture device 105A from a first position to a second position based on the position of the object in the first image and the direction the object is facing in the first image. The image capture device 105A can determine the direction the object is facing in the first image based on the positioning of multiple features of the object relative to each other within the first image. If the object is a person, the multiple features of the object can include at least one of a person's ears, a person's cheeks, a person's eyes, a person's eyebrows, a person's nose, a person's mouth, a person's chin, a person's appendages, or a combination thereof. For example, as regarding Figure 4As shown and discussed, the left distance between two features on the left side of the object (e.g., the left eye and left cheek of the object) can be compared with the right distance between two features on the right side of the object (e.g., the right eye and right cheek of the object). If the image capture device 105A determines that the left distance is equal to the right distance, or the difference between the left distance and the right distance is below a threshold, then the image capture device 105A determines that the object is facing the image capture device 105A. If the image capture device 105A determines that the left distance exceeds the right distance by at least the threshold amount, then the image capture device 105A determines that the object is facing to the right. If the image capture device 105A determines that the right distance exceeds the left distance by at least the threshold amount, then the image capture device 105A determines that the object is facing to the left.

[0167] The image capture device 105A can determine the direction the object is facing based on the movement of the object. More specifically, the image capture device 105A can receive a third image captured by the image sensor 130 that depicts the object. The image capture device 105A determines the direction of movement of the object based on the position of the object in the first image and the position of the object in the third image. The image capture device 105A determines the direction the object is facing in the first image based on the direction of movement of the object. For example, if the third image is captured after the first image is captured, and the image capture device 105A determines that the object appears to move from the first image to the third image in a particular direction within the photographic scene, then the image capture device 105A can determine that the object is facing that direction. Similarly, if the third image is captured before the first image is captured, and the image capture device 105A determines that the object appears to move from the third image to the first image in a particular direction within the photographic scene, then the image capture device 105A can determine that the object is facing that direction.

[0168] The visual difference between the first image and the second image can include an adjustment in the amount of negative space adjacent to the object in the direction the object is facing. The adjustment can be to increase the amount of negative space adjacent to the object in the direction the object is facing. For example, Figure 3A image 310 can be considered an example of the first image, where there is very little negative space in front of the object 305. Figure 3B image 320 can be considered an example of the second image, where compared to Figure 3ACompared with the image 310, there is more negative space in front of the object 305. In this example, the movement of the image capture device 105A is from the first position where the first image 310 is captured to the second position where the second image 320 is captured. The movement of the device can be a translational movement to the left of the first position. The adjustment can also be to reduce the amount of negative space adjacent to the object in the direction the object is facing. For example, if in the exemplary first image, the object is depicted very close to the edge of the frame, the visual difference generated based on the guidance will cause the object to be depicted slightly farther from the edge of the frame (e.g., closer to the center of the frame border in some cases) in the exemplary second image.

[0169] In some examples, the first image depicts a horizontal line. As Figure 7A and 7B shown in the images 710 and 740, the horizontal line depicted in the first image may not be horizontal. The visual difference between the first image and the second image can make the horizontal line in the second image horizontal. For example, Figure 7A the image 710 can be an example of the first image depicting a non-horizontal horizontal line. The indication 730 indicates that the image capture device 105A is to be rotated counterclockwise by approximately 15 degrees along the rolling rotation direction. The execution of the rotation indicated by the indication 730 before capturing the second image generates a visual difference between the first image and the second image, which makes the horizontal line in the second image horizontal. Similarly, Figure 7B the image 740 can be an example of the first image depicting a non-horizontal horizontal line. The indication 760 indicates that the image capture device 105A is to be rotated clockwise by approximately 20 degrees along the rolling rotation direction. The execution of the rotation indicated by the indication 760 before capturing the second image generates a visual difference between the first image and the second image, which makes the horizontal line in the second image horizontal. The image capture device 105A can receive sensor measurement data from one or more pose sensors of the image capture device 105A. The sensor measurement data from one or more pose sensors can be referred to as pose sensor measurement data. The image capture device 105A determines the pose of the image capture device 105A based on the sensor measurement data. The pose of the image capture device 105A can include the position of the image capture device 105A, the orientation of the image capture device 105A (e.g., pitch, roll, and / or yaw), or a combination thereof. Then, the movement of the image capture device 105A can be identified based on the sensor measurement data and / or based on the pose. In some aspects, one or more pose sensors include at least one of an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit, a global navigation satellite system (GNSS) receiver, and an altimeter.

[0170] The image capture device 105A may receive a third image of a scene captured by a second image sensor of the image capture device 105A. The first image of the scene and the third image of the scene may be captured within a time window that spans a period of time. For example, the time window may be one or more picoseconds, one or more nanoseconds, one or more milliseconds, one or more seconds, or a combination thereof. The third image is captured by the second image sensor within a threshold time of the image sensor capturing the first image. The second image sensor has a wider field of view than the image sensor. In some examples, the second image sensor receives light through a second lens, while the image sensor receives light through a first lens. The first lens has a wider viewing angle than the second lens. For example, the first lens may be Figure 9 a wide-angle lens, while the second lens is Figure 9 a normal lens. The guidance may be based on the depiction in the third image of a portion of the scene not depicted in the first image. For example, the image capture device 105A may identify that an object is at least partially outside the border of the first image based on the depiction of the object in the third image. For example, in Figure 9 , the view visible to the normal lens 910 may be an example of the first image, and the view visible to the wide-angle lens 920 is an example of the third image. In this example, most of the object 940 is outside the frame of the first image, but most of the object 940 is within the frame of the third image. The movement of the image capture device 105A may be determined based on the third image (the view visible to the wide-angle lens 920) in order to bring the object 940 within the border of the second image, which in this example is the next image captured using the image sensor corresponding to the normal lens. In some cases, the guidance may be based on a fourth image to be captured by the second image sensor after capturing the third image.

[0171] In some examples, the image capture device 105A may receive a third image of a scene captured by a second image sensor of the device within a threshold time of the image sensor capturing the second image. In some cases, the device may present both the second image and the third image, for example by displaying them side by side or sequentially. The user may choose to retain one of the second image and the third image and discard the other, or simply mark one of them as the primary image at this time and the other as an alternative secondary image.

[0172] In some cases, the guidance may instruct the device to remain stationary (e.g., between the capture of a first image and the capture of a second image). For example, if the object is stationary and well-positioned from the perspective of image composition, the guidance may instruct the device to remain stationary. Or, if the object is moving and will be better positioned at a specific future time point (or time range) from the perspective of image composition, the guidance may instruct the device to remain stationary and capture a second image at the specific future time point (or time range) (or within a threshold time near the specific future time point (or time range)).

[0173] In some aspects, the multiple training images include training images in which a second object shares one or more similarities with the object. In some cases, the second object may be the object. For example, the first object and the second object may be the same person, the same monument, the same building, or the same object. In some cases, the second object may be an object of the same type as the object. For example, the object may be a person, and the second object may be a different person. The object may be a building, and the second object may be a different building with similar dimensions. The object may be an object, and the second object may be a different object with similar dimensions. One or more similarities shared between the second object and the object include: one or more saliency values associated with the second object are within a predetermined range of one or more saliency values associated with the object.

[0174] One or more changes to one or more attributes indicated by the guidance may be based on one or more settings of one or more attributes used to capture the training images. The one or more attributes may include the pose of the image capture device 105A, in which case one or more changes to one or more attributes indicated by the guidance may be based on the pose of the image capture device that captured the training images during (or within the same time window as) the capture of the training images. The one or more attributes may include one or more image capture settings of the image capture device 105A, in which case one or more changes to one or more attributes indicated by the guidance may be based on one or more image capture settings used by the image capture device that captured the training images during the capture of the training images. The one or more attributes may include one or more image processing settings of the image capture device 105A, in which case one or more changes to one or more attributes indicated by the guidance may be based on one or more image processing settings used by the image capture device that captured the training images and applied to the training images during the capture of the training images. The visual difference between the first image and the second image may include the second image being more similar to the training image than the first image is to the training image. Thus, applying the guidance to create a visual difference may result in the second image being visually more similar to the training image than the first image is to the training image.

[0175] The image capture device can identify that a second object is depicted in the training image using at least one of feature detection, object detection, face detection, feature recognition, object recognition, face recognition, saliency mapping, another detection or recognition algorithm discussed herein, or a combination thereof.

[0176] The indication to move the image capture device 105A from the first position to the second position can include one or more position coordinates of the second position, such as latitude, longitude, and / or altitude coordinates. The indication to move the image capture device 105A from the first position to the second position can include a map that depicts at least one of the first position, the second position, and the path between the first position and the second position. The indication to move the image capture device 105A from the first position to the second position can include a direction from the first position to the second position, such as a walking direction, a driving direction, and / or a public transportation direction.

[0177] One or more changes to one or more attributes of the image capture device 105A can include applying image capture settings before the image sensor captures the second image. The image capture settings can correspond to at least one of zoom, focus, exposure time, aperture size, ISO, depth of field, analog gain, or aperture. A machine learning model can be used to generate the image capture settings.

[0178] The image capture device 105A can determine the image capture settings based on one or more image capture settings used to capture one or more of the plurality of training images. The image capture device 105A can output guidance by outputting an indicator that identifies one or more changes to one or more attributes of the device corresponding to the application of the image capture settings. The image capture device 105A can output guidance by automatically applying the image capture settings and thus automatically applying one or more changes to one or more attributes of the device.

[0179] In some cases, the image capture device 105A can receive the second image captured by the image sensor. One or more changes to one or more attributes of the image capture device 105A can include applying image processing settings to the second image during capture, at the time of capture, within a threshold time after capture, or some combination thereof. The image processing settings can correspond to at least one of brightness, contrast, saturation, gamma, levels, histogram, color adjustment, blur, sharpness, levels, curves, filtering, or cropping. A machine learning model can be used to generate the image processing settings.

[0180] The image capture device 105A can determine an image processing setting based on one or more image processing settings for processing one or more of a plurality of training images. The image capture device 105A can output an indication identifying the image processing setting and / or instructions for applying the image processing setting to the second image by during the capture of the second image, at the time of capture of the second image, within a threshold time after the capture of the second image, or some combination thereof. The image capture device 105A can output instructions by automatically applying the image processing setting to the second image during the capture of the second image, at the time of capture of the second image, within a threshold time after the capture of the second image, or some combination thereof.

[0181] In some cases, at least a subset of operation 1700 can be performed remotely by one or more web servers of a cloud service that performs image analysis (e.g., step 1710), trains a machine learning model, inputs the first image into the machine learning model (e.g., step 1715), uses the machine learning model to identify changes to an attribute (e.g., step 1720), generates and / or outputs instructions (e.g., operation 1720), or some combination thereof.

[0182] In some examples, the processes described herein (e.g., including operations 600, 800, 1000, 1200, 1400, 1600, 1700, and / or other processes described herein) can be performed by a computing device or apparatus. In one example, processes 600, 800, 1000, 1200, 1400, 1600, and / or 1700 can be performed by Figure 1 the image capture device 105A. In another example, a process including operations 600, 800, 1000, 1200, 1400, 1600, and / or 1700 can be performed by Figure 1 the image processing device 105B. A process including operations 600, 800, 1000, 1200, 1400, 1600, and / or 1700 can also be performed by Figure 1 the image capture and processing system 100. A process including operations 600, 800, 1000, 1200, 1400, 1600, and / or 1700 can be performed by a device having Figure 18Performed by the computing device of the computing device architecture 1800 shown. The computing device can include any suitable device, such as a mobile device (e.g., a mobile phone), a wireless communication device, a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or a smart watch, or other wearable devices), a server computer, a computing device of an autonomous vehicle or an autonomous vehicle, a robotic device, a television, a camera, a camera device, and / or any other computing device having the resource capabilities to perform the processes described herein (including the processes including operations 600, 800, 1000, 1200, 1400, 1600, and / or 1700). In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface can be configured to communicate and / or receive data based on Internet Protocol (IP) or other types of data.

[0183] The components of the computing device can be implemented with circuitry. For example, the components can include electronic circuitry or other electronic hardware and / or can be implemented using electronic circuitry or other electronic hardware, which can include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or can include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0184] The processes including operations 600, 800, 1000, 1200, 1400, 1600, and / or 1700 are shown as logical flowcharts, and the operations represent a sequence of operations that can be implemented with hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media, which, when executed by one or more processors, perform the operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a specific function or implement a specific data type. The order of the described operations is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0185] In addition, the processes including operations 600, 800, 1000, 1200, 1400, 1600, 1700 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) collectively executed on one or more processors, may be executed by hardware, or a combination thereof. As described above, the code may be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program including a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0186] Figure 18 is a diagram showing an example of a system for implementing certain aspects of the present technology. Specifically, Figure 18 an example of a computing system 1800 is shown, which can be any computing device, such as a component of a composed internal computing system, a remote computing system, a camera, or any combination thereof, where the components of the system communicate with each other using a connection 1805. The connection 1805 can be a physical connection using a bus or a direct connection to the processor 1810 (such as in a chipset architecture). The connection 1805 can also be a virtual connection, a network connection, or a logical connection.

[0187] In some embodiments, the computing system 1800 is a distributed system, where the functions described in the present disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent many such components, each component performing some or all of the functions described for the component. In some embodiments, the components can be physical or virtual devices.

[0188] The example system 1800 includes at least one processing unit (CPU or processor) 1810 and a connection 1805 that couples various system components including a system memory 1815 (e.g., read-only memory (ROM) 1820 and random access memory (RAM) 1825) to the processor 1810. The computing system 1800 may include a cache 1812 of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1810.

[0189] The processor 1810 may include any general-purpose processor and hardware services or software services (such as services 1832, 1834, and 1836 stored in the storage device 1830), which are configured to control the processor 1810 and dedicated processors in which software instructions are incorporated into the actual processor design. The processor 1810 can essentially be a completely independent computing system, including multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor can be symmetric or asymmetric.

[0190] To enable user interaction, the computing system 1800 includes an input device 1845, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, mobile input, voice, and so on. The computing system 1800 may also include an output device 1835, which can be one or more of a plurality of output mechanisms. In some cases, the multimode system can enable the user to provide multiple types of input / output to communicate with the computing system 1800. The computing system 1800 may include a communication interface 1840, which can generally govern and manage user input and system output. The communication interface can use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communication, including using audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, wireless signal transmission, low energy (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or combinations thereof. The communication interface 1840 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1800 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States-based Global Positioning System (GPS), the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. There is no limitation on operating on any particular hardware arrangement, and thus the basic features herein can be readily replaced by improved hardware or firmware arrangements developed.

[0191] The storage device 1830 can be a non-volatile and / or non-transitory and / or computer-readable memory device and can be a hard disk or other types of computer-readable media that can store data accessible by a computer, such as magnetic tape, flash memory cards, solid state storage devices, digital versatile discs, cassette tapes, floppy disks, flexible disks, hard disks, magnetic tapes, magnetic strips / ribbons, any other magnetic storage media, flash memory, memristor memory, any other solid state memory, compact disc read only memory (CD-ROM) discs, rewritable compact disc (CD) discs, digital video disc (DVD) discs, Blu-ray disc (BDD) discs, holographic discs, another optical medium, secure digital (SD) cards, micro secure digital (microSD) cards, A card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or combinations thereof.

[0192] The storage device 1830 may include software services, servers, services, etc., which, when the processor 1810 executes the code defining such software, cause the system to perform functions. In some embodiments, the hardware services that perform a particular function may include software components stored in a computer-readable medium and associated with the necessary hardware components (such as the processor 1810, the connection 1805, the output device 1835, etc.) to perform the function.

[0193] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include a non-transitory medium in which data can be stored and excludes carrier waves and / or transient electronic signals propagated wirelessly or by a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs), or digital versatile discs (DVDs), flash memory, memories, or memory devices. Code and / or machine-executable instructions may be stored on the computer-readable medium, and the code and / or machine-executable instructions may represent a process, a function, a subroutine, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Any suitable means, including memory sharing, message passing, token passing, network transmission, etc., may be used to pass, forward, or transmit information, arguments, parameters, data, etc.

[0194] In some embodiments, the computer-readable storage device, medium, and memory may include cables or wireless signals containing bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0195] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, those of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For clarity, in some cases, the present technology may be presented as separate functional blocks including functional blocks, which include devices, device components, steps or routines in a method implemented in software or a combination of software or hardware. In addition to those shown in the figures and / or described herein, additional components may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components so as not to obscure the embodiments with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and technologies may be shown without unnecessary details to avoid obscuring the embodiments.

[0196] The various embodiments may be described above as processes or methods depicted as flowcharts, process schematics, data flow diagrams, structural diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. When the operations of a process are completed, the process is terminated, but there may be additional steps not included in the figures. A process may correspond to a method, function, program, subroutine, subprogram, and so on. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0197] The processes and methods according to the above examples may be implemented using computer-executable instructions stored in or otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a particular function or a group of functions. Portions of the computer resources used may be accessed via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the methods according to the examples include magnetic or optical disks, flash memory, USB devices equipped with non-volatile memory, networked storage devices, and the like.

[0198] Devices implementing the disclosed processes and methods can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. A processor can execute the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein can also be embodied in peripheral devices or plug-in cards. As a further example, such functions can also be implemented on a circuit board between different chips or in different processes executed within a single device.

[0199] Instructions, the media for conveying these instructions, the computing resources for executing them, and the other structures for supporting these computing resources are exemplary components for providing the functions described in this disclosure.

[0200] In the foregoing description, aspects of the present application have been described with reference to specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while illustrative embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts can be otherwise implemented and used differently, and the appended claims are intended to be construed to cover such variations, unless limited by the prior art. The various features and aspects of the present application described above can be used alone or in combination. In addition, embodiments can be used in any number of environments and applications other than those described herein, without departing from the broader spirit and scope of this specification. Thus, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, methods are described in a particular order. It should be understood that in alternative embodiments, these methods can be performed in a different order than that described.

[0201] Those of ordinary skill in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.

[0202] In cases where a component is described as “configured to” perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operation, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operation, or any combination thereof.

[0203] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component, and / or that communicates, either directly or indirectly, with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0204] Claim language or other language that recites "at least one" of a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language that recites "at least one of A and B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" means: A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one" of a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.

[0205] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0206] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device with multiple uses including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques may be at least partially realized by a computer-readable data storage medium including program code, the program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), FLASH memory, magnetic or optical data storage medium, and the like. Additionally or alternatively, these techniques may be at least partially realized through a computer-readable communication medium that carries or conveys instructions or data structures in a form that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0207] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described herein may be provided in a dedicated software module or hardware module configured for encoding and decoding, or incorporated in a combined video encoder-decoder (CODEC).

Claims

1. An apparatus for guiding image capture, the apparatus comprising: One or more memory units storing instructions; and one or more processors that execute the instructions, wherein the execution of the instructions by the one or more processors causes the one or more processors to: Receive a first image of a scene captured by an image sensor; Identify an object depicted in the first image; Receive attitude sensor measurement data from one or more attitude sensors; Determine the attitude of the apparatus based on the attitude sensor measurement data; Input the first image into a machine learning model that is trained using a plurality of training images with the identified object; Use the machine learning model to identify one or more changes to one or more attributes associated with image capture, the one or more changes causing a visual difference between the first image and a second image to be captured by the image sensor after capturing the first image, wherein the one or more attributes include the attitude of the apparatus; and Output guidance indicating the one or more changes that produce the visual difference before the image sensor captures the second image.

2. The device according to claim 1, wherein The apparatus is at least one of a mobile device, a wireless communication device, and a camera.

3. The apparatus according to claim 1, wherein The apparatus includes a display configured to at least display the second image.

4. The apparatus according to claim 1, further comprising: One or more connectors coupled to the image sensor, wherein, The one or more processors receive the first image from the image sensor through the one or more connectors.

5. The apparatus according to claim 1, further comprising: The image sensor.

6. The device according to claim 1, wherein, Identifying an object depicted in the first image includes performing at least one of feature detection, object detection, feature recognition, object recognition, and generation of a saliency map.

7. The apparatus according to claim 6, wherein The object detection includes face detection, and The object recognition includes face recognition.

8. The apparatus according to claim 1, wherein, The execution of the instructions by the one or more processors causes the one or more processors to further: After outputting the guidance, receive the second image from the image sensor; and Output the second image, wherein outputting the second image includes at least one of displaying the second image using a display and transmitting the second image using a transmitter.

9. The device according to claim 1, wherein Identifying one or more changes to one or more attributes associated with image capture includes identifying a movement of the apparatus from a first position to a second position, wherein outputting the guidance includes outputting an indicator for moving the apparatus from the first position to the second position, the indicator identifying at least one of the translation direction of the movement, the translation distance of the movement, the rotation direction of the movement, and the rotation angle of the movement.

10. The apparatus according to claim 9, wherein Use the machine learning model to identify the second position.

11. The device according to claim 9, wherein, The indicator includes at least one of a visual indicator, an auditory indicator, and a vibration indicator.

12. The apparatus according to claim 9, wherein, The indicator includes one or more position coordinates of the second position.

13. The apparatus according to claim 9, wherein, The visual difference between the first image and the second image levels a horizontal line in the second image, wherein the horizontal line is not horizontal as depicted in the first image.

14. The apparatus according to claim 9, wherein Identifying the movement of the device from the first position to the second position is based on the attitude of the device, where the attitude of the device includes at least one of the position and the orientation of the device.

15. The apparatus according to claim 9, wherein, The execution of the instructions by the one or more processors causes the one or more processors to also: Determine the position of the object in the first image; And Based on at least one of the relative positioning of two features of the object and the direction of movement of the object between the first image and a third image captured by the image sensor, determine the direction the object faces in the first image, where identifying the movement of the device from the first position to the second position is based on the position of the object in the first image and the direction the object faces in the first image, where the visual difference between the first image and the second image includes an adjustment of the amount of negative space adjacent to the object in the direction the object faces.

16. The device according to claim 1, wherein, The execution of the instructions by the one or more processors causes the one or more processors to also: Receive a third image of the scene captured by a second image sensor, where the first image of the scene and the third image of the scene are captured within a time window, where the second image sensor has a wider field of view than the image sensor, and where the guidance is based on the depiction in the third image of a portion of the scene that is not depicted in the first image.

17. The apparatus according to claim 1, wherein, The guidance instructs the device to remain stationary between the capture of the first image and the capture of the second image.

18. The device according to claim 1, wherein, The plurality of training images includes training images depicting at least one of the object and a second object sharing one or more similarities with the object, where the one or more changes to the one or more attributes indicated by the guidance are based on one or more settings of the one or more attributes used to capture the training images.

19. The apparatus according to claim 18, wherein, The one or more similarities shared between the second object and the object include: one or more significance values associated with the second object are within a predetermined range of one or more significance values associated with the object.

20. The device according to claim 18, wherein, The visual difference between the first image and the second image includes: the second image is more similar to the training image than the first image is to the training image.

21. The device according to claim 1, wherein The one or more changes to the one or more attributes associated with image capture include: applying image capture settings before the image sensor captures the second image, where the image capture settings correspond to at least one of zoom, focus, exposure time, aperture size, ISO, depth of field, analog gain, and aperture level.

22. The apparatus according to claim 21, wherein, Outputting the guidance includes outputting an indicator that identifies the one or more changes to the one or more attributes associated with image capture corresponding to applying the image capture settings.

23. The apparatus according to claim 21, wherein, Outputting the guidance includes automatically applying the one or more changes to the one or more attributes associated with image capture corresponding to applying the image capture settings.

24. The device according to claim 1, wherein, The execution of the instructions by the one or more processors further causes the one or more processors to: Receive the second image captured by the image sensor, wherein the one or more changes to the one or more attributes associated with image capture include applying image processing settings to the second image, and wherein the image processing settings correspond to at least one of brightness, contrast, saturation, gamma, level, histogram, color adjustment, blur, sharpness, curve, filtering, and cropping.

25. A method for guiding image capture, the method comprising: Receiving a first image of a scene captured by an image sensor of an image capture device; Identifying an object depicted in the first image; Receiving pose sensor measurement data from one or more pose sensors; Determining the pose of the image capture device based on the pose sensor measurement data; Inputting the first image into a machine learning model that is trained using a plurality of training images with the identified object; Using the machine learning model to identify one or more changes to one or more attributes of the image capture device, the one or more changes causing a visual difference between the first image and a second image that will be captured by the image sensor after capturing the first image, and wherein the one or more attributes include the pose of the image capture device; and Outputting guidance indicating the one or more changes that produce the visual difference before the image sensor captures the second image.

26. The method according to claim 25, wherein, The method is performed by the image capture device, wherein the image capture device is at least one of a mobile device, a wireless communication device, and a camera.

27. The method according to claim 25, wherein Identifying the object depicted in the first image includes performing at least one of feature detection, object detection, feature recognition, object recognition, and generation of a saliency map.

28. The method according to claim 27, wherein The object detection includes face detection, and The object recognition includes face recognition.

29. The method according to claim 25, wherein, Identifying one or more changes to one or more attributes associated with image capture includes identifying a movement of the image capture device from a first position to a second position, and wherein outputting the guidance includes outputting an indicator for moving the image capture device from the first position to the second position, the indicator identifying at least one of the translation direction of the movement, the translation distance of the movement, the rotation direction of the movement, and the rotation angle of the movement.

30. The method according to claim 29, wherein, Using the machine learning model to identify the second position.

31. The method according to claim 29, wherein, The visual difference between the first image and the second image levels a horizontal line in the second image, wherein the horizontal line is not horizontal as depicted in the first image.

32. The method according to claim 29, further comprising: Determining the position of the object in the first image; and Determine the direction in which the object faces in the first image based on at least one of the relative positioning of two features of the object and the direction of movement of the object movement between the first image and a third image captured by the image sensor, wherein, Identifying the movement of the image capture device from the first position to the second position is based on the position of the object in the first image and the direction the object faces in the first image, wherein the visual difference between the first image and the second image includes an adjustment of the amount of negative space adjacent to the object in the direction the object faces.

33. The method according to claim 25, further comprising: Receive a third image of the scene captured by the second image sensor, wherein, The first image of the scene and the third image of the scene are captured within a time window, wherein the second image sensor has a wider field of view than the image sensor, and wherein the guidance is based on the depiction in the third image of a portion of the scene that is not depicted in the first image.

34. The method according to claim 25, wherein, The plurality of training images includes training images depicting at least one of the object and a second object sharing one or more similarities with the object, wherein the one or more changes to the one or more attributes indicated by the guidance are based on one or more settings of the one or more attributes used to capture the training images, and wherein the visual difference between the first image and the second image includes the second image being more similar to the training images than the first image is to the training images.

35. The method according to claim 34, wherein The one or more similarities shared between the second object and the object include: one or more significance values associated with the second object being within a predetermined range of one or more significance values associated with the object.

36. The method according to claim 25, wherein The one or more changes to the one or more attributes associated with image capture include: applying image capture settings before the image sensor captures the second image, wherein the image capture settings correspond to at least one of zoom, focus, exposure time, aperture size, ISO, depth of field, analog gain, and aperture scale.

37. The method according to claim 25, further comprising: Receive the second image captured by the image sensor, wherein, The one or more changes to the one or more attributes associated with image capture include applying image processing settings to the second image, wherein the image processing settings correspond to at least one of brightness, contrast, saturation, gamma, level, histogram, color adjustment, blur, sharpness, curve, filtering, and cropping.

38. An apparatus for guiding image capture, the apparatus including components for performing the method according to any one of claims 25-37.

39. A computer-readable medium storing code for guiding image capture, wherein the code is executable by a processor to perform the method according to any one of claims 25-37.

Citation Information

Patent Citations

  • Method and system for providing recommendation information related to photography

    US20190174056A1