Extended reality image capture recommendations using avatars

Customized machine learning and XR avatars assist users in selecting camera settings and poses, simplifying the capture of professional-quality images by providing intuitive alignment with desired results.

WO2025165347A1PCT designated stage Publication Date: 2025-08-07GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/013538
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Users face challenges in selecting and configuring appropriate camera settings and poses to achieve desired imaging results due to the complexity of available settings and the need to measure environment parameters, often lacking technical knowledge.

Method used

Utilizing customized machine learning models and extended reality techniques, users are guided through scanning an environment to capture imaging conditions, and XR avatars are rendered to indicate optimal camera settings and poses, allowing intuitive alignment with desired image results.

Benefits of technology

Simplifies the process of capturing professional-quality images by providing intuitive guidance through XR avatars, ensuring compatibility between poses and camera settings, thereby improving image quality without requiring extensive expertise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024013538_07082025_PF_FP_ABST
    Figure US2024013538_07082025_PF_FP_ABST
Patent Text Reader

Abstract

In described techniques, a request for recommended camera settings and a recommended pose for capturing an image of a subject in an environment may be received. An image of the environment may be received. The recommended camera settings and the recommended pose may be generated, based on the image of the environment and on the request. At least one avatar may be rendered with the recommended pose. Correspondence between a photographer and the avatar, and / or between the subject and the avatar, may be determined.
Need to check novelty before this filing date? Find Prior Art

Description

EXTENDED REALITY IMAGE CAPTURERECOMMENDATIONS USING AVATARSBACKGROUND

[0001] Historically, professional photographers and videographers have been trained to obtain desired imaging results, using available camera hardware and accessories. For example, photographers may be trained to select and configure a suitable aperture, white balance, and other camera settings, and to arrange desired lighting (e.g., lamps and reflectors). More recently, such camera settings may be reproduced and / or emulated using personal or wearable devices.SUMMARY

[0002] Described techniques enable a photographer to capture an image of an environment, and then receive, based on the image, a recommendation for camera settings for capturing an image of a subject within the environment. The recommendation may be made with respect to a pose of the photographer and / or a subject, including, e.g., a recommended position for the photographer and / or subject. The pose(s) may be conveyed in an extended reality (XR) context by rendering a photographer avatar and / or a subject avatar at the recommended position(s). Then, the photographer and / or subject may move to align with a corresponding avatar. Accordingly, a desired image may be captured in an intuitive manner, without relying on image capture expertise on the part of the photographer.

[0003] In a general aspect, a computer program product is tangibly embodied on a non-transitory computer-readable storage medium and comprises instructions. When executed by at least one computing device (e.g., by at least one processor of the computing device), the instructions are configured to cause the at least one computing device to receive a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment. The instructions, when executed by the at least one computing device, may further cause the at least one computing device to receive an image of the environment. The instructions, when executed by the at least one computing device, may further cause the at least one computing device to render an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0004] In another general aspect, a wearable device includes at least one frame for positioning the wearable device on a body of a user, at least one camera, at least one display, at least one processor, and at least one memory storing instructions. When executed, the instructions cause the at least one processor to receive a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment. When executed, the instructions cause the at least one processor to receive an image of the environment. When executed, the instructions cause the at least one processor to render an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0005] In another general aspect, a computer-implemented method includes receiving a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment, receiving an image of the environment, and rendering an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0006] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] FIG. 1 is a block diagram of a system for extended reality image capture recommendations using avatars.

[0008] FIG. 2 is a flowchart illustrating example operations of the system of FIG. 1.

[0009] FIG. 3 illustrates a more detailed example implementation of the system of FIG. 1.

[0010] FIG. 4A illustrates an example photographer avatar generated using the systems of FIGS. 1 and 3.

[0011] FIG. 4B illustrates an example subject avatar generated using the systems of FIGS. 1 and 3.

[0012] FIG. 4C illustrates an example subject with lighting recommendations generated using the systems of FIGS. 1 and 3.

[0013] FIG. 5 illustrates an example subject avatar based on a reference image.

[0014] FIG. 6 illustrates examples of posing landmarks that may be used to evaluate avatar correspondence with a photographer or subject.

[0015] FIG. 7 is a third person view of a user in an ambient computing environment.

[0016] FIGS. 8A and 8B illustrate front and rear views of an example implementation of a pair of smartglasses.DETAILED DESCRIPTION

[0017] Described systems and techniques are directed to assisting photographers and videographers in capturing desired and / or professional pictures and / or video, across a variety of conditions, contexts, settings, and types of desired pictures / videos, with a minimum of technical knowledge.

[0018] At least one technical problem solved by described techniques involves the large quantity and potential combinations of available camera settings, the selection of which in a particular environment may be daunting for many users. As a result, it is difficult for users to measure or otherwise quantify relevant environment parameters (e.g., light levels), and to then select, configure, and implement camera settings to obtain a desired image result in a desired context. In many cases, users may also be unaware of available camera settings and their effects, as well.

[0019] At least one technical solution to the above-referenced technical problem(s) includes the use of customized machine learning (ML) models and extended reality (XR) techniques to assist users in obtaining desired imaging results. For example, a user wearing XR glasses may be instructed to scan an image setting or environment by making a 360° turn while a camera of the XR glasses is activated. Imaging conditions within the image setting / environment may thus be captured, e.g., for input to an ML model, including, for example, lighting characteristics and objects present within the setting.

[0020] Desired image results or characteristics may also be received as inputs for the ML model, e g., as part of a request for camera settings and pose information. For example, GPS coordinates and / or a representative or desired image type may be specified. Verbal or textual descriptions may also be provided or selected. The ML model may thus output camera settings to obtain a desired image, either for selection / entry by a user (e.g., photographer) and / or by automatically configuring the camera settings.

[0021] In addition, a rendering engine may be configured to render a photographer avatar for the photographer and a subject avatar for a subject of the image to be captured, where the avatars are rendered as augmented reality assets, e.g., by XR glasses. The avatar(s) may be rendered in a pose(s), including a location, position, or orientation, determined to yield the desired image and associated image effects. For example, the photographer avatar may be rendered at a certain location, in a certain position, and facing a certain direction,while the subject avatar may be rendered in a recommended pose and associated position. Then, the user (photographer) and the subject may align themselves with their respective avatars. Once alignment is completed and confirmed, the desired image may be easily captured using the determined settings.

[0022] FIG. 1 is a block diagram of a system for extended reality image capture recommendations using avatars. In the example of FIG. 1, a head-mounted device (HMD) 102 is illustrated as being worn by a user, primarily referred to herein as a photographer 104. The photographer 104 may desire to capture an image of a subject 106. The subject 106 may also be wearing a HMD 109. As referenced above, and described in detail, below, the HMD 102 (and optionally the HMD 109) may be configured to provide a desired image capture of the subject 106.

[0023] For example, the HMD 102 may render a photographer avatar 104a, which indicates a virtual representative of the photographer 104. The HMD 102 may also render a subject avatar 106a, which indicates a virtual representative of the subject 106. As described below, an avatar may refer to any virtual representation of a person or other object, such as a 3D representation of the person / object that appears visually similar to the person / object being represented and / or that is labeled or otherwise indicated as representing the corresponding person / object, and that is rendered within an XR environment.

[0024] As described in detail, below, the photographer avatar 104a and the subject avatar 106a may be rendered as virtual objects within a real-world or physical environment 107. For example, as illustrated and described in more detail, below, the environment 107 may include, or be defined with respect to, various real- orld physical objects. The environment 107 may be indoors or outdoors, and may include one or more light sources and associated lighting characteristics. The environment 107 may be defined with respect to GPS coordinates, a depth / distance defined with respect to the HMD 102, or using other techniques to specify absolute or relative location(s).

[0025] The photographer avatar 104a and the subject avatar 106a may thus be rendered in the context of the environment 107. For example, the photographer avatar 104a and the subject avatar 1 6a may be rendered at a certain position(s) within the environment 107, such as next to, on, or under an object within the environment 107. The photographer avatar 104a and the subject avatar 106a may be rendered at specified, corresponding GPS coordinates.

[0026] Each of the photographer avatar 104a and the subject avatar 106a may be rendered in a corresponding pose. As used herein, the term pose refers generally to a bodyposition, location, action, attitude, orientation, or state that may be assumed by a person for purposes of image capture. For example, a pose may refer to a position such as sitting, standing, or lying down. A pose may refer to a relative or absolute placement or direction of a person’s arms, legs, or head. For example, a pose may refer to an object on which a person is standing, or a direction that a person is facing. A pose may also refer to a physical location at which a person is positioned. Various other examples of poses are provided below, or would be apparent.

[0027] A pose may be static or dynamic. Accordingly, the photographer avatar 104a and / or the subject avatar 106a may be static or dynamic in order to indicate a recommended pose. For example, when capturing video, the subject avatar 106a may be rendered dynamically, e.g., moving in a recommended fashion, to indicate recommended movements of the subject 106. For example, the subject avatar 106a may be rendered as changing location or body position, or any other pose aspect.

[0028] Such poses may relate to camera settings to perform a desired image capture. For example, one pose (including location) may cause ambient light (e.g., sunlight) to be incident on the subject 106, while, in another pose (including location), the ambient light may be behind the subject 106. Such differences in lighting generally correspond to differences in camera settings for image capture. Similarly, if the environment 107 includes a lake by a forest, a selected pose may have either the lake or the forest in the background, which also may correspond to differences in desirable camera settings for the HMD 102.

[0029] By rendering the photographer avatar 104a and the subject avatar 106a in an XR context, it is possible to convey all such pose information in a fast, intuitive manner, while also ensuring compatibility' between resulting poses and corresponding camera settings, as described in detail, below. For example, a need for verbal descriptions or physical gestures to be used to convey correct locations and poses (e.g., between the photographer 104 and the subject 106) may be reduced, minimized or eliminated. Rather, the photographer 104 and the subject 106 may each align themselves with their corresponding avatars 104a, 106a.

[0030] For example, as illustrated with respect to a second subject 108 and a second subject avatar 108a, the second subject 108 may move to a location of the second subject avatar 108a, and the HMD 102 may be configured to provide confirmation of alignment correspondence, or a correspondence indication, once the second subject 108 is fully or sufficiently aligned with the second subject avatar 108a. For example, the HMD 102 may calculate a degree or percentage of correspondence between a current pose of the second subject 108 and the rendered pose of the second subject avatar 108a, and correspondence maybe confirmed when the correspondence is higher than a threshold correspondence percentage. A correspondence indication may be provided that informs the photographer 104 and / or the subject(s) 106, 108 that sufficient correspondence alignment has been achieved. For example, a verbal confirmation may be displayed using the HMD 102. In other examples, another visual indication may be provided, such as a change in color or brightness of the relevant avatar(s).

[0031] In some examples, the HMD 102 may be used to confirm alignment (or correspondence) of the photographer 104 with the photographer avatar 104a, alignment of the subject 106 with the subject avatar 106a, and alignment of the second subject 108 with the second subject avatar 108a. In other examples, as shown in FIG. 1, the subject 106 may utilize a separate HMD 109, in which case the HMD 109 may be used to assist the subject 106 in aligning with the subject avatar 106a. For example, the HMD 102 may be in communication with the HMD 109, using suitable communications modules, as described and illustrated in more detail, below, with respect to FIG. 7.

[0032] In other examples, the photographer 104 may utilize the HMD 102 to render the photographer avatar 104a, the subject avatar 106a, and the second subject avatar 108a. Then, the photographer 104 may utilize the HMD 102 to achieve alignment with the photographer avatar 104a, while providing verbal or other instructions to the subject 106 and the second subject 108 to align with corresponding, respective avatars 106a, 108a.

[0033] In still other examples, the subject 106 may use the HMD 109 to render the photographer avatar 104a, the subject avatar 106a, and the second subject avatar 108a, and to confirm corresponding alignments (in conjunction with providing verbal instructions) of the photographer 104, the subject 106, and the second subject 108 with corresponding, respective avatars 104a. 106a, 108a. In such examples, the photographer 104 may use the HMD 102 or other suitable camera device (e.g., a smartphone) to perform subsequent image capture.

[0034] Thus, in general, it may be appreciated that at least one device (e.g., at least one HMD) may be used to render at least one avatar that enables at least one corresponding photographer and / or subject to assume a designated pose(s). Further, it may be appreciated that the at least one HMD may be used to confirm alignment therewith of a corresponding photographer and / or subject.

[0035] Additionally, as referenced above and described in more detail below, at least one HMD, e.g., the HMD 102 or the HMD 109, may be used to provide at least one camera setting that corresponds to the rendered avatar(s) and associated pose(s). For example, such camera setting(s) may be provided to the photographer 104 using an appropriate userinterface (UI), and / or may use a suitable interface to automatically configure suggested camera settings. As also referenced above, and described in detail, below, such camera setting(s) may be generated using a suitably trained ML model(s).

[0036] The HMD 102 may include, for example, any type of smartglasses or goggles, such as the smartglasses shown and described below with respect to FIGS. 7, 8A, and 8B. The HMD may also represent any other type of eyewear, as well as any headset, headband, hat, helmet or any other headwear that may be configured to provide the functionalities described herein.

[0037] The HMD 102 may therefore include any virtual reality (VR), augmented reality (AR), mixed reality (MR), or immersive reality (IR) device, generally referred to herein as an extended reality (XR) device, through which the photographer 104 may look to view the subject 106 and the second subject 108.

[0038] The subject 106 may include any physical (real-world) item or object that the photographer 104 may wish to view. The subject 106 may include, for example, a person, animal, or item. For example, the photographer 104 may wish to capture an artistic photograph of a movable item, such as a flower in a vase, and the subject avatar 106a may be rendered as an image of a flower in a vase at a recommended location (and with recommended camera settings) for image capture.

[0039] Any of the avatars 104a, 106a, 108a may be rendered using many different techniques. For example, when the subject 106 is a person, the subject avatar 106a may be rendered in the shape of a generic human image, an outline, or as an avatar based on an existing or captured image of the subject 106. Similarly, when the subject 106 is an animal or item, such as the flower in a vase referenced above, the subject avatar 106a may be a generic version or outline thereof, or may be based on a captured image of the subject 106.

[0040] Additionally, although the example of FIG. 1 is provided primarily with respect to the HMD 102 (or the HMD 109), it will be appreciated that additional or alternative devices (e.g., wearable or handheld) may also be used. For example, a smartphone may be used to render the subject avatar 106a in an augmented reality setting. Additional examples of devices that may be used, individually or in communication with one another, are provided below, e g., with respect to FIG. 7.

[0041] As show n in the exploded view' of FIG. 1 , the HMD 102 (and / or the HMD 109, or other wearable or handheld device(s)) may include a processor 110 (which may represent one or more processors), as well as a memory 112 (which may represent one or more memories (e.g., non-transitory computer readable storage media)). The processor 110may represent a Central Processing Unit (CPU) and / or a Graphical Processing Unit (GPU). More detailed examples of the HMD 102 and various associated hardware / software resources are provided below, e.g., with respect to FIGS. 7, 8A, and 8B.

[0042] The HMD 102 may include, or have access to, various sensors that may be used to detect, infer, designate, or determine aspects of the environment 107, e.g., of the photographer 104, the subject 106 and / or the second subject 108. For example, in FIG. 1, the HMD 102 includes a camera 114 and a depth sensor 116. For example, the camera 114 may represent any one or more standard RGB cameras, while the depth sensor 116 may represent any passive or active depth sensor. For example, the depth sensor 116 may represent a time- of-flight (ToF) camera or LiDAR sensor. In some implementations, relevant depth data may be determined in whole or in part by the camera 114.

[0043] The depth sensor 116 may provide depth data that may be used to generate a depth map, as described in more detail, below, which captures information characterizing relative depths of detected or rendered objects with respect to a defined perspective or reference point. For example, such a depth map may be used to define or determine a depth of the subject 106 or the subject 108, depths of the rendered avatars 104a, 106a, 108a, and depths of various other elements or aspects of the environment 107.

[0044] A GPS sensor 118 may be configured to provide relevant location coordinates. For example, the GPS sensor 118 may provide current GPS coordinates of the HMD 102, which may be used by the depth sensor 116 as a reference point for establishing relative depths in a field of view of the camera 114.

[0045] A light sensor 120 represents one or more sensors designed to capture and characterize aspects of lighting within the environment 107. For example, the light sensor 120 may capture an illuminance, intensity, direction, or color of light within the environment 107.

[0046] Although illustrated separately in FIG. 1 for ease of explanation and discussion, it will be appreciated that various elements or aspects of the HMD 102 may be combined, or may have overlapping functionalities. For example, the camera 114 may include depth and / or light sensing capabilities. For example, light sensing, such as illuminance sensing, may be used to gauge depth. Other hardware and software that may be included in the HMD 102, or other wearable or handheld device, are described below, with respect to FIGS. 7, 8A, and 8B.

[0047] Further in FIG. 1. photo advisor 125 refers to one or more software modules that utilize the above-described features and functions of the HMD 102 to provide the typesof extended reality image capture recommendations using avatars described herein. Although illustrated as a single module that includes separate, individual components within the context of the HMD 102, it will be appreciated that the photo advisor 125 may be implemented using two or more devices (e.g., the HMD 102 and the HMD 109, or the HMD 102 and a handheld device of the photographer 104).

[0048] For example, the photo advisor 125 may include, or utilize, a depth map generator 126. The depth map generator 126 should be understood to represent and illustrate depth map software stored using the memory 112 and executed using the processor 110, and configured to process depth-related data captured by the camera 114, the depth sensor 116, and / or the light sensor 120. As described below, the depth map generator 126 may be capable of determining a per-pixel depth of each pixel (and associated object or aspect) in an image frame captured by the camera 114. For example, such depth information may be captured and stored as a perspective image containing a depth value instead of a color value in each pixel. A depth map may be generated and stored, e.g., as a 2D array, or as a depth mesh (e.g., a realtime triangulated mesh). Examples of depth map generation and storage are provided in more detail, below.

[0049] The photo advisor 125 may further include a recommendation model 128. For example, the recommendation model 128 may represent a machine learning (ML) model trained to input environment information characterizing the environment 107, along with image preferences obtained from the photographer 104. and to output camera settings and other instructions / information used to capture desired images.

[0050] For example, environment information may be captured via the camera 114. For example, upon request by the photographer 104 for activation of the photo advisor 125, the recommendation model 128 may instruct the photographer 104 (e.g., using an available UI of the HMD 102) to scan the environment 107. For example, instructions may be provided to direct a field of view (FOV) of the HMD in a direction in which an image of the subject 106 is to be captured. In other examples, instructions may be provided to direct the FOV of the HMD across 180° or 360° around the photographer 104, to capture information regarding the environment 107.

[0051] In a specific example, instructions may be provided to “please turn in a full circle with camera on,” whereupon the photographer 104 may proceed to capture the requested video. The resulting video file may be input to the recommendation model 128, which may extract model for the recommendation model 128 therefrom. For example, such model inputs may include light levels and variations, light sources, recognized objects, and adepth map provided by the depth map generator 126.

[0052] Image preferences of the photographer 104 may also be received as inputs at the recommendation model 128. For example, image preferences can be expressed as a style or type of image desired, which may be expressed using a reference image. For example, the reference image may include a personal image provided by the photographer 104, or may include a publicly available image. Image preferences may include an initial or reference image of the subject 106, either captured contemporaneously or obtained from storage. Image preferences may be personalized to the photographer 104, and may be stored and customized over time as the photographer 104 uses the photo advisor 125.

[0053] Thus, the recommendation model 128 inputs environment information and image preferences, and outputs camera settings for the camera 114 together with poses for the photographer 104, the subject 106, and / or the second subject 108. Detailed examples of camera settings and poses are provided below. In general, camera settings may include any known or future camera settings, such as, for example, focal length (e.g., 35mm vs. 50mm), white balance, aperture, or shutter speed. Camera settings may also include more customized settings or groups of settings that are particular to the HMD 102, the camera 114, and / or the photographer 104.

[0054] Camera settings may also include generated, augmented, simulated, or modified environment factors that the photo advisor 125 may provide to obtain desired image results. For example, as described below with respect to FIG. 4C. virtual versions of lamps and / or reflectors, along with size / positioning information therefore, may be generated. The camera 114 may include a flash (e.g., one or more LEDs), and the camera settings generated by the recommendation model 128 may include parameterizations of such light sources (e.g., brightness levels and / or timing / duration of generated light).

[0055] The recommendation model 128 may also generate two or more sets of camera settings, for selection therebetween by the photographer 104. The sets of camera settings may be presented directly or indirectly to the photographer 104. For example, camera settings may be provided to the photographer 104 as a list using a suitable UI of the camera 114 and the HMD 102, and the list may then be adjusted as desired by the photographer 104, or a specific set of camera settings may be selected.

[0056] In other examples, the recommendation model 128 may provide two or more example images illustrating potential image results associated with different, corresponding camera settings. In such examples, the photographer 104 may obtain desired camera settings by selecting one of the provided example images.

[0057] The recommendation model 128 may also output recommended poses for the photographer 104, the subject 106, and / or the second subject 108, in conjunction with generated camera settings. As referenced above, such poses should be understood broadly to include any body position and location within the environment 107. As with the camera settings, the recommendation model 128 may propose two or more options for each recommended avatar, for selection therebetween, e.g.. by the photographer 104 and / or the subject 106.

[0058] Thus, the recommendation model 128 inputs environment information and image preferences, and outputs camera settings and poses. In the example of FIG. 1, the generated camera settings may be automatically implemented through the use of a settings controller 130, which may interface with the camera 114 to configure and parameterize operations of the camera 114 when capturing the desired image. Additionally, or alternatively, camera settings may be implemented partially or completely in a manual fashion, by instructing the photographer 104 and receiving corresponding selections from the photographer 104.

[0059] Meanwhile, the poses may be implemented using a rendering engine 132 and an avatar evaluator 134. For example, as described above, pose recommendations from the recommendation model 128 may be communicated to the rendering engine 132, which may then generate and render one or more of the avatars 104a, 106a. 108a. The avatar evaluator 134 may then capture and calculate a correspondence between each physical person (or object) and each corresponding avatar, as each person moves to assume each corresponding pose. For example, as described and illustrated with respect to the second subject 108 and the second subject avatar 108a, the avatar evaluator 134 may determine a degree of correspondence or alignment between the second subject 108 and the second subject avatar 108a, relative to a pre-determined threshold.

[0060] In order to provide and parameterize the recommendation model 128, a remote serv er 136 may be used to implement a training engine 138, which is configured to train the recommendation model 128 using available training data 140. Generally speaking, the training data 140 includes input / output pairs for corresponding training images, where the inputs include environment information and image preferences, and the outputs include camera settings and poses, of the types described above (where camera settings may include generated or simulated conditions, such as virtual lighting conditions). In other w ords, the training data 140 may include a combination of both real-world data of the photographer 104 and the subject 106, aligned with poses and virtual 3D avatars, and labeled withcorresponding camera parameters or settings (e.g., aperture, focal length, white balancing, or 6D0F positions). The training data 140 may include, for each training pair, a final image result representing an ideal or desired image result for corresponding input settings.

[0061] Then, the recommendation model 128 may include a neural network, such as a convolutional neural network (CNN), Unet, and / or a Transformer-based neural network (e.g., a multi-head attention based Transformer network), or any neural network that can be trained to provide the type of end-to-end solution described herein for providing recommended camera settings and poses to obtain desired, corresponding image results. As illustrated in FIG. 1, the recommendation model 128 may be trained remotely at the server 136, and deployed locally at the HMD 102 or other suitable device.

[0062] Once deployed, the recommendation model 128 may be further fine-tuned, trained, or configured to adapt to specific preferences of the photographer 104 (e.g., may be customized or personalized for the photographer 104). For example, the recommendation model 128 may capture implicit or explicit feedback from the photographer 104 that characterizes a desirability or suitability of a captured image.

[0063] FIG. 2 is a flowchart illustrating example operations of the system of FIG. 1. In the example of FIG. 2, operations 202-210 are illustrated as separate, sequential operations. However, in various example implementations, the operations 202-210 may be implemented in a different order than illustrated, in an overlapping or parallel manner, and / or in a nested, iterative, looped, or branched fashion. Further, various operations or suboperations may be included, omitted, or substituted.

[0064] In the example of FIG. 2, a request for recommended camera settings and a recommended pose for capturing an image of a subject in an environment may be received (202). For example, the photo advisor 125 may receive a request from the photographer 104 to capture an image of the subject 106, and / or the second subject 108. For example, the camera 114 may provide a UI which presents a selection option for use of the photo advisor 125.

[0065] An image of the environment may be received (204). For example, as noted above, the photo advisor 125 may use a suitable Ul to instruct the photographer 104 to capture at least one image, e.g., a video, of the environment 107. As also noted, the instructions may include various levels of detail, such as an instruction to capture a 180° or 360° field of view.

[0066] For example, the HMD 102 and / or the camera 114 may provide a suitable UI with selectable options for invoking and using functionalities of the photo advisor 125. The atleast one image may be captured using the camera 114, or using a camera of a separate device that is in communication with the HMD 102. As noted above, the instructions on image capture may include varying levels of detail, which may depend on a nature of the request for, or invocation of, the photo advisor 125. For example, the instructions may include a physical direction or object to include in the image capture.

[0067] The recommended camera settings and the recommended pose may be generated, based on the image of the environment and on the request (206) For example, the recommendation model 128 of FIG. 1 may receive the image of the environment 107 to generate the recommended camera settings and the recommended pose. Although not illustrated explicitly in FIG. 2, other inputs to the recommendation model 128 may be used, as well. For example, the request may include a description of characteristics of the image of the subject to be captured. For example, visual, verbal, and / or textual requests or preferences of the photographer 104 may be received as inputs to the recommendation model 128. For example, such requests or preferences may be received contemporaneously with the image currently being captured, or may be pre-stored by the photographer 104, or by the photo advisor 125. For example, the photographer 104 may provide a sample or reference image, including a sample image of the subject 106 and / or the second subject 108, or may verbally or textually describe desired image results. Resulting camera settings may be presented for selection by the photographer 104, and / or the settings controller 130 may automatically configure corresponding settings of the camera 114.

[0068] An avatar may be rendered with the recommended pose (208). For example, the rendering engine 132 may render any or all of the avatars 104a, 106a, 108a. The recommended pose may include a geographical location, which may be determined on an absolute basis using GPS coordinates, and / or may be determined on a relative basis with respect to a depth map of the depth map generator 126. The recommended pose may further include a body position or other characteristic(s), which may be visually represented by the corresponding avatar in the relevant XR environment.

[0069] Alignment with the avatar may be confirmed (210). For example, the avatar evaluator 134 may evaluate a current, real-time degree of alignment or correspondence between each of the photographer 104 and the photographer avatar 104a, the subject 106 and the subject avatar 106a, and the second subject 108 and the second subject avatar 108a.

[0070] As illustrated in FIG. 1, alignment of the subject 106 with the subject avatar 106a may initially be at or close to 0%, and the avatar evaluator 134 may display and update the alignment percentage, e g., on a suitable UI of the camera 114 as displayed using theHMD 102, as the subject 106 moves closer to, and otherwise adopts a pose of, the subject avatar 106a. Also with respect to FIG. 1, the avatar evaluator 134 may display a higher alignment percentage with respect to the second subject 108 and the second subject avatar 108a, e.g., 85%, because the second subject 108 is more closely aligned with the recommended pose of the second subject avatar 108a.

[0071] In the described manner, implementing improved camera settings and an improved pose to obtain a desired image result in a desired context can be simplified. Thereby, improved image results are possible.

[0072] FIG. 3 illustrates a more detailed example implementation of the system of FIG. 1. At block 302, the recommendation model 128 is trained remotely, e.g., at the server 136 of FIG. 1, and using the training data 140 and the training engine 138.

[0073] For example, as referenced above, the training data 140 may include captured images or video labelled with associated model inputs of environment information and / or image preferences, and associated model outputs of camera settings (including lighting conditions) and poses (including physical positions) of photographers and imaged subjects.

[0074] Such training data 140 may be obtained from external sources, such as image repositories in which included images have at least some of the just-referenced labels. Additionally, or alternatively, the training data 140 may be collected by capturing reference images in a variety of settings and contexts, including capturing the associated model inputs and model outputs at a time of image capture of training images.

[0075] For example, a captured training image may be labelled with depth information characterizing the relevant image capture device at an origin location defined as (0, 0, 0), with an image subject and any defined objects or light sources at corresponding (x, y. z) coordinates, e.g., expressed in meters. Available, example techniques for capturing and expressing depth information are referenced above, and described in more detail, below.

[0076] For example, as described with respect to the HMD 102 of FIG. 1, an image capture device may be used to generate depth maps, e.g., portrait depth information may be acquired using a Time of Flight (ToF) camera(s), RGB camera(s), and / or trained deep learning models, in combination with depth map(s) obtained using video see-through frames.

[0077] For example, available inputs at the HMD 102 (or other suitable image capture device) may include a camera position (x, y, z) and rotation (quaternion x, y, z, w), a scene depth map (e.g., stored as a 512x512, 16-bit array), subject segmentation (e.g., stored as a binary mask. 512x512 array), and video see-through RGB images (e.g., as a 512x512x3 RGB image).

[0078] In some implementations, a 3D depth mesh may be constructed. Then, when the imaged subject in the training data 140 includes a person, a human texture (RGB values) from video see-through images may be obtained. Such information may train and enable the recommendation model 128 to define, e.g., the subject avatar 106a at a recommended depth, which may then be projected or rendered by the rendering engine 132 onto a calculated depth mesh (triangle mesh), using generated depth values for the current environment 107.

[0079] Depth may be understood as a distance from a reference point. Depth can be distance and direction from a reference point. Each pixel in an image can have a depth. Therefore, each pixel in the image can have an associated distance (and optionally a direction) from a reference point. The reference point can be based on a position of a device (e.g., a mobile device, a tablet, a camera, and / or the like), as referenced above, including a camera. The reference point can be based on a global position (using, e.g., a global positioning system (GPS) of the device.

[0080] While the device moves (e.g., changes a perspective) in the real-world space, a depth of view with respect to the device can be tracked in the real-world space using, for example, device sensors and an application programming interface (API) configured to track the depth from the device using the sensors.

[0081] As referenced above, depth information may include a depth map having a depth value for each pixel in an image. A depth map can be an image associated with a color (e.g.. RGB, YUV, and / or the like) image or frame of a video. A depth map can be an image associated with a black and white, greyscale, and the like image or frame of a video. The depth map can store a distance value for each pixel in the image or frame. The depth value can have an 8-bit representation with values between, for example, 0 and 255, where 255 (or 0) represents the closest depth value and 0 (or 255) represents the most distant depth value. In an example implementation, the depth map can be normalized. For example, the depth map can be converted to a range of [0, 1], Normalization can include converting the actual range of depth values. For example, if the depth map includes values between 43 and 203, 43 would be converted to 0. 203 would be converted to 1 and the remaining depth values would be converted to values between 0 and 1.

[0082] A depth map can be generated by a camera including a depth (D) sensor. This ty pe of camera is sometimes called an RGBD camera. A depth map can be generated using an algorithm implemented as a post-processing function of a camera while capturing an image or frame of a video. For example, an image can be generated from a plurality (e.g., 3. 6, 9 or more) images captured by a camera. The algorithm can be implemented in anapplication programming interface (API) associated with capturing real-world space images (e.g., for augmented reality applications).

[0083] Depth information can be represented as a 3-D point cloud. In other words, a point cloud can be a depth map in three dimensions. A point cloud can be a collection of 3D points (X, Y, Z) that represent the external surface of the scene and can contain color information.

[0084] The depth information can include depth layers each having a number (e.g.. an index number or z-index number) indicating a layer order. The depth information can be a layered depth image (LDI) having multiple ordered depths for each pixel in an image. Color information can be color (e.g., RGB, YUV, and / or the like) for each pixel in an image. A depth image can be an image where each pixel represents a distance from the camera location.

[0085] Described implementations may use features that can run immediately on a variety of devices by focusing on real time depth map processing, and / or may use techniques requiring a persistent model of the environment generated with 3D surface reconstruction. For example, real-time depth maps provided suitable APIs may be obtained using only a single moving RGB camera to estimate depth. A dedicated depth camera, such as time-of- flight (ToF) cameras can instantly provide depth maps without any initializing camera motion. Additionally, data such as a live camera feed, phone position and orientation, and camera parameters including focal length, intrinsic matrix, extrinsic matrix, and projection matrix for each frame may be used to establish a mapping between the physical world and virtual objects, such as the various avatars 104a, 106a, 108a described herein.

[0086] Depth data may be stored in a low-resolution depth buffer (e.g., 160 120), which is a perspective camera image that contains a depth value instead of color in each pixel. Different categories of data structures may be used.

[0087] For example, a depth array may store depth in a 2D array of a landscape image with 16-bit integers on a CPU. Using phone orientation and maximum sensing range, depth may be accessed from any screen point or texture coordinates of the camera image.

[0088] Alternatively, a depth mesh is a real-time triangulated mesh generated for each depth map on both CPU and GPU. In contrast to traditional surface reconstruction with persistent voxels or triangles, depth mesh has little memory and compute overhead and can be generated in real time. Depth texture may be copied to the GPU from the depth array for per-pixel depth use cases in each frame.

[0089] Depth data may be mapped to the camera image and the real-world geometry,including adapting to changes in camera orientation, or conversion of points between local and global coordinate frames. Depth data may leverage localized depth, surface depth, or dense depth.

[0090] For example, localized depth uses a depth array to operate on a small number of points directly on the CPU, and may be used to compute physical distance, e.g., to the subject 106. Surface depth leverages the CPU or compute shaders on the GPU to create and update depth meshes in real time, thus enabling collision, physics, texture decal, geometry- aware shadows, and other effects. Dense depth may be copied to a GPU texture and used for rendering depth-aware effects with GPU-accelerated bilinear filtering in screen space.

[0091] Pixels in a color camera image may have depth values mapped to them, which is useful for the real-time projected avatar image operations described herein. Screen-space to / from world-space conversion may include the following operations.

[0092] For example, for a screen point p = [x, y], a corresponding depth value may be obtained from a depth array Dwxh (e.g., w = 120, h = 160, as in the example above). Then, the screen point may be re-projected to a camera-space vertex vPusing the camera intrinsic matrix K , where vp = D(p) I<1[p,l],

[0093] Given the camera extrinsic matrix C = [R|t], which consists of a 3x3 rotation matrix R and a 3x1 translation vector t, global coordinates gPin the world space may be determined as: gP= C [vp,l]. Hence, virtual objects, such as the rendered avatars described herein, may be rendered in a singular coordinate system of the environment 107 (and included real world objects).

[0094] In a reverse process, 3D points may be projected using the camera’s projection matrix P. Then the projected depth values may be normalized and the depth projection may be scaled to the size of the depth map wxh.

[0095] In examples using real-time depth meshes, a mesh refers to, e.g., a set of triangle surfaces that are connected to form a continuous surface, which is the most common representation of a 3D shape. A depth mesh provides the 3D vertex position and the normal vector of surface points to compute world-space texture coordinates. Game and graphics engines are optimized for handling mesh data and provide simple ways to transform, shade, and to detect interactions between shapes.

[0096] In example implementations, variations of screen-space depth meshing techniques may be used that rely on a densely tessellated quad in which each vertex is displaced based on a re-projected depth value. No additional data transfer between CPU and GPU is required during render time, making this method very efficient.

[0097] Thus, the training data 140 may be captured, and labelled with depth labels, using various aspects of the types of detailed depth information just described, or other suitable depth information, including depth information for image subjects as well as included objects and other landmarks that might be useful in providing pose recommendations. Similarly, other pose labels may be captured for inclusion in the training data 140, which characterize, e.g., a body position or orientation. Techniques for capturing and characterizing these types of pose information are described below, with respect to FIG. 6.

[0098] Also similarly, training images may be labelled with camera settings (including lighting conditions) used to capture the training images. As noted above, such camera settings may include aperture, focal length, and white balance parameters.

[0099] The recommendation model 128 may thus be trained using the labelled image data of the training data 140. For example, as noted above, the recommendation model 128 may be implemented as a transformer-based deep learning neural network.

[0100] For example, the training data 140 may be characterized in terms of input embeddings, which may be fed to existing transformer-based architectures with multi-head attention mechanisms. For example, input data / labels may be split into n-grams encoded as tokens, which may be converted to the types of vector embeddings just referenced. The multihead attention mechanism may then be used to contextualize the embeddings in the context of other tokens / embeddings, enabling key tokens to be amplified. Output labels (e.g.. associated embeddings thereof) may be used to compare with model outputs, so that the recommendation model 128 may be further optimized and tuned until acceptable levels of error are achieved. The recommendation model 128 may thus be deployed for use, e.g., in the HMD 102, as shown in FIG. 1.

[0101] Then, at block 304, a request is received for a recommendation of capture options for a currently viewed subject. For example, the photographer 104 may use a camera option of the camera 114 of the HMD 102 to activate the photo advisor 125. Then, at block 306, an instruction is issued to the photographer 104 to look in 360° to capture multi-view image data of the environment 107 and the subject 106.

[0102] At block 308, a reference image may be received. The reference image may cause the recommendation model 128 (e.g., transformer-based neural network, as just described) to converge to a specific pose displayed within the reference image. For example, when subsequent camera settings are generated for the reference image, the photographer 104 may be able to observe the generated camera settings, and thereby gain expertise in selectingcamera setings in future, similar contexts.

[0103] The reference image may be available within an image repository and associated with a physical (e.g., GPS-based) location of the environment 107. In other examples, the reference image may be captured by the photographer 104. For example, the photographer 104 may capture the subject 106 at a specific position / location in the reference image, so that the recommendation model 128 may recommend camera setings (including lighting conditions) and body positions / orientations that may provide improved image capture of the subj ect 106 at the defined location.

[0104] As described herein, other image preferences, in addition or in the alternative to the reference image, may be received verbally or textually from the photographer 104, perhaps from a suggested list of available image preferences. For example, the photographer 104 may input a preference for a portrait or artistic image capture.

[0105] At block 310, with reference to FIG. 1 , photographer avatar 104a and subj ect avatar 106a are generated by the recommendation model 128, with corresponding poses dictated by the captured environment video and any associated image preferences, such as those expressed by the reference image. As described, poses may include physical positions in the sense of physical locations, as well as body positions in the sense of expressions or orientations (e.g., sitting / standing / lying down, or direction relative to light sources).

[0106] Additionally at block 310, corresponding camera settings may be generated by the recommendation model 128. including the types of camera parameters mentioned above, as well as lighting conditions. For example, lighting conditions may specify a white light source at a specific location within the environment 107, or three-point lighting, or a specific color temperature. A more specific example of recommended lighting conditions is provided below with respect to FIG. 4C.

[0107] In specific examples, the recommendation model 128 may generate two or more (e.g., n) sample images, each including the top-n recommendations for camera setings and poses. Then, the photographer 104 may select an option from those suggested.

[0108] At block 312. the photographer avatar 104a may thus be rendered. At the same time, at the block 314. the subject avatar 106a may be rendered. As shown in FIGS. 4A and 4B, the avatars 104a, 106a may be rendered in a real-world context represented by the environment 107.

[0109] The avatars 104a. 106a may be rendered with varying levels of accuracy with respect to the corresponding photographer 104 and the subject 106. For example, the avatars 104a, 106a may be generated as generic human (or other suitable) outlines. In otherexamples, the avatars 104a, 106a may be based on a previous image capture of the photographer 104 and / or the subject 106.

[0110] At block 316, recommended camera settings may be configured and recommended lighting may be rendered. For example, the settings controller 130 may automatically configure settings of the camera 114 to match the recommended parameters, and / or may display options for selecting recommended parameters. Similar comments apply to the recommended lighting parameters, so that suitable virtual light sources may be rendered by the rendering engine 132.

[0111] At block 318, alignment between the photographer 104 and / or the subject 106 and corresponding avatars 104a, 106a may be computed. For example, as described and illustrated with respect to the second subject 108 and the second subject avatar 108a. the photographer 104 (and the subject 106) may move to assume a pose(s) of the photographer avatar 104a and the subject avatar 106a, respectively. Techniques for determining current photographer / subject poses for purposes of computing such alignments, are described below with respect to FIG. 6

[0112] Finally in FIG. 3, at block 320, a difference in color temperature and area light alignment may be determined. That is, for example, the recommended lighting may be generated to obtain a desired result. As the photographer 104 and / or the subject 106 move to align with their respective avatars 104a, 106a, any resulting differences between the recommended / desired results and actual lighting conditions at the recommended poses may be determined and accommodated.

[0113] FIG. 4A illustrates an example photographer avatar 402 generated using the systems of FIGS. 1 and 3, and corresponding to the photographer avatar 104a of FIG. 1. As illustrated, the photographer avatar 402 may be rendered in a crouching position at a specified depth / location.

[0114] As further illustrated in FIG. 4A, additional object avatars may be generated, such as a virtual camera 404. In this way, the photographer 104 may align corresponding objects for optimizing image capture. For example, the photographer 104 may wear the HMD 1 2, but the camera 114 may be a separate device in communication with the HMD 102. as described below with respect to FIG. 7. In such cases, the photographer 104 may align with the photographer avatar 402 of FIG. 4A, w hile also aligning the separate camera with the virtual camera 404.

[0115] FIG. 4B illustrates an example subject avatar 406a for a subject 406, generated using the systems of FIGS. 1 and 3. For example, the photographer 104 may view the subject406 through the HMD 102, as described with respect to FIG. 1. The HMD 102 may render the subject avatar 406a in the illustrated pose, e.g., standing up as opposed to the current seated pose of the subject 406.

[0116] A pose alignment 408 may be rendered in a display of the HMD 102, informing the photographer 104 that the subject 406 has a 10% alignment with the subject avatar 406a. For example, the 10% alignment may be determined based on proximity of GPS coordinates of the subject 406 relative to the subject avatar 406a, and / or on similarity’ in relative depths as described above.

[0117] Thus, the photographer 104 may provide verbal instructions to stand and take a step forward to achieve alignment, which may then be reflected in the displayed pose alignment 408. In other examples, as illustrated in FIG. 1 with respect to the HMD 109, the subject 406 may be wearing a separate HMD or may otherwise have visibility or feedback with respect to aligning with the subject avatar 406a.

[0118] FIG. 4C illustrates an example subject 412 with lighting recommendations generated using the systems of FIGS. 1 and 3. For example, the subject 412 may have previously aligned with a subject avatar (not shown in FIG. 4C), as described above with respect to FIG. 4B.

[0119] Recommended camera settings may include lighting information that includes a light source 414 and a plurality of reflectors, including a reflector 416 an a reflector 418, which may themselves be rendered / projected as providing a desired type, direction, and intensity of light. For example, the light source 414 may be positioned above and behind the subject 412. The reflector 416 may represent a reflector that also projects white light from virtual LED light sources. The reflector 418 may have different, specified reflection properties, including reflecting / absorbing specified colors / frequencies of light. Thus, using the positioning and depth techniques described above, the light source 414 and the reflectors 416, 418 may be positioned in virtually any manner that is determined by the recommendation model 128 to achieve a desired effect(s), including projections of realistic shadows / shading and reflections.

[0120] FIG. 5 illustrates an example subject avatar 506a based on a reference image 504. For example, an environment 502 may represent a location defined with respect to local landmarks or features, and / or relevant GPS coordinates. The reference image 504 may be accessed from a repository of available images, and may be tagged as having been captured in the context of the environment 502.

[0121] A photographer, not pictured in FIG. 5, may wish to capture an image of asubject, also not shown in FIG. 5, that is similar to the reference image 504. The reference image 504 may be provided to the photo advisor 125, e.g., to the recommendation model 128, as described with respect to FIGS. 1 and 3. Accordingly, the subject avatar 506a may be rendered with a recommended pose within the environment 502, which may mimic or be based on a pose of an original subject within the reference image.

[0122] In other examples, the recommended pose may be based on expressed preferences of the photographer and / or subject, or on cunent conditions within the environment 502. For example, a time of day and / or lighting / weather conditions may be different within the environment 502, as compared to the reference image 504. Nonetheless, described techniques may ensure that a captured image is optimized with respect to such conditions.

[0123] FIG. 6 illustrates examples of posing landmarks 602, labelled 0-32, that may be used to evaluate avatar correspondence with a photographer or subject. For example, the rendering engine 132 may use the posing landmarks 602 to render each of the various described avatars with poses recommended by the recommendation model 128. Then, the avatar evaluator 134 may be configured to receive a current image frame, e.g., from the camera 114, of a photographer or subject for whom the avatar is generated.

[0124] For example, the posing landmarks 602 may be rendered and tracked in image coordinates and in 3-dimensional world coordinates. For example, such coordinates may be related to the coordinate systems described above with respect to depth determinations, e.g., operations of the depth map generator 126.

[0125] With further reference to FIG. 1, the second subject avatar 108a may be rendered and the posing landmarks 602 may be included in the rendering, and may be reflective of the recommended pose. The second subject 108 may move to align with the second subject avatar 108a.

[0126] The avatar evaluator 134 may track a current pose of the second subject 108 using a variety of techniques. For example, the avatar evaluator 134 may receive a current image frame from the camera 114 and may use image segmentation techniques, referenced above, to identify the second subject therein. In some examples, video analysis / recognition techniques may be used to determine a current pose of the second subject 108, for comparison to the recommended pose of the second subject avatar 108a.

[0127] In other examples, additional information may be used to determine the pose of the second subject 108. For example, as described below with respect to FIG. 7. the second subject 108 may be wearing or cartying various other wearable / personal devices, such as aHMD, smartwatch, earbuds, or a smartphone. Such devices may be in communication with the HMD 102, so that the avatar evaluator 134 may use received information to facilitate pose determination. For example, such devices may include one or more inertial measurement units (IMUs), which are designed to measure force, angular rate, orientation, and other movement / positioning aspects, using, e.g., an accelerometer, gyroscope, and / or magnetometer. The avatar evaluator 134 may use IMU signals to determine a current body position and orientation, and other pose information, for determining correspondence of the second subject avatar 108 with the second subject avatar 108a.

[0128] Thus, described techniques assist users in taking professional-quality photos, even without any prior photography experience. Users may save time and effort, as processes for choosing poses, parameters, and lighting are automated. Users improve photography skills by receiving feedback on captured photos, and described techniques may be used in both the physical world and the virtual world. Moreover, although primarily described with respect to photographs / images, it will be appreciated that described techniques may be used in the context of videography, as well.

[0129] FIG. 7 is a third person view of a user 702 (analogous to the photographer 104 or the subject 106 of FIG. 1) in an ambient environment 7000, with one or more external computing systems shown as additional resources 752 that are accessible to the user 702 via a network 7200. FIG. 7 illustrates numerous different wearable devices that are operable by the user 702 on one or more body parts of the user 702. including a first wearable device 750 in the form of glasses worn on the head of the user, a second wearable device 754 in the form of ear buds worn in one or both ears of the user 702, a third wearable device 756 in the form of a watch worn on the wrist of the user, and a computing device 706 held by the user 702. In FIG. 7. the computing device 706 is illustrated as a handheld computing device but may also be understood to represent any personal computing device, such as a tablet or personal computer.

[0130] In some examples, the first wearable device 750 is in the form of a pair of smart glasses including, for example, a display, one or more images sensors that can capture images of the ambient environment, audio input / output devices, user input capability, computing / processing capability and the like. Additional examples of the first wearable device 750 are provided below, with respect to FIGS. 8A and 8B.

[0131] In some examples, the second wearable device 754 is in the form of an ear worn computing device such as headphones, or earbuds, that can include audio input / output capability, an image sensor that can capture images of the ambient environment 7000,computing / processing capability, user input capability and the like. In some examples, the third wearable device 756 is in the form of a smart watch or smart band that includes, for example, a display, an image sensor that can capture images of the ambient environment, audio input / output capability, computing / processing capability, user input capability and the like. In some examples, the handheld computing device 706 can include a display, one or more image sensors that can capture images of the ambient environment, audio input / output capability, computing / processing capability, user input capability, and the like, such as in a smartphone. In some examples, the example wearable devices 750, 754, 756 and the example handheld computing device 706 can communicate with each other and / or with external computing system(s) 752 to exchange information, to receive and transmit input and / or output, and the like. The principles to be described herein may be applied to other types of wearable devices not specifically shown in FIG. 7 or described herein.

[0132] The user 702 may choose to use any one or more of the devices 706, 750, 754, or 756, perhaps in conjunction with the external resources 752, to implement any of the implementations described above with respect to FIGS. 1-6. For example, the user 702 may use an application executing on the device 706 and / or the smartglasses 750 to execute the photo advisor 125 of FIG. 1.

[0133] As referenced above, the device 706 may access the additional resources 752 to facilitate the various image capture recommendation techniques described herein, or related techniques. In some examples, the additional resources 752 may be partially or completely available locally on the device 706. In some examples, some of the additional resources 752 may be available locally on the device 706, and some of the additional resources 752 may be available to the device 706 via the network 7200. As shown, the additional resources 752 may include, for example, server computer systems, processors, databases, memory storage, and the like. In some examples, the processor(s) may include training engine(s), transcription engine(s), translation engine(s), rendering engine(s), and other such processors. In some examples, the additional resources may include ML model(s), such as the recommendation model 128 used by the photo advisor 125 of FIG. 1.

[0134] The device 706 may operate under the control of a control system 760. The device 706 can communicate with one or more external devices, either directly (via wired and / or wireless communication), or via the network 7200. In some examples, the one or more external devices may include various ones of the illustrated wearable computing devices 750, 754, 756. another mobile computing device similar to the device 706, and the like. In some implementations, the device 706 includes a communication module 762 to facilitate externalcommunication. In some implementations, the device 706 includes a sensing system 764 including various sensing system components. The sensing system components may include, for example, one or more image sensors 765, one or more position / orientation sensor(s) 764 (including for example, an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetometer and other such sensors), one or more audio sensors 766 that can detect audio input, one or more image sensors 767 that can detect visual input, one or more touch input sensors 768 that can detect touch inputs, and other such sensors. The device 706 can include more, or fewer, sensing devices and / or combinations of sensing devices. Various ones of the communications modules may be used to control brightness settings among devices described herein, and various sensors may be used individually or together to perform the ty pes of gaze, depth, and / or brightness detection described herein.

[0135] Captured still and / or moving images may be displayed by a display device of an output system 772, and / or transmitted externally via a communication module 762 and the network 7200, and / or stored in a memory7770 of the device 706. The device 706 may include one or more processor(s) 774. The processors 774 may include various modules or engines configured to perform various functions. In some examples, the processor(s) 774 may include, e g., training engine(s), transcription engine(s), translation engine(s), rendering engine(s), and other such processors. The processor(s) 774 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 774 can be semiconductor-based including semiconductor material that can perform digital logic. The memory 770 may' include any type of storage device or non-transitory computer-readable storage medium that stores information in a format that can be read and / or executed by the processor(s) 774. The memory 770 may store applications and modules that, when executed by the processor(s) 774, perform certain operations. In some examples, the applications and modules may be stored in an external storage device and loaded into the memory7770.

[0136] Although not shown separately in FIG. 7, it will be appreciated that the various resources of the computing device 706 may be implemented in whole or in part within one or more of various wearable devices, including the illustrated smartglasses 750. earbuds 754, and smartwatch 756, which may be in communication with one another to provide the various features and functions described herein.

[0137] An example head mounted wearable device 800 in the form of a pair of smart glasses is shown in FIGS. 8A and 8B, for purposes of discussion and illustration. The example head mounted wearable device 800 includes a frame 802 having rim portions 803surrounding glass portion, or lenses 807, and arm portions 830 coupled to a respective rim portion 803. In some examples, the lenses 807 may be corrective / prescription lenses. In some examples, the lenses 807 may be glass portions that do not necessarily incorporate corrective / prescription parameters. A bridge portion 809 may connect the rim portions 803 of the frame 802. In the example shown in FIGS. 8 A and 8B, the wearable device 800 is in the form of a pair of smart glasses, or augmented reality glasses, simply for purposes of discussion and illustration.

[0138] In some examples, the wearable device 800 includes a display device 804 that can output visual content, for example, at an output coupler providing a visual display area 805, so that the visual content is visible to the user. In the example shown in FIGS. 8A and 8B. the display device 804 is provided in one of the two arm portions 830, simply for purposes of discussion and illustration. Display devices 804 may be provided in each of the two arm portions 830 to provide for binocular output of content. In some examples, the display device 804 may be a see through near eye display. In some examples, the displaydevice 804 may be configured to project light from a display source onto a portion of teleprompter glass functioning as a beamsplitter seated at an angle (e.g., 30-45 degrees). The beamsplitter may allow for reflection and transmission values that allow the light from the display source to be partially reflected while the remaining light is transmitted through. Such an optic design may allow a user to see both physical items in the world, for example, through the lenses 807, next to content (for example, digital images, user interface elements, virtual content, and the like) output by the display device 804. In some implementations, waveguide optics may be used to depict content on the display device 804.

[0139] The example wearable device 800, in the form of smart glasses as shown in FIGS. 8A and 8B, includes one or more of an audio output device 806 (such as, for example, one or more speakers), an illumination device 808, a sensing system 810, a control system 812, at least one processor 814, and an outward facing image sensor 816 (for example, a camera). In some examples, the sensing system 810 may include various sensing devices and the control system 812 may include various control system devices including, for example, the at least one processor 814 operably coupled to the components of the control system 812. In some examples, the control system 812 may include a communication module providing for communication and exchange of information between the wearable device 800 and other external devices. In some examples, the head mounted wearable device 800 includes a gaze tracking device 815 to detect and track eye gaze direction and movement. Data captured by the gaze tracking device 815 may be processed to detect and track gaze direction andmovement as a user input. In the example shown in FIGS. 8A and 8B, the gaze tracking device 815 is provided in one of two arm portions 830, simply for purposes of discussion and illustration. In the example arrangement shown in FIGS. 8A and 8B, the gaze tracking device 815 is provided in the same arm portion 830 as the display device 804, so that user eye gaze can be tracked not only with respect to objects in the physical environment, but also with respect to the content output for display by the display device 804. In some examples, gaze tracking devices 815 may be provided in each of the two arm portions 830 to provide for gaze tracking of each of the two eyes of the user. In some examples, display devices 804 may be provided in each of the two arm portions 830 to provide for binocular display of visual content.

[0140] The wearable device 800 is illustrated as glasses, such as smartglasses, augmented reality (AR) glasses, or virtual reality (VR) glasses. More generally, the wearable device 800 may represent any head-mounted device (HMD), including, e.g., goggles, helmet, or headband. Even more generally, the wearable device 800 and the computing device 706 may represent any wearable device(s). handheld computing device(s), or combinations thereof.

[0141] Use of the wearable device 800, and similar wearable or handheld devices such as those shown in FIG. 7, enables useful and convenient use case scenarios of implementations of FIGS. 1-6. For example, as shown in FIG. 8B, the display area 805 may be used to display subject 106 and / or environment 107 in the example of FIG. 1, or any of the various avatars 104a, 106a, 108a. More generally, the display area 805 may be used to provide any of the functionality described with respect to FIGS. 1-6 that may be useful in operating the photo advisor 125.

[0142] In a first example, referred to herein as Example 1, a computer program product is tangibly embodied on a non-transitory computer-readable storage medium and comprises instructions that, when executed by at least one computing device, are configured to cause the at least one computing device to: receive a request for recommended camera settings and a recommended pose for capturing an image of a subject in an environment; receive an image of the environment; and render an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0143] Example 2 includes the computer program product of Example 1 , wherein the recommended pose includes a recommended location, and the avatar is rendered at therecommended location within the environment.

[0144] Example 3 includes the computer program product of Example 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the avatar as a photographer avatar of a photographer capturing the image.

[0145] Example 4 includes the computer program product of Example 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the avatar as a subject avatar of the subject.

[0146] Example 5 includes the computer program product of Example 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: determine a correspondence between the avatar and a photographer capturing the image; and output a correspondence indication characterizing the correspondence.

[0147] Example 6 includes the computer program product of Example 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: determine a correspondence between the avatar and the subject; and output a correspondence indication characterizing the correspondence.

[0148] Example 7 includes the computer program product of any of Examples 1 to 6, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: process the image of the environment with a machine learning (ML) model trained to generate the recommended camera settings and the recommended pose.

[0149] Example 8 includes the computer program product of any of Examples 1 to 7, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: configure a camera for captunng the image of the subject in the environment with the recommended camera settings; and cause the camera to capture the image of the subject in the environment.

[0150] Example 9 includes the computer program product of any of Examples 1 to 8, wherein the avatar is rendered in a mixed reality setting provided by a head-mounted device (HMD).

[0151] Example 10 includes the computer program product of any of Examples 1 to 9, wherein the request includes a description of characteristics of the image of the subject to be captured.

[0152] In an eleventh example, referred to herein as Example 11, head-mounted device (HMD) comprising: at least one frame for positioning the HMD on a face of a user; at least one camera; at least one display; at least one processor; and at least one memory, the at least one memory storing a set of instructions, which, when executed, cause the at least one processor to: receive a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment; receive an image of the environment; and render an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0153] Example 12 includes the HMD of Example 11 , wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: determine a correspondence between the avatar and a photographer capturing the image; and output a correspondence indication characterizing the correspondence.

[0154] Example 13 includes the HMD of Example 11, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: determine a correspondence between the avatar and the subject; and output a correspondence indication characterizing the correspondence.

[0155] Example 14 includes the HMD of any of Examples 11 to 13, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: process the image of the environment with a machine learning (ML) model trained to generate the recommended camera settings and the recommended pose.

[0156] Example 15 includes the HMD of Examples 11 to 14, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to:configure the camera for capturing the image of the subj ect in the environment with the recommended camera settings: and cause the camera to capture the image of the subject in the environment.

[0157] Example 16 includes the HMD of any of Examples 11 to 15, wherein the request includes a description of characteristics of the image of the subject to be captured.

[0158] In a seventeenth example, referred to herein as Example 17, a computer- implemented method comprising: receiving a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment; receiving an image of the environment; and rendering an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

[0159] Example 18 includes the method of Example 17, further comprising: determining a correspondence between the avatar and a photographer capturing the image; and outputting a correspondence indication characterizing the correspondence.

[0160] Example 19 includes the method of Example 17, further comprising: determining a correspondence between the avatar and the subject; and outputting a correspondence indication characterizing the correspondence.

[0161] Example 20 includes the method of any of Examples 17 to 19, further comprising: rendering the avatar in motion to visually demonstrate the recommended pose.

[0162] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0163] These computer programs (also known as modules, programs, softw are, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms '‘machine-readable medium” “computer-readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0164] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid cry stal display) monitor, or LED (light emitting diode)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory' feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0165] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication netw orks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.

[0166] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication netw ork. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0167] A number of implementations have been described. Nevertheless, it w ill be understood that various modifications may be made without departing from the spirit and scope of the description and claims.

[0168] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components maybe added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.

[0169] Further to the descriptions above, a user is provided with controls allowing the user to make an election as to both if and when systems, programs, devices, networks, or features described herein may enable collection of user information (e.g., information about a user's social network, social actions, or activities, profession, a user’s preferences, or a user's current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that user information is removed. For example, a user’s identity may be treated so that no user information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over w hat information is collected about the user, how that information is used, and what information is provided to the user.

[0170] The computer system (e.g.. computing device) may be configured to wirelessly communicate with a netw ork server over a network via a communication link established with the network server using any known wireless communications technologies and protocols including radio frequency (RF), microw ave frequency (MWF), and / or infrared frequency (IRF) wireless communications technologies and protocols adapted for communication over the network.

[0171] In accordance with aspects of the disclosure, implementations of various techniques described herein may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Implementations may be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine-readable storage device, a computer-readable medium, a tangible computer-readable medium), for processing by, or to control the operation of, data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). In some implementations, a tangible computer-readable storage medium may be configured to store instructions that when executed cause a processor to perform a process. A computer program, such as the computer program(s) described above, may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed tobe processed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.

[0172] Specific structural and functional details disclosed herein are merely representative for the purposes of describing example implementations. Example implementations, however, may be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.

[0173] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the implementations. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including." when used in this specification, specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0174] It will be understood that when an element is referred to as being "coupled," "connected," or "responsive" to, or "on," another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may also be present. In contrast, when an element is referred to as being "directly coupled," "directly connected," or "directly responsive" to, or "directly on," another element, there are no intervening elements present. As used herein the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0175] Spatially relative terms, such as "beneath," "below," "lower," "above," "upper," and the like, may be used herein for ease of description to describe one element or feature in relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the term "below" can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 130 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly.

[0176] Example implementations of the concepts are described herein with reference to cross-sectional illustrations that are schematic illustrations of idealized implementations (and intermediate structures) of example implementations. As such, variations from theshapes of the illustrations as a result, for example, of manufacturing techniques and / or tolerances, are to be expected. Thus, example implementations of the described concepts should not be construed as limited to the particular shapes of regions illustrated herein but are to include deviations in shapes that result, for example, from manufacturing. Accordingly, the regions illustrated in the figures are schematic in nature and their shapes are not intended to illustrate the actual shape of a region of a device and are not intended to limit the scope of example implementations.

[0177] It will be understood that although the terms "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Thus, a "first" element could be termed a "second" element without departing from the teachings of the present implementations.

[0178] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0179] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover such modifications and changes as fall within the scope of the implementations. It should be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.

Claims

WHAT IS CLAIMED IS:

1. A computer program product, the computer program product being tangibly embodied on anon-transitory computer-readable storage medium and comprising instructions that, when executed by at least one computing device, are configured to cause the at least one computing device to: receive a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment; receive an image of the environment; and render an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

2. The computer program product of claim 1 , wherein the recommended pose includes a recommended location, and the avatar is rendered at the recommended location within the environment.

3. The computer program product of claim 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the avatar as a photographer avatar of a photographer capturing the image.

4. The computer program product of claim 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the avatar as a subject avatar of the subject.

5. The computer program product of claim 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: determine a correspondence between the avatar and a photographer capturing the image; and output a correspondence indication characterizing the correspondence.

6. The computer program product of claim 1 or 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: determine a correspondence between the avatar and the subject; and output a correspondence indication characterizing the correspondence.

7. The computer program product of any of claims 1 to 6, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: process the image of the environment with a machine learning (ML) model trained to generate the recommended camera settings and the recommended pose.

8. The computer program product of any of claims 1 to 7, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: configure a camera for captunng the image of the subject in the environment with the recommended camera settings.

9. The computer program product of any of claims 1 to 8, wherein the avatar is rendered in a mixed reality setting provided by a head-mounted device (HMD).

10. The computer program product of any of claims 1 to 9, wherein the request includes a description of characteristics of the image of the subject to be captured.

11. A head-mounted device (HMD) comprising: at least one frame for positioning the HMD on a face of a user; at least one camera; at least one display; at least one processor; and at least one memory, the at least one memory storing a set of instructions, which, when executed, cause the at least one processor to: receive a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment; receive an image of the environment; andrender an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

12. The HMD of claim 11, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: determine a correspondence between the avatar and a photographer capturing the image; and output a correspondence indication characterizing the correspondence.

13. The HMD of claim 11, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: determine a correspondence between the avatar and the subject; and output a correspondence indication characterizing the correspondence.

14. The HMD of any of claims 11 to 13, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: process the image of the environment with a machine learning (ML) model trained to generate the recommended camera settings and the recommended pose.

15. The HMD of any of claims 11 to 14. wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: configure the at least one camera for capturing the image of the subject in the environment with the recommended camera settings.

16. The HMD of any of claims 11 to 15, wherein the request includes a description of characteristics of the image of the subject to be captured.

17. A computer-implemented method comprising: receiving a request for a recommended camera setting and a recommended pose for capturing an image of a subject in an environment; receiving an image of the environment; and rendering an avatar with the recommended pose and the recommended camera setting based on the image of the environment and on the request.

18. The method of claim 17, further comprising: determining a correspondence between the avatar and a photographer capturing the image; and outputting a correspondence indication characterizing the correspondence.

19. The method of claim 17, further comprising: determining a correspondence between the avatar and the subject; and outputting a correspondence indication characterizing the correspondence.

20. The method of any of claims 17 to 19, further comprising: rendering the avatar in motion to visually demonstrate the recommended pose.

Citation Information

Patent Citations

  • Shooting guiding method and electronic equipment

    CN114095662A

  • Imaging apparatus

    US20140368716A1

  • Methods for adjusting characteristics of photographs captured by electronic devices and related program products

    US20210227128A1

  • Automatic camera guidance and settings adjustment

    US20210368094A1

  • Photographing method and terminal

    US20230018557A1