Synthetic image generation based on image segmentation

By combining image data from multiple cameras and different zoom levels with a machine learning system, a synthetic image preview frame is generated, which solves the problems of insufficient image processing efficiency and quality in existing technologies and achieves efficient image synthesis and augmented reality effects.

CN121264058APending Publication Date: 2026-01-02QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480037808.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-17
Filing Date
2024-06-07
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively utilize image data from different zoom levels to generate high-quality composite images, especially during the preview stage before image capture, resulting in insufficient image processing efficiency and quality.

Method used

By using image data from multiple cameras and different zoom levels, combined with a machine learning system, image segmentation and synthesis are performed to generate a composite image preview frame. The image is then adjusted and combined before receiving the captured image input, achieving accurate segmentation and synthesis of the foreground and background.

Benefits of technology

It improves the efficiency and quality of image processing in the image capture and preview stage, provides better image compositing effects, and enhances the image augmentation and extended reality experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121264058A_ABST
    Figure CN121264058A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing image data are provided. First image data of a scene may be obtained at a first zoom level and include at least a foreground portion and a background portion. A user input may be received indicating an adjustment to increase or decrease a zoom level of the foreground relative to the background, the adjustment corresponding to a second zoom level greater than or less than the first zoom level. Second image data of the scene may be obtained based on adjusting and using a second zoom level to include an adjusted foreground portion associated with the second zoom level. The adjusted foreground portion may be segmented from second image data of the scene. A composite image may be generated based on combining the segmented foreground portion from the second image data with at least a portion of the first image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in its entirety to image processing. For example, aspects of this disclosure relate to systems and techniques for generating synthetic images based on image segmentation performed on multiple images and / or multiple camera image previews. Background Technology

[0002] Many devices and systems allow a scene to be captured by generating images (or frames) and / or video data (including multiple frames). For example, a camera or a device that includes a camera can capture one or more images of a scene (e.g., a still image of the scene, one or more frames of video of the scene, etc.). In some cases, one or more images may be processed to perform one or more functions, may be output for display, may be output for processing and / or consumption by other devices, and for other purposes.

[0003] A common type of processing performed on images is image segmentation, which involves dividing image frames and video frames into multiple parts. For example, image frames and video frames can be segmented into foreground and background parts. In some examples, semantic segmentation can segment image frames and video frames into one or more segmentation masks based on object classification. For example, one or more pixels of an image frame and / or video frame can be segmented into categories such as people, hair, skin, clothing, houses, bicycles, birds, backgrounds, etc. Segmented image frames and video frames can then be used for a variety of applications. Applications using image segmentation are numerous, including, for example, computer vision systems, image enhancement and / or augmentation, image background replacement, extended reality (XR) systems, augmented reality (AR) systems, autonomous driving operations, and other applications. Summary of the Invention

[0004] The following provides a brief overview in relation to one or more aspects disclosed herein. Therefore, the following summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Thus, the sole purpose of the following overview is to present, in a simplified form, certain concepts relating to one or more aspects involving the mechanisms disclosed herein, prior to the detailed description that follows.

[0005] Systems, methods, apparatuses, and computer-readable media for image processing are disclosed. According to at least one exemplary example, a method for processing image data is provided. The method includes: obtaining first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; receiving user input instructing an adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than or less than the first zoom level; obtaining second image data of the scene based on the adjustment and using the second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; generating a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and generating a composite image based on combining at least a portion of the segmented foreground portion from the second image data of the scene with the first image data of the scene.

[0006] In another exemplary example, an apparatus for processing image data is provided. The apparatus includes at least one memory and at least one processor coupled to the at least one memory and configured to: acquire first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; receive user input instructing an adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than or less than the first zoom level; acquire second image data of the scene based on the adjustment and using the second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; generate a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and generate a composite image based on combining the segmented foreground portion from the second image data of the scene with at least a portion of the first image data of the scene.

[0007] In another exemplary example, a non-transitory computer-readable storage medium includes instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: obtain first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; receive user input instructing an adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than or less than the first zoom level; obtain second image data of the scene based on the adjustment and using the second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; generate a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and generate a composite image based on combining the segmented foreground portion from the second image data of the scene with at least a portion of the first image data of the scene.

[0008] In another exemplary example, an apparatus for processing image data is provided. The apparatus includes: components for acquiring first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; components for receiving user input instructing an adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than or less than the first zoom level; components for acquiring second image data of the scene based on the adjustment and using the second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; components for generating a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and components for generating a composite image based on combining the segmented foreground portion from the second image data of the scene with at least a portion of the first image data of the scene.

[0009] The aspects generally include, as described substantially with reference to the accompanying drawings and description and illustrated as shown in the drawings and description, methods, apparatus, systems, computer program products, non-transitory computer-readable media, user equipment, user gear, wireless communication equipment, and / or processing systems.

[0010] Some aspects include a device having a processor configured to perform one or more operations of any of the methods outlined above. Further aspects include a processing device for use in the device, the processing device being configured with processor-executable instructions to perform operations of any of the methods outlined above. Further aspects include a non-transitory processor-readable storage medium storing processor-executable instructions thereon, the processor-executable instructions being configured to cause the processor of the device to perform operations of any of the methods outlined above. Further aspects include a device having components for performing functions of any of the methods outlined above.

[0011] The features and technical advantages of the examples according to this disclosure have been summarized rather extensively above in order to provide a better understanding of the detailed description that follows. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily used as the basis for modifying or designing other structures for achieving the same purpose as this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and manner of operation) and the associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each figure in the drawings is provided for illustrative and descriptive purposes and not as a definition of limitation of the claims. The foregoing, as well as other features and aspects, will become more apparent upon reference to the following specification, claims, and appended drawings.

[0012] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim. Attached Figure Description

[0013] The accompanying drawings are provided to aid in describing various aspects of this disclosure, and are provided for illustrative purposes only and not for limiting the scope of the aspects. For a more detailed understanding of the foregoing features of this disclosure, a more specific description of the invention, briefly summarized above, can be obtained by referring to the aspects, some of which are illustrated in the drawings. However, it should be noted that the drawings illustrate only certain typical aspects of this disclosure and are therefore not to be considered as limiting its scope, as other equally valid aspects are permissible in this description. The same reference numerals in different drawings may identify the same or similar elements.

[0014] FIG. 1 Examples of system-on-chip (SoC) implementations based on some examples are illustrated;

[0015] FIG. 2A Examples of fully connected neural networks based on some examples are shown;

[0016] FIG. 2B Examples of locally connected neural networks are shown based on some examples;

[0017] FIG. 3A Examples of synthetic images generated based on some examples are provided, which are obtained from a first image captured at a first zoom level and from a second image captured at a second zoom level less than the first zoom level.

[0018] FIG. 3B Examples of synthetic images generated based on some examples are provided, which are obtained from a first image captured at a first zoom level and from a second image captured at a second zoom level greater than the first zoom level.

[0019] FIG. 4 Examples of image capture preview UIs, including at least two zoom adjustment features for adjusting the composite image in the image capture preview user interface (UI), are provided below.

[0020] FIG. 5 This is an illustration of an example image processing system for performing synthetic image generation based on the segmentation of multiple camera image frames, based on some examples;

[0021] FIG. 6 Examples of person or object segmentation from the foreground of a telephoto image frame are shown below;

[0022] FIG. 7 Examples of removing people or objects from the foreground of non-telephoto image frames are shown below;

[0023] FIG. 8 Examples of use based on some examples are shown FIG. 6 Foreground information segmented by people and using based FIG. 7 An example of a synthetic image generated by removing background information from the image;

[0024] FIG. 9 Examples are provided for previewing images before receiving commands to capture them. FIG. 8 An example of an image capture preview UI for composite images;

[0025] FIG. 10 Examples of image capture preview UI and foreground segmentation maps are provided, based on some examples, for use in compositing and / or adjusting composite images with a foreground portion translated from a first position to a second position;

[0026] FIG. 11Another example of an image capture preview UI that can be used to synthesize and / or adjust a synthesized image, based on some examples, is shown, wherein the image capture preview UI includes UI features for adjusting the zoom level of the foreground portion and UI features for adjusting the pan or position of the foreground portion;

[0027] FIG. 12 Examples of image capture preview UIs, based on some examples, are provided for synthesizing and / or adjusting a preview of a synthesized image based on one or more user input gestures corresponding to zoom level adjustment and / or position adjustment of the foreground portion;

[0028] FIG. 13 Examples of image capture preview UIs are provided, which can be used to synthesize and / or adjust the preview of a synthesized image based on zoom levels that adjust the background information obtained from a first image and / or zoom levels that adjust the foreground information obtained from a second image.

[0029] FIG. 14 Examples are shown below, illustrating how to adjust the UI by segmenting the foreground portion from a first image and synthesizing the segmented foreground portion with background information obtained from a second image.

[0030] FIG. 15A to FIG. 15D Examples of shadow matting and / or shadow compositing are shown, based on some examples;

[0031] FIG. 16A to FIG. 16C Examples of images corresponding to image completion and / or image restoration are shown, based on some examples;

[0032] FIG. 17 This is a flowchart illustrating an example of a process for processing image and / or video data, based on some examples; and

[0033] FIG. 18 This is a block diagram illustrating an example of a computing system used to implement some of the aspects described in this article. Detailed Implementation

[0034] Certain aspects and examples of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and examples may be applied independently, and some may be applied in combination. Specific details are set forth in the following description for purposes of explanation to provide a thorough understanding of various aspects of this application. However, it will be apparent that various aspects and examples may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0035] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0036] Image semantic segmentation is the task of generating segmentation results for image data frames such as still images or photographs. Video semantic segmentation is a type of image segmentation that involves generating segmentation results for one or more frames of a video (e.g., generating segmentation results for all or part of the image frames in a video). Image semantic segmentation and video semantic segmentation can be collectively referred to as "image segmentation" or "image semantic segmentation". The segmentation result may include one or more segmentation masks that are generated to indicate one or more locations, regions, and / or pixels within an image data frame that belong to a given semantic segment (e.g., a specific object, object class, etc.). For example, as further explained below, each pixel of the segmentation mask may include a value indicating the specific semantic segment (e.g., a specific object, object class, etc.) to which each pixel belongs.

[0037] In some examples, features can be extracted from image frames and used to generate one or more segmentation masks for the image frames based on the extracted features. In other cases, machine learning can be used to generate segmentation masks based on the extracted features. For example, a convolutional neural network (CNN) can be trained to perform semantic image segmentation by feeding many training images into the CNN and providing a known output (or label) for each training image. The known output for each training image may include a ground truth segmentation mask corresponding to a given training image.

[0038] In some cases, image segmentation can be performed to segment an image frame into a segmentation mask based on an object classification scheme (e.g., all pixels in a given semantic segment belong to the same category or class). For example, one or more pixels in an image frame can be segmented into categories such as person, hair, skin, clothing, house, bicycle, bird, background, etc. In some examples, the segmentation mask may include a first value for pixels belonging to a first category, a second value for pixels belonging to a second category, and so on. The segmentation mask may also include one or more categories for a given pixel. For example, the "person" category may have subcategories such as "hair," "face," or "skin," such that a set of pixels can be included in a first semantic segment with the "face" category and may also be included in a second semantic segment with the "person" category.

[0039] Segmentation masks can be used to apply one or more processing operations to an image data frame. For example, a system can perform image enhancement and / or image strengthening on an image data frame based on a semantic segmentation mask generated for the frame. In one example, a system can process certain portions of a frame with a specific effect, but may not apply that effect to portions of the frame corresponding to a specific class indicated by the segmentation mask of the frame. Image enhancement and strengthening processes may include, but are not limited to, personal beautification, such as skin smoothing or blemish removal; background replacement or blurring; providing extended reality (XR) experiences (e.g., virtual reality (VR), augmented reality (AR), or mixed reality (MR) experiences), etc. Semantic segmentation masks can also be used to manipulate certain objects or fragments in an image data frame, for example, by using a semantic segmentation mask to identify pixels in the image frame associated with the object or portion to be manipulated. In one example, background objects in a frame may be artificially blurred to visually separate them from a focused or foreground object of interest (e.g., a face) identified by a segmentation mask for that frame (e.g., an artificial bokeh effect may be generated based on the segmentation mask and applied), where the object of interest is not blurred. In some cases, segmentation information can be used to add visual effects to image data frames.

[0040] Some image capture devices may include multiple cameras, lenses, image sensors, and / or imaging systems. For example, smartphones and other mobile computing devices may include a first camera corresponding to a first focal length (e.g., a first zoom level), a second camera corresponding to a second focal length (e.g., a second zoom level), and so on. For example, a smartphone may include a first camera corresponding to a 1x zoom level, a second camera corresponding to a 3x zoom level (or other telephoto zoom levels greater than 1x), and / or a third camera corresponding to a 0.5x zoom level (or other wide-angle zoom levels less than 1x).

[0041] Images captured by different cameras and / or using different focal lengths (e.g., different zoom levels) can depict different views of the same scene. For example, a foreground portion may appear relatively larger in an image captured using a 3x telephoto zoom level than in an image captured using a 0.5x wide-angle zoom level. Similarly, a background portion may appear relatively larger in an image captured using a 3x telephoto zoom level than in an image captured using a 0.5x wide-angle zoom level. The foreground portion may be associated with one or more objects, such as foreground subjects (e.g., people, etc.). The background portion may be associated with one or more objects, such as background subjects (e.g., buildings, etc.). In some cases, the background portion may be associated with one or more objects that can be referred to as "background objects." Based on the camera position, the sizes of the various foreground and background portions, and / or the framing of the image, the foreground and background portions (e.g., their corresponding objects and / or subjects) may be fully captured in an image frame when using a first zoom level or focal length, and may be partially captured in an image frame when using a second zoom level or focal length.

[0042] In some cases, different zoom levels may be better suited for capturing different parts of a scene. For example, when capturing an image of a person standing near a tall building, a 0.5x or wide-angle image of the scene may include a full view of the building (e.g., the background portion) but only a relatively small view of the person (e.g., the foreground portion). A 3x or telephoto image of the scene may include a close-up or full view of the person but only a partial view of the building. Systems and techniques are needed for automatically compositing synthetic images using image data captured with different cameras and / or different focal lengths (e.g., different zoom levels). Systems and techniques are also needed for compositing and adjusting synthetic images using an image capture preview user interface (UI) of an imaging device, wherein the first frame corresponding to the synthetic image (e.g., a preview frame of the synthetic image) is generated and displayed on the UI before receiving input for capturing the final synthetic image (e.g., before receiving a command for capturing the image).

[0043] This document describes systems, apparatus, processes (also referred to as methods), and computer-readable media (collectively, "systems and techniques") for generating synthetic images based on image segmentation performed using a machine learning system on multiple images and / or multiple camera image preview frames (e.g., a first frame obtained before receiving input for capturing images). For example, the systems and techniques can be used to generate synthetic images comprising one or more foreground portions and one or more foreground portions captured using different corresponding focal lengths (e.g., also referred to as different zoom levels) and different corresponding cameras. For example, multiple images can be obtained using different cameras and different zoom levels (e.g., different focal lengths, different digital zooms, or cropping levels, etc.). In some aspects, each of the multiple images can be obtained using a corresponding camera, including those on smartphones, mobile computing devices, imaging devices, etc.

[0044] As used herein, a preview of a composite image may also be referred to as a “preview frame” and / or a “composite image preview frame.” In some aspects, a captured composite image (e.g., a final composite image) may also be referred to as a “composite image frame” and / or a “captured composite image frame.” In an exemplary example, a first frame may be captured and / or output before receiving input for capturing the frame. Input for capturing the frame may be received after capturing and / or outputting the first frame. A captured frame may be captured and / or output based on the input for capturing the frame, wherein the captured frame follows the first frame and the input for capturing the frame. In some aspects, the first frame is a preview frame corresponding to an image, and the captured frame is a captured frame corresponding to the image and / or the preview frame.

[0045] For example, systems and techniques can be used to generate or capture synthetic image frames based on receiving one or more user inputs corresponding to the zoom level of a foreground and / or background object corresponding to a synthetic image preview frame (e.g., a first synthetic image frame output before the input for the capture frame). The synthetic image preview frame can correspond to the captured synthetic image frame, and vice versa. In some aspects, the synthetic image preview frame and the captured synthetic image frame can correspond to one or more synthetic images generated based on segmentation of multiple image and / or multiple camera image previews using the aforementioned machine learning system. In some examples, the synthetic image frame corresponds to a synthetic image of the preview frame. In some cases, generating a synthetic image includes generating a preview frame corresponding to the synthetic image (e.g., a first synthetic image frame output before the input for the capture frame). The captured synthetic image frame can be the same as or similar to the preview frame (e.g., the captured synthetic image frame can be a second frame generated and output after the first frame and after receiving the input for the capture frame). In some cases, the captured synthetic image frame may differ from the preview frame.

[0046] In some aspects, systems and techniques can be used to independently control the zoom level and positioning of various objects in a scene as the user performs compositing. For example, the zoom level and / or positioning of various objects (e.g., people, pets, flowers, etc.) can be controlled or adjusted independently relative to each other and / or relative to the background of the composite image. In an exemplary example, the zoom level and / or positioning of various objects in the foreground or background of the composite image scene can be adjusted as the user performs compositing based on one or more user inputs provided by the user corresponding to adjustments made before initiating or triggering the capture of the composite image (e.g., before receiving a command for capturing the image). For example, the zoom level and / or positioning of various objects can be adjusted based on one or more user inputs to an image capture preview user interface (UI), and a corresponding composite image preview can be generated and displayed instantly on the image capture preview UI (e.g., using corresponding image capture previews from multiple cameras on the device). In some examples, image data can be captured using multiple (or all) cameras of the device and can be stored and used (e.g., in a gallery or image viewing application running on the device) to manipulate the zoom level and / or positioning of individual foreground portions or foreground elements after image capture.

[0047] The system and techniques can utilize image data acquired from multiple cameras, including those in smartphones, mobile computing devices, or other image capture devices used to implement a segmented zoom UI. The segmented zoom UI can be identical to the image capture preview UI. Multiple cameras with different zoom levels or focal lengths can be used to generate composite images and composite image previews (e.g., the first frame output before receiving input for the capture frame). Using multiple cameras to acquire simultaneous image data at different zoom levels or focal lengths can provide improved image quality for objects that have been enlarged or reduced, compared to using interpolation or other image resizing techniques to enlarge or reduce the size of objects.

[0048] One or more image segmentation machine learning networks can be used to find foreground portions (e.g., including one or more objects associated with the foreground portion, such as a foreground subject, which may include people, etc.) and separate the corresponding pixel or image data of the foreground portion. In some cases, the foreground portion and its corresponding shadow can be segmented from a foreground image data source. The foreground image data source can be image preview data or image capture data associated with a camera having a first focal length. Background image data can be obtained as image preview data or image capture data associated with a camera having a second focal length different from the first focal length.

[0049] For example, image preview data can be real-time preview data associated with and / or generated by a camera and / or image capture device. In some cases, image preview or image preview data can also be referred to as real-time image preview or real-time image preview data, respectively. In some examples, real-time image preview data can be obtained before image capture input and / or receiving commands for capturing an image (e.g., selection of capture UI elements, activation of image capture trigger, etc.). For example, image preview data may correspond to the first frame output before receiving input for capturing a frame. In some cases, real-time image preview data can be output using the display of the image capture device and used by the user (e.g., the image capture device) to view and / or synthesize an image before capturing it. For example, when attempting to capture a scene, real-time image preview data can be provided on the display of the image capture device. In some examples, real-time image preview data can be obtained after raw sensor data (e.g., collected by the sensors of the image capture device) undergoes various preprocessing stages such as de-mosaicing, denoising, etc. In some aspects, systems and techniques may include one or more preprocessing stages in and / or implemented as a real-time image preview data pipeline. For example, FIG. 5 One or more (or all) of processing blocks 510, 520, 530, 540, 550 and / or 560 can be used to generate real-time image preview data, wherein zoom segmentation is applied to increase or decrease the zoom level of the foreground portion of the image relative to the zoom level of the background portion of the image. As used herein, “real-time image preview data” may also be referred to as “image preview data”.

[0050] After segmenting the foreground portion from the foreground image data, image matting and image compositing can be performed to generate a composite image that combines the segmented foreground portion with the background image data. In some aspects, image completion and / or image inpainting can be used to remove the foreground portion from the background image data. The segmented foreground portion can then be combined with the background image data from which the foreground portion has been removed. In some examples, image harmonization can be performed to generate a composite image with color temperature, white balance, etc., that match between portions of the image data obtained from the foreground image data and portions of the image data obtained from the background image data.

[0051] Various aspects of this disclosure will be described with reference to the figures.

[0052] FIG. 1An example implementation of a System-on-a-Chip (SOC) 100 is illustrated, which may include a Central Processing Unit (CPU) 102 or a multi-core CPU configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., a weighted neural network), latency, frequency bin information, task information, and other information may be stored in a memory block associated with a Neural Processing Unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a Graphics Processing Unit (GPU) 104, a memory block associated with a Digital Signal Processor (DSP) 106, a memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from the program memory associated with the CPU 102 or from memory block 118.

[0053] SOC 100 may also include additional processing blocks tailored for specific functions, such as GPU 104, DSP 106, connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and multimedia processor 112 capable of, for example, detecting and recognizing gestures. In one implementation, the NPU is implemented within CPU 102, DSP 106, and / or GPU 104. SOC 100 may also include sensor processor 114, image signal processor (ISP) 116, and / or navigation module 120, which may include a global positioning system.

[0054] The SOC 100 may be based on the ARM instruction set. In one aspect of this disclosure, the instructions loaded into the CPU 102 may include code for searching in a lookup table (LUT) for a stored multiplication result corresponding to the multiplicative product of the input values ​​and filter weights. The instructions loaded into the CPU 102 may also include code for disabling the multiplier during the multiplication operation of the multiplicative product when a lookup table hit for the multiplicative product is detected. Additionally, the instructions loaded into the CPU 102 may include code for storing the calculated multiplicative product of the input values ​​and filter weights when a lookup table miss for the multiplicative product is detected.

[0055] The SOC 100 and / or its components may be configured to perform image processing using machine learning techniques according to the aspects of this disclosure discussed herein. For example, the SOC 100 and / or its components may be configured to perform semantic image segmentation according to the aspects of this disclosure. In some cases, the aspects of this disclosure can improve the accuracy and efficiency of semantic image segmentation by using neural network architectures such as transformers and / or shift window transformers when determining one or more segmentation masks.

[0056] Generally, machine learning (ML) can be considered a subset of artificial intelligence (AI). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks through pattern dependency and inference without explicit instructions. An example of an ML system is a neural network (also known as an artificial neural network), which can include groups of interconnected artificial neurons (e.g., neuron models). Neural networks can be used in a variety of applications and / or devices, such as image and / or video decoding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and more.

[0057] Individual nodes in a neural network mimic biological neurons by taking input data and performing simple operations on that data. The results of these simple operations on the input data are selectively passed to other neurons. Weights are associated with each vector and node in the network, and these values ​​constrain how the input data relates to the output data. For example, the input data for each node can be multiplied by its corresponding weight, and the products can be summed. The sum of the products can be adjusted with optional biases, and activation functions can be applied to the results to produce the node's output signal or "output activation" (sometimes called a feature map or activation map). The weights can initially be determined by an iterative stream of training data through the network (e.g., weights are established during training phases where the network learns how to identify a particular category based on the characteristics of its typical input data).

[0058] There are different types of neural networks, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Generative Adversarial Networks (GANs), Multilayer Perceptron (MLP) neural networks, Transformer Neural Networks, and so on. For example, a Convolutional Neural Network (CNN) is a feedforward artificial neural network. A CNN can comprise an assembly of artificial neurons, each with its own receptive field (e.g., a localized region of the input space) and collectively tiling the input space. RNNs work on the principle of storing the output of a layer and feeding that output back to the input to help predict the outcome of that layer. A GAN is a generative neural network that learns patterns in the input data so that the neural network model can generate new synthetic outputs, which are plausible from the original dataset. A GAN can comprise two neural networks operating together: a generative neural network that generates the synthetic output and a discriminative neural network that evaluates the authenticity of the output. In an MLP neural network, data is fed into the input layer, and one or more hidden layers provide an abstraction level to the data. The output layer can then be predicted based on this abstract data.

[0059] Deep learning (DL) is an example of machine learning techniques and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. Using multiple layers in a deep neural network allows for the progressive extraction of higher-level features from a given raw data input. For example, the output of the first layer of artificial neurons becomes the input of the second layer, the output of the second layer becomes the input of the third layer, and so on. The layers located between the input and output of the entire deep neural network are often called hidden layers. Hidden layers learn (e.g., are trained) by transforming intermediate inputs from previous layers into slightly more abstract and complex representations that can be provided to subsequent layers until the final or desired representation is obtained as the final output of the deep neural network.

[0060] As noted above, neural networks are examples of machine learning systems and can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes in the input layer, processed by hidden nodes in one or more hidden layers, and output is produced by output nodes in the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which can include artificial neurons (or nodes). Feature maps can include filters, kernels, etc. Nodes can include one or more weights used to indicate the importance of nodes in one or more layers. In some cases, deep learning networks may have a series of many hidden layers, where earlier layers are used to determine simple and low-level properties of the input, and later layers build a hierarchy of more complex and abstract properties.

[0061] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.

[0062] Deep learning architectures perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to identify features such as wheels, windshields, and others. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.

[0063] Neural networks can be designed to have multiple connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the recognition of higher-level concepts can aid in discerning specific lower-level features of the input.

[0064] The connections between layers in a neural network can be fully connected or locally connected. FIG. 2A An example of a fully connected neural network 202 is illustrated. In the fully connected neural network 202, neurons in the first layer can transmit their outputs to each neuron in the second layer, so that each neuron in the second layer will receive inputs from each neuron in the first layer. FIG. 2BAn example of a locally connected neural network 204 is illustrated. In the locally connected neural network 204, neurons in the first layer can connect to a limited number of neurons in the second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in the layer will have the same or similar connectivity pattern, but the connection strength can have different values ​​(e.g., 210, 212, 214, and 216). The connectivity pattern of locally connected networks can generate spatially dissimilar receptive fields in higher layers because higher-layer neurons in a given region can receive inputs that are trained to the properties of a restricted portion of the network's total input.

[0065] As previously mentioned, the systems and techniques described in this paper can be used based on machine learning systems (e.g., those referenced above). FIG. 1 , FIG. 2A and FIG. 2B The systems described herein perform image segmentation of multiple images and / or multiple camera image previews (e.g., a first frame output before receiving input for a capture frame) to generate a composite image. Generating a composite image may include generating a composite image preview frame (e.g., a first composite image frame output before receiving input for a capture frame), generating a composite image frame (e.g., a captured composite image frame generated after the first frame and / or after receiving input for a capture frame), and / or a combination of both. For example, generating a composite image may include generating one or more composite image preview frames and generating (e.g., capturing) a composite image frame based on one or more preview frames. In some aspects, multiple images may be acquired using different cameras and different zoom levels (e.g., different focal lengths, different digital zooms, or cropping levels, etc.). In some aspects, each of the multiple images may be acquired using a corresponding camera, including those on a smartphone, mobile computing device, imaging device, etc.

[0066] FIG. 3AAn example of synthetic image generation 300 is illustrated, wherein synthetic image 330 is generated based on foreground information obtained from a first image 320 captured at a first zoom level and background information obtained from a second image 310 captured at a second zoom level less than the first zoom level. For example, the example synthetic image generation 300 may correspond to a scene where a person (e.g., a foreground subject or other object associated with a foreground portion) is standing very close to a tall building (e.g., a background object or other object associated with a background portion). In this scene, image 310 is a wide-angle image obtained using a 1x zoom level, and image 320 is a telephoto image obtained using a 3x zoom level. The wide-angle image 310 can be obtained using a first camera, and the telephoto image 320 can be obtained using a second camera. The systems and techniques described herein can be used to generate synthetic image 330 to include foreground information of a person corresponding to the telephoto image 320 and background information of a building corresponding to the wide-angle image 310.

[0067] In some examples, image 310 may be a first image preview frame having a first zoom level (e.g., a first frame output before receiving input for capturing a frame), and image 320 may be a second image preview frame having a second zoom level different from the first zoom level (e.g., a second frame output before receiving input for capturing a frame). For example, image 310 may be a wide-angle image preview frame, and image 320 may be a telephoto image preview frame. In some aspects, composite image 330 may be a composite image preview frame that includes at least a portion of the first image preview frame 310 at the first zoom level and at least a portion of the second image preview frame 320 at the second zoom level.

[0068] In some examples, the composite image 330 is a composite image preview frame (e.g., image preview data) output to and / or displayed as image capture user interface (UI). For example, the composite image 330 may be a composite image preview frame displayed using the image capture UI of an image capture device or other electronic device (e.g., a smartphone, etc.). The composite image preview frame may be generated and / or output before receiving input for capturing the frame. In some cases, the composite image 330 may be a real-time preview (e.g., image preview data) output before receiving a command for capturing the image (e.g., receiving an image capture trigger, etc.). For example, the composite image 330 may be image preview data (e.g., including multiple composite image preview frames) output before receiving a command for capturing the image, where the command corresponds to user input to a shutter button on the image capture UI. The captured composite image frame may be generated (e.g., captured) based on receiving user input indicating a command for capturing the image (e.g., user input to a shutter button on the image capture UI). The captured composite image frame may be generated based on one or more of the corresponding composite image preview frames.

[0069] As previously noted, when attempting to capture a scene, image preview data can be provided on the display of the image capture device. In some examples, the image preview data can be obtained after the raw sensor data (e.g., collected by the sensors of the image capture device) undergoes various preprocessing stages such as demosaicing, denoising, etc. In some aspects, the systems and techniques described herein can include one or more preprocessing stages in and / or implemented as an image preview data pipeline. For example, FIG. 5 One or more (or all) of the processing blocks 510, 520, 530, 540, 550 and / or 560 can be used to generate image preview data, wherein zoom segmentation is applied to increase or decrease the zoom level of the foreground portion of the image relative to the zoom level of the background portion of the image.

[0070] In an exemplary example, the image preview data of the composite image (e.g., the composite image preview frame) can be changed and updated based on one or more user inputs for adjusting the composite image being previewed before capture. For example, the image preview data of the composite image can be changed based on user input for increasing or decreasing the zoom level of the foreground portion or foreground object relative to the main zoom level of the image (e.g., the zoom level of the background portion, etc.). The image preview data of the composite image can be changed based on user input for increasing or decreasing the zoom level of one or more background objects relative to the main zoom level of the image and / or relative to the adjusted zoom level of the foreground object. The image preview data of the composite image can also be changed based on user input for panning or repositioning one or more objects within the composite image preview frame (e.g., foreground objects and / or background objects at the main zoom level and / or at the corresponding adjusted zoom level).

[0071] When the image preview data of the composite image meets the user's expectations, the user can select a UI element corresponding to the command used to capture the image, and can obtain image capture data from one or more cameras of the user's image capture device for the captured composite image frame. Image capture data from each camera can be stored in the device storage of the user's image capture device. In some aspects, the image preview data of the composite frame can be similar to the captured composite image frame only before receiving the command for capturing the image. The captured composite image frame can be an image of higher visual quality than the composite image preview frame because the composite image preview frame is generated on-the-fly for output to the display of the user's image capture device (e.g., to guide or assist in image compositing and subsequent capture of the desired composite image), while the captured composite image frame can be processed without immediate latency constraints or requirements. In some aspects, the captured composite image frame can undergo processing pipelines used to generate the composite image preview frame (e.g., ...). FIG. 5 The image processing system 500 uses the same or similar processing pipeline. In some examples, a larger machine learning model can be used to generate the captured synthetic image frame, which includes a larger number of parameters, layers, etc., than those machine learning models used to generate the synthetic image preview frame. For example, the segmented foreground object determined from the captured image data for generating the synthetic image capture can be more accurate than the same segmented foreground object represented in the synthetic image preview frame. In another example, the inpainting region determined when generating the synthetic image capture can have fewer artifacts than the same inpainting region represented in and generated for the synthetic image preview frame.

[0072] In some cases, the difference in visual image quality between a synthetic image preview and a synthetic image capture can be based on the corresponding machine learning models used to perform various operations in the synthetic image generation processing pipeline (e.g., including...). FIG. 5 (The corresponding machine learning model and / or engine in the image processing system 500). For example, a corresponding machine learning model and / or engine can be used. FIG. 5 A relatively lightweight implementation of one or more of the image processing operations 510, 520, 530, 540, 550, and / or 560 is used to generate synthetic image preview frames. Corresponding to FIG. 5 A relatively strong or heavyweight implementation of one or more of the image processing operations 510, 520, 530, 540, 550, and / or 560 is used to generate synthetic image preview frames. In some aspects, it is possible to target FIG. 5 One or more (or all) of the image processing operations 510, 520, 530, 540, 550 and / or 560 use the same machine learning model to generate synthetic image preview frames and synthetic image capture frames.

[0073] In some cases, the composite image preview frame can be based on an image segmentation engine (e.g., FIG. 5 The lower-quality segmentation map generated by the image segmentation engine 510 is used to generate the composite image capture frame, and the composite image capture frame can be based on the image segmentation engine (e.g., FIG. 5 The high-quality segmentation map generated by the image segmentation engine 510 is used to generate the composite image preview frame. In some examples, the composite image preview frame can be based on the image completion and restoration engine (e.g., FIG. 5 The image completion and restoration engine (530) is used to generate relatively lightweight restoration techniques, and the synthesized image capture frames can be based on the image completion and restoration engine (e.g., FIG. 5 The image completion and restoration engine (530) implements relatively heavyweight restoration techniques to generate the composite image. In some cases, the composite image preview frame can be generated based on the overlay synthesis of segmented foreground objects with a second zoom level on a background image portion with a first zoom level, and the composite image capture frame can be based on... FIG. 3B The image compositing engine 540, the shadow compositing engine 550, and / or the image harmonization engine 560 are implemented to generate higher quality segmentation maps and larger ML models.

[0074] FIG. 3B Another example of synthesized image generation is illustrated, wherein synthesized image 380 is generated based on foreground information obtained from a first image 360 ​​captured at a first zoom level and background information obtained from a second image 370 captured at a second zoom level greater than the first zoom level. For example, FIG. 3AThe example of synthesized image generation can correspond to a scene where a person (e.g., a foreground subject) is standing far from tall buildings (e.g., background objects). In this scene, image 360 ​​is a wide-angle image obtained using a 1x zoom level, and image 370 is a telephoto image obtained using a 3x zoom level. A first camera can be used to obtain the wide-angle image 360, and a second camera can be used to obtain the telephoto image 370. The system and techniques described herein can be used to generate a synthesized image 380 to include foreground information of the person corresponding to the wide-angle image 360 ​​and background information of the buildings corresponding to the telephoto image 370.

[0075] In some examples, image 360 ​​may be a first image preview frame having a first zoom level (e.g., a first frame output before receiving input for capturing a frame), and image 370 may be a second image preview frame having a second zoom level different from the first zoom level (e.g., a second frame output before receiving input for capturing a frame). For example, image 360 ​​may be a wide-angle image preview frame, and image 370 may be a telephoto image preview frame. In some aspects, composite image 380 may be a composite image preview frame that includes at least a portion of the first image preview frame 360 ​​at the first zoom level and at least a portion of the second image preview frame 370 at the second zoom level.

[0076] In some examples, the composite image 380 is a composite image preview frame (e.g., a composite frame output before receiving input for a capture frame) that is output to and / or displayed as an image capture user interface (UI). For example, the composite image 380 may be a composite image preview frame displayed using the image capture UI of an image capture device or other electronic device (e.g., a smartphone, etc.). In some cases, the composite image 380 may be image preview data (e.g., a live preview) output before receiving a command for capturing an image (e.g., an image capture trigger, etc.). For example, the composite image 380 may be image preview data output before receiving a command for capturing an image (e.g., including multiple composite image preview frames), where the command corresponds to user input to a shutter button on the image capture UI. The captured composite image frame may be generated (e.g., captured) based on receiving user input indicating a command for capturing an image (e.g., user input to a shutter button on the image capture UI). The captured composite image frame may be generated based on one or more of the corresponding composite image preview frames.

[0077] In some respects, the same image capture UI can be used to generate and / or display images. FIG. 3B The composite image preview frame 330 (e.g., including the background portion associated with the wide-angle image preview frame and the foreground portion associated with the telephoto image preview frame) and FIG. 4The composite image preview frame 380 (e.g., including the background portion associated with the telephoto image preview frame and the foreground portion associated with the wide-angle image preview frame).

[0078] FIG. 4 An example of an image capture preview UI 400, according to some examples, includes at least two zoom adjustment features for adjusting a composite image in an image capture preview user interface (UI). For example, the image capture preview UI may include a first zoom adjustment feature and a second zoom adjustment feature (e.g., also referred to as corresponding first and second graphical user interface (GUI) elements) displayed in the image capture preview UI for adjusting a composite image preview frame (e.g., a composite frame output before receiving input for a capture frame). In some cases, UI 400 may correspond to a split zoom (e.g., "SZ") mode or feature of a camera application or GUI of a smartphone or other mobile computing device. In one exemplary example, the split zoom UI may include a first zoom adjustment feature (e.g., a first GUI or a first GUI element) that can receive user input for increasing or decreasing the main zoom value of the composite image preview frame depicted in UI 400. For example, the first zoom adjustment feature (e.g., a first GUI or a first GUI element) can be used to adjust the main zoom value corresponding to the composite image preview frame. FIG. 5 In the example UI 410, the main zoom value is shown as adjusted to the 1.5x zoom level.

[0079] The segmented zoom UI 400 may include a second zoom adjustment feature (e.g., a second GUI or second GUI element) that can receive user input for increasing or decreasing the zoom value of a segmented object within a composite image preview frame. In an exemplary example, the segmented object can be a foreground object or a background object. For example, the segmented object can be a person (e.g., a foreground object) within the main image preview. In example UIs 410 and 420, the second zoom adjustment feature (e.g., a second GUI element) corresponds to a "person zoom value". Example UI 410 corresponds to an initial state where the person zoom value and the main zoom value are equal (e.g., both are set to a 1.5x zoom level). In the initial state corresponding to UI 410, an image preview can be captured using a single camera with a 1.5x zoom level.

[0080] In the final state corresponding to Example UI 420, the person's zoom value is adjusted to a zoom level different from the 1.5x zoom level of the main image. For example, Example UI 420 shows the person's zoom value adjusted to a 3x zoom level. As will be described in more detail below, increasing or decreasing the person's zoom value can allow the system and technology to segment and remove the person from an image captured using the main zoom value (e.g., the image previewed in UI 410), and generate a composite image using the segmented image data of the person from the image captured using the person's zoom value (e.g., the image previewed in UI 420). As previously noted above, generating a composite image may include generating a composite image preview frame (e.g., a first frame output before receiving input for a capture frame) and / or generating (e.g., capturing) a composite image frame corresponding to one or more composite image preview frames (e.g., a captured frame output after the first frame and after receiving input for a capture frame).

[0081] FIG. 4 This is an illustration of an example image processing system 500 for performing synthetic image generation based on the segmentation of multiple camera image frames, according to some examples. In some aspects, the multiple camera image frames can be obtained from a first camera having at least a first focal length and a second camera having a second focal length. In some cases, the multiple camera image frames can be frames obtained before receiving input for capturing an image. For example, the multiple camera image frames can be preview frames obtained and / or output before receiving input for capturing an image. In some examples, the multiple camera image frames can be frames obtained after receiving input for capturing an image and / or based on receiving input for capturing an image. For example, the multiple camera image frames can be captured frames. A captured frame obtained after receiving input for capturing an image can correspond to a first frame obtained before receiving input for capturing an image (e.g., a captured frame can correspond to a preview frame).

[0082] A first camera can be used and a first zoom level (e.g., a first focal length) can be used to obtain the foreground source frame 502. A second camera can be used and a second zoom level (e.g., a second focal length) can be used to obtain the background source frame 504. The first camera and the second camera can be different from each other. In an exemplary example, the first camera and the second camera are included on the same computing device (e.g., an image capture device, etc.). The first zoom level and the second zoom level (e.g., a first focal length and a second focal length) can be different from each other. Image preview data corresponding to the foreground source frame 502 and the background source frame 504 (e.g., image frames output before receiving input for capturing images) can be at least partially based on the... FIG. 4 The example is obtained by segmenting one or more user inputs into the zoom UI 400. For example, foreground source frame 502 and background source frame 504 can be obtained based on... FIG. 4The main zoom level value and the human zoom level value are obtained from user input as depicted in the image.

[0083] In some cases, a first GUI or a first GUI element may be used to receive adjustments to the zoom level of a foreground object, and a second GUI or a second GUI element may be used to receive adjustments to the zoom level of a background object and / or the main zoom level. In some aspects, a second GUI element may be used to receive adjustments to the main zoom level, and a third GUI element may be used to receive adjustments to the zoom level of a background object. In some examples, one or more GUI elements for receiving user input indicating zoom level adjustments may be overlaid on (e.g., on top of or above the composite image preview frame) generated using the systems and techniques described herein (e.g., a frame obtained before receiving input for image capture). For example, an image capture preview interface may include a composite image preview frame and one or more zoom level adjustment GUIs overlaid on the composite image preview frame. In some cases, the generated composite image preview frame may include one or more GUI elements. In some examples, the generated composite image preview frame is separate from the corresponding overlay data of the corresponding GUI elements associated with rendering or displaying one or more GUI elements associated with zoom level adjustments.

[0084] In some cases, one or more of the GUI elements associated with zoom level adjustment can be collapsible. For example, FIG. 4 The "main zoom value" can be a default GUI overlay element displayed in conjunction with the composite image preview frame, and FIG. 4 The "Human Zoom Value" can be a GUI overlay element configured to collapse based on one or more additional user inputs. For example, additional user input can be received and used to begin displaying the "Human Zoom Value" GUI overlay element. In this example, the default state of the "Human Zoom Value" can be a collapsed state, where the "Human Zoom Value" GUI overlay element is not displayed by the image capture preview UI unless a specific user input is received instructing the display of the "Human Zoom Value" GUI overlay element on top of the composite image preview frame. In another example, the "Human Zoom Value" GUI overlay element can be displayed as an overlay on top of the composite image preview frame by default, and can receive additional user input instructing the collapse of the "Human Zoom Value" GUI overlay element. In such examples, the additional user input can cause the image capture preview UI to remove the "Human Zoom Value" GUI overlay element from the composite image preview frame displayed to the user.

[0085] In an illustrative example, background source frame 504 is shown as a live (e.g., instant) preview in an image capture preview interface (e.g., a frame output before receiving input for capturing the image), such as... FIG. 4Example UI 400. Background source frame 504 can correspond to FIG. 5 The "main zoom value". Foreground source frame 502 can be composited onto background source frame 504, where the composite image based on foreground source frame 502 and background source frame 504 is shown as a real-time (e.g., instantaneous) preview in the image capture preview interface (e.g., the composite frame output before receiving input for capturing the composite image). For example, it can be used... FIG. 4 The "human zoom value" is used to increase or decrease the relative size of the segmented foreground object 503. In some aspects, the background source frame 504 may be an image preview frame (e.g., a frame output before receiving input for capturing an image) captured using a first camera and a first zoom level (e.g., a first focal length), and the foreground source frame 502 may be an image preview frame (e.g., a frame output before receiving input for capturing an image) captured using a second camera and a second zoom level. The first camera and the second camera may be different from each other. The first zoom level and the second zoom level may be different from each other. In some aspects, the background source frame 504 and the foreground source frame 502 may include corresponding image data corresponding to the same scene and / or objects. For example, the foreground source frame 502 may include some (or all) of the scene and scene objects included in the background source frame 504 (e.g., in an example where the foreground source frame has a wider zoom level than the background source frame, the background source frame may depict a subset or portion of the scene in the foreground source frame). In another example, background source frame 504 may include some (or all) of the scene and scene objects included in foreground source frame 502 (e.g., in an example where the background source frame has a wider zoom level than the foreground source frame, the foreground source frame may depict a subset or portion of the scene in the background source frame).

[0086] As the relative size of the segmented foreground object 503 increases or decreases, the system and techniques can update the composite image preview displayed in the segmented zoom image capture preview interface. In some aspects, as the relative size of the segmented foreground object 503 increases or decreases, the specific camera or focal length used to obtain the foreground source frame 502 can be updated. For example, increasing the relative size of the segmented foreground object 503 can allow the system and techniques to obtain image data of the segmented foreground object 503 from a camera with a larger zoom level (e.g., a longer focal length). In some aspects, if an additional camera with a larger zoom level (longer focal length) is unavailable, the segmented zoom UI can remove UI features used to increase the zoom level. In another example, if an additional camera with a larger zoom level (longer focal length) is unavailable, the segmented zoom UI can perform digital zoom (e.g., cropping) using image preview data captured using the longest available focal length camera (e.g., an image frame output before receiving input for capturing the image).

[0087] Reducing the relative size of the segmented foreground object 503 allows the system and techniques to obtain segmented foreground object 503 image data from a camera with a smaller zoom level (e.g., a shorter focal length). For example, if the background source frame 504 is captured using a 1x zoom level, then FIG. 4 The foreground source frame 502 shown can correspond to a telephoto 3x zoom level. If the relative size of the segmented foreground object 503 is reduced (e.g., FIG. 6 If the “human zoom value” is reduced, then in some respects, the foreground source frame 502 image data can be updated as image preview data (e.g., an image frame output before receiving input for capturing an image) captured using a wide-angle camera with a zoom level of 0.5x (or other zoom levels less than 1x the main zoom level corresponding to the background source frame 504 image data).

[0088] In some aspects, the foreground source frame 502 image preview data (e.g., an image frame output before receiving input for capturing an image) can be provided to an image segmentation and / or image matting engine 510, which can be used to generate segmented foreground object image data 503. The segmented foreground object image data can also be referred to as "segmented foreground object result" and / or "segmented person result." The zoom level of the segmented foreground object result 503 is the same as the zoom level of the foreground source frame 502 (e.g., with...). FIG. 6 (The zoom level of the person is the same). For example, the zoom level of the foreground source frame 502 can be 3x, and the zoom level of the segmented foreground object result 503 can be 3x. In some aspects, the shadow matting engine 520 can be additionally used to generate the segmented foreground object result 503. The segmented foreground object result 503 can also be referred to as the segmented person result (e.g., in the example where the foreground object to be segmented is a person).

[0089] The following is for reference. FIG. 5 To describe further details of the segmentation result 503 ( FIG. 15A to FIG. 15D The result of the segmentation of 650 people can be compared with FIG. 5 The results of the segmentation were the same for 503 people. See below for reference. FIG. 4 To describe FIG. 7 Further details regarding shadow matting and / or the shadow matting engine 520. In some aspects, the system and techniques can generate segmented foreground object results 503 without using the shadow matting engine 520.

[0090] In some examples, the background source frame 504 image preview data (e.g., an image frame output before receiving input for capturing the image) can be provided to an image completion and / or restoration engine 530, which can be used to generate background-only image data 505. The background-only image data can also be referred to as the "background-only result." The zoom level of the background-only result can be the same as the zoom level of the background source frame 504 (e.g., with...). FIG. 7 (The "main zoom level" is the same). For example, the zoom level of the background source frame 504 can be 1x, and the zoom level of the background-only result 505 can also be 1x. In some aspects, the system and techniques can generate the background-only result 505 without using the image completion and / or restoration engine 530.

[0091] The following is for reference. FIG. 8 (To describe further details of background result 505). See below for reference. FIG. 8 Figure 16 and Figure 16 are used to describe further details of the image completion and / or repair engine 530.

[0092] In an exemplary example, image compositing engine 540 can generate a composite image 507 by combining the segmented foreground object result 503 with the background-only result 505. For example, the segmented foreground object result 503 can be added to the background-only result 505 to obtain the composite image 507. See below for reference. FIG. 9 To describe further details of the generation of the synthetic image (e.g., FIG. 5 and FIG. 15A to FIG. 15D The composite image 850 can be combined with FIG. 6 The composite image is the same as 507.

[0093] In some respects, the shadow compositing engine 550 can be additionally used to perform shadow compositing to generate more realistic shadow information in the composite image 507. See the following references. FIG. 4 To describe further details of shadow compositing.

[0094] In some respects, the image harmonization engine 560 can be used to generate an adjusted synthetic image 509 based on the white balance, color profile, hue, etc. of image data 503 of segmented foreground object results and image data 505 of background-only results (e.g., which are obtained using two separate cameras that can use different sensors, image capture parameters, etc.) that have undergone one or more image harmonization processes to match.

[0095] FIG. 5 Examples of person or object segmentation 600 from the foreground of a telephoto image frame are illustrated. In some aspects, the segmentation zoom UI 602 can be combined with... FIG. 5The split zoom UI 400 (and / or example UIs 410, 420) is the same as or similar to the split zoom UI 602. In the example of split zoom UI 602, the main zoom level (e.g., FIG. 5 The zoom level of the background source frame 504 image data is set to 1x, and the human zoom level (e.g., FIG. 5 The zoom level of the foreground source frame 502 image data is set to 3x.

[0096] In some respects, FIG. 6 The foreground source frame 502 image data can be the same as or similar to the foreground source frame 610 image data. A 3x zoom horizontal telephoto camera can be used to obtain the foreground source frame 610. In some examples, based on the segmentation zoom UI 602, the user increases the human zoom level to 3x, which can obtain 3x zoom telephoto image frame data 610 and its corresponding human segmentation map 620.

[0097] In an illustrative example, a segmentation machine learning network can be used to generate a 3x human segmentation map 620. The segmentation machine learning network can be used to obtain... FIG. 5 The foreground source frame 502 image data and the background source frame 504 image data are implemented using the same smartphone, mobile computing device, image capture device, etc. In some aspects, the 3x human segmentation image 620 may have the same pixel resolution as the 3x telephoto image data 610. Each pixel of the 3x human segmentation image 620 may include a value indicating whether the corresponding pixel in the 3x telephoto image data 610 is included or not included in the segmented human result. In some examples, the 3x human segmentation image 620 and the 3x segmented human result 650 may include only pixels corresponding to the image data of the human. In other examples, the 3x human segmentation image 620 and the 3x segmented human result 650 may include pixels corresponding to the image data of the human or the image data corresponding to the shadow cast by the human. FIG. 7 As illustrated, the segmentation information includes the shadows of the person. As previously noted above, the 3x segmented person result 650 can be obtained by combining the 3x telephoto image data 610 with the 3x person segmentation map 620. For example, in an exemplary example, the 3x segmented person result 650 can be obtained by multiplying the 3x telephoto image data 610 by the 3x person segmentation map 620. In some aspects, the 3x segmented person result 650 can be combined with... FIG. 7 The results of the segmentation of people were 503 identical or similar.

[0098] FIG. 6 Examples are given of removing 700 people or objects from the foreground of an image frame used as the background of a composite image, based on some examples. FIG. 4 Can be associated with FIG. 7 The split zoom UI 602 and / or FIG. 6The split zoom UI 400 (and sample UIs 410 and 420) is the same as or similar to the split zoom UI. FIG. 7 In the example, the main zoom level of the split zoom UI can be set to 1x, and the human zoom level can be set to 3x, as shown in the reference above. FIG. 5 As described in the split zoom UI 602.

[0099] In some respects, FIG. 5 Example foreground removal (e.g., removing a person or object 700) can be used FIG. 6 The background source frame 504 image data is used to perform the operation and can be obtained using a 1x zoom horizontal camera, which is included in the image data used for capturing. FIG. 5 Foreground source frame 502 image data (e.g., FIG. 6 The foreground source frame (610 image data) is on the same device as a 3x telephoto camera.

[0100] In an exemplary example, one or more machine learning (ML) and / or artificial intelligence (AI) models are used, corresponding to the foreground portion of interest (e.g., from...). FIG. 5 Image data of a person or other foreground object or subject segmented from the foreground frame 502 can be removed from the background source frame 504 image data. For example, one or more ML and / or AI models can be used to generate a 1x zoom level person removal result 720 based on removing people (e.g., foreground objects) from the 1x background image data 504. In some cases, models for generating... FIG. 7 The same segmentation machine learning model used to generate the 3x person segmentation map 620 and the 3x person segmentation result 650 removes people from 1x background image data. For example, one or more ML and / or AI models can be used to generate a 1x person segmentation map 710, which can be combined with background source frame 504 image data to generate a 1x subject removal result 720 in the background. In an illustrative example, the 1x person segmentation map 710 can be used... FIG. 5 The image matting / segmentation engine 510 is used to generate the image from the background source frame 504 (e.g., FIG. 5 The mask of the foreground object removed from the 1x human segmentation map. Using the 1x human segmentation map 710, the foreground portion of the background source frame 504 is removed to generate a 1x human removal result 720, which includes the missing regions (e.g., regions without pixel data) corresponding to the 1x human segmentation map 710. This can be used... FIG. 4 The image completion and restoration engine 530 fills in the missing or blank areas of the 1x human removal result 720 to generate a 1x background-only result 505. For example, FIG. 5The image completion and / or repair engine 530 can be used to generate image data for the missing parts of the image data in the 1x person removal result 720 by generating pixel data to fill the negative space corresponding to the removed person in the 1x background image based on the analysis of pixel information, semantic information, etc. of neighboring pixels that were not removed in the 1x person removal result 720 and / or the analysis of pixel information, semantic information, etc. of the overall 1x background image 504.

[0101] In some examples, the zoom level of 1x background-only result 505 can be the same as the zoom level of 1x background source frame 504 (e.g., with...). FIG. 5 (The "main zoom level" is the same). For example, the zoom level of the background source frame 504 can be 1x, and the zoom level of the background-only result 505 can also be 1x. In some aspects, the system and techniques can skip generating the 1x background-only result 505 by not utilizing the image completion and / or restoration engine 530. For example, the 1x background source frame 504 can be the same as... FIG. 8 The segmented foreground object result 503 is provided directly as input. FIG. 6 Image compositing engine 540.

[0102] FIG. 7 An example of generating 800 synthetic images is shown. For example, according to some examples, it is possible to use a method based on... FIG. 5 Foreground information segmented by people and using based FIG. 7 The background information removed by the person is used to generate the synthetic image 850. In an exemplary example, it can be based on the background information removed by the person. FIG. 6 and FIG. 5 The 1x background result 505 and FIG. 8 The three-fold segmentation of the human body result 650 is added to generate a composite image 850. As previously noted, the composite image 850 can be combined with... FIG. 5 The composite image 507 is identical or similar. A composite image 850 is generated by adding a 1x background-only result 505 and a 3x segmented person result 650 to include a 3x telephoto view of the person (e.g., a foreground object) within a 1x wide-frame view of the background scene. In some respects, FIG. 5 The composite image 850 and / or composite image generation 800 can be (respectively) used FIG. 9 Image compositing engine 540 and / or FIG. 8 The shadow compositing engine 550 is used to generate and / or execute shadows.

[0103] FIG. 9 Examples are provided for previewing an image before receiving a command to capture it (e.g., image capture triggered or image capture user input, etc.). FIG. 4An example of an image capture preview UI 900 for a composite image 850. For example, the composite image 850 may be output and / or displayed in the image capture preview UI 900 before a command for capturing an image is received. For example, one or more composite image preview frames may be generated and displayed in the image capture preview UI 900 based on corresponding image preview frames corresponding to foreground source preview frames and background source preview frames. The composite image 850 may be a composite image frame generated (e.g., captured) based on a command for capturing an image. In some aspects, the captured composite image frame may correspond to the composite image preview frame displayed in the image capture preview UI 900 when a command for capturing an image is received. In another example, the composite image 850 may be a composite image preview frame displayed in the image capture preview UI 900. In some aspects, FIG. 6 Image capture preview UI 900 can be used with FIG. 4 The split UI 400 and / or FIG. 10 One or more of the segmented UI 602 are the same or similar.

[0104] In the initial segmented UI view 910, the displayed image preview area corresponds to a 1x zoom level image of the scene (e.g., both foreground objects (e.g., people) and background objects (e.g., houses)), shown as being at the same 1x zoom level, using image preview data obtained from a 1x zoom wide-angle camera (e.g., image frames output before receiving input for capturing the image). This is based on increasing the zoom level of the person (e.g., increasing...). FIG. 4 User input (increased from "human zoom level" to 3x zoom level) previews the composite image 850 in the image preview area of ​​the segmented UI, as depicted in the final segmented UI view 950. In an illustrative example, the composite preview image 850 displayed in the segmented UI view 950 is generated in real-time using image preview frames corresponding to a 1x background frame and a 3x telephoto foreground frame. The image preview frames can be obtained as streaming image data from a 1x wide-angle image sensor and a 3x telephoto image sensor included in the computing device used to render the segmented UI view 950 and the image capture preview UI 900.

[0105] In some aspects, the image capture preview UI 900 may include image capture input elements or other UI elements for receiving commands for capturing images (e.g., triggering capture of the composite image output, etc.). For example, user input that selects or actuates the image capture / shutter button of the split UI view 950 can cause the system and technology to capture a 1x background frame and a 3x telephoto foreground frame at full resolution and generate a full-resolution composite image output corresponding to the composite image preview 850. In some examples, the 1x background frame image data and the 3x telephoto frame image data can be captured in parallel (e.g., simultaneously). In some examples, the 1x background frame image data and the 3x telephoto frame image data can be captured sequentially.

[0106] FIG. 6 Examples of usage that can be combined with FIG. 9 Segmented zoom UI 400 FIG. 10 The split zoom UI 602 and / or FIG. 10 An example of repositioning the foreground object in the split zoom UI 900, which is one or more identical or similar image capture preview UI 1010 (e.g., split zoom UI 1000).

[0107] For example, FIG. 8 The split zoom UI 1010 may include a main zoom level adjustment (e.g., shown here as set to 1x zoom level) and a human zoom level adjustment (e.g., shown here as set to 3x zoom level), as previously described. FIG. 9 The split zoom UI 1010 may additionally include a person (e.g., a foreground object) pan adjustment UI element, which can be used to pan or otherwise adjust the positioning of the person within the composite image preview displayed in the split zoom UI.

[0108] For example, the Segment Zoom UI 1010 corresponds to an initial composite image preview without applying translation adjustments to the segmented person. The initial composite image preview displayed by the Segment Zoom UI 1010 can be compared with... FIG. 10 and FIG. 6 The composite image preview is the same as or similar to 850.

[0109] For translation adjustment of UI elements (e.g., in) FIG. 8 In the example, user input (corresponding to the up arrow, down arrow, left arrow, and right arrow) can be used to generate a 3x translated human segmentation map 1030 using the 3x initial human segmentation map 1020. The 3x initial human segmentation map 1020 can be used with... FIG. 5The 3x human segmentation image 620 is the same as or similar to the original human segmentation image 1020. In some aspects, based on which arrow is pressed in the translation adjustment UI element, the segmented people in the original human segmentation image 1020 are moved 3x in the corresponding direction by appropriately filling and removing rows and / or columns. Based on the filling and removal of rows / columns, a 3x translated human segmentation image 1030 can be generated.

[0110] The 3x translated person segmentation image 1030 can be used to generate a translated composite image preview, as depicted in the image preview area of ​​the segmented zoom UI 1050. The translated composite image preview can be referenced as above. FIG. 11 The non-translated synthetic image preview 850 is generated as described (e.g., based on the inverted 3x translated human segmentation map 1030, multiplied by 1x human removal result, and using...). FIG. 10 The image synthesis engine 540 is combined with the 3x human segmentation results.

[0111] In one exemplary example, the systems and techniques described herein can be used to provide a split zoom UI (e.g., also referred to as a dual zoom UI) that can be used to individually change or otherwise adjust the zoom levels of the foreground and background of an image during the image preview stage (e.g., prior to final image capture based on receiving a command to capture the image, such as a user selection of a camera or shutter button included in the split zoom UI).

[0112] FIG. 12 Another example of an image capture preview UI 1100, which can be used to compose and / or adjust a composite image, is illustrated, based on some examples. This image capture preview UI includes UI features for adjusting the zoom level of a foreground object and UI features for adjusting the translation or position of the foreground object. The image capture preview UI 1100 can also be referred to as a split zoom UI and can be integrated with… FIG. 10 The split zoom UI is the same as or similar to 1000.

[0113] In the initial view 1110 of the segment zoom UI 1100, user input instructing the foreground object to be panned to the left is received. For example, the user input may correspond to the selection of a left-hand arrow on a panning UI element as described above. Based on the user input for the left-hand panning arrow, the updated view 1120 of the segment zoom UI 1100 can display a composite image preview, where the foreground object (e.g., a person) is panned to the left relative to background image data, which remains in its unpanned position. The panning adjustment can be performed in real-time in the preview area of ​​the segment zoom UI 1100, similar to or identical to the real-time (e.g., instantaneous) adjustment of the foreground and background segment zoom levels described above. In some examples, the selection of directional arrows (e.g., left, right, up, down) on the panning UI element may correspond to a predetermined panning amount. In some aspects, the directional arrows may be selected multiple times to increase the panning distance of the foreground object within the preview composite image. In another example, the length or distance of panning in a particular direction may be based on the length of time the user selects the corresponding directional panning arrow.

[0114] FIG. 11 Examples of an image capture preview UI 1200 are provided, illustrating how a composite image can be synthesized and / or adjusted based on one or more user input gestures corresponding to foreground object zoom level adjustment and / or foreground object position adjustment. The image capture preview UI 1200 may also be referred to as a split zoom UI and can be used with... FIG. 13 The split zoom UI 1000 and / or FIG. 10 One or more of the segmented zoom UI 1100 are the same or similar.

[0115] In the initial view 1210 of the split zoom UI 1200, user input is received to select a foreground object of interest. For example, the user input could be a touch input selecting a person as the foreground object of interest. Based on the user touch input selecting a person, the initial view 1210 of the split zoom UI 1200 can be updated to highlight the selected person in the foreground of the image preview.

[0116] In some aspects, users can use one or more touch-based gestures to adjust the composite image preview. For example, after selecting a person in the initial view 1210, a user can expand the person to zoom in (e.g., increase the zoom level of the selected person relative to the background image data) and / or pinch the person to shrink (e.g., decrease the zoom level of the selected person relative to the background image data), etc. In another example, after selecting a person in the initial view 1210, a user can pin and slide the person to place it in a desired position in the composite image preview. For example, the updated composite image preview displayed in the updated view 1220 of the split zoom UI 1200 could correspond to user input from expanding the selected person in the initial view 1210, where the zoom level of the selected person increases based on the user input.

[0117] FIG. 11 Examples of image capture preview UI 1300 are illustrated, which can be used to synthesize and / or adjust a composite image based on zoom levels that adjust background information obtained from a first image and / or foreground information obtained from a second image; the image capture preview UI 1300 may also be referred to as a segmented zoom UI, and can be used with... FIG. 12 Segmented zoom UI 1000 FIG. 12 The split zoom UI 1100 and / or FIG. 13 One or more of the segmented zoom UI 1200 are the same or similar.

[0118] In an exemplary example, the image capture preview UI 1300 may include a main zoom level adjustment (e.g., shown in initial view 1310 as set to 1x main zoom level), a person zoom level adjustment (e.g., shown in initial view 1310 as set to 2x zoom level), and a background zoom level adjustment (e.g., shown in initial view 1310 as set to 4x zoom level).

[0119] The image capture preview UI 1300's main zoom level adjustment and human zoom level adjustment can be compared with the above reference Figure 3 to... FIG. 14The described corresponding zoom level adjustments are the same or similar. In an exemplary example, background zoom level adjustment can be used to generate a composite image preview based on additionally segmenting and adjusting the zoom level of a background object (such as a house). For example, the zoom level of a selected background object (e.g., a house) can be adjusted separately from the zoom level of a selected foreground object (e.g., a person), and vice versa. The zoom level of a selected background object (e.g., a house) can also be adjusted separately from the zoom level of the main image frame (e.g., image data other than the person or the house). The zoom level of a selected foreground object (e.g., a person) can also be adjusted separately from the zoom level of the main image frame (e.g., image data other than the person or the house).

[0120] In the initial view 1310, background objects (e.g., houses) are segmented and combined into a 1x main image frame preview at a 4x zoom level, and foreground objects (e.g., people) are segmented and combined into a 1x main image frame preview at a 2x zoom level.

[0121] In updated view 1320, the background object (e.g., a house) is shown at the same 1x zoom level as the main image frame preview. In some examples, since the background object zoom level is the same as the main image frame zoom level, no segmentation of the background object is performed. For example, the preview image displayed in updated view 1320 can be generated by combining a 2x telephoto view of a segmented foreground person with a 1x image that includes the main scene view and the background house object.

[0122] In some examples, one or more of the segmented zoom adjustments described herein can be performed after image capture. For example, when a command to capture an image is received (e.g., a user selection of a camera or shutter button in the segmented zoom UI), image data can be captured using multiple different cameras and / or focal lengths associated with the computing device. In some examples, image data can be captured using each of multiple cameras included in the computing device, where each of the multiple cameras has a different focal length.

[0123] The composite image generated and stored in response to a command to capture an image can be generated and stored during image preview (e.g., before image capture, as previously mentioned above with reference to Figures 3 to 4). FIG. 15A to FIG. 15DThe system and techniques described herein can be used to generate updated composite images at a later time (e.g., after image capture) based on image data from all available cameras and focal lengths acquired simultaneously with the initial composite image. For example, additional image data from different focal lengths can be stored as part of the metadata of the initial composite image capture. In some aspects, at some point after the initial capture of the image data, the user can later adjust the zoom level of the foreground and / or background, and / or adjust the translation of foreground and / or background objects.

[0124] FIG. 5 Examples of image adjustment UI 1400, based on several examples, are illustrated, which can be used to segment foreground objects from a first image and composite the segmented foreground objects with background information obtained from a second image. For example, in the gallery view of a previously captured image, a user can long-press on a person in the photo and select the option to copy the person (or other selected foreground object) to a different photo. After selecting the "copy" option for the selected person or foreground object, the user can navigate to a different desired photo and long-press on the location where the user wants to paste the copied person or foreground object. Selecting the "paste" option causes the segmented person from the initial frame 1410 to be overlaid onto the composite frame 1420 or otherwise combined with that composite frame. In the example where the initial frame 1410 itself is a composite frame generated according to the system and techniques described herein, the selection of a person or other foreground object using the "copy" option in the initial frame 1410 can be based on previously generated segmentation information. In other examples, where the initial frame 1410 is not a composite frame generated using a segmented person taken at a different focal length than the background, selecting the "copy" option in the initial frame 1410 can trigger segmentation of the selected person from the initial frame 1410 using one or more segmentation machine learning networks, as previously described above.

[0125] FIG. 5 Example images 1500 are shown, corresponding to shadow matting and / or shadow compositing, based on some examples. These can be used... FIG. 5 The shadow matting engine 520 is used to perform shadow matting. It can be used... FIG. 5 The shadow compositing engine 550 is used to perform shadow compositing.

[0126] In some respects, shadow compositing and / or shadow matting can be similar to image compositing and / or image matting. In some cases, shadow compositing and / or shadow matting can be included in image compositing and / or image matting, or a subset thereof.

[0127] Image matting is the process of accurately cropping out the foreground from an image. For example, image matting can be used to extract the foreground from an image. FIG. 5 The foreground (e.g., a person) is cropped out from image 502 to generate FIG. 5The segmentation result is 503. Image compositing is the process of pasting a cut-out foreground into another image. For example, it can be done using... FIG. 5 Image compositing engine 540 is used to perform image compositing to... FIG. 5 The result of the division of people 503 and FIG. 5 The background image 505 combination is only.

[0128] In image segmentation, each pixel can be analyzed to determine whether it belongs to the foreground or background category of the image. However, this binary segmentation method (e.g., foreground pixels or background pixels) may not be able to handle natural scenes containing fine details (e.g., hair, fur, etc.). In some cases, scenes with fine details (such as hair and fur) can be segmented based on estimating the transparency value for each pixel of the foreground object. For example, without estimating the transparency value for each pixel of the segmented foreground object, parts of the background image may become trapped between the fine details of the foreground object during foreground segmentation, potentially resulting in unrealistic composite images. For instance, background pixels corresponding to a blue sky might become trapped between the fine details of the hair on the segmented foreground person's head, potentially producing unrealistic composite images when the segmented person is superimposed onto another image.

[0129] In an exemplary example, FIG. 6 The image segmentation engine 510 may further include or otherwise implement an image matting engine (e.g., a sub-engine or subsystem of engine 510). Image matting can be performed to estimate the foreground opacity of some (or all) pixels included in the foreground segmentation estimated by the image segmentation engine 510. Image matting can be used for more accurate segmentation of the foreground and background of an image.

[0130] In some respects, FIG. 5 The image segmentation engine 510 can generate a human segmentation map based on classifying each pixel as corresponding to a person or not (e.g., background pixels). The human segmentation map can be compared with... FIG. 5 The people segmented in Figure 620 are the same or similar.

[0131] FIG. 5 The image segmentation engine 510 can further generate a human matting mask. In both the human segmentation map and the human matting mask, pixels with a first value (e.g., equal to "1") belong to the foreground object (e.g., the person), and pixels with a second value (e.g., equal to "0") belong to the background.

[0132] A matting mask can be generated to include one or more pixels with values ​​between 0 and 1, indicating an estimated opacity value for the pixel. Pixels in the matting mask can each correspond to a corresponding pixel in a person segmentation map. For example, the matting mask can estimate opacity values ​​for pixels along the boundary between the foreground and background categories in the person segmentation map. Pixels in the matting mask with values ​​between 0 (e.g., background pixels) and 1 (e.g., foreground pixels) can be considered to belong to both the foreground and background portions of the segmentation, with their opacity based on the corresponding value in the matting mask. For example, a matting mask pixel with a value of 0.8 can be represented with 80% opacity in the segmented foreground (e.g., and also with 20% opacity in the background category). In some aspects, for FIG. 5 The segmentation result 503, using a matting mask, can be used to generate more realistic synthetic images based on a more gradual transition from the foreground to the background in the resulting composite image (e.g., such as...). FIG. 5 Synthetic images 507, 509, etc., and / or various other synthetic images described herein.

[0133] In some aspects, use FIG. 5 The complexity of image compositing performed by the image compositing engine 540 can be at least partially based on the complexity of image compositing performed by the engine 540. FIG. 5 The accuracy of the image matting mask generated or otherwise used by the image segmentation and image matting engine 510 varies. For example, if the image matting mask is more accurate, the complexity of the image compositing steps performed using the image compositing engine 540 can be reduced (e.g., multiplying the image matting mask with the foreground frame 502 and adding the result to the background frame 504 or 505). In some aspects, if the image matting mask is inaccurate and / or if the image segmentation engine 510 does not use the image matting mask, one or more post-processing steps can be applied to improve the realism of the boundary regions around the segmented result 503 in the synthesized image. For example, when compositing to... FIG. 15A to FIG. 15D When only the background frame 505 is used, blending can be used to smooth the transition around the boundary of the segmented result 503.

[0134] In some aspects, shadow matting and shadow compositing can be performed to generate composite images. For example, FIG. 15A The shadow matting engine 520 can be used with FIG. 5 The image segmentation and / or image matting engine 510 is used in conjunction to generate segmentation results 503 that further include shadow matting information.

[0135] FIG. 15B Four example image frames are depicted in connection with examples of shadow matting and shadow compositing. FIG. 5 The image frame is the background image frame, and can be compared with... FIG. 15CThe background image data is the same as or similar to 504. FIG. 15B The image frame is the foreground image frame, and can be compared with... FIG. 15D The foreground image data 502 is the same as or similar to the foreground image data. FIG. 15B The image frames are based on from FIG. 15C An example of a merged or composited result where only the shadows of the person are segmented from the foreground image frame. FIG. 15B The image frames are based on from FIG. 15C An example of a foreground image frame segmentation of people and their shadows, with the result of merging or compositing the shadows.

[0136] Images typically include one or more shadows or reflections. If segmentation is performed only on foreground people (e.g., if segmentation engine 510 does not segment the shadows of foreground people), the resulting synthesized image may be unrealistic. For example, FIG. 15D The image frames depict examples of merged or synthesized image results, where... FIG. 15B The person is segmented in the foreground image frame, but the person's shadow is not included in the segmentation. In this example, FIG. 15D The resulting synthetic image frames, based on a 3x segmentation of a person, may appear unrealistic due to the lack of appropriately sized and positioned shadows. FIG. 15D Image frames depicting examples of merged or synthesized image results, where both people and their shadows are derived from... FIG. 15D Segmentation is performed within the foreground image frame. This is based on including shadows in the image used for generation. FIG. 15A In the segmented human results from synthesized image frames, the synthesized results can appear more natural and realistic. Additionally, including shadows in the segmented human results can further enhance the effect. FIG. 15B The resulting composite image frame shows the impression of a person standing on the surface beneath their feet.

[0137] FIG. 7 The example composite image frame includes two shadows: corresponding to FIG. 5 The first shadow and the second shadow from the segmented person result in the background image, showing the location and relative size of the person (e.g., corresponding to the first shadow). FIG. 16A to FIG. 16C The positioning and relative size of the second shadow of the person in the foreground image.

[0138] Shadow masking can be used to remove shadows from foreground objects removed from an image. For example, FIG. 5 The 1x person removal result 710 is generated by removing the person from the foreground of a 1x background image frame and additionally by removing the person's shadow from the foreground of the 1x background frame. In an illustrative example, the person's shadow can be... FIG. 15D The shadow masking engine 520 removes shadows from the foreground of a 1x background image.

[0139] In some aspects, shadow matting can be performed to remove or cut out shadows from the foreground of an image, and shadow compositing can be performed to paste the removed shadows into the correct position in the composite image, such that the pasted shadows adapt to the texture and / or shape of the surface in which they are placed within the composite image. In some examples, shadow matting and shadow compositing can enhance the realism of the composite image, such as in examples where shadow matting is overlaid onto a background object in the composite image that has a different shape or orientation than the shadows in the foreground source image.

[0140] FIG. 5 Example image 1600 corresponding to image completion and / or image restoration is illustrated according to some examples. In some aspects, the image completion and restoration of Figure 16 can be combined with... FIG. 16A The image completion and restoration engine 530 performs the same or similar image completion and / or restoration. For example, image completion and restoration can be performed when the zoom level of the foreground (e.g., a person) is increased in a composite image preview. When the zoom level of the foreground (e.g., a person) is increased, the foreground segment from the higher zoom level image data will be placed on top of the smaller foreground in the current frame (e.g., the main or background frame) to generate a composite image preview.

[0141] In some cases, a larger (e.g., increased zoom level) segmentation of the foreground person will not be able to completely cover or hide the smaller foreground being replaced in the current frame. For example, FIG. 5 The segmented 3-fold person in the image frame does not fully cover the smaller foreground person from the 1-fold background frame, nor does it fully cover the shadow of the smaller foreground person from the 1-fold background frame.

[0142] In an exemplary example, FIG. 16B The image completion and inpainting engine 530 can be used to remove pixels corresponding to smaller people or other foreground objects from a 1x background frame, and then paste the segmented 3x person onto it to generate a composite image. In some aspects, when the zoom level of the foreground object (e.g., a person) is to be reduced, it may be necessary to erase the foreground object (e.g., a person) in the current frame before compositing the smaller segmented person into the frame. In some examples, the missing gaps from the removal of foreground objects and shadows from the current or background frame can be filled with interpolated pixel data or additional pixel data generated based on neighboring regions and / or other frames with missing information. For example, image inpainting and / or image completion can be implemented as a process of filling missing pixel locations based on known neighboring regions or patterns of the image.

[0143] FIG. 16C The image frame is the background frame of the image data, and can be compared with... FIG. 16B The background image data is the same as or similar to 504. FIG. 16CThe image frame depicts the background frame of image data in which foreground objects (e.g., people and shadows) are erased. For example, as described earlier above, foreground matting masks and shadow matting masks can be used to remove foreground objects (people and shadows). FIG. 16C Image frame depiction FIG. 5 The image frames that have already undergone restoration. In some respects, the restoration process can also be referred to as "image completion," and FIG. 5 The image frame can also be referred to as a "complete background-only image". FIG. 5 Image frames can be with FIG. 5 The completed image is only the same as or similar to the background image 505.

[0144] return FIG. 5 In some aspects of the discussion, the Image Harmonization Engine 560 can be used to perform image harmonization to improve the user experience. FIG. 5 The system 500 generates a composite image and / or a composite image preview with improved visual consistency. For example, image harmonization can improve the composite image by adjusting the appearance of the foreground image segmentation and the completed background image frames to make them compatible or otherwise consistent with each other (e.g., such as...). FIG. 17 The visual consistency of the synthesized image 509 is achieved through image harmonization. For example, image data acquired using different cameras may have different image capture properties and / or settings, such as color temperature values, automatic white balance (AWB) settings, etc. In some aspects, the system and techniques may use an image harmonization engine 560 to perform image harmonization to match the images combined into the synthesized image (e.g., such as...). FIG. 3A 507, Deharmonicized Composite Image FIG. 3B The color temperature and / or white balance settings of various parts of the image data (such as the harmonious composite image 509).

[0145] FIG. 4 This is a flowchart illustrating an example of a process 1700 for processing image and / or video data. At block 1702, process 1700 includes obtaining first image data of a scene, which is associated with a first zoom level and includes at least a foreground portion and a background portion. For example, the first image data may be the same as or similar to the following: FIG. 5 One of the corresponding images 310 and 320; FIG. 6 The corresponding one of the image data 360 and 370; FIG. 7 1.5 times the image data 410; FIG. 9 The corresponding one of the 3x image data 502 or the 1x image data 504; FIG. 3A The corresponding one of the 1x image data 602 or the 3x image data 610; FIG. 3A 1x the image data 504; FIG. 3A 1x image data 910; etc.

[0146] In some cases, the first image data may be a first frame obtained before receiving input for capturing the frame. For example, the first image data may be obtained before receiving input for capturing the frame. The first image data may also be obtained based on the input for capturing the frame before it is captured. For example, the first image data may be a preview frame obtained before receiving input for obtaining the captured frame. In some examples, the first image data may be associated with a composite image preview frame, and the captured frame may be a composite image capture frame.

[0147] In some cases, the first image data includes first image data acquired using a first camera having a first focal length (e.g., an image frame output before receiving input for capturing an image). For example, the first image data may be first image preview data acquired using a first camera having a first focal length (e.g., a first image preview frame). In some cases, the first image data is associated with a first camera having a first focal length corresponding to a first zoom level. In some examples, the first image data may be relatively wide-angle image data associated with a relatively wide-angle zoom level.

[0148] At box 1704, process 1700 includes receiving user input instructing an adjustment of the zoom level for increasing or decreasing the zoom level of the foreground portion relative to a background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than or less than the first zoom level. For example, the user input can be received using a user interface (UI) or graphical user interface (GUI) that is the same as or similar to one or more of the UIs and / or GUIs of Figures 3 through 16. In some examples, the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for increasing the zoom level of the foreground portion relative to a background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than the first zoom level. For example, the adjustment of the zoom level can increase the zoom level of the foreground portion, as referenced... FIG. 3B The composite image 330 is described. For example, both the foreground and background portions are... FIG. 12 The first zoom level is associated with 1x in the first image data 310. FIG. 12 In the composite image 330, the zoom level of the foreground portion is increased to three times the second zoom level, relative to the first zoom level of the background portion.

[0149] In another example, the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for decreasing the zoom level of the foreground portion relative to the background portion in the first image data, wherein the adjustment corresponds to a second zoom level that is less than the first zoom level. For example, the adjustment of the zoom level can decrease the zoom level of the foreground portion, as shown in the reference...FIG. 3A The composite image 380 is described.

[0150] In some cases, receiving user input instructing the adjustment of the zoom level of a foreground portion relative to a background portion in the first image data includes: receiving a first user input instructing the selection of a foreground portion from one or more foreground portions included in the first image data of the scene; and receiving a second user input instructing the adjustment of the zoom level of the selected foreground portion relative to a background portion (e.g., a non-selected portion of the first image data). For example, the first user input instructing the selection of the foreground portion may be combined with... FIG. 3B The example zoom UI 1210 depicts touch or long-press inputs that are the same or similar, and the second user input indicating the adjustment of the zoom level for increasing or decreasing the selected foreground portion can be the same as... FIG. 4 The increase in zoom level depicted in the example zoom UI 1220 is the same as or similar to that shown in the example zoom UI.

[0151] In some examples, user input instructing the user to adjust the zoom level relative to the background portion of the first image data, by increasing or decreasing the zoom level, is received in a graphical user interface (GUI). For example, the GUI may include a slider, where moving the slider in a first direction indicates an increase in the zoom level, and moving the slider in a second direction indicates a decrease in the zoom level. In some cases, the GUI includes multiple discrete step adjustments, each corresponding to a predetermined increase or decrease in the zoom level.

[0152] In some examples, process 1700 includes: receiving user input indicating an adjustment for increasing or decreasing the zoom level of a background portion relative to a foreground portion in first image data; and automatically determining a corresponding adjustment for increasing or decreasing the zoom level of the foreground portion. The corresponding adjustment for increasing or decreasing the zoom level of the foreground portion can be automatically determined relative to the user input indicating an adjustment for increasing or decreasing the zoom level of the background portion. A composite image can be generated based on the adjustment for increasing or decreasing the zoom level of the background portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the foreground portion.

[0153] In some examples, process 1700 includes automatically determining a corresponding adjustment for increasing or decreasing the zoom level of the background portion based on user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion. The composite image can be generated based on the adjustment for increasing or decreasing the zoom level of the foreground portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the background portion.

[0154] At box 1706, process 1700 includes acquiring second image data of a scene based on adjustments and using a second zoom level, the second image data including an adjusted foreground portion associated with the second zoom level. In some cases, the second image data may be a second frame acquired before receiving input for capturing a frame. For example, the first image data and the second image may be corresponding first and second frames acquired before receiving input for capturing a frame. The first and second frames may additionally be acquired based on receiving input for capturing a frame before capturing a frame. For example, the second image data may be a preview frame acquired before receiving input for acquiring the captured frame. In some examples, the first and second image data may be associated with a composite image preview frame, and the captured frame may be a composite image capture frame.

[0155] For example, the second image data can be the same as or similar to the following: FIG. 5 One of the corresponding images 310 and 320; FIG. 6 The corresponding one of the image data 360 and 370; FIG. 7 1.5 times the image data 410; FIG. 9 The corresponding one of the 3x image data 502 or the 1x image data 504; FIG. 3A The corresponding one of the 1x image data 602 or the 3x image data 610; FIG. 3A 1x the image data 504; FIG. 3B 1x image data 910; etc.

[0156] In some examples, the first image data can be FIG. 3B Image data 310, and the second image data can be FIG. 4 Image data 320. In some examples, the first image data may be... FIG. 5 The image data is 360, and the second image data can be... FIG. 5 Image data 370.

[0157] In some examples, the first image data can be FIG. 6 Image data 410. In some examples, the first image data may be... FIG. 6 Image data 504, and the second image data can be FIG. 7 Image data 502. In some examples, the first image data may be... FIG. 9 Image data 602, and the second image data can be FIG. 1 Image data 610. In some examples, the first image data may be... FIG. 18 Image data 504. In some examples, the first image data may be... FIG. 5 Image data 910.

[0158] In some aspects, the second image data includes second image data obtained using a second camera having a second focal length (e.g., an image frame output before receiving input for capturing an image). For example, the second image data may be second image preview data obtained using a second camera having a second focal length (e.g., a second image preview frame). In some cases, the second image data is associated with a second camera having a second focal length corresponding to a second zoom level. In some examples, the first camera and the second camera are different, and the first and second cameras are included in the imaging system of a computing device. For example, the first camera and the second camera may be different cameras included in the same computing device, smartphone, mobile computing device, user computing device, etc. In some cases, the first camera and the second camera are included in... FIG. 5 Example computing device 100 and / or FIG. 6 Examples of different cameras in the 1800 computing system or otherwise associated with them.

[0159] In some examples, the second image data is obtained based on adjustments to the zoom level of the foreground portion. For instance, adjusting the zoom level of the foreground portion involves increasing the zoom level, which can be achieved using a zoom level greater than the first zoom level (and a camera with a corresponding focal length). In another example, adjusting the zoom level of the foreground portion involves decreasing the zoom level, which can be achieved using a zoom level less than the first zoom level (and a camera with a corresponding focal length).

[0160] In some cases, obtaining second image data of the scene includes scaling first image data to obtain scaled first image data, wherein the scaled first image data is associated with a second zoom level. For example, the scaled first image data may include a scaled foreground portion corresponding to the foreground portion.

[0161] At box 1708, process 1700 includes generating a segmented foreground portion based on a foreground portion segmented and adjusted from second image data of the scene. For example, a segmentation machine learning network can be used to generate the segmented foreground portion. In some cases, a segmented foreground portion can be used... FIG. 8 The image segmentation and image matting engine 510 is used to generate segmented foreground parts. For example, the segmented foreground parts can be combined with... FIG. 6 The segmented foreground portion 503, FIG. 7 The foreground portion of the segment is 650. FIG. 10 The foreground portion of the segment is the same as or similar to 650.

[0162] In some cases, generating the foreground segmentation includes determining a segmentation map based on second image data, classifying each pixel in a plurality of pixels of the second image data as either foreground or non-foreground. For example, the segmentation map can be compared with... FIG. 5 Segmentation diagram 620 FIG. 5 Segmentation diagram 710 FIG. 5 The segmentation diagrams 12020 or 1030 are the same or similar.

[0163] In some cases, the segmented foreground portion can be generated by multiplying the segmented map by the second image data.

[0164] In some examples, a matting mask corresponding to the segmentation map can be generated, wherein the matting mask has transparency values ​​for at least a portion of the pixels in the segmentation map that are classified as foreground categories. In some cases, the segmented foreground portion can be generated based on combining the segmentation map with the matting mask.

[0165] In some cases, shadow matting information corresponding to the foreground portion can be determined based on second image data of the scene. For example, it can be used... FIG. 5 The shadow matting engine 520 determines the corresponding shadow based on second image data of the scene (e.g., foreground source image 502). FIG. 5 The shadow matting information is used to update the segmented foreground portion of the source image (e.g., the second image) 502, including the shadow matting information. In some examples, the shadow matting information can be used to update the segmented foreground portion to further include pixels of the second image data corresponding to the shadow of the foreground portion. For example, the segmented foreground portion 503 can be updated by the shadow matting engine 520 to include pixels of the second image data 502 corresponding to the shadow of the foreground portion.

[0166] At box 1710, process 1700 includes generating a synthetic image based on combining at least a portion of the segmented foreground portion from second image data of the scene with first image data of the scene. In some cases, it can be used FIG. 5 The image compositing engine 540 is used to generate composite images. In some cases, it can be additionally used... FIG. 3A Image completion and restoration engine 530 FIG. 3B Shadow compositing engine 550 and / or FIG. 4 One or more (or all) of the image harmonization engines 560 are used to generate synthetic images.

[0167] In some examples, the synthesized image can be combined with... FIG. 5 Composite image 330 FIG. 8 380 composite images FIG. 9 Composite image 420 FIG. 10 Composite images 507 or 509 FIG. 11850 composite images FIG. 12 950 composite images FIG. 13 Composite image 1010 or 1050 FIG. 14 Composite image 1110 or 1120 FIG. 3A Composite image 1220 FIG. 3B Composite image 1310 or 1320 FIG. 4 One or more of the composite images 1410 or 1420, composite images (c) or (d) of Figure 15 are the same or similar.

[0168] In some examples, generating the composite image includes generating a preview of the composite image. In some examples, generating the composite image includes outputting a first frame corresponding to the composite image before receiving input for capturing the composite image. For example, generating a preview of the composite image may include displaying a portion of the first image data composited with a portion of the second image data using an image capture user interface (UI). For example, a segmented, adjusted foreground portion having a second zoom level (obtained from the second image data or the second preview frame) may be composited with a background portion having a first zoom level (obtained from the first image data or the first preview frame).

[0169] In some cases, commands for capturing image frames (e.g., input for capturing frames) can be received for capturing a composite image, wherein receiving the command for capturing the image includes receiving user input to the image capture UI. In some examples, the input for capturing frames is a command for capturing a composite image and includes user input to the image capture UI. In some examples, the user input corresponds to a shutter button in the image capture UI. For example, the command for capturing the image could be user input corresponding to a shutter button in the image capture UI, such as in... FIG. 6 and FIG. 9 The shutter button is shown at the bottom center of each example UI 310, 320, 330, 360, 370, and 380; FIG. 10 The shutter button is shown at the bottom center of each example UI 410 and 420; FIG. 11 The shutter button is shown at the bottom center of the example UI 602; FIG. 12 The shutter button is shown at the bottom center of each example UI 910 and 950; FIG. 13 The shutter button is shown at the bottom center of each example UI 1010 and 1050; FIG. 5 The shutter button is shown at the bottom center of each example UI 1110 and 1120; FIG. 5 The shutter button is shown at the bottom center of each example UI 1210 and 1220; FIG. 7The shutter button is shown at the bottom center of each example UI1310 and 1320; and so on.

[0170] In some examples, generating the composite image includes outputting a first frame corresponding to the composite image and receiving input for capturing frames, where the input is received after the first frame is output. In some cases, the first frame may be a preview frame corresponding to the composite image. For example, the first frame may be a preview frame using... FIG. 5 The architecture 500 generates preview frames. Generating a composite image may additionally include outputting captured frames corresponding to the composite image based on receiving input for capturing frames (e.g., commands for capturing images, etc.). The captured frames can be generated using... FIG. 7 The architecture 500 generates composite image capture frames. In some examples, the first frame is a preview frame corresponding to the composite image, and the captured frame is the composite image. The captured frame may be different from the first frame.

[0171] In some cases, the composite image is output before receiving a command for capturing the image. In some cases, a preview of the composite image is output, and a command for capturing a composite image frame corresponding to the preview is received. In some cases, a preview of the composite image is output, and user input instructing adjustment of the zoom level of the foreground portion is received based on the preview of the composite image. In some cases, user input instructing adjustment of one or more of a first zoom level or a second zoom level is received based on the preview of the composite image. In some cases, the composite image is displayed in a preview, wherein the preview includes the composite image and at least a first graphical user interface (GUI) associated with receiving user input instructing adjustment of the zoom level of the foreground portion to increase or decrease.

[0172] In some cases, the preview includes a first GUI overlaid on the composite image. In some examples, the first GUI is collapsible within the preview. In some examples, process 1700 further includes: receiving user input on one or more of the preview or the first GUI; and collapsing the first GUI based on the user input, wherein collapsing the first GUI includes removing the overlay of the first GUI from the preview.

[0173] In some examples, box 1710 also includes removing pixels corresponding to the foreground portion from the first image data based on segmentation information of the foreground portion in the first image data. For example, based on... FIG. 5 The segmentation information of the foreground portion in the same or similar first image data as the 1x main body segmentation image 710 can be obtained from... FIG. 7 and FIG. 5 The first image data 504 removes the foreground portion. Removing the foreground portion from the first image data may include using an image completion engine (e.g., ...). FIG. 5Image completion and restoration engine 530) generates and FIG. 5 Subject removal results in the same or similar backgrounds (720). In some cases, removing the foreground portion from the first image data may include using an image completion engine (e.g., FIG. 7 The image completion and repair engine 530 generates repaired first image data, wherein each removed pixel in the pixels corresponding to the foreground portion of the first image data is replaced with the corresponding repaired pixel.

[0174] In some examples, an image completion engine can be used to generate first image data for restoration, where each removed pixel in the pixels corresponding to the foreground portion of the first image data is replaced with the corresponding restored pixel. For example, the image completion engine can be used with... FIG. 13 The image completion and restoration engine 530 is the same as or similar to the original. In some cases, the first image data to be restored can be the same as... FIG. 13 and FIG. 13 The background result is identical or similar to 505 only if it is 1x the original. In some cases, generating a synthetic image includes: generating an inverted segmentation map based on a segmentation map corresponding to the segmented foreground portion inverted from the second image data; and adding the segmented foreground portion to the product of the inverted segmentation map and the repaired first image data.

[0175] In some cases, the synthesized image includes background image data of the scene associated with a first zoom level and foreground portion image data corresponding to an adjustment of the second zoom level, as well as user input indicating the adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion in the first image data.

[0176] In some cases, additional user input can be received indicating an adjustment to the zoom level for increasing or decreasing the background portion relative to the foreground portion in the first image data. Third image data of the scene can be obtained based on the additional adjustment and using a third zoom level corresponding to the additional adjustment, the third image data including at least the adjusted background portion associated with the third zoom level. The third zoom level may differ from at least one (or both) of the first and second zoom levels. A segmented background portion can be generated based on segmenting the background portion from the third image data of the scene. A composite image can be generated based on combining a segmented foreground portion from the second image data of the scene, a segmented background portion from the third image data of the scene, and a portion of the first image data of the scene. For example, the first zoom level may correspond to... FIG. 11 The x1 main zoom level, the second zoom level can correspond to FIG. 11 The x2 zoom level, and the third zoom level can correspond to FIG. 11 x4 background zoom level.

[0177] In some examples, additional user input can be received, indicating adjustments to the positioning of the foreground portion. For example, the additional user input could be an indication of... FIG. 10 The translation input is adjusted to the same or similar translation adjustment as the left translation adjustment. In some cases, a synthetic image can be generated further based on the translated segmented foreground portion according to additional user input, wherein the segmented foreground portion is partially translated relative to the first image data of the scene. For example, the synthetic image can be compared with... FIG. 10 The translated synthetic image 1120 is the same as or similar to it, and can correspond to FIG. 10 The non-translated synthetic image 1110. In some cases, the translated foreground segmentation is based on a translated segmentation map generated from additional user input corresponding to the segmented foreground portion and indications of adjustments to the positioning of the foreground portion. For example, the segmented foreground portion can be... FIG. 18 The initial body segmentation map x3 is the same as or similar to 1020, and the translated segmentation map can be the same as... FIG. 18 The subject segmentation image 1030 is the same as or similar to the x3 translation. In some cases, the translated composite image can be the same as or similar to the translated composite image 1050, and can correspond to... FIG. 5 Non-translated synthetic image 1010.

[0178] In some cases, a first GUI may be used to receive user input instructing adjustments to the zoom level of a foreground portion relative to the background portion in the first image data. A second GUI may be used to receive user input instructing adjustments to the zoom level of a background portion relative to the foreground portion in the first image data. In some cases, the composite image may also include a first GUI and a second GUI. For example, the composite image may be displayed in a preview (e.g., the composite image may be a preview frame and / or a first frame output before receiving input for capturing an image), wherein the preview includes the composite image, the first GUI, and the second GUI. In some cases, displaying the composite image in a preview includes: outputting preview image data corresponding to the composite image; overlaying the first GUI onto the preview image data; and overlaying the second GUI onto the preview image data. In some examples, one or more of the first GUI or the second GUI are collapsible within the preview. In some cases, the composite image may also include a third GUI element. The third GUI element may include a capture icon associated with capturing the composite image (e.g., associated with receiving input for capturing an image).

[0179] In some examples, the processes described herein (e.g., process 1700 and / or any other processes described herein) may be performed by a computing device, apparatus, or system. In one example, process 1700 may be performed by a device having… ​The computing device architecture 1800 is used to perform the computing device or system described herein. The computing device, apparatus, or system may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a connected watch or smartwatch, or other wearable devices), a server computer, a computing device for autonomous vehicles or autonomous vehicles, a robotic device, a laptop computer, a smart TV, a camera, and / or any other computing device with the resource capability to perform the processes described herein (including process 1700 and / or any other processes described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0180] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0181] Process 1700 is illustrated as a logic flowchart, the operations of which represent a sequence of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.

[0182] Additionally, process 1700 and / or any other process described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0183] ​ Example computing device architecture 1800 illustrates example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1800 may implement... ​ The components of the computing device architecture 1800 are shown to communicate electrically with each other using a connection 1805 (such as a bus). The example computing device architecture 1800 includes a processing unit (CPU or processor) 1810 and a computing device connection 1805 that couples various computing device components, including computing device memories 1815 (such as read-only memory (ROM) 1820 and random access memory (RAM) 1825), to the processor 1810.

[0184] The computing device architecture 1800 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1810. The computing device architecture 1800 may copy data from memory 1815 and / or storage device 1830 to cache 1812 for fast access by the processor 1810. In this way, the cache can provide performance improvements by avoiding latency for the processor 1810 while waiting for data. These and other engines can control or be configured to control the processor 1810 to perform various actions. Other computing device memory 1815 may also be used. Memory 1815 may include various different types of memory with different performance characteristics. The processor 1810 may include any general-purpose processor and hardware or software services configured to control the processor 1810 (such as services 1 1832, 2 1834, and 3 1836 stored in storage device 1830), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1810 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.

[0185] To enable user interaction with the computing device architecture 1800, the input device 1845 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The output device 1835 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some instances, a multimodal computing device allows the user to provide multiple types of input to communicate with the computing device architecture 1800. The communication interface 1840 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0186] Storage device 1830 is non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape cassettes, flash memory cards, solid-state storage devices, digital universal optical discs, magnetic tape cartridges, random access memory (RAM) 1825, read-only memory (ROM) 1820, and combinations thereof. Storage device 1830 may include services 1832, 1834, 1836 for controlling processor 1810. Other hardware or software modules or engines are envisioned. Storage device 1830 may be connected to computing device connection 1805. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 1810, connection 1805, output device 1835, etc., to perform that function.

[0187] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0188] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components, or a particular aspect, or example. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0189] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects and examples can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring aspects and examples in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects and examples.

[0190] The aspects and examples described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While a flowchart may describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0191] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0192] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as flash memory), memory or memory devices, magnetic disks or optical discs, flash memory, USB devices provided with non-volatile memory, network storage devices, compressed optical discs (CDs) or digital versatile optical discs (DVDs), any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, engines, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0193] In some aspects and examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.

[0194] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0195] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0196] In the foregoing description, aspects of this application have been described with reference to specific aspects and examples thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, while illustrative aspects and examples of this application have been described in detail herein, it should be understood that the inventive concept can be implemented and employed in a variety of other ways, and the appended claims are intended to be construed as including these variations, unless limited by prior art. Various features and aspects of the above applications may be used individually or in combination. Furthermore, aspects and examples may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative aspects and examples, the methods may be performed in a different order than that described.

[0197] Those skilled in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein can be represented by less than or equal to ("<"), respectively. ") and greater than or equal to (" The symbol '(')' is used to replace the existing description without deviating from its scope.

[0198] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0199] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0200] The various exemplary logic blocks, modules, engines, circuits, and algorithm steps described in conjunction with the aspects and examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0201] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0202] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0203] The claim language or other language that expresses "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that expresses "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language that expresses "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repeating information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0204] Claims or other languages ​​that specify "at least one processor, the at least one processor is configured to," "at least one processor is configured to," "one or more processors, the one or more processors are configured to," or "one or more processors are configured to," indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim stating "at least one processor is configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks of operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim stating "at least one processor is configured to: X, Y, and Z" could mean that any single processor can perform only at least one subset of operations X, Y, and Z.

[0205] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0206] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).

[0207] The exemplary aspects of this disclosure include:

[0208] Aspect 1. A method comprising: obtaining first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; receiving user input instructing an adjustment of the zoom level relative to the background portion included in the first image data, by increasing or decreasing the foreground portion; obtaining second image data of the scene based on the adjustment and using a second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; generating a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and generating a composite image based on combining the segmented foreground portion from the second image data of the scene with at least a portion of the first image data of the scene.

[0209] Aspect 2. The method according to aspect 1, the method further comprising: receiving a command for capturing an image frame corresponding to the synthesized image.

[0210] Aspect 3. The method according to any one of Aspects 1 to 2, wherein generating the composite image comprises: outputting a first frame corresponding to the composite image; and receiving input for capturing a frame, wherein the input is received after outputting the first frame.

[0211] Aspect 4. The method according to aspect 3, wherein generating the composite image further comprises: outputting a captured frame corresponding to the composite image based on the input received for capturing the frame.

[0212] Aspect 5. The method according to aspect 4, wherein the first frame is a preview frame corresponding to the composite image, and wherein the captured frame is the composite image.

[0213] Aspect 6. The method according to any one of Aspects 4 to 5, wherein the captured frame is different from the first frame.

[0214] Aspect 7. The method according to any one of Aspects 1 to 6, wherein generating the composite image comprises: outputting a first frame corresponding to the composite image; and receiving input for capturing a frame, wherein the input is received after outputting the first frame.

[0215] Aspect 8. The method according to aspect 7, wherein: the first image data includes first image data obtained using a first camera having a first focal length; and the second image data includes second image data obtained using a second camera having a second focal length.

[0216] Aspect 9. The method according to aspect 8, wherein the first image data and the second image data are obtained before receiving input for capturing frames.

[0217] Aspect 10. The method according to aspect 9, wherein the first image data is associated with a preview frame obtained using the first camera, and wherein the second image data is associated with a preview frame obtained using the second camera.

[0218] Aspect 11. The method according to any one of Aspects 8 to 10, wherein outputting the first frame comprises: displaying a portion of the first image data synthesized with a portion of the second image data using an image capture user interface (UI).

[0219] Aspect 12. The method according to aspect 11, wherein the input for capturing the frame is a command for capturing the composite image and includes user input for the image capture UI.

[0220] Aspect 13. The method according to aspect 12, wherein the user input corresponds to the shutter button of the image capture UI.

[0221] Aspect 14. The method according to any one of aspects 1 to 13, the method further comprising outputting the synthesized image before receiving a command for capturing an image.

[0222] Aspect 15. The method according to any one of aspects 1 to 14, the method further comprising: outputting a preview of the composite image; and receiving a command for capturing a composite image frame corresponding to the preview of the composite image.

[0223] Aspect 16. The method according to any one of Aspects 1 to 15, the method further comprising: outputting a preview of the composite image; and receiving, based on the preview of the composite image, the user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion.

[0224] Aspect 17. The method according to aspect 16, the method further comprising: receiving user input indicating an adjustment for increasing or decreasing one or more of the first zoom level or the second zoom level based on the preview of the synthesized image.

[0225] Aspect 18. The method according to any one of Aspects 1 to 17, the method further comprising: displaying the composite image in a preview, wherein the preview includes the composite image and at least a first graphical user interface (GUI) associated with receiving user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion.

[0226] Aspect 19. The method according to aspect 18, wherein the preview includes the first GUI superimposed on the composite image.

[0227] Aspect 20. The method according to any one of Aspects 18 to 19, wherein the first GUI is collapsible within the preview.

[0228] Aspect 21. The method according to aspect 20, the method further comprising: receiving user input on one or more of the preview or the first GUI; and collapsing the first GUI based on the user input, wherein collapsing the first GUI includes removing the overlay of the first GUI from the preview.

[0229] Aspect 22. The method according to any one of Aspects 1 to 21, wherein generating the foreground portion of the segmentation comprises: determining a segmentation map based on the second image data, classifying each pixel of a plurality of pixels in the second image data as either a foreground category or a non-foreground category; and multiplying the segmentation map by the second image data.

[0230] Aspect 23. The method according to aspect 22, the method further comprising: generating a matting mask corresponding to the segmentation map, wherein the matting mask has transparency values ​​for at least a portion of a plurality of pixels of the segmentation map classified as the foreground category; and generating a foreground portion of the segmentation based on combining the segmentation map with the matting mask.

[0231] Aspect 24. The method according to any one of Aspects 1 to 23, the method further comprising: determining shadow matting information corresponding to the shadow of the foreground portion based on the second image data of the scene; and updating the segmented foreground portion using the shadow matting information to further include pixels of the second image data corresponding to the shadow of the foreground portion.

[0232] Aspect 25. The method according to any one of Aspects 1 to 24, the method further comprising: removing pixels corresponding to the foreground portion from the first image data based on segmentation information of the foreground portion in the first image data; and generating repaired first image data using an image completion engine, wherein each removed pixel among the pixels corresponding to the foreground portion in the first image data is replaced with a corresponding repaired pixel.

[0233] Aspect 26. The method according to aspect 25, wherein generating the synthetic image comprises: generating an inverted segmentation map based on inverting a segmentation map corresponding to the segmented foreground portion from the second image data; and adding the segmented foreground portion to the product of the inverted segmentation map and the repaired first image data.

[0234] Aspect 27. The method according to aspect 26, wherein the synthesized image includes background image data of the scene associated with the first zoom level and foreground portion image data corresponding to the adjustment of the second zoom level, as well as the user input indicating the adjustment.

[0235] Aspect 28. The method according to any one of Aspects 1 to 27, wherein the second image data is obtained based on the adjustment for increasing or decreasing the zoom level relative to the background portion.

[0236] Aspect 29. The method according to any one of Aspects 1 to 28, wherein: the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for increasing the zoom level of the foreground portion relative to the background portion included in the first image data. And the second zoom level is greater than the first zoom level.

[0237] Aspect 30. The method according to any one of Aspects 1 to 29, wherein: the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for decreasing the zoom level of the foreground portion relative to the background portion included in the first image data. And the second zoom level is smaller than the first zoom level.

[0238] Aspect 31. The method according to any one of Aspects 1 to 30, wherein receiving the user input instructing the adjustment of the zoom level of the foreground portion relative to the background portion comprises: receiving a first user input instructing the selection of a foreground portion from one or more foreground portions included in the first image data of the scene; and receiving a second user input instructing the adjustment of the zoom level of the selected foreground portion relative to the background portion.

[0239] Aspect 32. The method according to any one of Aspects 1 to 31, the method further comprising: receiving user input instructing an additional adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data; obtaining third image data of the scene based on the additional adjustment and using a third zoom level corresponding to the additional adjustment, the third image data including at least the adjusted background portion associated with the third zoom level; generating a segmented background portion based on segmenting the adjusted background portion from the third image data of the scene; and generating the composite image based on combining the segmented foreground portion from the second image data of the scene with the segmented background portion from the third image data of the scene and a portion of the first image data of the scene.

[0240] Aspect 33. The method according to any one of Aspects 1 to 32, the method further comprising: receiving additional user input indicating an adjustment of the positioning of the foreground portion; and further generating the composite image based on translating the segmented foreground portion based on the additional user input, wherein the segmented foreground portion is translated relative to the portion of the first image data of the scene.

[0241] Aspect 34. The method according to aspect 33, wherein translating the segmented foreground portion is based on generating a translated segmentation map corresponding to the segmented foreground portion and the additional user input indicating the adjustment of the positioning of the foreground portion.

[0242] Aspect 35. The method according to any one of Aspects 1 to 34, wherein: the first image data is associated with a first camera having a first focal length corresponding to the first zoom level; and the second image data is associated with a second camera having a second focal length corresponding to the second zoom level.

[0243] Aspect 36. The method according to aspect 35, wherein the first camera is different from the second camera, and wherein the first camera and the second camera are included in an imaging system of a computing device.

[0244] Aspect 37. The method according to any one of Aspects 1 to 36, wherein obtaining the second image data of the scene comprises: scaling the first image data to obtain scaled first image data, wherein the scaled first image data is associated with the second zoom level, and wherein the scaled first image data includes a scaled foreground portion corresponding to the foreground portion.

[0245] Aspect 38. The method according to any one of aspects 1 to 37, wherein the user input indicating the adjustment of the zoom level of the foreground portion is received in a graphical user interface (GUI).

[0246] Aspect 39. The method according to aspect 38, wherein the GUI includes a slider, wherein moving the slider in a first direction indicates an increase in the zoom level, and moving the slider in a second direction indicates a decrease in the zoom level.

[0247] Aspect 40. The method according to any one of Aspects 38 to 39, wherein the GUI includes a plurality of discrete pace adjustments, each of the plurality of discrete pace adjustments corresponding to an increase or decrease in the configuration of the zoom level.

[0248] Aspect 41. The method according to any one of Aspects 1 to 40, the method further comprising: receiving, using a first graphical user interface (GUI), the user input indicating an adjustment of the zoom level of the foreground portion; and receiving, using a second GUI, the user input indicating an adjustment of the zoom level of the background portion.

[0249] Aspect 42. The method according to aspect 41, wherein the synthesized image further includes the first GUI and the second GUI.

[0250] Aspect 43. The method according to aspect 42, the method further comprising: displaying the composite image in a preview, wherein the preview includes the composite image, the first GUI and the second GUI.

[0251] Aspect 44. The method according to aspect 43, wherein displaying the composite image in the preview comprises: outputting preview image data corresponding to the composite image; overlaying the first GUI onto the preview image data; and overlaying the second GUI onto the preview image data.

[0252] Aspect 45. The method according to any one of Aspects 43 to 44, wherein one or more of the first GUI or the second GUI are collapsible within the preview.

[0253] Aspect 46. The method according to any one of aspects 43 to 45, wherein the synthesized image further comprises a third GUI element.

[0254] Aspect 47. The method according to aspect 46, wherein the third GUI element includes a capture icon associated with capturing the composite image.

[0255] Aspect 48. The method according to any one of Aspects 1 to 47, the method further comprising: receiving user input indicating an adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data; automatically determining a corresponding adjustment for increasing or decreasing the zoom level of the foreground portion, wherein the corresponding adjustment is automatically determined relative to the user input indicating the adjustment for increasing or decreasing the zoom level of the background portion; and generating the composite image based on the adjustment for increasing or decreasing the zoom level of the background portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the foreground portion.

[0256] Aspect 49. The method according to any one of Aspects 1 to 48, the method further comprising: automatically determining a corresponding adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data based on the user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion; and generating the composite image based on the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the foreground portion and the zoom level of the background portion.

[0257] Aspect 50. An apparatus for processing image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain first image data of a scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; receive user input instructing an adjustment of the zoom level for increasing or decreasing the foreground portion relative to the background portion included in the first image data; obtain second image data of the scene based on the adjustment and using a second zoom level, the second image data including at least the adjusted foreground portion associated with the second zoom level; generate a segmented foreground portion based on segmenting the adjusted foreground portion from the second image data of the scene; and generate a composite image based on combining the segmented foreground portion from the second image data of the scene with at least a portion of the first image data of the scene.

[0258] Aspect 51. The apparatus according to aspect 50, wherein the at least one processor is further configured to: receive a command for capturing an image frame corresponding to the synthesized image.

[0259] Aspect 52. The apparatus according to any one of Aspects 50 to 51, wherein, in order to generate the composite image, the at least one processor is configured to: output a first frame corresponding to the composite image; and receive input for capturing a frame, wherein the input is received after the output of the first frame.

[0260] Aspect 53. The apparatus according to aspect 52, wherein, in order to generate the composite image, the at least one processor is further configured to: output a captured frame corresponding to the composite image, wherein the captured frame is output based on the input for capturing the frame.

[0261] Aspect 54. The apparatus according to aspect 53, wherein the first frame is a preview frame corresponding to the composite image, and wherein the captured frame is the composite image.

[0262] Aspect 55. The apparatus according to any one of aspects 53 to 54, wherein the captured frame is different from the first frame.

[0263] Aspect 56. The apparatus according to any one of aspects 50 to 55, wherein, in order to generate the composite image, the at least one processor is configured to: output a first frame corresponding to the composite image; and receive input for capturing a frame, wherein the input is received after the output of the first frame.

[0264] Aspect 57. The apparatus according to aspect 56, wherein: the first image data includes first image data obtained using a first camera having a first focal length; and the second image data includes second image data obtained using a second camera having a second focal length.

[0265] Aspect 58. The apparatus according to aspect 57, wherein the first image data and the second image data are obtained before receiving input for capturing frames.

[0266] Aspect 59. The apparatus according to aspect 58, wherein the first image data is associated with a preview frame obtained using the first camera, and wherein the second image data is associated with a preview frame obtained using the second camera.

[0267] Aspect 60. The apparatus according to any one of aspects 57 to 59, wherein, in order to output the first frame, the at least one processor is configured to: display a portion of the first image data synthesized with a portion of the second image data using an image capture user interface (UI).

[0268] Aspect 61. The apparatus according to aspect 60, wherein the input for capturing a frame is a command for capturing the composite image and includes user input for the image capture UI.

[0269] Aspect 62. The apparatus according to aspect 61, wherein the user input corresponds to a shutter button of the image capture UI.

[0270] Aspect 63. The apparatus according to any one of aspects 50 to 62, wherein the at least one processor is further configured to output the composite image before receiving a command for capturing an image.

[0271] Aspect 64. The apparatus according to any one of aspects 50 to 63, wherein the at least one processor is further configured to: output a preview of the composite image; and receive a command for capturing a composite image frame corresponding to the preview of the composite image.

[0272] Aspect 65. The apparatus according to any one of aspects 50 to 64, wherein the at least one processor is further configured to: output a preview of the composite image; and receive, based on the preview of the composite image, the user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion.

[0273] Aspect 66. The apparatus according to aspect 65, wherein the at least one processor is further configured to: receive user input indicating an adjustment for increasing or decreasing one or more of the first zoom level or the second zoom level based on the preview of the composite image.

[0274] Aspect 67. The apparatus according to any one of aspects 50 to 66, wherein the at least one processor is further configured to: display the composite image in a preview, wherein the preview includes the composite image and at least a first graphical user interface (GUI) associated with receiving user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion.

[0275] Aspect 68. The apparatus according to aspect 67, wherein the preview includes the first GUI superimposed on the composite image.

[0276] Aspect 69. The apparatus according to any one of Aspects 67 to 68, wherein the first GUI is foldable within the preview.

[0277] Aspect 70. The apparatus according to aspect 69, wherein the at least one processor is further configured to: receive user input to one or more of the preview or the first GUI; and collapse the first GUI based on the user input, wherein collapsing the first GUI includes removing the overlay of the first GUI from the preview.

[0278] Aspect 71. The apparatus according to any one of aspects 50 to 70, wherein, in order to generate the segmented foreground portion, the at least one processor is configured to: determine a segmentation map based on the second image data, classifying each pixel of a plurality of pixels in the second image data as either a foreground category or a non-foreground category; and multiply the segmentation map by the second image data.

[0279] Aspect 72. The apparatus according to aspect 71, wherein the at least one processor is further configured to: generate a matting mask corresponding to the segmentation map, wherein the matting mask has transparency values ​​for at least a portion of a plurality of pixels of the segmentation map classified as the foreground category; and generate a foreground portion of the segmentation based on combining the segmentation map with the matting mask.

[0280] Aspect 73. The apparatus according to any one of Aspects 50 to 72, wherein the at least one processor is further configured to: determine shadow matting information corresponding to the shadow of the foreground portion based on the second image data of the scene; and update the segmented foreground portion using the shadow matting information to further include pixels of the second image data corresponding to the shadow of the foreground portion.

[0281] Aspect 74. The apparatus according to any one of Aspects 50 to 73, wherein the at least one processor is further configured to: remove pixels corresponding to the foreground portion from the first image data based on segmentation information of the foreground portion in the first image data; and generate repaired first image data using an image completion engine, wherein each removed pixel of the pixels corresponding to the foreground portion in the first image data is replaced with a corresponding repaired pixel.

[0282] Aspect 75. The apparatus according to aspect 74, wherein generating the synthetic image comprises: generating an inverted segmentation map based on inverting a segmentation map corresponding to the segmented foreground portion from the second image data; and adding the segmented foreground portion to the product of the inverted segmentation map and the repaired first image data.

[0283] Aspect 76. The apparatus according to aspect 75, wherein the synthesized image includes background image data of the scene associated with the first zoom level and foreground portion image data corresponding to an adjustment of the second zoom level, as well as the user input indicating the adjustment.

[0284] Aspect 77. The apparatus according to any one of aspects 50 to 76, wherein the second image data is obtained based on the adjustment for increasing or decreasing the zoom level relative to the background portion.

[0285] Aspect 78. The apparatus according to any one of Aspects 50 to 77, wherein: the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for increasing the zoom level of the foreground portion relative to the background portion included in the first image data. And the second zoom level is greater than the first zoom level.

[0286] Aspect 79. The apparatus according to any one of Aspects 50 to 78, wherein: the adjustment for increasing or decreasing the zoom level of the foreground portion is an adjustment for decreasing the zoom level of the foreground portion relative to the background portion included in the first image data. And the second zoom level is smaller than the first zoom level.

[0287] Aspect 80. The apparatus according to any one of aspects 50 to 79, wherein, in order to receive the user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion relative to the background portion, the at least one processor is configured to: receive a first user input indicating a foreground portion selected from one or more foreground portions included in the first image data of the scene; and receive a second user input indicating an adjustment for increasing or decreasing the zoom level of the selected foreground portion relative to the background portion.

[0288] Aspect 81. The apparatus according to any one of Aspects 50 to 80, wherein the at least one processor is further configured to: receive user input instructing an additional adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data; obtain third image data of the scene based on the additional adjustment and using a third zoom level corresponding to the additional adjustment, the third image data including at least the adjusted background portion associated with the third zoom level; generate a segmented background portion based on segmenting the adjusted background portion from the third image data of the scene; and generate the composite image based on combining the segmented foreground portion from the second image data of the scene with the segmented background portion from the third image data of the scene and a portion of the first image data of the scene.

[0289] Aspect 82. The apparatus according to any one of Aspects 50 to 81, wherein the at least one processor is further configured to: receive additional user input indicating an adjustment of the positioning of the foreground portion; and further generate the composite image based on translating the segmented foreground portion based on the additional user input, wherein the segmented foreground portion is translated relative to the portion of the first image data of the scene.

[0290] Aspect 83. The apparatus according to aspect 82, wherein, in order to translate the segmented foreground portion, the at least one processor is configured to generate a segmentation map corresponding to the segmented foreground portion and the additional user input indicating the adjustment of the positioning of the foreground portion.

[0291] Aspect 84. The apparatus according to any one of Aspects 50 to 83, wherein: the first image data is associated with a first camera having a first focal length corresponding to the first zoom level; and the second image data is associated with a second camera having a second focal length corresponding to the second zoom level.

[0292] Aspect 85. The apparatus according to aspect 84, wherein the first camera is different from the second camera, and wherein the first camera and the second camera are included in the imaging system of the computing device.

[0293] Aspect 86. The apparatus according to any one of Aspects 50 to 85, wherein, in order to obtain the second image data of the scene, the at least one processor is configured to: scale the first image data to obtain scaled first image data, wherein the scaled first image data is associated with the second zoom level, and wherein the scaled first image data includes a scaled foreground portion corresponding to the foreground portion.

[0294] Aspect 87. The apparatus according to any one of aspects 50 to 86, wherein the user input indicating the adjustment of the zoom level of the foreground portion is received in a graphical user interface (GUI).

[0295] Aspect 88. The apparatus according to aspect 87, wherein the GUI includes a slider, wherein moving the slider in a first direction indicates an increase in the zoom level, and moving the slider in a second direction indicates a decrease in the zoom level.

[0296] Aspect 89. The apparatus according to aspect 88, wherein the GUI includes a plurality of discrete timing adjustments, each of the plurality of discrete timing adjustments corresponding to an increase or decrease in the configuration of the zoom level.

[0297] Aspect 90. The apparatus according to any one of aspects 50 to 89, wherein the at least one processor is further configured to: receive, using a first graphical user interface (GUI), the user input indicating the adjustment of the zoom level of the foreground portion; and receive, using a second GUI, the user input indicating the adjustment of the zoom level of the background portion.

[0298] Aspect 91. The apparatus according to aspect 90, wherein the synthesized image further includes the first GUI and the second GUI.

[0299] Aspect 92. The apparatus according to aspect 91, wherein the at least one processor is further configured to: display the composite image in a preview, wherein the preview includes the composite image, the first GUI, and the second GUI.

[0300] Aspect 93. The apparatus according to aspect 92, wherein, in order to display the composite image in the preview, the at least one processor is configured to: output preview image data corresponding to the composite image; overlay the first GUI onto the preview image data; and overlay the second GUI onto the preview image data.

[0301] Aspect 94. The apparatus according to any one of aspects 92 to 93, wherein one or more of the first GUI or the second GUI are collapsible within the preview.

[0302] Aspect 95. The apparatus according to any one of aspects 92 to 94, wherein the synthesized image further comprises a third GUI element.

[0303] Aspect 96. The apparatus according to aspect 95, wherein the third GUI element includes a capture icon associated with capturing the composite image.

[0304] Aspect 97. The apparatus according to any one of Aspects 50 to 96, wherein the at least one processor is further configured to: receive user input indicating an adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data; automatically determine a corresponding adjustment for increasing or decreasing the zoom level of the foreground portion, wherein the corresponding adjustment is automatically determined relative to the user input indicating the adjustment for increasing or decreasing the zoom level of the background portion; and generate the composite image based on the adjustment for increasing or decreasing the zoom level of the background portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the foreground portion.

[0305] Aspect 98. The apparatus according to any one of Aspects 50 to 97, wherein the at least one processor is further configured to: automatically determine a corresponding adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data based on receiving user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion; and generate the composite image based on the automatically determined corresponding adjustment for increasing or decreasing the zoom level of the foreground portion and the zoom level of the background portion.

[0306] Aspect 99. A non-transitory computer-readable storage medium comprising instructions stored thereon, the instructions causing the at least one processor, when executed by at least one processor, to perform an operation according to any one of aspects 1 to 49, and / or 103.

[0307] Aspect 100. A non-transitory computer-readable storage medium comprising instructions stored thereon, the instructions causing the at least one processor, when executed by at least one processor, to perform an operation according to any one of aspects 50 to 98 and / or 104.

[0308] Aspect 101. An apparatus comprising one or more components for performing operations according to any one of aspects 1 to 49 and / or 103.

[0309] Aspect 102. An apparatus comprising one or more components for performing operations according to any one of aspects 50 to 98 and / or 104.

[0310] Aspect 103. The method according to any one of Aspects 1 to 49, wherein the adjustment for increasing or decreasing the zoom level of the foreground portion relative to the background portion corresponds to a second zoom level greater than the first zoom level or a second zoom level less than the first zoom level.

[0311] Aspect 104. The apparatus according to any one of Aspects 50 to 98, wherein the adjustment for increasing or decreasing the zoom level of the foreground portion relative to the background portion corresponds to a second zoom level greater than the first zoom level or a second zoom level less than the first zoom level.

Claims

1. A method, the method comprising: Obtain first image data of the scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; Receive user input instructing an adjustment of the zoom level of the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than the first zoom level or a second zoom level less than the first zoom level; Based on the adjustment, a second image data of the scene is obtained using the second zoom level, the second image data including at least the foreground portion of the adjustment associated with the second zoom level; The segmented foreground portion is generated based on the segmentation of the adjusted foreground portion from the second image data of the scene; as well as A synthetic image is generated by combining the segmented foreground portion of the second image data from the scene with at least a portion of the first image data from the scene.

2. The method according to claim 1, further comprising: The output corresponds to the first frame of the synthesized image; Receive a command for capturing an image frame corresponding to the synthesized image, wherein the first frame is output before receiving the command for capturing the image frame; and Based on the received command for capturing the image frame, output the captured frame corresponding to the synthesized image.

3. The method according to claim 2, wherein: The first frame is a preview frame corresponding to the synthesized image; The captured frame is the synthetic image; and The captured frame is different from the first frame.

4. The method according to claim 1, wherein: The first image data includes first image data obtained using a first camera with a first focal length; and The second image data includes second image data obtained using a second camera with a second focal length.

5. The method of claim 4, wherein generating the synthesized image comprises: The image capture user interface (UI) outputs a first frame corresponding to the synthesized image, wherein the first frame includes a portion of the first image data synthesized with a portion of the second image data; as well as Receive input for capturing frames, wherein the input is received after the first frame is output and includes user input for the image capture UI.

6. The method according to claim 1, further comprising: Output a preview of the synthesized image; as well as Based on the preview of the synthesized image, the user input indicating the adjustment for increasing or decreasing the zoom level of the foreground portion is received; or The user input, based on the preview of the synthesized image, is used to indicate adjustments for increasing or decreasing one or more of the first zoom level or the second zoom level.

7. The method of claim 1, wherein the second image data is obtained based on the adjustment for increasing or decreasing the zoom level relative to the background portion.

8. The method of claim 1, wherein receiving the user input instructing the adjustment of the zoom level of the foreground portion relative to the background portion comprises: Receive first user input instructing the selection of a foreground portion from one or more foreground portions included in the first image data of the scene; as well as Receive a second user input instructing the adjustment of the zoom level of the selected foreground portion relative to the background portion.

9. The method according to claim 1, further comprising: The user input, which indicates the adjustment of the zoom level for increasing or decreasing the foreground portion, is received using a first graphical user interface (GUI). as well as The second GUI receives user input instructing users to adjust the zoom level of the background portion, either increasing or decreasing it.

10. A non-transitory computer-readable medium having instructions stored thereon, the instructions causing the one or more processors, when executed, to: Obtain first image data of the scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; Receive user input instructing an adjustment of the zoom level of the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than the first zoom level or a second zoom level less than the first zoom level; Based on the adjustment, a second image data of the scene is obtained using the second zoom level, the second image data including at least the foreground portion of the adjustment associated with the second zoom level; The segmented foreground portion is generated based on the segmentation of the adjusted foreground portion from the second image data of the scene; as well as A synthetic image is generated by combining the segmented foreground portion of the second image data from the scene with at least a portion of the first image data from the scene.

11. An apparatus for processing image data, the apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, the at least one processor being configured to: Obtain first image data of the scene, the first image data being associated with a first zoom level and including at least a foreground portion and a background portion; Receive user input instructing an adjustment of the zoom level of the foreground portion relative to the background portion included in the first image data, wherein the adjustment corresponds to a second zoom level greater than the first zoom level or a second zoom level less than the first zoom level; Based on the adjustment, a second image data of the scene is obtained using the second zoom level, the second image data including at least the foreground portion of the adjustment associated with the second zoom level; The segmented foreground portion is generated based on the segmentation of the adjusted foreground portion from the second image data of the scene; as well as A synthetic image is generated by combining the segmented foreground portion of the second image data from the scene with at least a portion of the first image data from the scene.

12. The apparatus of claim 11, wherein the at least one processor is further configured to: Receive a command for capturing an image frame corresponding to the synthesized image.

13. The apparatus according to claim 12, wherein, In order to generate the synthesized image, the at least one processor is configured to: The output corresponds to the first frame of the composite image, wherein the first frame is output before receiving the command for capturing the image frame corresponding to the composite image.

14. The apparatus according to claim 13, wherein, In order to generate the synthesized image, the at least one processor is further configured to: The output corresponds to a captured frame of the synthesized image, wherein the captured frame is output based on the received command for capturing the image frame.

15. The apparatus of claim 14, wherein the first frame is a preview frame corresponding to the composite image, and wherein the captured frame is the composite image.

16. The apparatus of claim 14, wherein the captured frame is different from the first frame.

17. The apparatus according to claim 11, wherein, In order to generate the synthesized image, the at least one processor is configured to: The output corresponds to the first frame of the synthesized image; and Receive input for capturing frames, wherein the input is received after the output of the first frame.

18. The apparatus according to claim 17, wherein: The first image data includes first image data obtained using a first camera with a first focal length; and The second image data includes second image data obtained using a second camera with a second focal length.

19. The apparatus of claim 18, wherein the first image data and the second image data are obtained prior to receiving input for capturing a frame.

20. The apparatus of claim 19, wherein the first image data is associated with a preview frame obtained using the first camera, and wherein the second image data is associated with a preview frame obtained using the second camera.

21. The apparatus according to claim 18, wherein, In order to output the first frame, the at least one processor is configured to: The image capture user interface (UI) is used to display a portion of the first image data that is composited with a portion of the second image data.

22. The apparatus of claim 21, wherein the input for capturing the frame is a command for capturing the composite image and includes user input for the image capture UI.

23. The apparatus of claim 22, wherein the user input corresponds to a shutter button of the image capture UI.

24. The apparatus of claim 11, wherein the at least one processor is further configured to: Output a preview of the synthesized image; and The user input, based on the preview of the synthesized image, indicates the adjustment for increasing or decreasing the zoom level of the foreground portion.

25. The apparatus of claim 24, wherein the at least one processor is further configured to: The user input, based on the preview of the synthesized image, is used to indicate adjustments for increasing or decreasing one or more of the first zoom level or the second zoom level.

26. The apparatus of claim 11, wherein the at least one processor is further configured to: The composite image is displayed in a preview, wherein the preview includes the composite image and at least a first graphical user interface (GUI) associated with receiving user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion.

27. The apparatus of claim 26, wherein the preview includes the first GUI superimposed on the composite image.

28. The apparatus of claim 11, wherein the at least one processor is further configured to: Based on the segmentation information of the foreground portion in the first image data, pixels corresponding to the foreground portion are removed from the first image data; and The image completion engine generates the first image data for repair, wherein each removed pixel in the pixels corresponding to the foreground portion of the first image data is replaced with the corresponding repair pixel.

29. The apparatus of claim 11, wherein the second image data is obtained based on the adjustment for increasing or decreasing the zoom level relative to the background portion.

30. The apparatus according to claim 11, wherein, In order to receive the user input instructing the adjustment of the zoom level relative to the background portion or the foreground portion, the at least one processor is configured to: Receive first user input instructing the selection of a foreground portion from one or more foreground portions included in the first image data of the scene; as well as Receive a second user input instructing the adjustment of the zoom level of the selected foreground portion relative to the background portion.

31. The apparatus of claim 11, wherein the at least one processor is further configured to: Receive user input instructing additional adjustments to the zoom level of the background portion relative to the foreground portion included in the first image data, either increasing or decreasing the zoom level. Based on the additional adjustment, a third image data of the scene is obtained using a third zoom level corresponding to the additional adjustment, the third image data including at least a background portion of the adjustment associated with the third zoom level; The segmented background portion is generated by segmenting the adjusted background portion from the third image data of the scene; as well as The composite image is generated by combining the segmented foreground portion of the second image data from the scene with the segmented background portion of the third image data from the scene and a portion of the first image data from the scene.

32. The apparatus of claim 11, wherein the at least one processor is further configured to: Receive additional user input instructing adjustments to the positioning of the foreground portion; and The composite image is further generated by translating the segmented foreground portion based on the additional user input, wherein the segmented foreground portion is translated relative to the portion of the first image data of the scene.

33. The apparatus according to claim 11, wherein, In order to obtain the second image data of the scene, the at least one processor is configured to: The first image data is scaled to obtain scaled first image data, wherein the scaled first image data is associated with the second zoom level, and wherein the scaled first image data includes a scaled foreground portion corresponding to the foreground portion.

34. The apparatus of claim 11, wherein the user input indicating the adjustment for increasing or decreasing the zoom level of the foreground portion is received in a graphical user interface (GUI).

35. The apparatus of claim 34, wherein the GUI includes a slider, wherein moving the slider in a first direction indicates an increase in the zoom level, and moving the slider in a second direction indicates a decrease in the zoom level.

36. The apparatus of claim 35, wherein the GUI includes a plurality of discrete timing adjustments, each of the plurality of discrete timing adjustments corresponding to an increase or decrease in the configuration of the zoom level.

37. The apparatus of claim 11, wherein the at least one processor is further configured to: The user input, which indicates an adjustment for increasing or decreasing the zoom level of the foreground portion, is received using a first graphical user interface (GUI); and The second GUI receives user input instructing users to adjust the zoom level of the background portion, either increasing or decreasing it.

38. The apparatus of claim 11, wherein the at least one processor is further configured to: Receive user input instructing the user to adjust the zoom level of the background portion relative to the foreground portion included in the first image data, either by increasing or decreasing the zoom level. Automatically determine a corresponding adjustment for increasing or decreasing the zoom level of the foreground portion, wherein the corresponding adjustment is automatically determined relative to the user input indicating the adjustment for increasing or decreasing the zoom level of the background portion; and The composite image is generated based on the adjustment of the zoom level for increasing or decreasing the background portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level for the foreground portion.

39. The apparatus of claim 11, wherein the at least one processor is further configured to: Automatically determine a corresponding adjustment for increasing or decreasing the zoom level of the background portion relative to the foreground portion included in the first image data, based on the user input indicating an adjustment for increasing or decreasing the zoom level of the foreground portion; and The composite image is generated based on the adjustment of the zoom level for increasing or decreasing the foreground portion and the automatically determined corresponding adjustment for increasing or decreasing the zoom level for the background portion.