Post-capture photo viewpoint selection and refinement

An image correction model addresses visual distortions in reoriented images by processing them with additional data, improving the accuracy and realism of 3D photo effects generated from monocular images.

WO2025250142A1PCT designated stage Publication Date: 2025-12-04GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/032091
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing image reorientation models introduce visual distortions such as stretching, blurring, and unrealistic textures when generating new viewpoints from monocular images, leading to inaccurate and unrealistic depictions of scenes.

Method used

An image correction model, typically a neural network, is used to process reoriented images generated by an image reorientation model to remove visual distortions, utilizing additional data like near-duplicate images, defect masks, and auxiliary data to improve accuracy and realism.

Benefits of technology

The image correction model enhances the quality of 3D photo effects by accurately representing scenes from new viewpoints without apparent degradation, using fewer computational resources and training iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024032091_04122025_PF_FP_ABST
    Figure US2024032091_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A method includes determining a viewpoint modification of a first viewpoint from which an input image represents a scene, and determining, based on the input image and the viewpoint modification, a reoriented image that (i) represents the scene from a second viewpoint that differs from the first viewpoint and (ii) includes a visual distortion of the scene. The visual distortion may be associated with the viewpoint modification. The method also includes processing the input image and the reoriented image using an image correction model configured to remove visual distortions associated with viewpoint modifications, and generating, using the image correction model and based on processing the input image and the reoriented image, an output image that includes a correction of at least part of the visual distortion in the reoriented image. The output image may represent the scene from the second viewpoint. The method additionally includes outputting the output image.
Need to check novelty before this filing date? Find Prior Art

Description

Post-Capture Photo Viewpoint Selection and RefinementBACKGROUND

[0001] Machine learning models may be used to process various types of data, including images, to generate various desirable outputs. Improvements in the machine learning models and / or the training thereof allow the models to carry out the processing of data faster, to utilize fewer computing resources for the processing, and / or to generate outputs that are of relatively higher quality.SUMMARY

[0002] An input image may represent a scene from a first viewpoint. An image reorientation model may be configured to generate, based on the input image, a reoriented image that represents the scene from a second viewpoint. Thus, the image reorientation model may be used to generate a three-dimensional (3D) photo effect based on a monocular input image. In some cases, the reoriented image may include one or more visual distortions introduced by the image reorientation model in the course of generating visual content for the second viewpoint. An image correction model may be configured to process the reoriented image and remove therefrom at least part of the one or more visual distortions, thereby generating an output image that more accurately and / or consistently represents the scene from the second viewpoint. The image correction model may thus improve the quality of 3D photo effects available for input images and allow for larger viewpoint changes without apparent degradation in the resulting image quality.

[0003] In a first example embodiment, a method may include determining a viewpoint modification of a first viewpoint from which an input image represents a scene. The method may also include determining, based on the input image and the viewpoint modification, a reoriented image that (i) represents the scene from a second viewpoint that differs from the first viewpoint and (ii) includes a visual distortion of the scene. The visual distortion may be associated with the viewpoint modification. The method may additionally include processing the input image and the reoriented image using an image correction model configured to remove visual distortions associated with viewpoint modifications. The method may further include generating, using the image correction model and based on processing the input image and the reoriented image, an output image that includes a correction of at least part of the visual distortion in the reoriented image. The output image may represent the scene from the second viewpoint. The method may yet further include outputting the output image.

[0004] In a second example embodiment, a system may include a processor and a non- transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations in accordance with the first example embodiment.

[0005] In a third example embodiment, a system may include a processor configured to perform operations in accordance with the first example embodiment.

[0006] In a fourth example embodiment, a non-transitory computer-readable medium may have stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations in accordance with the first example embodiment.

[0007] In a fifth example embodiment, a system may include various means for carrying out each of the operations of the first example embodiment.

[0008] These, as well as other embodiments, aspects, advantages, and alternatives, will become apparent to those of ordinary skill in the art by reading the following detailed description, with reference where appropriate to the accompanying drawings. Further, this summary and other descriptions and figures provided herein are intended to illustrate embodiments by way of example only and, as such, that numerous variations are possible. For instance, structural elements and process steps can be rearranged, combined, distributed, eliminated, or otherwise changed, while remaining within the scope of the embodiments as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 illustrates a computing device, in accordance with examples described herein.

[0010] Figure 2 illustrates a computing system, in accordance with examples described herein.

[0011] Figure 3 illustrates an image processing system, in accordance with examples described herein.

[0012] Figure 4A illustrates an example input image, in accordance with examples described herein.

[0013] Figure 4B illustrates an example reoriented image, in accordance with examples described herein.

[0014] Figure 4C illustrates an example output image, in accordance with examples described herein.

[0015] Figure 5 illustrates a training system, in accordance with examples described herein.

[0016] Figure 6 illustrates a flow chart, in accordance with examples described herein.

[0017] Figure 7 illustrates a flow chart, in accordance with examples described herein.DETAILED DESCRIPTION

[0018] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as being an “example,” “exemplary,” and / or “illustrative” is not necessarily to be construed as preferred or advantageous over other embodiments or features unless stated as such. Thus, other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein.

[0019] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0020] Further, unless context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be generally viewed as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.

[0021] Additionally, any enumeration of elements, blocks, or steps in this specification or the claims is for purposes of clarity. Thus, such enumeration should not be interpreted to require or imply that these elements, blocks, or steps adhere to a particular arrangement or are carried out in a particular order. Unless otherwise noted, figures are not drawn to scale.I. Overview

[0022] An image reorientation model may be configured to generate new viewpoints of a scene depicted in a monocular image. The image reorientation model may include a machine learning model, such as a neural network. The image reorientation model may be alternatively referred to as a three-dimensional (3D) photo model because it can create a 3D or 3D-like representation of the scene by generating a plurality of new viewpoints of the monocular image. However, some image reorientation models may introduce visual distortions into reoriented images (representing the new viewpoints) generated thereby. These visual distortions may include stretching, blurring, unrealistic textures, unrealistic object shapes, introduction of noise, and / or inadvertent deletion of salient features, among other possibilities.

[0023] Thus, the new viewpoints might not, in some cases, accurately and / or realistically depict aspects of the scene. The visual distortions may result from (i) an input image to the image reorientation model lacking image data for regions of the scene revealed in the new viewpoints and / or (ii) shortcomings in the architecture and / or training of the image reorientation model, among other possibilities. Accordingly, different scenes and / or different image reorientation models may lead to different visual distortions. Thus, it is desirable to avoid and / or correct such visual distortions in reoriented images generated by image reorientation models.

[0024] Accordingly, an image processing system may utilize an image correction model configured to remove and / or correct at least some of the visual distortions present in the reoriented images generated by the image reorientation model, thereby generating output images that include fewer or no visual distortions. The image correction model may include a machine learning model, such as a neural network. The image correction model may be trained to remove visual distortions based on training on reoriented images generated using one or more image reorientation models, and may thus be configured to remove various types of visual distortions that the one or more image reorientation models are likely and / or expected to introduce into reoriented images at inference time. Accordingly, the image correction model may improve the accuracy, realism, and / or plausibility of reoriented images that represent new viewpoints. Further, using a separate image correction model to correct the visual distortions may be more computationally efficient, use less memory, involve fewer training iterations, and / or result in a smaller overall image processing system than increasing a size and / or complexity of the image reorientation model in an attempt to avoid the visual distortions altogether.

[0025] The image correction model may be configured to generate an output image based on processing an input image that represents a scene from a first viewpoint and a reoriented image that represents the scene from a second viewpoint that is different from the first viewpoint. The second viewpoint of the reoriented image may be selected using, for example, an interactive user interface that displays various viewpoints of the scene (as generated by the image reorientation model) based on user input to the interactive user interface. The output image may represent the scene from the second viewpoint, but might not include some or all of the visual distortions present in the reoriented image. Processing the input image by the image correction model may provide information about how aspects of the scene look without visual distortions, thus allowing the image correction to model this normal appearance in generating visual corrections.

[0026] The image correction model may additionally be configured to process other data that provides additional information about the scene and which may be useful in correcting the visual distortions. This other data may include a near-duplicate image of the scene, a defect mask, auxiliary data, and / or a textual prompt, among other possibilities. The near-duplicate image may be another image of the scene that differs from the input image in its perspective and / or capture time, and thus provides a visually diverse representation of aspects of the scene. The defect mask may be generated by the image reorientation model, and may indicate portions of the reoriented image that have been generated and / or modified by the image reorientation model. The auxiliary data may include a depth map (e.g., generated by a monocular depth model), a feature map (e.g., generated by a convolutional neural network), a 3D model of a subject of the input image, and / or a semantic map, each of which may correspond to the input image, the reoriented image, and / or the near-duplicate image. The textual prompt may specify, for example, which portion(s) of the scene include visual distortions and are thus to be corrected.

[0027] The image correction model may be trained using training samples, each of which includes a training input image, a training reoriented image illustrating a viewpoint modification relative to the input image, and a ground-truth output image having a same viewpoint as the training reoriented image. The training reoriented image may be generated by the image reorientation model and may include one or more training visual distortions, which the image correction model may attempt to remove when generating a training output image. The ground-truth output image may lack the one or more training visual distortions, and may thus be compared to the training output image to quantify how well the image correction model has managed to remove the one or more training distortions from the training reoriented image. Respective input images and ground-truth output images of the training samples may be generated using cameras arranged in different physical poses relative to one another, thus providing examples of how different viewpoint changes relate to various types and / or extents of visual distortions. In some cases, the image correction model may additionally be trained using a training near-duplicate image, a training defect mask, training auxiliary data, and / or a training textual prompt, among other possibilities.II. Example Computing Devices and Systems

[0028] Figure 1 illustrates an example computing device 100. Computing device 100 is shown in the form factor of a mobile phone. However, computing device 100 may be alternatively implemented as a laptop computer, a tablet computer, and / or a wearable computing device, among other possibilities. Computing device 100 may include variouselements, such as body 102, display 106, and buttons 108 and 110. Computing device 100 may further include one or more cameras, such as front-facing camera 104 and rear-facing camera 112.

[0029] Front-facing camera 104 may be positioned on a side of body 102 typically facing a user while in operation (e.g., on the same side as display 106). Rear-facing camera 112 may be positioned on a side of body 102 opposite front-facing camera 104. Referring to the cameras as front and rear facing is arbitrary, and computing device 100 may include multiple cameras positioned on various sides of body 102.

[0030] Display 106 could represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some examples, display 106 may display a digital representation of the current image being captured by front-facing camera 104 and / or rear-facing camera 112, an image that could be captured by one or more of these cameras, an image that was recently captured by one or more of these cameras, and / or a modified version of one or more of these images. Thus, display 106 may serve as a viewfinder for the cameras. Display 106 may also support touchscreen functions that may be able to adjust the settings and / or configuration of one or more aspects of computing device 100.

[0031] Front-facing camera 104 may include an image sensor and associated optical elements such as lenses. Front-facing camera 104 may offer zoom capabilities or could have a fixed focal length. In other examples, interchangeable lenses could be used with front-facing camera 104. Front-facing camera 104 may have a variable mechanical aperture and a mechanical and / or electronic shutter. Front-facing camera 104 also could be configured to capture still images, video images, or both. Further, front-facing camera 104 could represent, for example, a monoscopic, stereoscopic, or multiscopic camera. Rear-facing camera 112 may be similarly or differently arranged. Additionally, one or more of front-facing camera 104 and / or rear-facing camera 112 may be an array of one or more cameras.

[0032] Computing device 100 could be configured to use display 106 and front-facing camera 104 and / or rear-facing camera 112 to capture images of a target object. The captured images could be a plurality of still images or a video stream. The image capture could be triggered by activating button 108, pressing a softkey on display 106, or by some other mechanism. Depending upon the implementation, the images could be captured automatically at a specific time interval, for example, upon pressing button 108, upon appropriate lighting conditions of the target object, upon moving computing device 100 a predetermined distance, or according to a predetermined capture schedule.

[0033] Figure 2 is a simplified block diagram showing some of the components of an example computing system 200. By way of example and without limitation, computing system 200 may be a cellular mobile telephone (e.g., a smartphone), a computer (such as a desktop, notebook, tablet, server, or handheld computer), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a gaming console, a robotic device, a vehicle, or some other type of device. Computing system 200 may represent, for example, aspects of computing device 100.

[0034] As shown in Figure 2, computing system 200 may include communication interface 202, user interface 204, processor 206, data storage 208, and camera components 224, all of which may be communicatively linked together by a system bus, network, or other connection mechanism 210. Computing system 200 may be equipped with at least some image capture and / or image processing capabilities. It should be understood that computing system 200 may represent a physical image processing system, a particular physical hardware platform on which an image sensing and / or processing application operates in software, or other combinations of hardware and software that are configured to carry out image capture and / or processing functions.

[0035] Communication interface 202 may allow computing system 200 to communicate, using analog or digital modulation, with other devices, access networks, and / or transport networks. Thus, communication interface 202 may facilitate circuit-switched and / or packet-switched communication, such as plain old telephone service (POTS) communication and / or Internet protocol (IP) or other packetized communication. For instance, communication interface 202 may include a chipset and antenna arranged for wireless communication with a radio access network or an access point. Also, communication interface 202 may take the form of or include a wireline interface, such as an Ethernet, Universal Serial Bus (USB), or High- Definition Multimedia Interface (HDMI) port, among other possibilities. Communication interface 202 may also take the form of or include a wireless interface, such as a Wi-Fi, BLUETOOTH®, global positioning system (GPS), or wide-area wireless interface (e.g., WiMAX or 3 GPP Long-Term Evolution (LTE)), among other possibilities. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used over communication interface 202. Furthermore, communication interface 202 may comprise multiple physical communication interfaces (e.g., a Wi-Fi interface, a BLUETOOTH® interface, and a wide-area wireless interface).

[0036] User interface 204 may function to allow computing system 200 to interact with a human or non-human user, such as to receive input from a user and to provide output to theuser. Thus, user interface 204 may include input components such as a keypad, keyboard, touch-sensitive panel, computer mouse, trackball, joystick, microphone, and so on. User interface 204 may also include one or more output components such as a display screen, which, for example, may be combined with a touch-sensitive panel. The display screen may be based on CRT, LCD, LED, and / or OLED technologies, or other technologies now known or later developed. User interface 204 may also be configured to generate audible output(s), via a speaker, speaker jack, audio output port, audio output device, earphones, and / or other similar devices. User interface 204 may also be configured to receive and / or capture audible utterance(s), noise(s), and / or signal(s) by way of a microphone and / or other similar devices.

[0037] In some examples, user interface 204 may include a display that serves as a viewfinder for still camera and / or video camera functions supported by computing system 200. Additionally, user interface 204 may include one or more buttons, switches, knobs, and / or dials that facilitate the configuration and focusing of a camera function and the capturing of images. It may be possible that some or all of these buttons, switches, knobs, and / or dials are implemented by way of a touch-sensitive panel.

[0038] Processor 206 may comprise one or more general purpose processors - e.g., microprocessors - and / or one or more special purpose processors - e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, application-specific integrated circuits (ASICs), and / or tensor processing units (TPUs). In some instances, special purpose processors may be capable of image processing, image alignment, and merging images, among other possibilities. Data storage 208 may include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic storage, and may be integrated in whole or in part with processor 206. Data storage 208 may include removable and / or non-removable components.

[0039] Processor 206 may be capable of executing program instructions 218 (e.g., compiled or non-compiled program logic and / or machine code) stored in data storage 208 to carry out the various functions described herein. Therefore, data storage 208 may include a non-transitory computer-readable medium, having stored thereon program instructions that, upon execution by computing system 200, cause computing system 200 to carry out any of the methods, processes, or operations disclosed in this specification and / or the accompanying drawings. The execution of program instructions 218 by processor 206 may result in processor 206 using data 212.

[0040] By way of example, program instructions 218 may include an operating system 222 (e.g., an operating system kernel, device driver(s), and / or other modules) and one or moreapplication programs 220 (e.g., camera functions, address book, email, web browsing, social networking, audio-to-text functions, text translation functions, and / or gaming applications) installed on computing system 200. Similarly, data 212 may include operating system data 216 and application data 214. Operating system data 216 may be accessible primarily to operating system 222, and application data 214 may be accessible primarily to one or more of application programs 220. Application data 214 may be arranged in a file system that is visible to or hidden from a user of computing system 200.

[0041] Application programs 220 may communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs may facilitate, for instance, application programs 220 reading and / or writing application data 214, transmitting or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on.

[0042] In some cases, application programs 220 may be referred to as “apps” for short. Additionally, application programs 220 may be downloadable to computing system 200 through one or more online application stores or application markets. However, application programs can also be installed on computing system 200 in other ways, such as via a web browser or through a physical interface (e.g., a USB port) on computing system 200.

[0043] Camera components 224 may include, but are not limited to, an aperture, shutter, recording surface (e.g., photographic film and / or an image sensor), lens, shutter button, infrared projectors, and / or visible-light projectors. Camera components 224 may include components configured for capturing of images in the visible-light spectrum (e.g., electromagnetic radiation having a wavelength of 380 - 700 nanometers) and / or components configured for capturing of images in the infrared light spectrum (e.g., electromagnetic radiation having a wavelength of 701 nanometers - 1 millimeter), among other possibilities. Camera components 224 may be controlled at least in part by software executed by processor 206.III. Example Image Processing System

[0044] Figure 3 illustrates an example image processing system 300. Image processing system 300 may include image reorientation model 308, auxiliary model 312, and image correction model 320. Image processing system 300 may be implemented using software, hardware, or a combination thereof. Image processing system 300 may be implemented using computing device 100, computing system 200, a client device (e.g., a mobile device), a server device, and / or a combination thereof.

[0045] Image processing system 300 may be configured to generate output image 322 based on input image 302 and viewpoint modification 304. In some cases, image processing system 300 may be configured to generate output image 322 further based on near-duplicate image 306, among other possible inputs. Output image 322 may represent the contents of input image 302 reoriented according to viewpoint modification 304, and may lack some or all visual distortions caused by this reorientation. Thus, image processing system 300 may be configured to generate new viewpoints of input image 302 and correct any distortions associated with (e.g., introduced and / or caused by) performance of viewpoint modification 304.

[0046] Input image 302 (e.g., a monocular red-green-blue (RGB) image) may represent a scene from a first viewpoint. The scene may include, for example, a person, a landscape, a vehicle, and / or an animal, among other possibilities. In some cases, the scene may include a foreground (e.g., a person), a background (e.g., a landscape behind the person), and / or one or more other distinct and / or separable layers. In some implementations, image processing system 300 may be configured to uncrop (e.g., using an ML-based image uncropping model) input image 302 before providing input image 302 as input to image reorientation model 308 (thus allowing image reorientation model 308 having to generate visual content for at least some portions of the scene that are not represented in input image 302). Viewpoint modification 304 may represent a change from the first viewpoint to a second viewpoint that differs from the first viewpoint. For example, viewpoint modification 304 may include a translation and / or rotation relative to the scene. Viewpoint modification 304 may be obtained by way of a user interface (UI) (e.g., from a user), and / or may be determined (e.g., suggested) by image processing system 300.

[0047] As one example, the UI may include a UI component that is rotatable and / or translatable (e.g., relative to input image 302 and / or an interactive representation of the scene) to define viewpoint modification 304. For example, the UI component may represent a position of an observer (e.g., camera) relative to input image 302 and / or the interactive representation of the scene. Thus, rotation and / or translation of the UI component may modify the observer’s viewpoint of the scene represented by input image 302. The UI component may be moveable in discrete increments that, when sufficiently dense, approximate and / or simulate a continuous range of movements relative to the viewpoint of input image 302. As the UI component is manipulated and / or interacted with, image processing system 300 may be configured to generate (e.g., using image reorientation model 308, as discussed below) at least a partial preview of how input image 302 will look after a corresponding viewpoint modification. Thispartial preview may be displayed using the UI, thus providing visual feedback on how various interactions with the UI component will affect output image 322.

[0048] As another example, viewpoint modification 304 may be selected from a plurality of candidate viewpoint modifications. The plurality of candidate viewpoint modifications may be determined by image processing system 300 and may represent predetermined combinations of rotations and / or translations relative to the first viewpoint. For example, the predetermined combinations may include combinations of horizontal translation and / or rotation at predetermined increments (e.g., -20°, -15°, -10°, -5°, +5°, +10°, +15°, and +20° of rotation relative to the first viewpoint) and combinations of vertical translation and / or rotation at predetermined increments. Thus, the plurality of candidate viewpoints may represent a sampling (e.g., a uniform sampling) of viewpoint modifications that can be achieved using image processing system 300.

[0049] Image reorientation model 308 may be configured to generate reoriented image 310 based on input image 302 and viewpoint modification 304. Reoriented image 310 may represent the scene from the second viewpoint. Thus, input image 302 and reoriented image 310 may represent aspects of the same underlying scene, but may do so from different viewpoints. Accordingly, input image 302 and reoriented image 310 may visually differ from one another in various ways. For example, reoriented image 310 may include visual content (e.g., generated by an inpainting model of image reorientation model 308) corresponding to portions of the scene that have been revealed (i.e., become disoccluded) due to viewpoint modification 304, and may lack visual content corresponding to portions of the scene that have been obstructed / occluded due to viewpoint modification 304. In some implementations, image reorientation model 308 may be configured to generate a plurality of reoriented images 310 to represent previews of the visual effects of different viewpoint modifications 304. Viewpoint modification 304 may be represented in a reference frame and / or coordinate system utilized by image reorientation model 308.

[0050] Reoriented image 310 may include one or more visual distortions of the scene. For example, the one or more visual distortions may include stretching of parts of the scene, blurring or parts of the scene, tearing / separation of parts of an object in the scene (e.g., the hand of a foreground subject ending up in the background and floating on its own), physically implausible appearance of parts of the scene, parts of the scene appearing inconsistent with the visual content of input image 302, merging of a background object with part of the foreground (e.g., part of the background appearing to get stuck in a foreground subject’s face), inaccurate object boundaries that appear ragged or unnatural, inaccurate representations of fine layeringdetails (e.g., matting of individual hairs), missing and / or inaccurate inpainting and / or outpainting, “cardboard cutout” type artifacts (e.g., due to parts of the scene not properly deforming due to errors and / or imperfections in modeling 3D properties of objects), among other possibilities. The visual distortions may alternatively be referred to as visual defects and / or visual artifacts, among other possibilities. These visual distortions may be (inadvertently) introduced into reoriented image 310 by image reorientation model 308. For example, image reorientation model 308 might not be able to generate distortion-free image data for some viewpoint modifications, scene types, and / or visual features, among other possibilities. The type and / or extent of the one or more visual distortions may depend on an architecture of image reorientation model 308, a size of image reorientation model 308, and / or training data used to train image reorientation model 308, among other potential causes.

[0051] Image reorientation model 308 may include a generative machine learning (ML) model, which may include one or more neural networks. In some implementations, image reorientation model 308 may be configured to generate an intermediate representation of the scene based on input image 302 to facilitate generation of reoriented image 310. The intermediate representation may include one or more 3D polygon meshes, one or more point clouds, one or more 3D Gaussians (e.g., as discussed in the paper titled “3D Gaussian Splatting for Real-Time Radiance Field Rendering,” authored by Kerbl et al., and published as arXiv:2308.04079, which is hereby incorporated by reference), one or more multi-plane images (e.g., as discussed in the paper titled “Single- View View Synthesis with Multiplane Images,” authored by Tucker et al., and published as arXiv:2004.11364, which is hereby incorporated by reference), and / or one or more neural radiance fields (e.g., as discussed in the paper titled “ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image,” authored by Sargent et al., and published as arXiv:2310.17994, which is hereby incorporated by reference), among other possibilities. For example, image reorientation model 308 may include aspects of the 3D photo system discussed in U.S. Patent Application 17 / 907,529, which is hereby incorporated by reference as if set forth herein in its entirety.

[0052] In other implementations, image reorientation model 308 may be configured to generate reoriented image 308 without relying on an intermediate representation (e.g., without relying on a 3D model or representation) of the scene. For example, image reorientation model 308 may include a viewpoint-conditioned generative model (e.g., a diffusion model) that has been trained to generate (e.g., directly or incrementally) reoriented images based on corresponding input images and various viewpoint modifications. The viewpoint-conditionedgenerative model may alternatively be referred to as a controllable-viewpoint generative model.

[0053] In some implementations, the viewpoint-conditioned generative model may be configured to generate one or more incremental reoriented images based on incremental viewpoint modifications, where the incremental viewpoint modifications collectively form (e.g., sum to) viewpoint modification 304, and where the incremental reoriented images allow the viewpoint-conditioned generative model to incrementally go from the first viewpoint of input image 302 to the second viewpoint of reoriented image 310. The viewpoint-conditioned generative model may include, for example, an image generation diffusion model (e.g., as discussed in the paper titled “Zero-l-to-3: Zero-shot One Image to 3D Object,” authored by Liu et al., and published as arXiv:2303.11328, which is hereby incorporated by reference) and / or a video generation diffusion model (e.g., as discussed in the paper titled “Lumiere: A Space-Time Diffusion Model for Video Generation,” authored by Bar-Tai et al., and published as arXiv:2401.12945, which is hereby incorporated by reference), among other possibilities.

[0054] In some implementations, image reorientation model 308 may also be configured to generate defect mask 318 corresponding to reoriented image 310. Defect mask 318 may indicate portions of reoriented image 310 that are likely and / or expected to contain visual distortions. For example, defect mask 318 may indicate portions of reoriented image 310 that include (i) generated visual content that is not part of input image 302 and / or (ii) modified visual content that is based on modifications of visual content that is part of input image 302. As one example, defect mask 318 may be based on and / or may include one or more inpainting masks that indicate (i) parts of reoriented image 310 where image reorientation model 308 performed inpainting and / or outpainting, and / or (ii) parts of reoriented image 310 where image correction model 320 is to perform inpainting and / or outpainting.

[0055] Image correction model 320 may be configured to generate output image 322 based on reoriented image 310 and input image 302. Specifically, image correction model 320 may be configured to remove from reoriented image 310 at least part of the visual distortion added thereto by image reorientation model 308. Thus, output image 322 may represent the scene from the second viewpoint, and may do so more accurately and / or consistently than reoriented image 310 at least in that output image 322 may lack at least part of the visual distortion presented in reoriented image 310. In some cases, output image 322 may lack most or all of the visual distortion, and may thus provide a substantially distortion-free representation of the scene from the second viewpoint. Thus, output image 322 may provide a representationof the scene that more accurately represents the scene than reoriented image 310 and / or is more consistent with the physical properties of scene features than reoriented image 310.

[0056] In some cases, image correction model 320 may have been trained to remove the types of visual distortions that image reorientation model 308 is (inadvertently) configured to introduce into reoriented images generated thereby. Thus, as image reorientation model 308 changes (e.g., in its architecture, size, training, etc.), thereby potentially changing the type and / or extent of visual distortions introduced thereby, image correction model 320 may be retrained to account for such changes. Accordingly, image correction model 320 may be viewed as being matched to image reorientation model 308 and / or the particular visual distortions of image reorientation model 308.

[0057] In other cases, image correction model 320 may have been trained to correct a plurality of different types of visual distortions configured to be introduced into reoriented images by a plurality of different types and / or instances of image reorientation model 308. For example, image correction model 320 may be trained to correct a plurality of different types of visual distortions configured to be introduced into reoriented images by a plurality of different architectures of image reorientation model 308, each of which may have been trained using a different training data set. Accordingly, image correction model 320 may be viewed as being agnostic to the type of image reorientation model 308 and / or different types visual distortions.

[0058] In generating output image 322, input image 302 may provide image correction model 320 with a distortion-free representation of the scene. That is, input image 302 may provide a baseline reference relative to which image correction model 320 may judge whether a given part of reoriented image 310 (i) includes a distortion or (ii) accurately represents the actual and / or likely content of a corresponding part of the scene. Thus, providing input image 302 as an input to image correction model 320 may allow image correction model 320 to more accurately and / or completely correct visual distortions in reoriented image 310.

[0059] In some implementations, image correction model 320 may be configured to generate output image 322 further based on near-duplicate image 306. Near-duplicate image 306 may be a near-duplicate of input image 302 in that these two images may represent aspects of the same scene from different viewpoints. Specifically, near-duplicate image 306 may include first image data corresponding to a first portion of the scene and second image data corresponding to a second portion of the scene. The first portion of the scene may be represented in both input image 302 and near-duplicate image 306, and the second portion of the scene may be represented in near-duplicate image 306 but not in input image 302. Thus, near-duplicate image 306 may include both redundant and non-redundant image data, therebyproviding an alternative and / or visually diverse representation of the scene, which may assist image correction model 320 with correcting visual distortions present in reoriented image 310.

[0060] In some implementations, image correction model 320 may be configured to generate output image 322 further based on defect mask 318. Since defect mask 318 may indicate portions of reoriented image 310 that are likely and / or expected to contain visual distortions, defect mask 318 may provide image correction model 320 with a suggestion as to which portions of reoriented image 310 are to be corrected.

[0061] In some implementations, image processing system 300 may include auxiliary model 312 configured to generate auxiliary data 314. Auxiliary model 312 may be configured to generate auxiliary data 314 based on input image 302, reoriented image 310, and / or nearduplicate image 306, among other possibilities. Auxiliary data 314 may be used by image correction model 320 to assist with generating output image 322. Auxiliary data 314 may include information indicative of a physical structure of the scene, semantics of the scene, and / or other properties of the scene, and may thus provide image correction model 320 with cues and / or prior knowledge as to which parts of reoriented image 310 might include visual distortions.

[0062] As one example, auxiliary model 312 may include a depth model (e.g., monocular depth model) configured to generate a depth map. Thus, auxiliary data 314 may include one or more depth maps corresponding to input image 302, reoriented image 310, and / or near-duplicate image 306. The one or more depth maps may allow image correction model 320 to identify, for example, parts of reoriented image 310 associated with depth values that are inconsistent with depth values of corresponding parts of input image 302.

[0063] As another example, auxiliary model 312 may include a 3D subject model configured to generate a 3D representation of a subject (e.g., person) present in input image 302. Thus, auxiliary data 314 may include the 3D representation of the subject. For example, auxiliary model 312 may include and / or be based on the PHORHUM model described in the paper titled “Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing,” authored by Alldieck et al., and published as arXiv:2204.08906, which is hereby incorporated by reference. In some cases, a plurality of different types of 3D subject models corresponding to different types of subjects may be used.

[0064] As a further example, auxiliary model 312 may include a feature map model configured to generate a feature map. Thus, auxiliary data 314 may include one or more feature maps corresponding to input image 302, reoriented image 310, and / or near-duplicate image 306. For example, auxiliary model 312 may include a convolutional neural network configuredto generate one or more convolutional feature maps. The one or more feature maps may represent visual features (e.g., lines, edges, blobs, combinations thereof, etc.) identified by the feature map model, and may be useful to image correction model 320 in identifying visual distortions in reoriented image 310.

[0065] As a yet further example, auxiliary model 312 may include a semantic model configured to generate a semantic map. Thus, auxiliary data 314 may include one or more semantic maps corresponding to input image 302, reoriented image 310, and / or near-duplicate image 306. The one or more semantic maps may represent classifications of different parts of the scene, as represented by a corresponding image. For example, a semantic map corresponding to input image 302 may assign, to each respective pixel of a plurality of pixels of input image 302, a corresponding value that represents a classification (e.g., person, dog, cat, tree, sky, landscape, etc.) of a visual feature represented by the respective pixel. By representing the classifications of different parts of the scene, the one or more semantic maps may assist image correction model 320 with identifying types of visual features and / or transitions between different types of visual features, some of which may be associated with visual distortions in reoriented image 310.

[0066] Image correction model 320 may include a generative machine learning model. For example, image correction model 320 may include and / or be based on a generative latent diffusion model (LDM) architecture and / or a generative adversarial network (GAN) architecture, among other possibilities. Additionally or alternatively, image correction model 320 may include, be based on, and / or utilize the techniques associated with, among other possibilities, (i) the Imagen model, as discussed in the paper titled “Photorealistic Text-to- Image Diffusion Models with Deep Language Understanding,” authored by Saharia et al., and published as arXiv:2205.11487, (ii) the ControlNet model, as discussed in the paper titled “Adding Conditional Control to Text-to-Image Diffusion Models,” authored by Zhang et al., and published as arXiv:2302.05543, (iii) the InstructPix2Pix model, as discussed in the paper titled “InstructPix2Pix: Learning to Follow Image Editing Instructions,” authored by Brooks et al., and published as arXiv:2211.09800, and / or (iv) the RealFill model, as discussed in the paper titled “RealFill: Reference-Driven Generation for Authentic Image Completion,” authored by Tang et al. and published as arXiv:2309.16668, each of which is hereby incorporated by reference.

[0067] Accordingly, in some implementations, image correction model 320 may be configured to generate output image 322 further based on noise image 316. For example, when image correction model 320 includes a diffusion model, noise image 316 may be iterativelydenoised by the diffusion model as part of the process of generating output image 322. Thus, in some implementations, a plurality of different noise images 316 may be used to generate a plurality of different output images 322, thus providing a variety of possible corrections of reoriented image 310. Providing multiple different output images 322 that vary in how the visual distortions are corrected therein may increase the likelihood of at least one image providing an accurate, extensive, and / or complete correction of the visual distortions. The plurality of different output images 322 may be displayed to a user, and the user may select one or more of these images for output and / or usage (e.g., based on the user’s preference for how the visual distortions have been corrected, accuracy of the corrections, and / or completeness of the corrections, among other factors).

[0068] In some implementations, image correction model 320 may be configured to generate output image 322 further based on textual prompt 324. Textual prompt 324 may instruct image correction model 320 to correct visual distortions in reoriented image 310. In some cases, textual prompt 324 may be predetermined. For example, textual prompt 324 may include a general instruction to “Remove visual distortions from the reoriented image.” Textual prompt 324 may include a name and / or type of image reorientation model 308 (e.g., “The reoriented image was generated by a model that uses a polygon mesh.”), thus providing information about the types of visual distortions likely to be present in reoriented image 310. In other cases, textual prompt 324 may be specified by a user. For example, the user may write textual prompt 324 to describe portions of reoriented image 310 that include the visual distortion and are thus to be corrected by image correction model 320. For example, textual prompt 324 may include a specific instruction to “Remove visual distortions around the face of the subject of the reoriented image.”

[0069] In some implementations, image correction model 320 may include and / or be used in connection with a super-resolution model to generate output image 322. For example, a first (distortion correction) model of image correction model 320 may be configured to generate an intermediate image that has a first resolution and corrects at least part of the visual distortions present in reoriented image 310. A second (super-resolution) model of image correction model 320 may be configured to process the intermediate image and, based thereon, generate output image 322 having a second resolution that is greater than the first resolution. Thus, in some cases, distortion correction and upsampling may be performed using two different models. In other cases, both the distortion correction and the upsampling may be performed jointly using one model. For example, image correction model 320 maysimultaneously remove the distortions and upsample the image data in the process of generating output image 322.IV. Example Images

[0070] Figure 4A illustrates an example input image 402, which provides one example of input image 302. Input image 402 represents, from a first viewpoint, cup 400 filled with coffee and positioned on top of table 404. Cup 400 may be considered to be the subject and / or foreground of input image 402, while table 404 and regions visible beyond table 404 may be considered to be the background of input image 402.

[0071] Figure 4B illustrates an example reoriented image 410, which provides one example of reoriented image 410. Reoriented image 410 represents, from a second viewpoint different from the first viewpoint, cup 400 filled with coffee and positioned on top of table 404. The second viewpoint corresponds to a downward translation of a camera that captured input image 402 (e.g., a negative displacement along the y-axis of computing device 100 illustrated in Figure 1) and / or a rotation along a pitch axis of the camera (e.g., a rotation about the x-axis of computing device 100 illustrated in Figure 1).

[0072] Reoriented image 410 includes a plurality of visual distortions, including visual distortion 430, visual distortion 432, and visual distortion 434, each of which may be associated with the change from the first viewpoint to the second viewpoint performed by image reorientation model 308. Visual distortion 430 includes an incorrect change in a height of cup 400. Specifically, cup 400 appears taller in reoriented image 410 than it realistically should. Visual distortion 432 includes an incorrect shape of a portion of the handle of cup 400. Specifically, the handle of cup 400 appears to connect to a body of cup 400 in a physically unrealistic manner. Visual distortion 434 includes an incorrect position and shape of the edge of table 404. Specifically, the edge of table 404 is visible where it should be obscured by cup 400, and appears to give table 404 an irregular, rather than circular, shape.

[0073] Figure 4C illustrates an example output image 422, which provides one example of output image 322. Output image 422 represents, from the second viewpoint, cup 400 filled with coffee and positioned on top of table 404. Output image 422 does not include most or all of visual distortions 430, 432, and 434, and thus more accurately and / or realistically represents cup 400, table 404, and any surrounding objects than reoriented image 410. First, a height of cup 400 in output image 422 is smaller than a height of cup 400 in reoriented image 410, which provides a more accurate and / or realistic depiction of cup 400 after the viewpoint change. Second, the shape of the portion of the handle of cup 400 in output image 422 appears more realistically connected to the body of cup 400 than in reoriented image 410. Third, the edge oftable 404 is correctly obscured by cup 400 in output image 422 and thus appears to give table 404 a circular shape, unlike reoriented image 410.V. Example Training System

[0074] Figure 5 illustrates training system 500 configured to train image correction model 320 based on training data samples. Training system 500 may include viewpoint calculator 508, image reorientation model 308, image correction model 320, loss function(s) 516, and model parameter adjuster 520. Training system 500 may be implemented using software, hardware, or a combination thereof. Training system 500 may be implemented using computing device 100, computing system 200, a client device (e.g., a mobile device), a server device, and / or a combination thereof.

[0075] An example training data sample may include training input image 502, groundtruth output image 504, and, in some cases, training sample metadata 506. Training system 500 may be configured to generate, based on the training data samples, a trained version of image correction model 320 that is configured to remove visual defects from reoriented images generated by image reorientation model 308 and, in some cases, trained versions of one or more other components of image processing system 300.

[0076] Training input image 502 may represent a training scene from a first training viewpoint. The training scene may be a phy si cal / real -world training scene. Training input image 502 may be analogous to input image 302, but may be processed at training time rather than at inference time. Ground-truth output image 504 may represent the training scene from a second training viewpoint that is different from the first training viewpoint. In some cases (e.g., when the training scene includes moving objects), training input image 502 and ground-truth output image 504 may be captured at substantially the same time using two or more cameras. The two or more cameras may be separated from one another by a known distance and / or a known rotation. Training sample metadata 506 may represent properties of the two or more cameras, including (i) the intrinsic parameters of each camera, (ii) the distance between the two or more cameras, and / or (iii) the rotation of one camera relative to the other camera(s), among others. In other cases (e.g., when the training scene includes only static objects), training input image 502 and ground-truth output image 504 may be captured at different times using a single camera. In further cases, training input image 502 and ground-truth output image 504 may represent multiple rendered viewpoints of a synthetic training scene.

[0077] Different training data samples may be captured with different relative spatial arrangements of the cameras in order to represent different types, directions, and / or magnitudes of potential viewpoint modifications. That is, in order to train image correction model 320 tocorrect visual distortions resulting from a wide range of potential viewpoint modifications, the training data samples may be obtained to represent a wide range of actual viewpoint modifications. For example, the training data samples may include image pairs that represent horizontal translation, vertical translation, yaw rotation, roll rotation, pitch rotation, and / or combinations thereof. Thus, in some cases, a maximum value of viewpoint modification 304 along a given axis may be based on and / or correspond to a maximum viewpoint modification along the given axis represented by the training data samples. For example, a maximum horizontal translation of viewpoint modification 304 may be based on and / or correspond to a maximum horizontal translation represented by the training data samples, thus allowing image correction model 320 to operate with respect to viewpoint modifications that have been represented by the training data samples.

[0078] The training data samples may be captured using a camera rig that allows the two or more cameras to be placed in different spatial arrangements relative to one another. For example, the camera rig may include one or more horizontal bars that allow two cameras to be horizontally separated from one another by, for example, 5 centimeters to 60 centimeters, and one or more vertical bars that allow two cameras to be vertically separated from one another by, for example, 5 centimeters to 60 centimeters. The camera rig may include a plurality of holders configured to hold computing devices (e.g., smartphones) that include the camera devices, thus allowing the training data samples to be obtained from the same or similar types of cameras as input images 302. Each holder may be vertically and / or horizontally translatable along and / or rotatable about one or more axes.

[0079] In some cases, prior to image capture, one or more of the computing devices may be provided with metadata that represents the relative arrangement of the cameras on the camera rig. The cameras on the camera rig may be synchronized with one another such that each camera captures a respective image at approximately, substantially, and / or exactly the same time. For example, the cameras may be synchronized using the operations and / or techniques discussed in the paper titled “Wireless Software Synchronization of Multiple Distributed Cameras,” authored by Ansari et al., and published as arXiv: 1812.09366, which is hereby incorporated by reference. Each time a training sample is captured, the computing devices on the camera rig may be configured to generate and / or store corresponding training sample metadata in association with the training sample.

[0080] Viewpoint calculator 508 may be configured to determine training viewpoint modification 510 based on training input image 502, ground-truth output image 504, and, in some cases, training sample metadata 506. Image reorientation model 308 may be configuredto generate training reoriented image 512 based on training viewpoint modification 510 and training input image 502. Image correction model 320 may be configured to determine training output image 514 based on training input image 502 and training reoriented image 512. Training viewpoint modification 510, training reoriented image 512, and training output image 514 may be analogous to, respectively, viewpoint modification 304, reoriented image 310, and output image 322, but may be determined at training time rather than at inference time.

[0081] In some implementations, image correction model 320 may be configured to determine training output image 514 further based on training near-duplicate image (analogous to near-duplicate image 306), training auxiliary data (analogous to auxiliary data 314), training defect mask (analogous to defect mask 318), training noise image (analogous to noise image 316), and / or training textual prompt (analogous to textual prompt 324), among other possibilities, as represented by arrow 524.

[0082] Training viewpoint modification 510 may be determined such that a viewpoint of training reoriented image 512 is approximately, substantially, and / or exactly equal to the second training viewpoint of ground-truth output image 504. Thus, ground-truth output image 504 may provide a distortion-free representation of the training scene from the second viewpoint, thereby allowing training system 500 to quantify how effectively image correction model 320 is able to remove visual distortions from training reoriented image 512, which also represents the scene from the second viewpoint. Accordingly, viewpoint calculator 508 may be configured to determine a transformation that relates training input image 502 and groundtruth output image 504 in a reference frame and / or coordinate system used by image reorientation model 308.

[0083] In some cases, training sample metadata 506 may represent the first viewpoint of training input image 502 and / or the second training viewpoint of ground-truth output image 504 using metric values (e.g., centimeters), while image reorientation model 308 may be configured to process viewpoint modifications expressed using relative values (i.e., unitless values that correspond to, but do not directly represent, physical distances). Accordingly, viewpoint calculator 508 may be configured to transform the physical viewpoint modification between training input image 502 and ground-truth output image 504, as represented using metric values by training sample metadata 506, into training viewpoint modification 510, represented using relative values along a relative value scale of image reorientation model 308.

[0084] Specifically, viewpoint calculator 508 may be configured to determine, based on training input image 502 and / or ground-truth output image 504, one or more relative depth values associated with the training scene. The one or more relative depth values may bedetermined using a depth model that is also utilized by image reorientation model 308, thus expressing the one or more relative depth values using the relative value scale of image reorientation model 308. Viewpoint calculator 508 may also be configured to determine, based on training input image 502, ground-truth output image 504, and / or training sample metadata 506 (which represents a metric distance between a first camera position associated with training input image 502 and a second camera position associated with ground-truth output image 504), one or more metric depth values associated with the training scene. Viewpoint calculator 508 may further be configured to determine a mapping between relative depth and metric depth based on the one or more relative depth values and the one or more metric depth values.

[0085] The mapping may be expressed as a function that scales relative depth values to metric depth values, or vice versa. For example, viewpoint calculator 508 may be configured to determine the function by solving an optimization problem such that the mapping from relative depth to metric depth, or vice versa, is consistent and / or accurate across a plurality of points in the training scene. Thus, viewpoint calculator 508 may be configured to improve and / or maximize the consistency and / or accuracy of the mapping for the scene as a whole, rather than for individual points therein. The function may be linear or nonlinear. Viewpoint calculator 508 may then apply the mapping to the viewpoint modification between training input image 502 and ground-truth output image 504, represented using metric values by training sample metadata 506, thus determining training viewpoint modification 510 that is represented using the relative value scale of image reorientation model 308.

[0086] Additionally or alternatively, viewpoint calculator 508 may be configured to use an iterative search process (e.g., grid search) to determine training viewpoint modification 510. For example, viewpoint calculator 508 may select a plurality of candidate training viewpoint modifications, and transform training input image 502 (possibly with the assistance of image reorientation model 308) using each of the plurality of candidate training viewpoint modifications to generate a plurality of candidate output images. Viewpoint calculator 508 may also be configured to compare ground-truth output image 504 to each of the plurality of candidate output images to identify a candidate output image that most closely matches groundtruth output image 504. For example, viewpoint calculator 508 may be configured to generate, for each respective candidate output image, a corresponding color consistency value that represents how well the pixel values of pixels of the respective candidate output image match the pixel values of corresponding pixels (e.g., corresponding in pixel space) of ground-truth output image 504. Viewpoint calculator 508 may determine training viewpoint modification 510 based on a particular candidate training viewpoint modification that corresponds to thecandidate output image that most closely matches ground-truth output image 504. For example, training viewpoint modification 510 may be equal to the particular candidate training viewpoint modification.

[0087] To further improve the accuracy of training viewpoint modification 510, the search process may be repeated one or more times based on an additional plurality of candidate training viewpoint modifications that are near and / or similar to the particular candidate training viewpoint modification resulting from a prior iteration (e.g., within a smaller grid surrounding the particular candidate training viewpoint modification may be searched). Specifically, viewpoint calculator 508 may transform training input image 502 using each of the additional plurality of candidate training viewpoint modifications to generate an additional plurality of candidate output images, compare ground-truth output image 504 to each of the additional plurality of candidate output images to identify a particular candidate output image that most closely matches ground-truth output image 504, and determine training viewpoint modification 510 based on a given candidate training viewpoint modification that corresponds to the particular candidate output image. The iterative search process may conclude when, for example, the corresponding color consistency value of the particular candidate output image meets or exceeds a threshold color consistency value.

[0088] In some implementations, viewpoint calculator 508 may be configured to use feature matching, optical flow, 2D-3D correspondences, and / or pose estimation techniques, among other possibilities, to determine training viewpoint modification 510.

[0089] Loss function(s) 516 may be configured to determine loss value 518 based on training output image 514, ground-truth output image 504, and / or one or more intermediate values used in determination thereof. For example, loss function(s) 516 may be configured to determine loss value 518 based on a difference (e.g., a mean square error, a mean absolute error, etc.) between training output image 514 and ground-truth output image 504. Loss function(s) 516 may additionally or alternatively be configured to compare a latent representation of training output image 514 to a latent representation of ground-truth output image 504 using a perceptual loss function. In implementations that utilize diffusion models, loss function(s) 516 may additionally or alternatively be configured to compare training noise added by a forward diffusion process to noise detected by image correction model 320 (e.g., when image correction model 320 and / or aspects thereof are configured to predict the noise rather than explicitly predict a denoised image).

[0090] Model parameter adjuster 520 may be configured to determine updated model parameters 522 based on loss value 518. Specifically, updated model parameters 522 may beselected such that, during a subsequent iteration of processing of training input image 502, training output image 514 more closely matches ground-truth output image 504. Updated model parameters 522 may include one or more updated parameters of any trainable component of image correction model 320. In some implementations, image correction model 320 may be trained while the parameters of one or more other trainable models of image processing system 300 are held fixed. In other implementations, image correction model 320 may be trained jointly with the one or more other trainable models of image processing system 300, and thus updated model parameters 522 may also include parameters for the one or more other trainable components.

[0091] Model parameter adjuster 520 may be configured to determine updated model parameters 522 by, for example, determining a gradient of loss function(s) 516. Based on this gradient and loss value 518, model parameter adjuster 520 may be configured to select (e.g., using gradient descent) updated model parameters 522 that are expected to reduce loss value 518, and thus improve a performance of image correction model 320 and / or image processing system 300. After applying updated model parameters 522 to image correction model 320, the operations discussed above may be repeated to compute another instance of loss value 518 and, based thereon, another instance of updated model parameters 522 may be determined and applied to image correction model 320 to further improve the performance thereof. Such training of image correction model 320 may be repeated until, for example, loss value 518 is reduced to below a target loss value.VI. Additional Example Operations

[0092] Figure 6 illustrates a flow chart of operations related to correcting visual distortions associated with viewpoint modifications of an input image. The operations may be carried out by computing device 100, computing system 200, and / or image processing system 300, among other possibilities. The embodiments of Figure 6 may be simplified by the removal of any one or more of the features shown therein. Further, these embodiments may be combined with features, aspects, and / or implementations of any of the previous figures or otherwise described herein.

[0093] Block 600 may involve determining a viewpoint modification of a first viewpoint from which an input image represents a scene.

[0094] Block 602 may involve determining, based on the input image and the viewpoint modification, a reoriented image that (i) represents the scene from a second viewpoint that differs from the first viewpoint and (ii) includes a visual distortion of the scene. The visual distortion may be associated with the viewpoint modification.

[0095] Block 604 may involve processing the input image and the reoriented image using an image correction model configured to remove visual distortions associated with viewpoint modifications.

[0096] Block 606 may involve generating, using the image correction model and based on processing the input image and the reoriented image, an output image that includes a correction of at least part of the visual distortion in the reoriented image. The output image may represent the scene from the second viewpoint.

[0097] Block 608 may involve outputting the output image.

[0098] In some examples, a defect mask may be determined that represents one or more portions of the scene depicted in the reoriented image and for which the reoriented image includes visual content associated with one or more visual distortions resulting from the viewpoint modification. The output image may be generated further based on processing the defect mask by the image correction model.

[0099] In some examples, the defect mask may be based on an inpainting mask that indicates that the input image lacks image data for a portion of the one or more portions of the scene.

[0100] In some examples, a near-duplicate image may be obtained that includes (i) first image data corresponding to a first portion of the scene and (ii) second image data corresponding to a second portion of the scene. The first portion of the scene may be represented by the input image and the second portion of the scene might not be represented by the input image. The output image may be generated further based on processing the nearduplicate image by the image correction model.

[0101] In some examples, a depth map corresponding to at least one of the input image or the reoriented image may be determined. The output image may be generated further based on processing the depth map by the image correction model.

[0102] In some examples, a feature map corresponding to at least one of the input image or the reoriented image may be determined. The output image may be generated further based on processing the feature map by the image correction model.

[0103] In some examples, a three-dimensional (3D) subject model that represents a subject present in the scene may be determined. The output image may be generated further based on processing the 3D subject model by the image correction model.

[0104] In some examples, a semantic map corresponding to at least one of the input image or the reoriented image may be determined. The output image may be generated further based on processing the semantic map by the image correction model.

[0105] In some examples, the reoriented image may be determined by an image reorientation model. The visual distortion of the scene may be generated by the image reorientation model as part of the determination of the reoriented image. The image correction model may have been trained based on outputs of the image reorientation model to remove a plurality of types of visual distortions generated by the image reorientation model.

[0106] In some examples, the viewpoint modification may include one or more of (i) a translation relative to the first viewpoint or (ii) a rotation relative to the first viewpoint.

[0107] In some examples, generating the output image may include generating, using the image correction model and based on processing the input image and the reoriented image, an intermediate image that has a first resolution and includes the correction of the at least part of the visual distortion in the reoriented image. Generating the output image may also include generating, using a super-resolution model and based on processing the intermediate image and the input image thereby, the output image. The output image may have a second resolution that is greater than the first resolution.

[0108] In some examples, generating the output image may include generating a plurality of output images based on a plurality of different noise images. Each respective image of the plurality of output images (i) may be generated by the image correction model based on a corresponding noise image of the plurality of different noise images and (ii) may include a corresponding correction of at least a corresponding part of the visual distortion in the reoriented image. Outputting the output image may include outputting the plurality of output images to provide a plurality of different corrections of the visual distortion in the reoriented image.

[0109] In some examples, determining the viewpoint modification may include causing the input image to be displayed using a user interface, and obtaining, by way of the user interface, a user input representing the viewpoint modification.

[0110] In some examples, causing the input image to be displayed using the user interface may include generating an interactive representation of the input image. The interactive representation of the input image may be updated based on the user input to visually represent the viewpoint modification.[OHl] In some examples, the reoriented image may be determined using one or more of a three-dimensional polygon mesh, a point cloud, a three-dimensional Gaussian, a multiplane image, a neural radiance field, or a viewpoint-conditioned generative model.

[0112] In some examples, determining the viewpoint modification may include determining a plurality of candidate viewpoint modifications, causing a plurality of candidatereoriented images corresponding to the plurality of candidate viewpoint modifications to be displayed using a user interface, and obtaining, by way of the user interface, a selection of the reoriented image from the plurality of candidate reoriented images.

[0113] In some examples, the image correction model may include a generative machine learning model.

[0114] In some examples, the image correction model may have been trained by a training process that includes obtaining a training sample. The training sample may include (i) a training input image representing a training scene from a first training viewpoint, (ii) a ground-truth output image representing the training scene from a second training viewpoint that differs from the first training viewpoint, and (iii) a training reoriented image that represents the training scene from the second training viewpoint and includes a training visual distortion of the training scene. The training process may also include generating, using the image correction model and based on the training input image and the training reoriented image, a training output image by correcting at least part of the training visual distortion in the training reoriented image. The training process may additionally include determining a loss value based on comparing the training output image to the ground-truth output image, and adjusting one or more parameters of the image correction model based on the loss value.

[0115] In some examples, obtaining the training sample may include determining, based on the first training image and the second training image, a training viewpoint modification that relates the first training viewpoint to the second training viewpoint. Obtaining the training sample may also include determining, based on the training input image and the training viewpoint modification, the training reoriented image. The training visual distortion may be associated with the training viewpoint modification.

[0116] In some examples, determining the training viewpoint modification may include determining, based on one or more of the training input image or the ground-truth output image, one or more relative depth values associated with the training scene. Determining the training viewpoint modification may also include determining, based on (i) the training input image, (ii) the ground-truth output image, and (iii) a metric distance between a first camera position associated with the training input image and a second camera position associated with the ground-truth output image, one or more metric depth values associated with the training scene. Determining the training viewpoint modification may further include determining, based on the one or more relative depth values and the one or more metric depth values, a mapping between relative depth and metric depth. Determining the training viewpoint modification may yet further include determining, based on the mapping and the metric distance between the firstcamera position and the second camera position, a transformation between the first camera position and the second camera position. The transformation may be expressed in relative distance units used by a three-dimensional model in connection with generating the training reoriented image.

[0117] In some examples, the viewpoint modification may be smaller than a maximum viewpoint modification that corresponds to a maximum distance and / or rotation between training cameras used to generate training samples for training the image correction model.

[0118] Figure 7 illustrates a flow chart of operations related to training an image correction model. The operations may be carried out by computing device 100, computing system 200, and / or training system 500, among other possibilities. The embodiments of Figure 7 may be simplified by the removal of any one or more of the features shown therein. Further, these embodiments may be combined with features, aspects, and / or implementations of any of the previous figures or otherwise described herein. For example, the operations of Figure 7 may be used in combination with any of the examples discussed in connection with Figure 6.

[0119] Block 700 may involve obtaining a training sample. The training sample may include (i) a training input image representing a training scene from a first training viewpoint, (ii) a ground-truth output image representing the training scene from a second training viewpoint that differs from the first training viewpoint, and (iii) a training reoriented image that represents the training scene from the second training viewpoint and includes a training visual distortion of the training scene. The training visual distortion may be associated with a training viewpoint modification from the first training viewpoint to the second training viewpoint.

[0120] Block 702 may involve generating, using an image correction model and based on the training input image and the training reoriented image, a training output image by correcting at least part of the training visual distortion in the training reoriented image.

[0121] Block 704 may involve determining a loss value based on comparing the training output image to the ground-truth output image.

[0122] Block 706 may involve adjusting one or more parameters of the image correction model based on the loss value to configure the image correction model to remove visual distortions associated with viewpoint modifications.

[0123] Block 708 may involve outputting the image correction model.

[0124] In some examples, obtaining the training sample may include determining, based on the first training image and the second training image, a training viewpoint modification that relates the first training viewpoint to the second training viewpoint. Obtainingthe training sample may also include determining, based on the training input image and the training viewpoint modification, the training reoriented image.

[0125] In some examples, determining the training viewpoint modification may include determining, based on one or more of the training input image or the ground-truth output image, one or more relative depth values associated with the training scene. Determining the training modification may also include determining, based on (i) the training input image, (ii) the ground-truth output image, and (iii) a metric distance between a first camera position associated with the training input image and a second camera position associated with the ground-truth output image, one or more metric depth values associated with the training scene. Determining the training modification may additionally include determining, based on the one or more relative depth values and the one or more metric depth values, a mapping between relative depth and metric depth. Determining the training modification may further include determining, based on the mapping and the metric distance between the first camera position and the second camera position, a transformation between the first camera position and the second camera position. The transformation may be expressed in relative distance units used by a three-dimensional model in connection with generating the training reoriented image.VII. Conclusion

[0126] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as illustrations of various aspects. Many modifications and variations can be made without departing from its scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims.

[0127] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying figures. In the figures, similar symbols typically identify similar components, unless context dictates otherwise. The example embodiments described herein and in the figures are not meant to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0128] With respect to any or all of the message flow diagrams, scenarios, and flow charts in the figures and as discussed herein, each step, block, and / or communication can represent a processing of information and / or a transmission of information in accordance with example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be executed out of order from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. Further, more or fewer blocks and / or operations can be used with any of the message flow diagrams, scenarios, and flow charts discussed herein, and these message flow diagrams, scenarios, and flow charts can be combined with one another, in part or in whole.

[0129] A step or block that represents a processing of information may correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a block that represents a processing of information may correspond to a module, a segment, or a portion of program code (including related data). The program code may include one or more instructions executable by a processor for implementing specific logical operations or actions in the method or technique. The program code and / or related data may be stored on any type of computer readable medium such as a storage device including random access memory (RAM), a disk drive, a solid state drive, or another storage medium.

[0130] The computer readable medium may also include non-transitory computer readable media such as computer readable media that store data for short periods of time like register memory, processor cache, and RAM. The computer readable media may also include non-transitory computer readable media that store program code and / or data for longer periods of time. Thus, the computer readable media may include secondary or persistent long term storage, like read only memory (ROM), optical or magnetic disks, solid state drives, compactdisc read only memory (CD-ROM), for example. The computer readable media may also be any other volatile or non-volatile storage systems. A computer readable medium may be considered a computer readable storage medium, for example, or a tangible storage device.

[0131] Moreover, a step or block that represents one or more information transmissions may correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions may be between software modules and / or hardware modules in different physical devices.

[0132] The particular arrangements shown in the figures should not be viewed as limiting. It should be understood that other embodiments can include more or less of each element shown in a given figure. Further, some of the illustrated elements can be combined or omitted. Yet further, an example embodiment can include elements that are not illustrated in the figures.

[0133] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purpose of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method comprising: determining a viewpoint modification of a first viewpoint from which an input image represents a scene; determining, based on the input image and the viewpoint modification, a reoriented image that (i) represents the scene from a second viewpoint that differs from the first viewpoint and (ii) includes a visual distortion of the scene, wherein the visual distortion is associated with the viewpoint modification; processing the input image and the reoriented image using an image correction model configured to remove visual distortions associated with viewpoint modifications; generating, using the image correction model and based on processing the input image and the reoriented image, an output image that includes a correction of at least part of the visual distortion in the reoriented image, wherein the output image represents the scene from the second viewpoint; and outputting the output image.

2. The computer-implemented method of claim 1, further comprising: determining a defect mask that represents one or more portions of the scene depicted in the reoriented image and for which the reoriented image includes visual content associated with one or more visual distortions resulting from the viewpoint modification, wherein the output image is generated further based on processing the defect mask by the image correction model.

3. The computer-implemented method of claim 2, wherein the defect mask is based on an inpainting mask that indicates that the input image lacks image data for a portion of the one or more portions of the scene.

4. The computer-implemented method of any of claims 1-3, further comprising: obtaining a near-duplicate image that includes (i) first image data corresponding to a first portion of the scene, wherein the first portion of the scene is represented by the input image, and (ii) second image data corresponding to a second portion of the scene, wherein the second portion of the scene is not represented by the input image, and wherein the output image is generated further based on processing the near-duplicate image by the image correction model.

5. The computer-implemented method of any of claims 1-4, further comprising one or more of: determining a depth map corresponding to at least one of the input image or the reoriented image, wherein the output image is generated further based on processing the depth map by the image correction model; determining a feature map corresponding to at least one of the input image or the reoriented image, wherein the output image is generated further based on processing the feature map by the image correction model; determining a three-dimensional (3D) subject model that represents a subject present in the scene, wherein the output image is generated further based on processing the 3D subject model by the image correction model; or determining a semantic map corresponding to at least one of the input image or the reoriented image, wherein the output image is generated further based on processing the semantic map by the image correction model.

6. The computer-implemented method of any of claims 1-5, wherein the reoriented image is determined by an image reorientation model, wherein the visual distortion of the scene is generated by the image reorientation model as part of the determination of the reoriented image, and wherein the image correction model has been trained based on outputs of the image reorientation model to remove a plurality of types of visual distortions generated by the image reorientation model.

7. The computer-implemented method of any of claims 1-6, wherein the viewpoint modification comprises one or more of (i) a translation relative to the first viewpoint or (ii) a rotation relative to the first viewpoint.

8. The computer-implemented method of any of claims 1-7, wherein generating the output image comprises: generating, using the image correction model and based on processing the input image and the reoriented image, an intermediate image that has a first resolution and includes the correction of the at least part of the visual distortion in the reoriented image; andgenerating, using a super-resolution model and based on processing the intermediate image and the input image thereby, the output image, wherein the output image has a second resolution that is greater than the first resolution.

9. The computer-implemented method of any of claims 1-8, wherein: generating the output image comprises generating a plurality of output images based on a plurality of different noise images, each respective image of the plurality of output images is (i) generated by the image correction model based on a corresponding noise image of the plurality of different noise images and (ii) includes a corresponding correction of at least a corresponding part of the visual distortion in the reoriented image, and outputting the output image comprises outputting the plurality of output images to provide a plurality of different corrections of the visual distortion in the reoriented image.

10. The computer-implemented method of any of claims 1-9, wherein determining the viewpoint modification comprises: causing the input image to be displayed using a user interface; and obtaining, by way of the user interface, a user input representing the viewpoint modification.

11. The computer-implemented method of claim 10, wherein causing the input image to be displayed using the user interface comprises generating an interactive representation of the input image, and wherein the method further comprises: updating, based on the user input, the interactive representation of the input image to visually represent the viewpoint modification.

12. The computer-implemented method of claim 1-11, wherein the reoriented image is determined using one or more of a three-dimensional polygon mesh, a point cloud, a three- dimensional Gaussian, a multi-plane image, a neural radiance field, or a viewpoint-conditioned generative model.

13. The computer-implemented method of any of claims 1-9, wherein determining the viewpoint modification comprises: determining a plurality of candidate viewpoint modifications;causing a plurality of candidate reoriented images corresponding to the plurality of candidate viewpoint modifications to be displayed using a user interface; and obtaining, by way of the user interface, a selection of the reoriented image from the plurality of candidate reoriented images.

14. The computer-implemented method of any of claims 1-13, wherein the image correction model comprises a generative machine learning model.

15. The computer-implemented method of any of claims 1-14, wherein the image correction model has been trained by: obtaining a training sample comprising (i) a training input image representing a training scene from a first training viewpoint, (ii) a ground-truth output image representing the training scene from a second training viewpoint that differs from the first training viewpoint, and (iii) a training reoriented image that represents the training scene from the second training viewpoint and includes a training visual distortion of the training scene; generating, using the image correction model and based on the training input image and the training reoriented image, a training output image by correcting at least part of the training visual distortion in the training reoriented image; determining a loss value based on comparing the training output image to the groundtruth output image; and adjusting one or more parameters of the image correction model based on the loss value.

16. The computer-implemented method of claim 15, wherein obtaining the training sample comprises: determining, based on the first training image and the second training image, a training viewpoint modification that relates the first training viewpoint to the second training viewpoint; and determining, based on the training input image and the training viewpoint modification, the training reoriented image, wherein the training visual distortion is associated with the training viewpoint modification.

17. The computer-implemented method of claim 16, wherein determining the training viewpoint modification comprises:determining, based on one or more of the training input image or the ground-truth output image, one or more relative depth values associated with the training scene; determining, based on (i) the training input image, (ii) the ground-truth output image, and (iii) a metric distance between a first camera position associated with the training input image and a second camera position associated with the ground-truth output image, one or more metric depth values associated with the training scene; determining, based on the one or more relative depth values and the one or more metric depth values, a mapping between relative depth and metric depth; and determining, based on the mapping and the metric distance between the first camera position and the second camera position, a transformation between the first camera position and the second camera position, wherein the transformation is expressed in relative distance units used by a three-dimensional model in connection with generating the training reoriented image.

18. A computer-implemented method comprising: obtaining a training sample comprising (i) a training input image representing a training scene from a first training viewpoint, (ii) a ground-truth output image representing the training scene from a second training viewpoint that differs from the first training viewpoint, and (iii) a training reoriented image that represents the training scene from the second training viewpoint and includes a training visual distortion of the training scene, wherein the training visual distortion is associated with a training viewpoint modification from the first training viewpoint to the second training viewpoint; generating, using an image correction model and based on the training input image and the training reoriented image, a training output image by correcting at least part of the training visual distortion in the training reoriented image; determining a loss value based on comparing the training output image to the groundtruth output image; adjusting one or more parameters of the image correction model based on the loss value to configure the image correction model to remove visual distortions associated with viewpoint modifications; and outputting the image correction model.

19. A system comprising a processor configured to perform the method of any of claims 1-18.

20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a computing device, cause the computing device to perform the method of any of claims 1-18.

Citation Information

Patent Citations

  • Method for controlling a refrigerator and non-transitory computer usable medium having computer-readable instructions embodied therein for same

    US10041718B2

  • Single image 3D photography with soft-layering and depth-aware inpainting

    EP4150560B1

  • Layered view synthesis system and method

    WO2023235273A1