Real-time facial repair and reillumination in video using facial
By combining a neural network model with an autoencoder and a distortion classifier, the image restoration system solves the problem of poor video quality in low-light environments, achieves efficient and fast image restoration, and generates high-quality video streams.
Patent Information
- Application Number
- CN202480008555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-22
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-05
Smart Images

Figure CN120604261A_ABST
Abstract
Description
Background Art
[0001] In recent years, in digital communication (particularly in the field of video streaming), significant hardware and software progress has been arranged. For example, individuals increasingly participate in remote conferences that rely on video conferencing tools. Although existing systems are improving, due to both internal and external factors, they often stream poor-quality videos. For example, if an individual is in an environment with poor lighting, such as a dark room or a room with poor lighting conditions, the video quality will be impaired, making it difficult for the individual to see. In particular, lack of light on the subject increases the blur, noise, distortion and artifact in the video image. In addition, even in an ideal environment, lower-quality hardware components (such as, poor-quality webcams) can produce substandard video and images. As a result, existing systems must consume a large amount of computer resources to attempt to correct the low-quality image problem, which also results in increased waiting time and delay. In addition, in this case, despite these and other efforts, existing systems can usually not provide high-quality video streaming and face other problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The detailed description provides additional specificity and detail for one or more implementations through the use of the accompanying drawings, as briefly described below.
[0003] Figure 1A to Figure 1B An example flow chart providing an overview for implementing an image restoration system to restore and improve image quality is shown, according to one or more implementations.
[0004] Figures 2A to 2B An example computing environment for implementing an image restoration system in accordance with one or more implementations is shown.
[0005] Figures 3A to 3C An example block diagram for training and utilizing an image restoration machine learning model is shown in accordance with one or more implementations.
[0006] Figures 4A to 4B An example block diagram illustrating additional details of an image inpainting machine learning model according to one or more implementations is shown.
[0007] Figures 5A to 5B An example processing flow for generating real and synthetic data for training an image restoration machine learning model is shown in accordance with one or more implementations.
[0008] Figure 6 Example image results are shown for comparing an image inpainting system according to one or more implementations with other existing systems.
[0009] Figure 7
[0014] An example series of acts for generating, enhancing, repairing, and relighting a digital image according to one or more implementations is shown.
[0010] Figure 8 Example components included within a computer system are shown.
[0011] Specific implementation method
[0012] The present disclosure describes an image restoration system that accurately and efficiently generates high-quality images captured under low-quality and / or low-light environmental conditions. For example, for users participating in a video stream in a low-light environment, the image restoration system improves the quality of the image by dynamically relighting the user's face, and further enhances the image quality so that, among other benefits, other users watching the video stream are unaware of the user's adverse environmental conditions. For example, in some cases, the image restoration system simulates the addition of physical light to illuminate the subject captured by the camera. In addition, the image restoration system provides improved image accuracy that is superior to existing systems, while also being more efficient and significantly faster than existing systems.
[0013] For context, the image restoration system provides image enhancement, and in particular, facial enhancement technology that repairs and restores high-quality images captured in low-quality and / or low-light environmental conditions. In addition to the technical benefits of improved accuracy and efficiency described in detail below, enhancing facial quality in videos and images significantly improves the user experience in many applications. These applications include video conferencing, mobile applications, various displays, and cameras. The image restoration system corrects defects caused by different conditions that affect video quality, including lighting / exposure (e.g., dark rooms, windows, and lights), blur issues (e.g., camera losing focus, people moving), distance from the camera that reduces facial quality, different camera resolutions, and many more real-world scenarios.
[0014] More specifically, the image restoration system generates and utilizes an image restoration machine learning model to improve the quality of low-quality images by relighting and restoring the images in real time. In various instances, the image restoration machine learning model is a neural network that corrects various problems, such as low light, reflected colored light, image distortion, image noise, blur, poor exposure, and other related issues, thereby producing an enhanced image that accurately reflects the original scene. In many instances, the image restoration system generates an image restoration machine learning model that is significantly smaller, more efficient, and more accurate than existing systems. In various implementations, the image restoration machine learning model is implemented by pairing an autoencoder model with a distortion classifier model.
[0015] For example, the image restoration system identifies an image that includes the user's face and the image background, and detects (and in some cases, crops) the face within the image. The system then utilizes a face restoration machine learning model to generate a light-enhanced facial image. To achieve this, the system combines an autoencoder and a distortion classifier within the face restoration machine learning model. Furthermore, the system generates an enhanced digital image by combining the light-enhanced facial image with the image background. Finally, the enhanced digital image is provided for display on a computing device, such as the user's client device and / or the devices of the other user(s) participating in the video call.
[0016] It is important to note that the following document primarily discusses image inpainting systems in the context of relighting faces within images. However, methods and techniques similar to those described in this document can be used to relight other objects within digital images, including avatars generated from a user's face. For example, for poorly illuminated objects, the image inpainting system utilizes an object image inpainting machine learning model with similar autoencoders and distortion classifiers to relight and improve the appearance of objects within digital images.
[0017] Implementations of the present disclosure are directed to solving one or more of the above-mentioned problems and other problems in the art. For example, various systems, computer-readable media, and methods utilize an image restoration system to relight, restore, and / or reconstruct low-quality images into high-quality, well-lit images. In particular, the image restoration system utilizes an image restoration machine learning model with an autoencoder combined with the output from a distortion classifier to generate highly accurate images. Furthermore, the image restoration machine learning model utilizes a model architecture that results in significantly faster processing times than conventional systems.
[0018] Specifically, the image restoration system offers several technical benefits in terms of computational accuracy and efficiency compared to existing computing systems. The image restoration system provides benefits and addresses issues associated with relighting, repairing, and restoring low-quality images captured in harsh environments. Specifically, the image restoration machine learning model provides highly accurate images by utilizing an autoencoder combined with the output from a distortion classifier. Furthermore, the image restoration system employs a model architecture that yields significantly faster speeds than traditional systems, providing practical applications across a variety of industries.
[0019] As described above, the image restoration system improves accuracy and efficiency compared to existing systems. For illustration, the image restoration system generates and / or utilizes an image restoration machine learning model, which is a lightweight but effective machine learning model that enhances different types of distortions in digital images, such as light, downsampling, and noise. Furthermore, in many implementations, the image restoration machine learning model operates in real time.
[0020] In detail, the image restoration system generates an image restoration machine learning model that efficiently balances accuracy (e.g., facial quality) with computational cost. Many existing systems sacrifice accuracy, such as facial quality, for computational efficiency, and vice versa. For example, some existing systems are designed to restore extremely low-quality images by applying noise reduction and downsampling, which may destroy and distort the image. In addition, the goal of these existing systems is to take consumer photos where the subject is small and far away. Therefore, when the subject's face occupies most of the image, these existing systems overcorrect and produce inaccurate images.
[0021] Thus, the image restoration system provides an improved balance between accuracy and latency compared to existing systems. For example, in various implementations, the image restoration system generates an image restoration machine learning model that utilizes a model architecture having a space-to-depth layer (space2depth) on the encoder side of an autoencoder and a depth-to-space (depth2space) layer on the decoder side of the autoencoder, the space-to-depth layer (space2depth) followed by a convolutional neural network (CNN) layer, and the depth-to-space layer followed by a dense layer. In this way, the model architecture of the image restoration machine learning model provides increased accuracy at a low computational cost.
[0022] More specifically, in one or more implementations, by utilizing spatial-to-depth layers and depth-to-spatial layers, the image restoration system maintains lossless spatial dimensions when performing data reduction and data expansion. In contrast, most existing systems use classical downsampling and upsampling layers (e.g., bilinear, nearest neighbor), which results in information loss (e.g., inaccurate results). Furthermore, the image restoration system operates at significantly faster speeds than existing systems on small-capacity computing devices.
[0023] Some existing systems attempt to identify and correct third-party deblurred, denoised, or poorly illuminated images. However, these systems are limited to a small number of degradation types. In contrast, image restoration systems provide a full refiner model that corrects distortions caused by downsampling, blurring, exposure, light variations, poor lighting, color distortion, chroma degradation, Gaussian noise, and JPEG compression. In practice, image restoration systems combine the outputs from a distortion classifier into an autoencoder to correct the range of distortions in an image.
[0024] Furthermore, in many implementations, the distortion classifier and the autoencoder are trained together to maximize their collaborative contribution to each other in improving image accuracy. In other words, in many instances, the image restoration system trains the distortion classifier together with the autoencoder to further guide the autoencoder's decoder on how to recover from each specific distortion. As a result, the accuracy and efficiency of the image restoration machine learning model are improved. Furthermore, due to its targeted training, the image restoration system handles realistic, real-world situations better than existing systems. In fact, the image restoration system is trained based on real-world scenarios (e.g., video calls).
[0025] To further illustrate, in various implementations, the image restoration system avoids image artifacts and makes skin and facial textures appear more realistic by combining the output from the distortion classifier into the autoencoder. For example, generative adversarial network (GAN) priors often add image artifacts and make skin and facial textures appear unrealistic. Therefore, instead of using only face and GAN priors as in existing systems, the image restoration system incorporates the distortion class prior from the distortion classifier for use by the generator (i.e., decoder) of the autoencoder.
[0026] As explained in the foregoing discussion, the present disclosure utilizes various terms to describe the features and advantages of one or more implementations described. For example, as used herein, the term "digital image" (or, simply, "image") refers to one or more digital graphics files that, when rendered, display one or more pixels. In many cases, an image includes at least one face, while in some implementations, an image includes one or more objects. In addition, an image includes an image background that includes non-face pixels (or, non-object pixels) of the image that includes a face (or, object). In addition, an image can be part of a sequence of images, such as an image frame in a video or part of a sequence of images captured at different times.
[0027] As used in this document, the term "relighting" refers to adjusting the light displayed in an image, typically by adding a simulated light source. In some instances, relighting an image provides a computer-based solution that is equivalent to simulating the addition of physical light to illuminate a subject captured by a camera.
[0028] Furthermore, the term "machine learning model" refers to a computer model or computer representation that can be trained (e.g., optimized) based on input to approximate an unknown function. For example, a machine learning model may include, but is not limited to, an autoencoder model, a distortion classification model, a neural network (e.g., a convolutional neural network or a deep learning model), a decision tree (e.g., a gradient boosted decision tree), a linear regression model, a logistic regression model, or a combination of these models (e.g., an image restoration machine learning model including an autoencoder model (referred to as autoencoder) and a distortion classification model (referred to as distortion classification)).
[0029] As another example, the term "neural network" refers to a machine learning model that includes interconnected artificial neurons that communicate and learn to approximate complex functions, generating outputs based on multiple inputs provided to the model. For example, a neural network includes an algorithm (or set of algorithms) that employs deep learning techniques and utilizes training data to adjust network parameters and model high-level abstractions in the data. There are various types of neural networks, such as convolutional neural networks (CNNs), residual learning neural networks, recurrent neural networks (RNNs), generative neural networks, generative adversarial neural networks (GANs), and single-shot detection (SSD) networks.
[0030] As another example, the term "synthesized image" or "generated image" refers to an image produced by a system or model. For example, an image restoration system creates synthetic images to train an image restoration machine learning model. In some cases, the images are synthesized from some or all of a training image dataset.
[0031] As used herein, the terms "object mask," "segmentation mask," or "image mask" (or, simply, "mask") refer to an indication of a plurality of pixels within an image. In particular, image restoration systems utilize image masks to isolate pixels in a segmented region from other pixels in the image (e.g., to segment a face from the background). An image mask can be square, circular, or other closed shape.
[0032] Additional details regarding example implementations of the image restoration system are discussed in conjunction with the following figures. For example, Figure 1A to Figure 1B An example flow chart providing an overview for implementing an image restoration system to restore and improve image quality according to one or more implementations is shown. In particular, Figure 1A Included is a series of acts 100 in which an image restoration system enhances a user's face during a video call by utilizing a face restoration machine learning model. Figure 1B The overall process of generating an enhanced digital image from a digital image using a face restoration machine learning model and other components is shown.
[0033] To illustrate, in Figure 1AIn FIG, a sequence of actions 100 includes action 102, which captures a user's face from a digital image during a video in which the face is poorly illuminated. For example, the image inpainting system detects, captures, and tracks the user's face within one or more video frames or another image. As shown, images are often poorly illuminated, resulting in low-quality characteristics such as noise, blur, and distortion. As described above, the image inpainting system can similarly perform the sequence of actions 100 with respect to a target object other than a face.
[0034] As shown, the series of actions 100 includes an action 104 of using a face inpainting machine learning model to generate a light-enhanced facial image that relights and refines the face. For example, the image inpainting system utilizes a face inpainting machine learning model that includes an autoencoder and a distortion classifier to generate a light-enhanced facial image that shows a well-lit face from a captured image of a dimly lit face. Generating and utilizing the face inpainting machine learning model is further described below.
[0035] The series of acts 100 includes an act 106 of replacing the user's face in the digital image with a light-enhanced facial image. For example, in various implementations, the image restoration system combines the light-enhanced facial image of the well-illuminated face with the image background of the digital image. In this way, the image restoration system selectively targets the user's face for enhancement, thereby conserving computing resources. Additional details regarding the insertion of the enhanced face are provided below in conjunction with subsequent figures.
[0036] The sequence of actions 100 also includes an action 108 of repeating the sequence of actions in real time with other digital images during the video to show a well-lit face. Furthermore, in most cases, the image restoration system employs a face restoration machine learning model to enhance several images or video frames from a video stream in real time. The architecture and objective training of the face restoration machine learning model enables the image restoration system to perform image enhancement in real time, even on devices with limited computing resources (including NPUs), such as network processing unit-based devices.
[0037] As mentioned earlier, Figure 1B Provides a general overview of how image restoration systems utilize face restoration machine learning models and other components to create improved digital images from original digital images. Figure 1B , a digital image 112 is shown, which can originate from a variety of sources, such as a video stream or a captured image. The digital image is typically obtained from a live or real-time image, but the digital image can also include stored images, such as those from video playback. For example, in the case of a shared video call with other users, the digital image 112 would be a video frame.
[0038] As shown, digital image 112 includes the user's face and an image background. Often, digital image 112, and in particular the user's face, is poorly illuminated due to poor lighting around the user (e.g., insufficient lighting, poorly placed lighting, or reflections of colored light from the device screen onto the user's face). In some implementations, the hardware of the camera device capturing the user may be insufficient to capture a high-quality image of the user in a poorly lit environment.
[0039] In some implementations, digital image 112 displays the faces of multiple users. For example, multiple users can participate in a video call from the same location and can be included in the same image. In various implementations, digital image 112 shows an object (e.g., a target object) instead of or in addition to the user's face.
[0040] As shown in the figure, Figure 1B A face tracking model 114 is included. For example, the face tracking model 114 identifies the presence of a face (e.g., a user's face) in the digital image 112, detects the location of the face within the digital image 112, tracks the face, and / or crops the user's face. In some implementations, the face tracking model 114 is a machine learning model and / or a neural network.
[0041] In various implementations, facial tracking model 114 crops the detected face to generate a cropped facial image 116 of the user's face and the image background surrounding the face. In alternative implementations, facial tracking model 114 does not crop the user's face. In some instances, facial tracking model 114 determines that the user's face does not need to be cropped (e.g., the user's face occupies a large portion of the image).
[0042] In one or more implementations, the facial tracking model 114 detects and tracks the user's face. For example, the facial tracking model 114 tracks the face between a series of images (e.g., between consecutive video frames). Facial tracking can include vertical and horizontal movement as well as movement toward or away from the camera. In some instances, the facial tracking model 114 also tracks when the user turns their face or when a face disappears from one image and reappears in a subsequent image.
[0043] Furthermore, facial tracking model 114 can detect and track multiple faces within an image. For example, if digital image 112 includes two faces, facial tracking model 114 can track each of the two faces separately and generate a cropped facial image of each face. Similarly, an image restoration system can use an object tracking model to track and capture a target object instead of, or in addition to, capturing a face.
[0044] By utilizing the facial tracking model 114, the image restoration system improves overall efficiency. For example, using the image restoration machine learning model on smaller target images requires less computer processing. Furthermore, tracking allows for more efficient processing between frames, and allows the light-enhanced image to transition smoothly between frames, rather than appearing jumpy and glitchy.
[0045] In some implementations, the image restoration system tracks the user's body relative to the user's face to enhance the lighting of the entire user, not just their face. To achieve this, the image restoration system uses an additional tracking model and / or background segmentation model to generate a cropped image of the user's body. In some implementations, the image restoration system utilizes a scale factor generated for the light-enhanced facial image to generate the body-enhanced image. Furthermore, the image restoration system can enhance the user's body and face separately and intelligently combine the two to generate a composite image.
[0046] Figure 1B 1 shows how the image restoration system works. The system takes a cropped facial image 116 and provides it to a facial restoration machine learning model 120, which includes an autoencoder 122 and a distortion classifier 124. The model generates a light-enhanced facial image 126 and a facial image mask 128 that separates the light-enhanced face from the image background within the light-enhanced facial image 126 (or, the cropped facial image 116). Figures 3A to 3C Additional details about the face restoration machine learning model 120 can be found in .
[0047] Figure 1B Also included is an image blending model 130. For example, the image inpainting system utilizes the image blending model 130 to blend the light-enhanced facial image 126 with the background in the cropped facial image 116 (or some or all of the image background of the digital image 112) to generate a relighted cropped facial image 132. For example, the image blending model 130 utilizes the facial image mask 128 to separate the relighted facial portion from the background portion on the light-enhanced facial image 126 and / or the cropped facial image 116 before combining the relighted face with the image background. In various implementations, the image blending model 130 utilizes alpha blending or other blending techniques to make the relighted face and the image background appear natural.
[0048] As shown, the image inpainting system generates an enhanced digital image 134 from the digital image 112 and the relighted cropped facial image 132. For example, the image inpainting system replaces the portion corresponding to the user's face (and surrounding area) with the relighted cropped facial image 132. In some implementations, the image inpainting system generates the enhanced digital image 134 directly from the digital image 112 by using the light-enhanced facial image 126, the facial image mask 128, and the image blending model 130 to generate the enhanced digital image 134 directly from the digital image 112, rather than indirectly generating the relighted cropped facial image 132.
[0049] As described above, in various instances, the image restoration system relights both the user's face and body. For example, the image restoration system uses the facial restoration machine learning model 120 or another image restoration machine learning model to generate an enhanced image of the user's body and a corresponding body or person image mask. In addition, the image restoration system utilizes the image blending model 130 to blend the user's body with the image background of the digital image 112 (or a cropped portion) to produce a modified digital image. The image blending model 130 is then used to blend the light-enhanced facial image 126 with the modified digital image to directly or indirectly generate an enhanced digital image 134.
[0050] Following a general overview of the image restoration system, additional details are provided regarding the components and elements of the image restoration system. For illustration, Figures 2A to 2B An example computing environment and architecture diagram of an image restoration system are provided. Specifically, Figures 2A to 2B An example computing environment is shown for implementing an image restoration system in accordance with one or more implementations.
[0051] To illustrate, Figure 2A A computing environment 200 is shown that includes a client device 202 and a server device 208 that are connected to each other via a network 212. The client device 202 includes a digital communication system 204 and an image restoration system 206. The server device 208 includes an image relighting server system 210. Figure 8 Provides additional details about these and other computing devices. Additionally, Figure 8 Additional details regarding a network (eg, network 212 as shown) are also provided.
[0052] Although Figure 2A An example arrangement and configuration of computing environment 200 is shown, but other arrangements and configurations are possible. For example, computing environment 200 may include additional client devices that communicate with each other. As another example, image restoration system 206 may be entirely located on server device 208, which facilitates video calls between client devices.
[0053] As mentioned above, Figure 2A Client device 202 in FIG. 1 includes a digital communication system 204. In various implementations, digital communication system 204 manages digital image communication between computing devices. For example, digital communication system 204 facilitates video calls or other video streaming between computing devices. In some implementations, digital communication system 204 manages the capture, storage, and / or access of digital images, which may include a user's face and / or a target object.
[0054] As shown, the digital communication system 204 includes an image restoration system 206. In some implementations, the image restoration system 206 is located outside the digital communication system 204 (e.g., on the client device 202 or on another device). Generally, the image restoration system 206 utilizes an image restoration machine learning model to accurately and efficiently generate high-quality, well-lit images from low-quality, poorly lit images under various conditions. In addition to selectively relighting the image, the image restoration system 206 also restores and recaptures the user's face (or other object) from noise and distortion caused by poorly lit environments and / or poorly functioning cameras. In combination Figure 2B and subsequent figures provide additional details regarding the image restoration system 206 .
[0055] Additionally, the computing environment 200 includes a server device 208 having an image relighting server system 210. In various implementations, the image relighting server system 210 provides a version of the image inpainting system 206 to the client device 202. In some implementations, the image relighting server system 210 receives a video feed from the client device 202, generates a light-enhanced video feed (e.g., a set of images), and provides the light-enhanced video feed to the computing devices of other participants and / or the client device 202. In some implementations, the image relighting server system 210 trains and updates an image inpainting machine learning model (e.g., a face inpainting machine learning model) offline and provides it to the client device 202 so that the image inpainting system 206 can apply the image inpainting machine learning model.
[0056] Figure 2B 2 shows an image restoration system 206 within a digital communication system 204 on a computing device 201. The computing device 201 can be used as a client device 202, a server device 208, or another computing device. Figure 2BAs shown, the image restoration system 206 includes various components and elements implemented in hardware and / or software. For example, the image restoration system 206 includes a digital image manager 222, a face / object tracking manager 224, an image restoration model manager 226, an image blending manager 228, and a storage manager 230 including a digital image 232, an image restoration machine learning model 234 having an autoencoder 236 and a distortion classifier 238, and other image models 240.
[0057] Generally, digital image manager 222 manages the capture, storage, access, and other management of digital images 232. In various examples, face / object tracking manager 224 detects faces or objects in digital images 232 and generates cropped images. In some cases, face / object tracking manager 224 utilizes one of other image models 240, such as a face / object tracking model, to track faces or objects between a series of images.
[0058] In many implementations, the image restoration model manager 226 generates, trains, and utilizes the image restoration machine learning model 234. In some implementations, the image restoration model manager 226 also uses an image model among other image models, such as a facial feature extraction model, the output of which is provided to the distortion classifier 238. Furthermore, in many implementations, the output of the distortion classifier 238 is combined midway into the autoencoder 236, as further described below. In some implementations, the image restoration model manager 226 generates an avatar based on the user's face.
[0059] Furthermore, in various implementations, the image blending manager 228 blends the light-enhanced image from the image inpainting machine learning model 234 with the corresponding digital image to generate an enhanced image in which the user's face is clear and bright. In these implementations, the image blending manager 228 also uses a corresponding segmentation mask also generated by the image inpainting machine learning model 234 to generate the enhanced image. In some examples, the image inpainting system 206 uses one of other image models, such as an image blending model.
[0060] Figures 3A to 3C An example block diagram for training and utilizing an image restoration machine learning model according to one or more implementations is shown. Specifically, Figure 3A shows an example of training an image restoration machine learning model, while Figure 3B Shown is the use of a trained image inpainting machine learning model. Figure 3C Provides an example architecture for an image restoration machine learning model.
[0061] As shown in the figure, Figure 3AIt includes training data 302, image restoration machine learning model 310, and loss model 340. In addition, Figure 3A A facial feature extraction model 330 is included, which may be present when the image inpainting machine learning model 310 is a face inpainting machine learning model. In the case where the image inpainting machine learning model 310 is an object inpainting machine learning model, an object feature extraction model may be added when advantageous. In some implementations, the image inpainting machine learning model 310 involves a GAN and / or a U-Net.
[0062] As shown, training data 302 includes training images 304 and ground truth re-lighted images 306. Training data 302 may include images of faces or objects in varying lighting environments, and corresponding ground truth images of faces in well-lit environments. In some instances, training data 302 includes real and / or synthetic images. Figures 5A to 5B Provides further details on generating synthetic images for training data or creating training data from real images.
[0063] As shown, the image restoration machine learning model 310 includes an autoencoder 312 and a distortion classifier 322. The autoencoder 312 includes an encoder 314, a decoder 315, and a connection layer 318. In many instances, the decoder 315 acts as a generator that generates an image by reconstructing the feature vector of the processed input image.
[0064] The image restoration machine learning model 310 is designed as a hybrid model that utilizes depth-to-spatial layers and dense layers to handle various distortions more efficiently. The model architecture of the image restoration machine learning model 310 includes spatial to deep on the encoder 314 (followed by CNN layers) and depth to spatial on the decoder 315 (followed by dense layers). Both depth-to-spatial and depth-to-spatial operations ensure lossless spatial dimensions when reducing or expanding / increasing data during processing. In contrast, existing systems use traditional downsampling and upsampling layers, such as bilinear and nearest neighbor, resulting in information loss. The image restoration machine learning model 310 strikes a balance between high accuracy and low computational cost, and runs much faster than existing models. The following is combined with Figure 4A Provides more information about depth-to-spatial and depth-to-spatial operations.
[0065] The image restoration system 206 utilizes training data 302, including training images 304, to provide input to an autoencoder 312 and a distortion classifier 322 via an encoder 314. The encoder 314 processes the training data 302 and generates an encoded feature vector 316 (e.g., a latent vector), which is sent to a connection layer 318. Simultaneously, the distortion classifier 322 processes the training data 302 to generate a distortion classification 324 (e.g., a distortion class prior), which is also provided to the connection layer 318. The encoded feature vector 316 and the distortion classification 324 are concatenated at the connection layer 318, resulting in an improved input to the autoencoder 312.
[0066] To improve the efficiency of the distortion classifier 322, the facial feature extraction model 330 pre-processes the training data 302 by generating facial features of the faces in the input image. The image restoration system 206 provides the pre-trained features to the facial feature extraction model 330, and the facial feature extraction model 330 outputs the extracted facial features as input to the distortion classifier 322. This approach enables the image restoration machine learning model 310 to generate the distortion classification 324 from the extracted facial features, rather than relying solely on the input image.
[0067] As previously described, the encoded feature vector 316 is concatenated or supplemented with the distortion classification 324 at the connection layer 318. The connection layer 318 then provides the modified encoder output 320 to the decoder 315, which generates a light-enhanced image 326. Furthermore, in various examples, the decoder 315 generates a corresponding instance of a segmentation mask 328 for the light-enhanced image 326. Furthermore, in some implementations, the decoder 315 generates an avatar of the user's face, where the avatar is based on the light-enhanced version of the user's face. The loss model 340 receives the light-enhanced image 326 and / or the segmentation mask 328 for training purposes.
[0068] Many existing systems use only face and GAN priors (e.g., generated pre-trained face GAN features), which can result in artifacts and unrealistic faces and facial textures. The image restoration system 206 addresses this problem by incorporating a classifier, such as a distortion classifier 322, to learn and send a signal to the autoencoder 312 about the type of distortion applied to each input image. The distortion classification 324 of the input image is integrated into the autoencoder 312 and applied to the encoder 314 when generating the light-enhanced image 326. Thus, the image restoration machine learning model 310 integrates the distortion classifier 322, which predicts the type of degradation in the input image and provides this category information (i.e., the distortion classification 324) as a priori information in the autoencoder 312.
[0069] Regarding the distortion classifier 322, as described above, the image restoration system 206 improves upon the existing system by incorporating the distortion classification from the distortion classifier 322 into the autoencoder 312. The distortion classifier 322 is trained to recognize various types of distortion, including noise, blur, shake, exposure, low light level, light hue, color, chroma, and image resizing. It can also detect and signal distortion from screen lighting, such as white or colored light reflected from the user's face.
[0070] By utilizing the distortion class prior, the image restoration system 206 generates and provides a distortion classification 324 to preserve facial texture and user identity (e.g., preserve the identity of the person without adding any artifacts to the face or changing the texture of the face). In some implementations, the distortion classifier 322 is a lightweight CNN-based classifier that learns the type of distortion to apply to each image. The image restoration system 206 combines these distortion types into the encoder features of the autoencoder 312 to generate a more accurate version of the light-enhanced image 326.
[0071] In many cases, such as when training the image restoration machine learning model 310, the image restoration system 206 utilizes a degradation model. Typically, the degradation model generates low-light and / or distorted images in real time for training. For example, the degradation model synthetically and / or randomly applies multiple degradation techniques to high-quality (HQ) images during training, so that the image restoration machine learning model 310 learns to restore degraded low-quality (LQ) images to be as close as possible to the corresponding HQ images. In various implementations, the degradation model uses smoother but more diverse types of distortion to train the image restoration machine learning model 310.
[0072] In various implementations, the degradation model is a full refiner model that integrates downsampling, blur, exposure / light variation, illumination (color distortion), chroma degradation, Gaussian noise, and JPEG compression in the autoencoder 312. To illustrate, in some implementations, the degradation model follows the following formula:
[0073]
[0074] In this formula, ↓, ★, η, C, N, k, and JPEG represent downsampling, exposure, color dithering (e.g., lighting, contrast, saturation, hue), chroma, noise, blur, and JPEG compression, respectively. In addition, in various implementations, the image restoration system 206 randomly samples r, e, γ, δ, σ, and q.
[0075] Algorithm 1 provided below illustrates example steps of the degradation model used by the image restoration system 206 .
[0076]
[0077]
[0078] Algorithm 1-Degradation Algorithm
[0079] Algorithm 1 shows the input to the image restoration system 206. It consists of a list of high-quality (HQ) images for each batch. dist (e.g., the percentage of images that will be distorted for each batch), and the distortion range for each distortion type (r, e, γ, δ, σ, and q). For each image, the image restoration system 206 first checks whether the distortion percentage limit has been reached for the batch. If the limit has not been reached, the system applies reduction, exposure changes to each RGB channel, color dithering (e.g., by simulating different lighting on a face or object), chroma, additive white Gaussian noise, Gaussian blur convolution, and compresses the image using JPEG operations. If the image exceeds the Percent dist If the image inpainting system 206 uses the same image without any distortion as the target image, the image inpainting system 206 allows the autoencoder 312 to see both LQ and HQ images during training, increasing robustness and ensuring that the image is not over-enhanced, which may produce unnecessary and over-compensated unnatural artifacts.
[0080] Furthermore, Algorithm 1 shows that the image restoration system 206 can train the distortion classifier 322 to predict the type of distortion present in the image using three main categories, namely noise, blur, and exposure. As previously described, the image restoration system 206 trains the distortion classifier 322 in parallel with the autoencoder 312.
[0081] The degradation model is used not only to apply degradation to each HQ image, but also to generate ground truth class labels (ground truth relighted image 306) for the distortion classifier based on the type of distortion present. Distortion classes typically include noise (e.g., Gaussian, JPEG, and chroma), blur (Gaussian blur and shrinkage), and exposure.
[0082] With respect to the decoder 315, in various implementations, the decoder 315 is an image generator that processes the modified encoder output 320 (encoder / classifier cascade output) and applies a series of depth-to-space operations (e.g., rearranging data from depth (channels) to space (weights and height)), followed by a dense block. The image restoration system 206 utilizes two CNN layers (with channel sizes of 3 and 1) in the decoder 315 to produce a light-enhanced image 326 and a segmentation mask 328.
[0083] Furthermore, in various implementations, the image restoration system 206 preserves the identity of the person (e.g., the model does not change the identity of the person). For example, the image restoration system 206 trains the model to not introduce additional artifacts that could change the identity of the person in terms of age, gender, skin color, makeup, or other facial features.
[0084] In some implementations, the light-enhanced image 326 and the segmentation mask 328 have a cropped face (or, a cropped object). To achieve this, the image restoration system 206 uses facial segmentation to allow the model to focus only on enhancing facial areas while predicting facial boundaries.
[0085] With respect to training, the loss model 340 compares the light-enhanced image 326 and / or the segmentation mask 328 generated by the image inpainting machine learning model 310 with the ground-truth re-illuminated image 306 corresponding to the training image 304 provided to the image inpainting machine learning model 310 to determine the amount of loss or error. In some implementations, the loss model 340 discards background information (e.g., using the segmentation mask 328) and focuses only on enhancing the facial region of the light-enhanced image 326 while also improving how the facial region is segmented.
[0086] To determine the amount of loss, the loss model 340 can use one or more loss functions. The amount of loss is provided back to the image restoration machine learning model 310 as feedback 354 to adjust the weights, parameters, layers and / or nodes of the model. Loss functions or types include pixel-by-pixel loss 342, feature loss 344, texture information loss 346, adversarial loss 348, classification loss 350, and segmentation loss 352. In addition, the image restoration system 206 trains the distortion classifier 322 in parallel with the autoencoder 312 to better train the decoder 315 on how to remove specific types of real-world distortions. The image restoration machine learning model 310 is trained in an end-to-end manner via backpropagation until the model converges or meets another training criterion.
[0087] Image restoration system 206 utilizes various loss functions. For example, pixel-wise loss 342 is a reconstruction loss type that measures the error between the light-enhanced image 326 (i.e., the predicted image) and the ground-truth re-illuminated image 306. Feature loss 344 is a perceptual loss type that measures the error between the predicted image and the ground-truth image in the high-level feature maps of the facial feature extraction network. Texture information loss 346 is a style loss type that measures the error between the predicted image and the matrix representing the ground-truth image.
[0088] In addition, adversarial loss 348 represents an adversarial loss type that measures the error value of the loss from the generator (e.g., decoder 315) of the GAN model. Classification loss 350 measures the error value of the cross entropy loss with respect to distortion classifier 322, which predicts the type of distortion applied to the image. Segmentation loss 352 represents a dice loss type that measures the error value of the overlap between the predicted mask and the true value mask. In some implementations, the image restoration system 206 may also use a color enhancement loss.
[0089] In various implementations, the image restoration system 206 trains the image restoration machine learning model 310 to be computationally efficient. For example, the image restoration system 206 incorporates synthetic distortions that simulate real-world distortions into the training data 302 and uses less severe degradation scales than existing systems, which requires less computation to achieve more accurate results. Furthermore, compared to existing systems, the image restoration system 206 utilizes a reduced set of parameters, allowing the image restoration system 206 to operate in real time across various computing devices (e.g., 80-100 frames per second of inference on a low-power neural processing unit (NPU)).
[0090] Once the image restoration machine learning model 310 has been trained, it enables the image restoration system 206 to achieve various purposes. For example, it enhances noisy, blurred, or low-quality faces, while also restoring faces under different lighting and exposure conditions. In some cases, the image restoration machine learning model 310 generates segmentations of facial regions, which is useful for post-production purposes such as real-time video editing and streaming. In several implementations, instead of blending the light-enhanced face or person with the original image background, the image restoration system 206 combines the light-enhanced face or person with a different background. Thus, the image restoration system 206 eliminates the need for a green screen type background and can also work in real time.
[0091] To illustrate the functionality of the image inpainting system 206 using a trained image inpainting machine learning model for inferring an input image into an enhanced relighted image, Figure 3B A visual representation is provided. Specifically, the diagram depicts the image restoration system 206 providing a cropped image 370 to an image restoration machine learning model 310, which in this case is a face restoration machine learning model. The image restoration machine learning model 310 uses an autoencoder 312 in conjunction with distortion classification from a distortion classifier 322 to generate a light-enhanced cropped image 372 and a corresponding image mask 374. Furthermore, the image restoration system 206 utilizes an image blending model 376 to generate a blended, re-illuminated cropped image 378, which is used to create the enhanced image, as previously described.
[0092] In various implementations, the encoder 314 of the autoencoder 312 receives the output from the degradation model (LQ image) and passes through several shuffling layers (spatial to deep) followed by dense blocks. The spatial to deep layers rearrange the data from space (weights and height) to depth (channels), allowing lossless spatial dimensionality increase. As previously mentioned, the spatial to deep layers are computationally efficient.
[0093] In some implementations, the distortion classifier 322 predicts the type of degradation present in the input. The image restoration system 206 uses this additional information to help the decoder 315 restore and enhance the image. Thus, in various implementations, as previously described, the image restoration system 206 first extracts LQ image features from the facial feature extraction model 330 and passes them along with one or more labels generated by the degradation model to the distortion classifier 322. Furthermore, the output of the distortion classifier 322 (i.e., the distortion classification 324) is concatenated with the final output of the encoder (i.e., the encoded feature vector 316).
[0094] Figure 3C An exemplary architecture 380 of an image restoration machine learning model 310 is shown. Notably, the autoencoder 312 includes various spatial to depth layers (shown as S2D) in the encoder 314, while the decoder 315 includes a depth to spatial layer (shown as D2S). In addition, the autoencoder 312 includes various other layers, including a decoder that receives a cropped image (e.g., X ~ ) of the input layer, dense layer, convolution layer, batch normalization (BN) layer, rectified linear activation function (ReLU) layer, pooling layer, and cascade layer (e.g., in connection layer 318). The distortion classifier 322 also includes a convolution layer, a BN layer, a ReLU layer, a pooling layer, a dense layer and / or another neural network layer.
[0095] Figures 4A to 4B An example block diagram illustrating additional details of an image inpainting machine learning model in one or more implementations is shown. Figure 4A Corresponding to the depth-to-spatial layer. The spatial-to-depth layer in the encoder rearranges the data from depth (channels) to space (weights and height) while maintaining the lossless spatial dimension during data reduction. The depth-to-spatial layer in the decoder performs the opposite function (rearranging the data from space to depth).
[0096] exist Figure 4A , a visual example of a depth-to-spatial layer that rearranges data from depth to space is shown. The example includes depth 402 and space 404. In various implementations, the depth-to-spatial layer performs an operation that outputs a copy of the input tensor, moving values from the depth dimension to the height and width dimensions in a spatial block.
[0097] Figure 4BAn example of a dense block layer 410 (e.g., a dense block module) within an autoencoder 312 is shown. The dense block layer 410 includes a convolutional (conv) layer, a batch normalization (BN) layer, a rectified linear activation function (ReLU) layer, and several residual connections. The dense block layer 410 typically follows the spatial-to-depth layer in the encoder 314 and the depth-to-spatial layer in the decoder 315. In various implementations, each dense block layer 410 extracts rich local features via dense residual connections.
[0098] Turning to the next accompanying figure, Figures 5A to 5B An example process flow is shown for generating real and synthetic data for training an image restoration machine learning model in one or more implementations. Figure 5A describes the process of generating training data from real images, while Figure 5B The process of generating training data from synthetic images is shown.
[0099] As previously described, the image restoration system 206 generates training data to ensure that the image restoration machine learning model 310 is trained to enhance low-quality images under real-world conditions without making the training too broad, which would hinder the model's ability to handle non-real-world conditions it is unlikely to encounter. By doing so, the image restoration system 206 keeps the model lightweight, small, and efficient.
[0100] Figure 5A It shows how the image inpainting system 206 starts with a real image 502 and applies various image distortions to generate a distorted training image 506. The image inpainting system 206 also generates a ground truth image 508 from the real image 502. In many instances, the ground truth image is the original input image.
[0101] The image distortion model 504, which includes various types of distortion functions (e.g., based on size, exposure, noise, blur, color, and color dithering), distorts the real image 502 into a distorted training image 506. The image inpainting system 206 can utilize additional or alternative types of distortion functions.
[0102] In some implementations, the image restoration system 206 applies colored lighting with varying degrees of illumination, contrast, saturation, and hue to the real image 502. This method generates training data that trains the image restoration machine learning model to detect and remove unwanted colored lighting from the input image when enhancing the image through relighting and other processes.
[0103] exist Figure 5B In
[0045] , the image restoration system 206 is described as creating synthetic images for training data. For illustration, Figure 5BAn act of generating a synthesized face is included 512. In various implementations, the image restoration system 206 generates the synthesized face using three-dimensional modeling or rendering software that can be randomized with various hairstyles, facial features, and accessories.
[0104] also, Figure 5B An action 514 is included in which multiple digital images of a composite face are captured from multiple light sources. For example, the image restoration system 206 generates composite light sources such as ambient light, artificial light, and device screen light, which can vary in color, illumination, color temperature, bulb type, and time of day. These light sources can be located anywhere to illuminate the subject's face.
[0105] Using random selections of light sources and lighting positions, the image restoration system 206 generates multiple sets of images for a single face. In some implementations, the selections are weighted to select one light source type more than another. In alternative implementations, the image restoration system 206 generates one or more images that include light from multiple light sources having the same or different positions.
[0106] As shown in the figure, Figure 5B The process includes combining the multiple images in the set into a combined composite image 516. In this way, the image restoration system 206 generates a composite image that simulates the user's real-world scene. For example, once the individual light source images are rendered, the image restoration system 206 combines the multiple images in the set using a weighted summation to generate multiple composite images with multi-source lighting conditions. The process 516 may also involve further processing the summed image to add effects such as noise, chromatic aberration, color, and lens distortion, such as Figure 5A As described in .
[0107] As also shown, action 518 involves capturing a real image of the synthesized face, such as direct natural light, a ring light, or another light source that a real-world user would use to capture their face on camera. In various implementations, the image restoration system 206 associates the real image with corresponding multi-source illumination images of the same subject's face and adds it to the training data.
[0108] Throughout this document, it has been mentioned that the image restoration system 206 provides significant benefits and improvements over existing systems. For example, in Figure 6 In , we can see the comparison between the image restoration system and other existing systems. Specifically, Figure 6 A first enhanced face set 602 (eg, women) and a second enhanced face set 604 (eg, men) are included, as well as an input image 606 and a ground truth image 612 for comparison.
[0109] As also shown in the figure, compared with the prior art system 610, Figure 6 Shown are enhanced versions of two faces based on image inpainting system 206, referred to as "image inpainting model 608." Prior art system 610 is a multi-degradation model focused on face inpainting and has been published within the past two years. As shown in the image inpainting model 608 image, image inpainting system 206 improves low-light noise and blur compared to the enhanced images generated by prior art system 610. Furthermore, image inpainting system 206 produces more natural and realistic results. In other evaluations, image inpainting system 206 was found to do a better job of removing color lighting from faces.
[0110] The researchers also conducted multiple performance evaluations, including peak signal-to-noise ratio (PSNR), pixel-by-pixel metrics such as structural similarity (SSIM), and learned perceptual image patch similarity (LPIPS). These evaluations revealed that the image restoration system 206 performed with the highest accuracy level, the fastest speed (2X-4X times faster than other evaluated models), and the smallest model size (4X-20X smaller than other evaluated models).
[0111] Now turn Figure 7 , which depicts an example flow chart overview of a series of actions 700 utilizing the image restoration system 206 according to one or more implementations. Specifically, Figure 7
[0014] An example series of acts for generating, enhancing, repairing, and relighting a digital image according to one or more implementations is shown.
[0112] Although Figure 7 Actions according to one or more implementations are shown, but alternative implementations may omit, add, reorder, and / or modify any of the actions shown. Figure 7 The actions of may be performed as part of a method (e.g., a computer-implemented method). Alternatively, a non-transitory computer-readable medium may include a program that, when executed by a processing system including a processor, causes a computing device to perform Figure 7 In another implementation, a system (e.g., a processing system including a processor) may execute Figure 7 action.
[0113] As shown, the series of actions 700 includes an action 710 of detecting a face within an image. For example, in an example implementation, action 710 involves detecting a face within a digital image, the digital image comprising the face and an image background. In some implementations, action 710 includes detecting a face within a digital image by identifying the face within the digital image using a face tracking model; cropping the face within the digital image to generate a cropped image; providing the face image to an image inpainting machine learning model (a face or object inpainting machine learning model); and / or tracking the face across a set of consecutive digital images.
[0114] As further shown, the series of actions 700 includes an action 720 of utilizing an image inpainting machine learning model to generate a light-enhanced facial image of a face. For example, in an example implementation, action 720 involves utilizing an image inpainting machine learning model that includes an autoencoder and a distortion classifier to generate a light-enhanced facial image of a face within a digital image. In one or more implementations, action 720 includes utilizing the image inpainting machine learning model to generate a light-enhanced facial image from the facial image by combining the outputs of the distortion classifier and the encoder as input to a generator. In some implementations, action 720 includes improving lighting on the face, inpainting low-quality facial features to higher-quality facial features, reducing blur, and reducing noise.
[0115] In some implementations, the image inpainting machine learning model generates an enhanced digital image that improves the lighting on the light-enhanced facial image relative to the digital image and inpaints low-quality facial features into higher-quality facial features in the light-enhanced facial image. In various cases, the image inpainting machine learning model rearranges data in an autoencoder to maintain lossless spatial dimensions. For example, generating a light-enhanced face includes rearranging data in an autoencoder to maintain lossless spatial dimensions. In some implementations, the autoencoder utilizes spatial-to-depth and depth-to-space rearrangement neural network layers to maintain lossless spatial dimensionality changes.
[0116] In various implementations, act 720 includes generating, with an encoder, a feature vector based on a face within the digital image, a light-enhanced facial image; generating, with a distortion classifier, a distortion classification of the face within the image as part of generating the light-enhanced facial image; and / or generating, with a generator, the light-enhanced facial image from the feature vector and the distortion classification as part of generating the light-enhanced facial image. In some implementations, the distortion classifier generates the distortion classification to indicate an amount of noise distortion, blur distortion, exposure distortion, and / or light distortion in a face detected within the digital image.
[0117] In one or more implementations, act 720 includes generating a light-enhanced facial image having an image size that matches the image size of the cropped image and / or generating a facial image mask that separates non-facial pixels from facial pixels in the light-enhanced facial image. For example, the image inpainting machine learning model generates a light-enhanced facial image having image dimensions that match (e.g., are the same or substantially the same as) the original facial image and includes a facial image mask that identifies facial pixels in the light-enhanced facial image. In some instances, generating the enhanced digital image includes blending the light-enhanced facial image with an image background of the digital image using the facial image mask. In some cases, the digital image depicts colored light shining on the face, and the image inpainting machine learning model generates the light-enhanced facial image to remove the colored light depicted in the digital image. In some implementations, detecting a face within the digital image includes detecting colored light shining on the face and / or generating the light-enhanced facial image includes removing the effects of the colored light included in the digital image. Thus, if there is excessive light on the face, or if there is camera color imbalance, exposure, camera noise, etc., the inpainting machine learning model enhances the lighting of the face.
[0118] As further shown, the series of actions 700 includes generating an enhanced image with an image background 730. For example, in an example implementation, action 730 involves generating an enhanced digital image (with a re-lit face) by combining the light-enhanced facial image with the original image background.
[0119] As further shown, the series of acts 700 includes displaying the augmented image on the computing device at act 740. For example, in an example implementation, act 740 involves providing the augmented digital image for display on the computing device.
[0120] In some implementations, the series of actions 700 includes additional actions. For example, in some implementations, the additional actions of segmenting a person or body part from a digital image to generate an image of the person or body, the person or body part being connected to a face; generating a modified digital image with a re-lit person or body part by combining the image of the person or body enhancement with an image background; and generating an enhanced digital image with a re-lit face and a re-lit person or body part by combining the light-enhanced facial image with the modified digital image. In various implementations, the series of actions 700 includes segmenting a body part from a digital image to generate a body image, the body part being connected to a face and generating a body-enhanced image using an image inpainting machine learning model. In one or more implementations, generating the enhanced digital image also includes combining the body-enhanced image with the light-enhanced facial image and the image background.
[0121] In some examples, the series of actions 700 also includes actions of generating a modified digital image by blending the person or body-enhanced image with the image background using a first set of blending weights and / or generating an enhanced digital image by blending the light-enhanced facial image with the modified digital image using a second set of blending weights. In various implementations, the first set of blending weights is different from the second set of blending weights.
[0122] In one or more implementations, the series of actions 700 includes tracking a face on a digital video having a digital image set, the digital image set including the digital image; generating a face-enhanced digital image set based on the digital image set using an image inpainting machine learning model, wherein the face-enhanced digital image set includes the enhanced digital image; and providing the face-enhanced digital image set as a face-enhanced digital video for display on a computing device.
[0123] In various implementations, the series of actions 700 includes generating an image inpainting machine learning model by training a distortion classifier and an autoencoder in parallel to improve the accuracy of the generator, wherein the image inpainting machine learning model uses real digital images and synthetic digital images. In some instances, generating the synthetic digital image includes generating a synthetic face; capturing a plurality of digital images, each digital image shining light from a different light source on the synthetic face; and combining the plurality of digital images into a combined digital image to produce the synthetic digital image.
[0124] In some implementations, generating an image restoration machine learning model includes utilizing a loss model function that includes a pixel-wise loss, a feature loss, a texture information loss, an adversarial loss, a classification loss, and / or a segmentation loss.
[0125] Furthermore, in one or more implementations, the series of actions 700 includes the additional actions of detecting or segmenting an object from a digital image having an image background to generate an image of the object; generating a light-enhanced image of the object (from the object image) using an object relighting neural network including an autoencoder and a distortion classifier; generating an enhanced digital image with the relighted object by combining the light-enhanced object image with the image background, wherein the enhanced digital image has improved lighting on the object with respect to the digital image and maintains the same image background as the digital image; and providing the enhanced digital image for display on a computing device.
[0126] Furthermore, in some implementations, generating a light-enhanced object image includes generating a distortion classification (from the object image) using a distortion classifier of the object relighting neural network; generating a feature vector for the object image using an encoder network of the object relighting neural network; and generating a light-enhanced object image from a combination of the distortion classification and the feature vector using a generator of the object relighting neural network.
[0127] In various implementations, the series of actions 700 also includes tracking the object across a digital image set, the digital image set including the digital images; generating an object-enhanced digital image set based on the digital image set using an object relighting neural network, wherein the object-enhanced digital image set includes the enhanced digital images; and providing the object-enhanced digital image set for display on a computing device as an object-enhanced digital video.
[0128] In this disclosure, "network" is defined as one or more data links that can transmit electronic data between computer systems, modules, and other electronic devices. Networks can include public networks such as the Internet and private networks. When information is transmitted or provided to a computer via a network or another communication connection (hardwired, wireless, or a combination of hardwired and wireless), the computer views the connection as a transmission medium. Transmission media can include networks and / or data links that carry computer-executable instructions or data structures and can be accessed by general-purpose or special-purpose computers. The scope of computer-readable media includes the above combinations.
[0129] Additionally, the network described herein may represent a network or combination of networks (such as the Internet, an intranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more such networks) through which one or more computing devices can access the image restoration system. In practice, the network described herein may include one or more networks that use one or more communication platforms or technologies to transmit data. For example, the network may include the Internet or other data links that enable the transmission of electronic data between various client devices and components (e.g., server devices and / or virtual machines thereon) of the cloud computing system.
[0130] Furthermore, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures may be automatically transferred from a transmission medium to a non-transitory computer-readable storage medium (device), and vice versa. For example, computer-executable instructions or data structures received over a network or data link may be cached in random access memory (RAM) within a network interface module (NIC) and then ultimately transferred to the computer system RAM and / or a less volatile computer storage medium (device) at the computer system. Thus, it should be understood that a non-transitory computer-readable storage medium (device) may be included in a computer system component that also (or even primarily) utilizes a transmission medium.
[0131] Computer-executable instructions include, for example, instructions and data that, when executed by a processor, cause a general-purpose computer, a special-purpose computer, or a dedicated processing device to perform a specific function or group of functions. In some implementations, computer-executable instructions are executed by a general-purpose computer to convert the general-purpose computer into a special-purpose computer that implements the components of the present disclosure. Computer-executable instructions can be, for example, binary code, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described using language specific to structural features and / or method actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. On the contrary, the features and actions described are disclosed as example forms of implementing the claims.
[0132] Figure 8 800. The computer system 800 may be used to implement various computing devices, components, and systems described herein. As used herein, a "computing device" refers to an electronic component that performs operations based on a programmed instruction set. A computing device includes a group of electronic components, client devices, server devices, and the like.
[0133] In various implementations, computer system 800 represents one or more of the above-mentioned client devices, server devices, or other computing devices. For example, computer system 800 can refer to various types of network devices capable of accessing data on a network, a cloud computing system, or another system. For example, a client device can refer to a mobile device such as a mobile phone, a smart phone, a personal digital assistant (PDA), a tablet computer, a laptop computer, or a wearable computing device (e.g., a headset or a smart watch). A client device can also refer to a non-mobile device such as a desktop computer, a server node (e.g., from another cloud computing system), or another non-portable device.
[0134] Computer system 800 includes a processing system having a processor 801. Processor 801 may be a general-purpose single-chip or multi-chip microprocessor (e.g., an Advanced Reduced Instruction Set Computer (RISC) Machine (ARM)), a dedicated microprocessor (e.g., a Digital Signal Processor (DSP)), a microcontroller, a programmable gate array, etc. Processor 801 may be referred to as a central processing unit (CPU). Although processor 801 is shown as Figure 8 A single processor is depicted in the computer system 800, but in an alternative configuration, a combination of processors (e.g., an ARM and DSP) could be used.
[0135] The computer system 800 also includes a memory 803 in electronic communication with the processor 801. The memory 803 can be any electronic component capable of storing electronic information. For example, the memory 803 can be implemented as random access memory (RAM), read-only memory (ROM), magnetic disk storage media, optical storage media, flash memory devices in RAM, on-board memory included in the processor, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, etc., including combinations thereof.
[0136] Instructions 805 and data 807 may be stored in memory 803. Instructions 805 may be executed by processor 801 to implement some or all of the functionality disclosed herein. Executing instructions 805 may include using data 807 stored in memory 803. Any of the various examples of modules and components described herein may be implemented in part or in whole as instructions 805 stored in memory 803 and executed by processor 801. Any of the various examples of data described herein may be data 807 stored in memory 803 and used during execution of instructions 805 by processor 801.
[0137] The computer system 800 may also include one or more communication interfaces 809 for communicating with other electronic devices. The one or more communication interfaces 809 may be based on wired communication technology, wireless communication technology, or both. Some examples of the one or more communication interfaces 809 include a universal serial bus (USB), an Ethernet adapter, a wireless adapter operating according to the Institute of Electrical and Electronics Engineers (IEEE) 802.11 wireless communication protocol, A wireless communication adapter, and an infrared (IR) communication port.
[0138] The computer system 800 may also include one or more input devices 811 and one or more output devices 813. Some examples of the one or more input devices 811 include a keyboard, a mouse, a microphone, a remote control device, buttons, a joystick, a trackball, a touchpad, and a light pen. Some examples of the one or more output devices 813 include speakers and a printer. A specific type of output device that is typically included in the computer system 800 is a display device 815. The display device 815 used with the implementations disclosed herein can utilize any suitable image projection technology, such as a liquid crystal display (LCD), a light emitting diode (LED), gas plasma, electroluminescence, etc. A display controller 817 may also be provided for converting data 807 stored in the memory 803 into text, graphics, and / or moving images (as appropriate) that are displayed on the display device 815.
[0139] The various components of the computer system 800 may be coupled together via one or more buses, which may include a power bus, a control signal bus, a status signal bus, a data bus, etc. For clarity, the various buses are shown in FIG. Figure 8 It is illustrated as bus system 819.
[0140] Those skilled in the art will appreciate that the present invention can be practiced in a network computing environment having many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, and the like. The present disclosure can also be practiced in a distributed system environment in which local and remote computer systems linked (by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) over a network both perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0141] The techniques described herein may be implemented in hardware, software, firmware, or any combination thereof, unless expressly described as being implemented in a particular manner. Any features described as modules, components, etc. may also be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a non-transient processor-readable storage medium comprising instructions that, when executed by at least one processor, perform one or more of the methods described herein. Instructions may be organized into routines, programs, objects, components, data structures, etc. that may perform specific tasks and / or implement specific data types and may be combined or distributed as needed in various implementations.
[0142] Computer-readable media can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Thus, by way of example, implementations of the present disclosure may include at least two distinct types of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0143] As used herein, non-transitory computer-readable storage media (devices) may include RAM, ROM, EEPROM, CD-ROM, solid-state drives (SSD) (e.g., RAM-based), flash memory, phase-change memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer.
[0144] The steps and / or actions of the methods described herein may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is required for proper operation of the described method, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims.
[0145] The term "determining" encompasses a variety of actions, and thus, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a data repository, or another data structure), ascertaining, etc. Furthermore, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), etc. Furthermore, "determining" may include resolving, selecting, choosing, establishing, etc.
[0146] The terms "comprising," "including," and "having" are intended to be inclusive and mean that there may be additional elements in addition to the listed elements. Furthermore, it should be understood that reference to "one implementation" or "an implementation" of the present disclosure is not intended to be interpreted to exclude the existence of additional implementations that also incorporate features. For example, any element or feature described with respect to an implementation herein may be combined with any element or feature of any other implementation described herein, where compatible.
[0147] The present disclosure may be implemented in other specific forms without departing from its spirit or characteristics. The described implementation is to be considered illustrative rather than restrictive. The scope of the present disclosure is indicated by the appended claims rather than by the preceding description. Changes within the meaning and range of equivalents of the claims are to be included within their scope.
Claims
1. A computer-implemented method comprising: detecting a face within a digital image (112), the digital image (112) comprising the face and an image background; generating a light-enhanced facial image (126) of the face using an image inpainting machine learning model comprising an autoencoder (122, 236, 312) and a distortion classifier (124, 238, 322); generating an enhanced digital image (134) by combining the light-enhanced facial image with the image background; and The enhanced digital image (134) is provided for display on a computing device (201).
2. The computer-implemented method of claim 1 , wherein generating the enhanced digital image comprises: Lighting on the face is improved, low-quality facial features are repaired to higher-quality facial features, blur is reduced, and noise is reduced.
3. The computer-implemented method of claims 1-2, wherein generating the light-enhanced facial image further comprises: generating an encoded feature vector based on the face within the digital image; generating a distortion classification for the face within the digital image; as well as The light-enhanced facial image is generated based on the encoded feature vector and the distortion classification.
4. The computer-implemented method of claim 3 , wherein generating the distortion classification comprises: An amount of noise distortion, blur distortion, exposure distortion, and light distortion is indicated for the face detected in the digital image.
5. The computer-implemented method of claims 1 to 4, wherein detecting the face within the digital image comprises: identifying the face in the digital image; cropping the face in the digital image to generate a cropped image; as well as The face is tracked across a set of consecutive digital images.
6. The computer-implemented method of claim 5 , wherein generating the light-enhanced facial image comprises: generating the light-enhanced facial image to have an image size matching an image size of the cropped image; as well as A facial image mask is generated that separates non-facial pixels from facial pixels in the light-enhanced facial image.
7. The computer-implemented method of claim 6, wherein generating the enhanced digital image comprises: The light-enhanced facial image is blended with the image background of the digital image using the facial image mask.
8. The computer-implemented method of claims 1 to 7, wherein: Detecting the face within the digital image includes: detecting colored light impinging on the face; and Generating the light-enhanced facial image includes removing the effects of the colored light included in the digital image.
9. The computer-implemented method of claims 1 to 8, wherein generating the light-enhanced facial image comprises: The data is rearranged in the autoencoder to maintain lossless spatial dimensionality.
10. The computer-implemented method of claims 1 to 9, further comprising: segmenting a body part from the digital image to generate a body image, the body part being connected to the face; generating a body-enhanced image using the image restoration machine learning model; as well as Wherein generating the enhanced digital image further comprises: combining the body-enhanced image with the light-enhanced facial image and the image background.
11. The computer-implemented method of claim 10, wherein generating the body-augmented image further comprises: A scaling factor generated for the light-enhanced facial image is utilized.
12. The computer-implemented method of claims 1 to 11, further comprising: tracking the face across a digital video having a digital image set including the digital image; generating a facially enhanced digital image set from the digital image set using the image restoration machine learning model, wherein the facially enhanced digital image set includes the enhanced digital image; as well as The set of facially enhanced digital images is provided for display as a facially enhanced digital video on the computing device.
13. A system comprising: An image restoration machine learning model comprising a distortion classifier and an autoencoder (122, 236, 312) having an encoder (314) and a generator; a processor (801); and Computer memory (803), the computer memory (803) including instructions (805) that, when executed by the processor (801), cause the system to perform operations comprising: detecting a face within a digital image (112), the digital image (112) comprising the face and an image background; generating a light-enhanced facial image (126) of the face using the image inpainting machine learning model based on combining the outputs of the distortion classifier and the encoder (314) as inputs to the generator; and An enhanced digital image (134) is generated by combining the light-enhanced facial image (126) with the image background.
14. The system of claim 13, wherein the instructions further comprise: The image inpainting machine learning model is generated by training the distortion classifier and the autoencoder in parallel to improve the accuracy of the generator, wherein the image inpainting machine learning model uses both real digital images and synthetic digital images.
15. The system of claim 13, wherein the instructions further comprise generating a composite digital image by: Generate synthetic faces; capturing a plurality of digital images, each digital image shining light from a different light source onto the composite face; and The plurality of digital images are combined into a combined digital image to generate the composite digital image.
16. The system of claim 13, wherein the instructions further comprise: The image restoration machine learning model is generated by utilizing a loss model function, wherein the loss model function includes pixel-by-pixel loss, feature loss, texture information loss, adversarial loss, classification loss, or segmentation loss.
17. The system of claim 13, wherein the distortion classifier generates a distortion classification to indicate an amount of noise distortion, blur distortion, exposure distortion, or light distortion in the face detected in the digital image.
18. A computer-implemented method comprising: detecting an object within a digital image (112), the digital image (112) including an image background; generating a light-enhanced image of the object using an object relighting neural network (212) comprising an autoencoder (122, 236, 312) and a distortion classifier (124, 238, 322); generating an enhanced digital image (134) having a re-illuminated object by combining the light-enhanced object image with the image background, wherein the enhanced digital image (134) improves illumination on the object compared to the digital image (112) and maintains an image background that is the same or substantially similar to the digital image (112); as well as The enhanced digital image (134) is provided for display on a computing device (201).
19. The computer-implemented method of claim 18, wherein generating the light-enhanced object image comprises: generating a distortion classification using the distortion classifier of the object relighting neural network; generating a feature vector using an encoder network of the object relighting neural network; The light-enhanced object image is generated using a generator of the object relighting neural network based on a combination of the distortion classification and the feature vector.
20. The computer-implemented method of claim 18, further comprising: tracking the object across a set of digital images, the set of digital images including the digital image; generating a set of object-enhanced digital images from the set of digital images using a generator of the object relighting neural network, wherein the set of object-enhanced digital images includes the enhanced digital images; as well as The object-augmented digital image is provided for display as an object-augmented digital video on the computing device.