Image replacement repair
The repair technique trained by deep learning models can accurately replace closed captions in images or videos, solving the problems of visual incongruity and high resource consumption in existing technologies, and achieving efficient and accurate image replacement and content customization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-13
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies often result in visual inconsistencies or inaccuracies when replacing closed captions in images or videos, affecting user experience and consuming significant resources.
A deep learning model is used to train the inpainting model by masking regions of specific types of objects. The loss function is used to focus on the detailed reconstruction within the masked regions, and the masked and reverse-masked images are combined to generate accurate replacement images.
It enables accurate replacement of closed captions and other content after the final image rendering, reducing resource consumption, improving user experience, and supporting dynamic updates and customized content.
Smart Images

Figure CN114514560B_ABST
Abstract
Description
Technical Field
[0001] This document describes a method for replacing parts of an image using impainting. Background Technology Summary of the Invention
[0002] Generally, an innovative aspect of the subject matter described in this specification can be embodied in a computer-implemented method for replacing objects in an image. This method may include identifying a first object at a location within a first image; masking a target region based on the first image and the location of the first object to generate a masked image; generating a second image different from the first image based on the masked image and a repair machine learning model; training the repair machine learning model using the difference between a target region of a training image and the content of an image generated at a location corresponding to the target region of the training image; generating a third image based on the masked image and the second image; and adding a new object different from the first object to the third image.
[0003] These and other embodiments may each optionally include one or more of the following features.
[0004] In some implementations, a loss function is used to train the instigation machine learning model, which represents the difference between the target region of the training image and the content of the generated image at the location corresponding to the target region of the training image.
[0005] In some implementations, the first image is a video frame.
[0006] In some implementations, generating a third image based on a masking image and a second image includes masking an inverse target region based on the second image and a location corresponding to a target region in the first image to generate an inverse masking image, and generating a third image based on the masking image and the inverse masking image. The inverse target region may include a region of the second image outside the location corresponding to the target region in the first image. Masking the inverse target region based on the second image and the location corresponding to the target region in the first image to generate an inverse masking image may include at least some content of the second image within the target region and exclude at least some content of the second image outside the target region. Generating a third image may include combining the inverse masking image with the masking image.
[0007] In some implementations, the method includes extrapolating a fourth image based on a third image, wherein each of the first, second, third, and fourth images is a frame of a video.
[0008] Other embodiments of this aspect include corresponding systems, apparatuses, and computer programs configured to perform the actions of the method and encoded on computer storage devices.
[0009] Specific embodiments of the topics described in this document can be implemented to achieve one or more of the following advantages. In some environments, there has been no prior method to replace image portions of a file, such as the final rendering of a video or image, without creating a region around the portion to be replaced that is visually inconsistent with the original video or image. For example, closed captions in the final rendering of a video cannot be simply altered or replaced to provide translation, correction, or updating because previous methods typically replaced the region around the original closed caption with a solid shape of a single color, or attempted to remove the original closed caption using remediation techniques and render the translated caption over the frame or removed caption. These previous methods are discordant and / or inaccurate and negatively impact the user's experience when viewing the video or image, thereby affecting the user's reaction to the video or image content. Generally, existing solutions for remediation are trained using the entire image and may not be able to effectively reconstruct smaller areas of the image due to variations in structure and detail on the image; this drawback is addressed by the techniques, devices, and systems discussed herein. The technique described in this paper enables a system to provide an improved solution for image replacement by masking the region to be replaced, using a repair machine learning model (also known as a repair model) trained on regions within the image surrounding a specific type of object to repair (e.g., predictively create or reconstruct) a proper subset of an image smaller than the image (e.g., corresponding to the masked region), and rendering a replacement image on the repaired region.
[0010] This new system supports replacing portions of images in the final rendered video or images with natural and synthetic backgrounds. Because current inpainting model training techniques use images with highly variable visual characteristics, existing inpainting methods produce inaccurate results when applied to images with regions varying in structure, color, texture, and other visual characteristics. This new image replacement method uses a deep learning model trained on reconstructing a limited region around an image containing a specific type of object. For example, when the target content is a simple 3D shape, the new system can use an inpainting model trained to reconstruct the region around the simple 3D shape, and when the target content is text, the new system can use an inpainting model trained to reconstruct the region around the text. Because the inpainting model is trained to focus on a limited region around the target object (e.g., within a masked region), the model can more accurately reconstruct details and variations within that region. More specifically, the loss function of the inpainting model can focus on the portion of the newly generated image corresponding to the masked region, rather than allowing regions outside the masked region to dominate the loss function, since content outside the masked region should be easily reproducible, resulting in very small differences between those portions of the original and newly generated images. Furthermore, the restoration model can be trained for specific types of objects, allowing it to be trained for reconstruction around and within complex objects with multiple transparent areas; text with specific styles and spacing; specific images representing logos; and other specific applications.
[0011] This system provides an effective solution that allows for alteration and / or customization of videos and images based on the original images after final rendering. It improves the accuracy of masking and removing target regions from images using inpainting methods by implementing a novel solution that trains the inpainting model using a tailored portion of test data to enhance the accuracy of reconstructing regions with various visual characteristics. Furthermore, the system allows for dynamic content updates based on specific characteristics, including user preferences. For example, if a user is watching a video with English audio and Spanish closed captions, and the user's preferred language is Chinese, the system allows for alteration and re-rendering of the video to include Chinese closed captions. In some implementations, the system can automatically detect desired changes or customizations. For example, the system can determine to replace English closed captions with generated Chinese closed captions based on user information indicating that the user's preferred language is Chinese. In some implementations, the system is provided with the changes or customizations to be performed. For example, the system can receive an instruction that a set of Chinese closed captions is to be replaced and that the object to be replaced is a set of Spanish closed captions. Additionally, non-text content can be replaced as needed. For example, a video can be modified to replace items deemed unsuitable for the current audience or that violate the terms of service for distributing the video (e.g., excessive graphic content). In some implementations, the system allows content to move relative to its frame. For example, a marker in the top layer of content against a blurred background can be repositioned to make the background content visible.
[0012] This approach allows for the customization of images and videos based on factors such as specific applications, user preferences, and geographic location. Furthermore, the improved removal and re-rendering process reduces the overall resource consumption by allowing individual video items to be adapted and used for different applications and users, thus providing high efficiency in video and image content creation. This allows for reduced bandwidth for transmitting video across the network, as only one video item needs to be sent to the node, and reduces the processing resources required to render more than one video item, since the aspects of a video item modified by the system are significantly fewer than a whole new video item being rendered (and then sent to the node). The system provides video designers, developers, and content providers with a way to adjust image and video content after the design, editing, and post-processing stages, which previously required significant resource expenditure.
[0013] Furthermore, this approach offers the advantage of allowing video creators to modify only a portion of a video project. It allows video projects to be customized for specific applications, user preferences, and geographic locations. In addition to reducing the amount of resources required to customize and render videos for each request, the system also reduces the overall resource consumption by allowing individual video projects to be adapted and used for different applications and users.
[0014] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the specification, drawings, and claims. Attached Figure Description
[0015] Figure 1 This is a block diagram of an example environment for improved image replacement and restoration.
[0016] Figure 2 The data flow of the improved image replacement and restoration process is described.
[0017] Figure 3 This is a diagram of an example machine learning process used to train an image replacement and restoration model.
[0018] Figure 4 This is a flowchart of an example process for improving image replacement and repair.
[0019] Figure 5 This is a block diagram of an example computing system.
[0020] The same reference numerals and names in different figures denote the same elements. Detailed Implementation
[0021] This document describes methods, systems, and devices for improving content replacement in images and videos, enabling a single image or video to be used in multiple applications, modified to conform to local specifications and standards, comply with the terms of service for image / video distribution, and adapted to the preferences of different users.
[0022] In some implementations, the system allows the final rendered image or video to be changed upon request and customized based on the context to which the image or video will be presented (e.g., the geographic region to which the image will be distributed and / or audience information). In some cases, video creators who wish to adapt video projects for specific clients can use the techniques discussed here.
[0023] For example, the system can receive a final rendered video from a content provider. This video includes images that will be adapted based on the viewer's location. In this example, the video is a public service announcement to be used throughout the United States, and the images could be a sign displayed in the corner of each frame of the video. This sign could be adjusted to the viewer's state. The system can detect the sign as a target, its position within the frame, the target's solid shape, and determine the bounding box surrounding the target. The bounding box can be determined based on parameters including target accuracy granularity and target image quality, among others. For each frame, the system can then generate a mask to remove the sign from the portion within the frame's bounding box and a "reverse" mask for the portion outside the frame's bounding box. The mask for the portion within the frame's bounding box is applied to the frame to create a masked image with the content within the bounding box removed. The masked image is fed to a remediation model trained to focus on the region surrounding the bounding box. The remediation model is used to generate a second, different frame that includes both the content inside and outside the bounding box. Therefore, the remediation model used to generate the content is suitable for this specific application. Inverse masking of the frame portion outside the bounding box is applied to a second, different frame to create an inverse mask image, where the content outside the bounding box is removed. The system can then render a new frame by compositing the mask image and the inverse mask image. In other words, the portion of the original image outside the bounding box can be combined with the portion of the generated image inside the bounding box to create a new image that does not include the target. A replacement object is identified and can be placed in the new frame at the same location as the target in the original frame. In some implementations, the replacement object is placed in the new frame at a different location relative to the target's location in the original frame. The system can use the replacement object to render the new frame and use the new frame to render the entire video to produce the final video to be presented in response to the request.
[0024] Figure 1 This is a block diagram of an example environment 100 for improved image replacement and restoration. Example environment 100 includes a network 102, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. Network 102 connects an electronic document server 104 (“Electronic Document Server”), user equipment 106, and a digital component distribution system 110 (also referred to as DCDS 110). Example environment 100 may include many different electronic document servers 104 and user equipment 106.
[0025] User equipment 106 is an electronic device capable of requesting and receiving resources (e.g., electronic documents) via network 102. Example user equipment 106 includes personal computers, mobile communication devices, and other devices that can send and receive data via network 102. User equipment 106 typically includes user applications, such as web browsers, to facilitate sending and receiving data via network 102; however, local applications executed by user equipment 106 can also facilitate sending and receiving data via network 102.
[0026] One or more third parties 150 include content providers, product designers, product manufacturers, and other parties involved in the design, development, marketing, or distribution of videos, products, and / or services.
[0027] An electronic document is data that presents a set of content on user device 106. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, search results pages, and feed sources. Native applications (e.g., “apps”), such as those installed on mobile, tablet, or desktop computing devices, are also examples of electronic documents. Electronic document 105 (“electronic Doc”) may be provided to user device 106 by electronic document server 104. For example, electronic document server 104 may include a server hosting a publisher’s website. In this example, user device 106 may initiate a request for a given publisher’s web page, and electronic document server 104 hosting the given publisher’s web page may respond to the request by sending machine hypertext markup language (HTML) code that initiates the rendering of the given web page on user device 106.
[0028] Electronic documents can include a variety of content. For example, electronic document 105 can include static content (e.g., text or other specified content) that is inherent to the electronic document itself and / or does not change over time. Electronic documents can also include dynamic content that can change over time or based on each request. For example, the publisher of a given electronic document may maintain a data source for populating portions of the electronic document. In this example, a given electronic document may include tags or scripts that, when user device 106 processes (e.g., renders or executes) the given electronic document, cause user device 106 to request content from the data source. User device 106 integrates the content obtained from the data source into the presentation of the given electronic document to create a composite electronic document that includes content obtained from the data source.
[0029] In some cases, a given electronic document may include a digital content tag or digital content script referencing DCDS 110. In these cases, when user equipment 106 processes the given electronic document, user equipment 106 executes the digital content tag or digital content script. Execution of the digital content tag or digital content script configures user equipment 106 to generate a request 108 for digital content, which is sent to DCDS 110 via network 102. For example, the digital content tag or digital content script enables user equipment 106 to generate a packetized data request that includes header and payload data. Request 108 may include data such as the name (or network location) of the server requesting the digital content, the name (or network location) of the requesting device (e.g., user equipment 106), and / or information that DCDS 110 can use to select the digital content provided in response to the request. User equipment 106 sends request 108 to the server of DCDS 110 via network 102 (e.g., a telecommunications network).
[0030] Request 108 may include data specifying the characteristics of an electronic document and the locations where digital content can be presented. For example, data specifying references to an electronic document (e.g., a webpage) that will present digital content (e.g., a URL), available locations on the electronic document that can be used to present the digital content (e.g., digital content slots), the size of the available locations, the position of the available locations within the presentation of the electronic document, and / or the media types eligible for presentation in these locations may be provided to DCDS 110. Similarly, data specifying keywords (“document keywords”) for selecting an electronic document or entities (e.g., people, places, or things) referenced by the electronic document may also be included in Request 108 (e.g., as payload data) and provided to DCDS 110 to identify digital content items eligible for presentation with the electronic document.
[0031] Request 108 may also include data related to other information, such as information already provided by the user, geographic information indicating the state or region from which the request was submitted, or other information that provides context for the environment in which the digital content will be displayed (e.g., the type of device that will display the digital content, such as a mobile device or a tablet). The information provided by the user may include demographic data of the user of user device 106. For example, demographic information may include characteristics such as age, gender, geographic location, education level, marital status, household income, occupation, hobbies, social media data, and whether the user owns specific items.
[0032] For the situations discussed here where systems collect or may use personal information about users, users may be given the opportunity to control whether programs or features collect personal information (e.g., information about a user's social networks, social actions or activities, occupation, user preferences, or user's current location), or to control whether and / or how content that may be more relevant to the user is received from content servers. Furthermore, some data may be anonymized in one or more ways before being stored or used, thereby removing personally identifiable information. For example, a user's identity may be anonymized, making it impossible to determine the user's personally identifiable information, or the user's geographic location may be generalized to the location information obtained (e.g., city, zip code, or state), making it impossible to determine the user's specific location. Therefore, users can control how content servers collect and use information about them.
[0033] Data specifying characteristics of user equipment 106 may also be provided in request 108, such as information identifying the model of user equipment 106, the configuration of user equipment 106, or the size (e.g., physical size or resolution) of the electronic display (e.g., a touchscreen or desktop monitor) presenting the electronic document. Request 108 may be transmitted, for example, over a packet network, and request 108 itself may be formatted as packet data with a header and payload data. The header may specify the destination of the packet, and the payload data may include any information discussed above.
[0034] In response to receiving request 108 and / or using the information included in request 108, DCDS 110 selects digital content to be presented with a given electronic document. In some implementations, DCDS 110 is implemented in a distributed computing system (or environment) including, for example, a server and a group of interconnected computing devices that, in response to request 108, identify and distribute digital content. This group of computing devices operates together to identify a set of digital content eligible for presentation in an electronic document from a corpus of millions or more available digital content. For example, millions or more available digital content may be indexed in a digital component database 112. Each digital content index entry may reference the corresponding digital content and / or include distribution parameters (e.g., selection criteria) that regulate the distribution of the corresponding digital content.
[0035] In some implementations, the digital components from the digital component database 112 may include content provided by a third party 150. For example, the digital component database 112 may receive photographs of public intersections from a third party 150 that uses machine learning and / or artificial intelligence to navigate public streets. In another example, the digital component database 112 may receive specific questions from a third party 150 that provides services to cyclists, and the third party 150 may request responses from users. Furthermore, the DCDS 110 may render video content, including, for example, content from the video rendering processor 130 and video content stored in the digital component database 112 provided by the video rendering processor 130.
[0036] The identification of eligible digital content can be segmented into multiple tasks, and then these tasks can be distributed among the computing devices in the group of multiple computing devices. For example, different computing devices can each analyze different parts of the digital component database 112 to identify various digital content with distribution parameters that match the information included in request 108.
[0037] DCDS 110 aggregates the results received from the group of multiple computing devices and uses information associated with the aggregated results to select one or more instances of digital content to be provided in response to request 108. In turn, DCDS 110 can generate and send response data 114 (e.g., digital data representing a response) via network 102, which enables user equipment 106 to integrate the selected set of digital content into a given electronic document, such that the selected set of digital content and the content of the electronic document are presented together on the display of user equipment 106.
[0038] Image replacement processor 120 receives or acquires images and videos and replaces specific portions of the images or videos. Image replacement processor 120 includes object processor 122, restoration model 124, and rendering processor 126. Object processor 122 processes the image or video to detect objects and generates bounding boxes around the detected objects. Restoration model 124 uses machine learning techniques to reconstruct specific portions of the image or video within, for example, the bounding boxes generated by object processor 122. Rendering processor 126 renders the altered image and / or video.
[0039] For ease of explanation, object processor 122, repair model 124, and rendering processor 126 are... Figure 1The image replacement processor 120 is shown as a separate component of the image replacement processor 120. The image replacement processor 120 can be implemented as a single system on a non-transitory computer-readable medium. In some embodiments, one or more of the object processor 122, the repair model 124, and the rendering processor 126 can be implemented as an integrated component of a single system. The image replacement processor 120, its components object processor 122, repair model 124, and rendering processor, as well as their respective operations and outputs, will be described in further detail below.
[0040] The repair model 124 can be a statistical and / or machine learning model. For example, the repair model 124 can be a generative adversarial network (GAN) model trained using a loss function focused on a finite region of the image (e.g., within a masked region or at a specified distance from the masked region). In some implementations, in addition to a GAN model, the repair model 124 can include any of a variety of models, such as decision trees, linear regression models, logistic regression models, neural networks, classifiers, support vector machines, inductive logic programming, ensemble models (e.g., using techniques such as bagging, boosting, random forests, etc.), genetic algorithms, Bayesian networks, etc., and can be trained using various methods, such as deep learning, association rules, inductive logic, clustering, maximum entropy classification, learned classification, etc. In some examples, the machine learning model can use supervised learning. In some examples, the machine learning model can use unsupervised learning. The training of the repair model 124 is detailed below. Figure 3 Detailed description.
[0041] The techniques described below enable the system to efficiently edit and render videos, regardless of complexity.
[0042] Figure 2 It shows Figure 1 Example data stream 200 for an improved image replacement and restoration process in an example environment. The operation of data stream 200 is performed by various components of system 100. For example, the operation of data stream 200 may be performed by image replacement processor 120 communicating with DCDS 110 and / or user equipment 106.
[0043] The process begins in step A, where image replacement processor 120 receives final render 202. Final render 202 can be an image or video. For example, image replacement processor 120 may receive final image render 202 from, for example, a content provider of a third party 130 or from DCDS 110. In some embodiments, image replacement processor 120 receives final render 202 from, for example, a user of a user device 106. In some embodiments, image replacement processor 120 retrieves final render 202 from a storage medium such as a database. For example, image replacement processor 120 retrieves final render 202 from digital component database 112. In some embodiments, image replacement processor 120 may receive instructions to retrieve final render 202.
[0044] The process continues to step B, where the image replacement processor 120 detects the target object to be replaced and generates a bounding box 206 around the target object. For example, the object processor 122 may detect the target object 204 and its object type. In this particular example, the target object 204 is a text label. In other implementations, the target object may be other types of content, such as text, open shapes, closed shapes, complex shapes with transparency parameters, simple shapes, moving shapes and / or changing shapes, as well as other types of objects.
[0045] Object processor 122 can also generate a bounding box 206 surrounding target object 204. In some embodiments, the bounding box 206 is determined based on predetermined parameters. For example, the bounding box 206 may be determined as a rectangular shape surrounding the target object, with a margin of at least a predetermined size between the bounding box 206 and the target object 204. For example, object processor 122 can determine the bounding box 206 around the target object 204 by detecting the edges of the target object 204 and following the margin parameters. For example, the margin parameters may be provided by the user of user device 106, content provider, DCDS 110, and / or third party 130.
[0046] The object processor 122 also detects the position of the target object 204 within the final rendering 202. For example, the object processor 122 can determine the position of the target object 204 based on its x and y positions. In some implementations, the object processor 122 can determine the position of the target object 204 relative to a reference point (e.g., a reference pixel position and the area occupied by the target object). In some implementations, the object processor 122 can determine the absolute position of the target object 204.
[0047] The process continues to step C, where the image replacement processor 120 reconstructs the content data of the final rendered region 202 within the bounding box 206. First, in substep C-1, the image replacement processor 120 generates a mask for the region of the final rendered region 202 within the bounding box 206 to remove the content of the final rendered region 202 within the bounding box 206. The image replacement processor 120 then applies the mask to the final rendered region 202 to generate a masked image 208 that does not include the content of the final rendered region 202 within the bounding box 206. In some implementations, the object processor 122 generates the mask and applies it to the final rendered region 202 to generate the masked image 208.
[0048] Next, in sub-step C-2, the repair model 124 can be achieved by uniformly sampling r∈[r min ,r max To reconstruct the content within the region surrounding the target object (e.g., the region within bounding box 206), where r min and r max These are the minimum and maximum aspect ratios, respectively. After repairing model 124, uniform sampling can be implemented. Where a min and a max These are the minimum and maximum coverage areas, respectively; x max and y max These are the width and height of the image, respectively; m is the desired margin around the edge of the target object. The inpainting model 124 can then be set to w = rh to conform to the selected aspect ratio. In some implementations, the selected aspect ratio is the aspect ratio of the original image. In other implementations, the aspect ratio can be selected based on user input or other parameters. The inpainting model 124 can then uniformly sample x. corner ∈[m,x max -(m+w)] and y corner ∈[m,y max -(m+h)].
[0049] Using the sampled data, the inpainting model 124 can reconstruct content within the region of the final rendered 202 inside the bounding box 206 without the target object 204. The inpainting model 124 can perform, for example, structural or geometric inpainting, texture inpainting, and / or a combination of structural and texture inpainting. This is because the inpainting model 124 is capable of adapting images with structural properties ranging from complex, variable textures and boundaries to simple edges and structural attributes. For example, the inpainting model 124 can generate a second image 210 using a masked image that includes the content within the region of the final rendered 202 reconstructed by the inpainting model 124. In other words, the second image can be a reconstruction of the final rendered 202 generated by the inpainting model 124 using the masked image 208 as input.
[0050] In some embodiments, the second image 210 includes the content of the final render 202 outside the bounding box 206, as represented in the final render 202. For example, the content of the final render 202 outside the bounding box 206 can be used in the second image 210 exactly as it was input, or it can be reproduced by the repair model 124 exactly as it was input to the repair model 124. In some embodiments, the second image 210 includes content representing the content of the final render 202 outside the bounding box 206 generated by the repair model 124. For example, the repair model 124 can generate the content of the final render 202 outside the bounding box 206, and the generated content may not represent its content exactly as it was input into the final render 202. The second image 210 may have the same size as the final render 202. In some embodiments, the second image 210 may have a different size than the final render 202. For example, the second image 210 may be a scaled version of the final render 202.
[0051] The process continues to step D, where image replacement processor 120 generates a "reverse" mask image. First, in sub-step D-1, image replacement processor 120 generates a "reverse" mask for the region of the second image 210 outside the location of bounding box 212. Image replacement processor 210 may generate bounding box 212 at the same location and scale as bounding box 206 in the final render 202. Bounding box 212 may contain reconstructed content generated by instigator 124. Instigator 124 may, for example, generate a reverse mask for the region of the second image 210 outside bounding box 212 to remove the region of the second image 210 outside bounding box 212. Instigator 124 then applies the reverse mask to the second image 210 to create a reverse mask image 214. In some implementations, object processor 122 generates a mask and applies it to the second image 210 to generate the reverse mask image 214.
[0052] In sub-step D-2, the repair model 124 can then create a composite of the masking image 208 and the inverse masking image 210 to generate a third image representing a version of the final render 202 that does not include the target object 204. For example, the repair model 124 can use alpha blending to combine the masking image 208 and the inverse masking image 210 into a third image 216 that does not include the target object 204. Other techniques for combining the masking image 208 and the inverse masking image 210 can also be used to create the third image 216, which includes a portion of each of the masking image 208 and the inverse masking image 210.
[0053] The process continues to step E, where DCDS 110 receives a request 108 for content from user equipment 106. Request 108 is sent from user equipment 106 to DCDS 110 when the client device interacts with digital content. For example, if a user of user equipment 106 clicks a link to download a shopping app, the link may cause user equipment 106 to send request 108 to DCDS 110. Request 108 may include interaction tracking data from client equipment 106. For example, request 108 may include tracking data such as indications of the interaction, the digital content that user equipment 106 interacted with, and an identifier that uniquely identifies user equipment 106. In some implementations, request 108 includes indications of the digital content provider and the location of the destination server hosting the requested resource.
[0054] DCDS 110 processes the request and forwards the request for a specific image or video, including customized information, to image replacement processor 120. For example, DCDS 110 can use the above-mentioned... Figure 1 The described content selection process selects a specific video and determines customization information to provide to the image replacement processor 120. In some implementations, the DCDS 110 determines the customization information based on the context in which the image or video will be presented at the user equipment 106. In some implementations, the DCDS 110 determines the customization information based on information from, for example, a third party 130 that provides content to the DCDS 110.
[0055] The process continues to step F, where image replacement processor 120 adds a replacement object to third image 216 or a composite image. Replacement object 218 can be of the same type as or a different type from target object 204. For example, replacement object 218 can be an image, while target object 204 is text. In some implementations, replacement object 218 can be selected based on custom information. In some implementations, replacement object 218 is provided to image replacement processor 120. Object processor 122 can add replacement object 218 at the same position as target object 204 in the final render 202. In some implementations, object processor 122 can add replacement object 218 at a different position relative to the position of target object 204 in the final render 202. In a particular example, object processor 122 can move target object 204 to a different position in third image 216 relative to the position of target object 204 in the final render 202.
[0056] The process continues to step G, where the image replacement processor 120 re-renders the image with the added object to generate a new image. For example, the rendering processor 126 may create a new render 220 by compositing the third image 216 and the replacement object 218. In some implementations, the rendering processor 126 may use alpha blending technology to create the new render 220.
[0057] In some implementations, the image replacement processor 120 can create a new render 220 by creating a composite of the second image 210 and the replacement object 218. For example, the image replacement processor 120 may omit step D.
[0058] The process continues to step H, where image replacement processor 120 provides new image or video to DCDS 110 for presentation to user equipment 106 along with or as response 114. For example, image replacement processor 120 provides new rendering 220 to DCDS 110 for presentation to user equipment 106 along with response 114. As described above, in addition to the requested electronic document, response data 114 may also indicate image or video content re-rendered by image replacement processor 120. In response to DCDS 110 receiving request 108 and determining that the distribution parameters are satisfied based on the received distribution parameters and the user data indicated in request 108, DCDS 110 sends response data 114 to user equipment 106.
[0059] Figure 3 This is a diagram of an example block diagram of system 300 used for training an intelligent garment system. For example, system 300 can be used to train a repair model 124, as described above. Figure 1 As stated above.
[0060] In this specific example, sample 202 is provided to training module 210 as input to train the GAN-based inpainting model. Sample 202 can be a positive sample (i.e., a sample with an inpainting region mask having a threshold accuracy level) or a negative sample (i.e., a sample with an inpainting region having an accuracy level less than the threshold). Sample 202 can include ground truth regions, or regions determined to have a threshold accuracy level relative to the actual correct regions from the original image.
[0061] Ground truth indicates the actual, correct regions from the original image. For example, sample 202 may include a training set where actual regions are removed from the original image and provided to training module 210. In some implementations, ground truth regions can be generated and provided to training module 210 as sample 202 by generating a restored region mask using restoration model 124 and verifying that the restored region mask has a threshold level of accuracy relative to the actual correct regions from the original image. In some implementations, the accuracy of the restored region mask can be manually verified by humans. Restoration region masks can be automatically detected and labeled by extracting data from a data storage medium containing verified restored region masks.
[0062] The ground-real future motion can be associated with specific inputs of sample 202, making these inputs labeled with ground-realistic data. Using the ground-realistic labels, training module 210 can use sample 202 and the labels to validate the model output of insulation model 124 and continue training the model to improve the accuracy of the model in generating insulation region masks.
[0063] Training module 210 uses target region loss function 212 to train instigation model 124. Training module 110 uses target region loss function 212 to train instigation model 124 to generate instigation region masks. Target region loss function 212 may consider variables such as predicted texture and / or predicted structure, as well as other variables. For example, target region loss function 212 may focus on regions surrounding a target object, such as the region within the bounding box 206 surrounding target object 204, as per [the context of the missing text]. Figure 2 What is depicted.
[0064] For example, the repair model 124 can use a loss function, such as specifying a loss function as follows:
[0065]
[0066] L(G,D)=E x,y [logD(x,y)]+E x [log(1-D(x,G(x)))]+λE x,y [‖yG(x)‖1]
[0067] The loss function represents the difference between the target region of the first image and the content of the second image at the location corresponding to the target region of the first image. Because the content outside the masked region is highly reproducible and, in some implementations, can be precisely copied, resulting in very small differences between those parts of the original image and the newly generated image, the system's loss function focuses on the portion of the newly generated image corresponding to the masked region, rather than allowing the region outside the masked region to dominate the loss function.
[0068] Training module 210 can train inpainting model 124 manually, or the process can be automated. Training module 210 uses loss function 212 and samples 202 labeled with ground truth regions to train inpainting model 124 to understand what is important to the model. For example, training module 210 can train inpainting model 124 by optimizing the model based on the loss function. In some embodiments, optimizing the model based on the loss function involves minimizing the result of the loss function, which minimizes the difference between the target region of the training image and the content of the generated image at the location corresponding to the target region of the training image. In some embodiments, training module 210 can use a threshold for the difference, such as a maximum percentage difference. For example, training module 210 can train inpainting model 124 to reconstruct content at locations corresponding to target regions in the training image with a maximum difference of 5%, 10%, or 20%, and other percentages. In some embodiments, training module 210 can train inpainting model 124 to reconstruct content at locations corresponding to target regions in the training image with a threshold number of pixels, such as a maximum number of different pixels or a minimum number of similar or identical pixels. Training module 210 allows the repair model 124 to learn by changing the weights applied to different variables, thereby emphasizing or de-emphasizing the importance of variables within the model. By changing the weights applied to variables within the model, training module 210 allows the model to learn which types of information (e.g., which textures, structural components, boundaries, etc.) should be weighted more heavily to produce a more accurate repair model.
[0069] Samples and variables can be weighted based on, for example, feedback from the image replacement processor 120. For instance, if the image replacement processor 120 collects feedback data indicating that the loss function 212 is producing results that do not have threshold accuracy for a particular type of target object, the inpainting model 124 can weight the particular configuration more heavily than other configurations. The inpainted regions that meet the threshold accuracy level generated by the inpainting model 124 can be stored as positive samples 202.
[0070] The repair model 124 can use various types of models, including general models that can be used for all types of target objects and attributes of the original image or the final render 202, as well as custom models that can be used for a specific subset of target objects or original images that share a set of features, and can dynamically adjust the model based on the type of the detected target object or original image. For example, the classifier can use a base GAN for all target objects and then adjust the model for each object.
[0071] The repair model 124 can adapt a large number of target objects and generate masks for them. In some implementations, the training process of the repair model 124 is a closed-loop system that receives feedback information and dynamically updates the repair model network based on that feedback information.
[0072] Figure 4 This is a flowchart of an example process 400 for improving image replacement and restoration. In some implementations, process 400 may be performed by one or more systems. For example, process 400 may be performed by... Figure 1-3 The process 400 is implemented using an image replacement processor 120, a DCDS 110, a user equipment 106, and a third party 140. In some embodiments, the process 400 may be implemented as instructions stored on a non-transitory computer-readable medium, and when these instructions are executed by one or more servers, they may cause one or more servers to perform the operations of the process 400.
[0073] The process 400 for replacing an object in an image begins with identifying a first object at a location within the first image (402). For example, the image replacement processor 120 receives a final render 202 and identifies a target object 204 at a specific location within the first image. The object processor 122 of the image replacement processor 120 can perform the identification. The object processor 122 can also identify the object type of the target object 204 and determine the location of the target object 204 in the final render 202. The object processor 122 can generate a bounding box 206 around the target object 204.
[0074] Process 400 continues by masking the target region based on the positions of the first image and the first object to generate a masked image (404). For example, image replacement processor 120 can generate a mask based on the final render 202 to remove the region within the bounding box 206. Object processor 122 or repair model 124 can perform the above-mentioned... Figure 1-3 The mask generation and application process described above.
[0075] Process 400 continues to generate a second image different from the first image based on the masked image and a repair machine learning model, which is trained using a loss function representing the difference between the target region of the training image and the content of the generated image at the location corresponding to the target region of the training image (406). For example, repair model 124 can generate a second image 210 based on the final rendering 202, as described above regarding... Figure 1-3 As described above, the repair model 124 is trained to focus on regions that are subsets of the entire image and surrounding target objects of a specific type.
[0076] Process 400 continues to generate a third image (408) based on the masking image and the second image. For example, image processor 120 can generate an inverse mask image based on the second image and the inverse mask, as described above. Figure 1-3 The image processor 120 can then generate a third image by synthesizing masked regions and inverse masking. For example, the inpainting model 124 or the object processor 122 can generate a third image 216.
[0077] Process 400 continues by adding new objects to the third image (410). For example, image processor 120 can generate a new render 220 by adding a replacement object 218, different from the target object 206, to the third image 216. Render processor 126 can use, for example, alpha blending techniques to create a new render 222. In some embodiments, image processor 120 adds the replacement object 218 to the third image 216 at a relative position to the target object 206 within the final render 202. In some embodiments, image processor 120 adds the replacement object 218 to different positions within the third image 216. In some embodiments, the replacement object 218 is the same object as the target object 206, and process 400 repositions the target object 206 relative to its position within the final render 202.
[0078] Figure 5 This is a block diagram of an example computer system 500 that can be used to perform the operations described above. System 500 includes a processor 510, memory 520, storage device 530, and input / output device 540. Each of components 510, 520, 530, and 540 can be interconnected, for example, using a system bus 550. Processor 510 is capable of processing instructions that execute within system 500. In one embodiment, processor 510 is a single-threaded processor. In another embodiment, processor 510 is a multi-threaded processor. Processor 510 is capable of processing instructions stored in memory 520 or on storage device 530.
[0079] The memory 520 stores information within the system 500. In one embodiment, the memory 520 is a computer-readable medium. In one embodiment, the memory 520 is a volatile memory cell. In another embodiment, the memory 520 is a non-volatile memory cell.
[0080] Storage device 530 provides high-capacity storage for system 500. In one embodiment, storage device 530 is a computer-readable medium. In various other embodiments, storage device 530 may include, for example, a hard disk drive, an optical disk drive, a storage device shared by multiple computing devices over a network (e.g., a cloud storage device), or some other high-capacity storage device.
[0081] Input / output device 540 provides input / output operations for system 500. In one embodiment, input / output device 540 may include one or more network interface devices, such as Ethernet cards, serial communication devices (e.g., RS-232 ports), and / or wireless interface devices (e.g., 802.11 cards). In another embodiment, input / output device may include a driver device configured to receive input data and send output data to other input / output devices (e.g., keyboards, printers, and display device 560). However, other embodiments may also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[0082] Although already Figure 5 An example processing system is described herein, but implementations of the subjects and functional operations described herein may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations of one or more of them.
[0083] An electronic document (for simplicity, it will be referred to as a document) does not necessarily correspond to a file. A document can be stored as a part of a file that contains other documents, as a single file dedicated to the document in question, or as multiple collaborative files.
[0084] Embodiments of the subject matter and operations described in this specification can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more computer program instruction modules encoded on a computer storage medium (or medium) for execution by or control of the operation of a data processing apparatus. Alternatively or additionally, program instructions can be encoded on artificially generated propagated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these, or included therein. Furthermore, while the computer storage medium is not a propagated signal, it can be a source or destination of computer program instructions encoded in artificially generated propagated signals. The computer storage medium can also be one or more separate physical components or media (e.g., multiple CDs, discs, or other storage devices) or included therein.
[0085] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0086] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, systems-on-a-chip, or a combination thereof. The apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement various computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.
[0087] A computer program (also known as a program, software, software application, script, or code) can be written in any programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment. A computer program may, but does not need to, correspond to a file in a file system. A program may be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located in one place or distributed across multiple locations and interconnected through a communication network.
[0088] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuits, and the devices can be implemented as special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0089] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for performing actions according to instructions and one or more storage devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), and so on. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory can be supplemented or incorporated by dedicated logic circuitry.
[0090] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, that the user can use to provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's client device in response to a request received from a web browser.
[0091] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes backend components, such as a data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface or a web browser through which a user can interact with embodiments of the subject matter described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).
[0092] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact via a communication network. The client-server relationship arises from computer programs running on their respective computers and involves a client-server relationship. In some embodiments, the server transmits data (e.g., HTML pages) to a client device (e.g., to display data to a user interacting with the client device and to receive user input from that user). Data generated at the client device (e.g., the result of user interaction) can be received from the client device at the server.
[0093] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features characteristic of particular embodiments of a particular invention. Some features described in the context of independent embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in some combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0094] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0095] Therefore, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and the desired result can still be obtained. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A computer-implemented method for replacing an object in an image, comprising: identifying a first object at a location within a first image and generating a target region surrounding the first object; masking the target region based on the first image and the location of the first object to produce a masked image; generating a second image different from the first image based on the masked image and a inpainting machine learning model, the inpainting machine learning model trained using a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image; generating a third image based on the masked image and the second image; and adding a new object different from the first object to the third image in response to a request from a user device, wherein the new object is selected based on customization information associated with the request and placed at the same location as the first object. The first image is a video frame.
2. The computer-implemented method of claim 1, wherein, The inpainting machine learning model is trained using a loss function of a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image.
3. The computer-implemented method of claim 1, wherein, Generating the third image based on the masked image and the second image comprises:
4. The computer-implemented method of claim 1, wherein, masking an inverse target region based on the second image and the location corresponding to the target region of the first image to produce an inverse masked image; and generating the third image based on the masked image and the inverse masked image. The inverse target region includes a region of the second image outside of the location corresponding to the target region of the first image.
6. The computer-implemented method of claim 4, wherein masking the inverse target region based on the second image and the location corresponding to the target region of the first image to produce the inverse masked image comprises generating the inverse masked image from the second image, the inverse masked image including at least some content of the second image within the target region and not including at least some content of the second image outside of the target region.
5. The computer-implemented method of claim 4, wherein, Generating the third image comprises compositing the inverse masked image with the masked image.
8. The computer-implemented method of claim 1, further comprising extrapolating a fourth image based on the third image, wherein each of the first image, the second image, the third image, and the fourth image is a video frame.
7. The computer-implemented method of claim 4, wherein, 9. A system for replacing an object in an image, comprising: and one or more memory elements including instructions that, when executed, cause the one or more processors to perform operations comprising: one or more processors; identifying a first object at a location within a first image and generating a target region surrounding the first object; masking the target region based on the first image and the location of the first object to produce a masked image; generating a second image different from the first image based on the masked image and a inpainting machine learning model, the inpainting machine learning model trained using a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image; generating a third image based on the masked image and the second image; and adding a new object different from the first object to the third image in response to a request from a user device, wherein the new object is selected based on customization information associated with the request and placed at the same location as the first object. adding a new object to the third image different from the first object in response to a request from a user device, wherein the new object is selected based on customization information associated with the request and is placed in the same location as where the first object is located.
10. The system of claim 9, wherein, The inpainting machine learning model is trained using a loss function representing a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image.
11. The system of claim 9, wherein, The first image is a video frame.
12. The system of claim 9, wherein, Generating the third image based on the masked image and the second image includes: masking an inverse target region based on the second image and a location corresponding to the target region of the first image to produce an inverse masked image; and generating the third image based on the masked image and the inverse masked image.
13. The system of claim 12, wherein, The inverse target region includes a region of the second image that is outside of the location corresponding to the target region of the first image.
14. The system of claim 12, wherein, Masking the inverse target region based on the second image and the location corresponding to the target region of the first image to produce the inverse masked image includes generating the inverse masked image from the second image that includes at least some content of the second image within the target region and does not include at least some content of the second image outside of the target region.
15. The system of claim 12, wherein, Generating the third image includes compositing the inverse masked image with the masked image.
16. The system of claim 9, the operations further comprising extrapolating a fourth image based on the third image, wherein each of the first, second, third, and fourth images are video frames.
17. A non-transitory computer storage medium encoded with instructions that, when executed by a distributed computing system, cause the distributed computing system to perform operations comprising: identifying a first object at a location within a first image and generating a target region around the first object; masking the target region based on the first image and the location of the first object to produce a masked image; generating a second image different from the first image based on the masked image and an inpainting machine learning model, the inpainting machine learning model trained using a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image; generating a third image based on the masked image and the second image; and adding a new object to the third image different from the first object in response to a request from a user device, wherein the new object is selected based on customization information associated with the request and is placed in the same location as where the first object is located.
18. The non-transitory computer storage medium of claim 17, wherein the inpainting machine learning model is trained using a loss function representing a difference between a target region of a training image and content of an image generated at a location corresponding to the target region of the training image.
19. The non-transitory computer storage medium of claim 17, wherein, The first image is a video frame.
20. The non-transitory computer storage medium of claim 17, wherein generating the third image based on the masked image and the second image includes: masking an inverse target region based on the second image and a location corresponding to the target region of the first image to produce an inverse masked image; and Based on the masking image and the inverse masking image, a third image is generated.
Citation Information
Patent Citations
A method for removing station captions and subtitles in an image based on a deep neural network
CN109472260A
Automatic personalized story generation for visual media
US20190200050A1