Semantic base image copy
The system uses generative AI to create images by combining descriptive captions and visual embeddings, addressing the challenge of maintaining source image integrity and introducing creative variations, ensuring effective brand representation and product depiction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-05-20
- Publication Date
- 2026-07-23
Smart Images

Figure 2026524572000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for copying an image using generative artificial intelligence in a manner that uniquely changes the image while maintaining the semantic context and visual quality of the source image.
Background Art
[0002] For the purpose of generally presenting the context of the present disclosure, the background art provided herein is described. The achievements of the inventor(s) recited herein are not expressly or implicitly recognized as prior art to the present disclosure, either within the scope described in this background art section or in aspects of this specification that may not be eligible as prior art at the time of filing and in other cases.
[0003] In recent years, remarkable progress has been made in the fields of image generation and image modification. In particular, generative artificial intelligence (AI) models are beginning to be widely used in both personal and commercial fields. In some cases, it is desirable to generate an inaccurate "copy" of a source image—that is, an image that is similar to or influenced by the source image in some respects, but is original and different from the source image. For example, in digital advertising, advertisers may want to expand their catalog of available ads and / or improve the performance of existing ads by developing a single creative / asset / image into a number of new creatives / asset / images. However, generative AI "copies" often fail to strike a good balance between (1) maintaining the visual properties of the source image (e.g., maintaining the look and feel of the source image), (2) maintaining the semantic context of the source image (e.g., maintaining the content depicted in the source image at an explanatory / conceptual level), and (3) being meaningfully different from the source image (e.g., the copy is a creative variation of the source image and not merely a trivial variation of the source image). In the context of digital advertising, for example, original creative copy generated by generative AI may fail to capture the look and feel associated with a particular advertiser (e.g., what is expressed in the source image through color and style), fail to accurately represent the advertised product, or fail to differ from the original creative to a degree sufficient to create a significant difference in performance compared to the original (e.g., a different level of customer or user engagement). [Overview of the Initiative]
[0004] In the disclosed technology, the system generates a new image by semantically copying a source image. As used herein, “copying” a source image generally means a new image generated in some way derived from the source image, but not an exact copy of the source image, nor necessarily, but (possibly) any part or pixel of the source image is reused or duplicated. The system of the disclosure generates a new image from a source image (i.e., generates a semantic copy) by (1) applying the source image to a first generative artificial intelligence (AI) model that generates a descriptive caption for the source image to generate a text prompt; (2) generating a visual embedding based on the source image (e.g., a visual embedding of either the source image or a processed version of the source image); and (3) using a second generative AI model to generate a new image based on both the text prompt and the visual embedding.
[0005] By generating a new image based on a text prompt that is a descriptive caption generated directly from the source image (or at least a text prompt derived from such a caption), the system provides a "semantic" copy of the source image. That is, the new image adheres tightly to the source image at a conceptual / descriptive level. Furthermore, since the system also generates the new image using the visual embedding of the source image, the new image adheres tightly to the visual characteristics of the source image. Moreover, text prompts derived from the source image itself are less likely to conflict with the visual embedding of the source image (compared to, for example, manually generated text prompts), and therefore less likely to create a strange, confusing, or unappealing new image.
[0006] Furthermore, using visual embeddings in conjunction with text prompts increases the likelihood that new images will depict something truly relevant to the source image. In other words, visual embeddings help the second (image-generating) generative AI model avoid hallucination regarding what should be represented at a conceptual level. For example, if the text prompt "mobile phone in a car" is derived from a source image showing a mobile phone mounted on a car's dashboard, using a visual embedding of the source image simultaneously prevents the generative AI model from generating a new image, for example, one where a huge mobile phone is mounted on the roof of the car and outside the car. Thus, the visual embedding derived from the source image serves as a foundation for the generative AI model. Using text prompts and visual embeddings together helps ensure that the generated image is based on the source image and, at the same time, introduces diversity that can be introduced through variations of the text prompt prior to the visual embedding and / or image processing steps (such as cropping), thereby preserving important concepts.
[0007] The technology of this disclosure also offers other technical advantages. One advantage arises from the fact that text prompts contain information that is not strictly present in the source image itself. As a result, text prompts provide a higher level of entropy than when using visual embeddings without text prompts, even though they are derived from the source image. In other words, text prompts provide a second generative AI model with a wider creative / imaginative space to create new images.
[0008] As another example, converting a source image into a text prompt using a first (captioning) generative AI model facilitates modifications to the source image compared to making such modifications based solely on visual embeddings or via other image processing means. This is because generative AI models (e.g., large-scale language models) are relatively adept at making modifications in semantic space (e.g., adding, removing, or changing the state or position of objects), and because the disclosed technique converts the source image into a descriptive / captioned text, such modifications can be made efficiently and relatively easily. In some embodiments, the disclosed system makes such modifications using a third generative AI model by modifying or "mutating" the text prompt (along with the visual embeddings) before the system applies the text prompt to a second generative AI model.
[0009] Those skilled in the art will also find other advantages apparent by reading this disclosure and examining the corresponding drawings.
[0010] In one embodiment, a semantic-based image copying method includes generating a text prompt using one or more processors. Generating a text prompt includes applying a source image to a first generative AI model to generate a descriptive caption for the source image. The method also includes generating a visual embedding based on the source image using one or more processors, and generating a new image based on the text prompt and the visual embedding using a second generative AI model with one or more processors.
[0011] In other embodiments, the system includes one or more processors and one or more memories that store instructions, which, when executed by one or more processors, cause one or more processors to (1) generate a text prompt, the generation of the text prompt includes applying a source image to a first generative AI model to generate a descriptive caption for the source image, (2) generate a visual embedding based on the source image, and (3) use a second generative AI model to generate a new image based on the text prompt and the visual embedding.
[0012] In another embodiment, one or more non-temporary computer-readable media, when executed by one or more processors, stores instructions causing one or more processors to (1) generate a text prompt, the generation of which includes applying a source image to a first generative AI model to generate a descriptive caption for the source image; (2) generate a visual embedding based on the source image; and (3) use a second generative AI model to generate a new image based on the text prompt and the visual embedding. [Brief explanation of the drawing]
[0013] [Figure 1] A block diagram of an exemplary system that can implement semantic-based image copying technology is shown. [Figure 2] Figure 1 illustrates an exemplary process for semantic-based copying of a source image, which can be implemented by the computing system shown. [Figure 3A] Figure 1 illustrates an exemplary process for semantic-based copying of a source image with feedback, which can be implemented by the computing system shown. [Figure 3B] Figure 1 illustrates an exemplary process for semantic-based copying of a source image with feedback, which can be implemented by the computing system shown. [Figure 4]An example of a semantic copy that can be created by the computing system in Figure 1 is shown using the processes in Figure 2, Figure 3A, or Figure 3B. [Figure 5] An example of a copy that can be created by an alternative embodiment that uses text prompts to generate a new image but does not use visual embedding is shown. [Figure 6] A flowchart illustrating an exemplary method for semantic-based copying of source images is shown. [Modes for carrying out the invention]
[0014] Figure 1 is a block diagram of an exemplary system 100 that can implement semantic-based image copying / generation technology. The exemplary system 100 includes a computing system 102, a client device 104, a content provider 106 (e.g., a content provider server), and a network 110. The computing system 102 is remote from the client device 104 and the content provider 106 and is communicatively coupled to the client device 104 and the content provider 106 via the network 110. In some embodiments, the system 100 does not include the client device 104 and / or the content provider 106.
[0015] Network 110 may be a single communication network (e.g., the Internet), and in some embodiments, it may also include one or more additional networks. As just one example, network 110 may include a cellular network, the Internet, and a server-side local area network (LAN). On the other hand, while Figure 1 shows only a single client device 104 and a single content provider 106, it will be understood that computing system 102 may also communicate with roughly the same number of other client devices (e.g., millions) as client device 104, and / or with roughly the same number of other content providers (e.g., thousands) as content provider 106.
[0016] In general, the computing system 102 can perform image copying / generation services (for example, to a provider such as content provider 106). As stated above, the term “copy” in this specification generally refers to a new image generated in some way derived from the source image, but not an exact copy of the source image, nor does it necessarily reuse or duplicate any part / pixel of the source image.
[0017] In the context of digital advertising or marketing, for example, a computing system 102 may use existing images from a content provider, such as content provider 106, to generate new images that the content provider can use in additional digital advertisements. In such an example, the new / additional images can be used to offer a wider variety of images / advertisements, and their performance can then be measured (e.g., based on click-through rates, conversion rates, etc.) to determine which images / advertisements are most effective. In another example of digital advertising, the new / additional images may have a different aspect ratio than the original images, making the new images suitable for ad slots with different aspect ratio constraints (e.g., web pages or mobile applications). In particular, the techniques described herein (e.g., in relation to Figures 2, 3A, 3B, and 6) can change the aspect ratio of the source image in a more seamless manner than conventional techniques (e.g., adding cropping to protruding area detection).
[0018] As another example, computing system 102 may generate new images / copy (e.g., images of explanatory materials) intended to facilitate viewer understanding, the performance of which is measured by a mechanism that determines what percentage of viewers who see the images take the correct action. Other contexts are also possible; however, for the sake of clarity and consistency, this disclosure primarily uses examples related to embodiments / contexts of digital advertising.
[0019] The client device 104 is generally configured to access information resources (e.g., web pages and / or user interfaces of mobile applications or other applications) that can present images generated by the computing system 102. For example, the computing system 102 may generate digital advertisements that include (or consist entirely of) semantic copy as discussed herein. The computing system 102 or other computing systems may provide digital advertisements to users of the client device 104 and / or other similar client devices using appropriate techniques such as conducting auctions (e.g., keyword bidding by advertisers, auctions based on relevance metrics, etc.). Digital advertisements may be provided in slots on web pages visited by the user, and / or in slots on application user interfaces displayed to the user.
[0020] Content provider 106 can generally delegate or request computing system 102 to generate one or more images and / or provide source images on which the image generation is based. For example, content provider 106 may be a digital advertiser providing digital advertising images for each of several offered products or services as part of one or more advertising campaigns owned or managed by content provider 106. Other examples include source images being screenshots of web pages hosted by content provider 106, screenshots of mobile applications provided by content provider 106, etc.
[0021] The computing system 102 includes a network interface 120, a processor 122, and memory 124. The network interface 120 includes hardware, firmware, and / or software configured to enable the computing system 102 to exchange electronic data with client devices 104 and other similar client devices (and possibly content providers such as 106) via the network 110. For example, the network interface 120 may include a wired or wireless router and a modem. The processor 122 may be a single processor (e.g., a central processing unit (CPU)) or may include multiple processors (e.g., multiple CPUs, or one or more CPUs and one or more graphics processing units (GPUs)). The computing system 102 may be a single computing device (such as a server) located in a single location, or may include multiple coordinated computing devices that are either located in the same location or distributed remotely.
[0022] Memory 124 is a computer-readable, non-transitory memory unit or device, or a collection of such units / devices, which may include persistent and / or non-persistent memory components. Memory 124 stores instructions executable by processor 122 to perform various operations, including instructions of various software applications and data generated and / or used by such applications. In the exemplary system 100 of FIG. 1, memory 124 stores instructions of semantic copy generator 130, which includes captioner 140, prompt mutator 142, image processor 144, visual embedder 146, and image generator 148.
[0023] Memory 124 can also store generative artificial intelligence (AI) models. Specifically, in the exemplary system 100 of FIG. 1, memory 124 stores a first generative AI model 150 used by captioner 140, a second generative AI model 152 used by image generator 148, and a third generative AI model 154 used by prompt mutator 142. In other embodiments, the third generative AI model 154 is not included in system 100. More generally, in some embodiments, it is understood that memory 124 may omit one or more modules / elements shown in FIG. 1, such as prompt mutator 142 and / or image processor 144. In some embodiments, memory 124 may include one or more additional modules / elements not shown in FIG. 1, such as a module that facilitates delivering images (e.g., digital advertisements) to a user of a device such as client device 104. In some embodiments, the first generative AI model 150, the second generative AI model
[0024] The client device 104 can be or include any fixed, mobile, or portable computing device with wired and / or wireless communication capabilities (e.g., smartphones, tablet computers, laptop computers, desktop computers, smart wearable devices such as smart glasses or smartwatches, computers in vehicle head units, etc.). In the exemplary embodiment of FIG. 1, the client device 104 includes a network interface 160, a processor 162, a memory 164, and a display 166. The processor 162 may be a single processor or include multiple processors.
[0025] The memory 164 includes one or more computer-readable non-transitory storage units or devices that can include persistent and / or non-persistent memory components. The memory 164 stores instructions executable by the processor 162 to perform various operations, including instructions for various software applications and data generated and / or used by such applications.
[0026] In the exemplary system 100 of Figure 1, memory 164 stores at least application 170. Generally, application 170 is executed by processor 162 and provides one or more user interfaces via display 166, through which the user interface(s) enable the user to access information resources that may include images (semantic copies) generated by computing system 102. For example, application 170 may be a web browser application, and images generated by computing system 102 may be included in content slots of web pages that the user visits and is presented on display 166. In a more specific example, the images may be digital advertisements generated by computing system 102, then selected by computing system 102 (or another computing system) for insertion into content slots and provided to client device 104. In other embodiments, application 170 is a dedicated application (e.g., a "mobile app"), and images generated by computing system 102 are included in content slots of user interfaces presented by application 170 on display 166.
[0027] The display 166 includes hardware, firmware, and / or software configured to allow the user to view the visual output of the client device 104, and may use any preferred display technology (e.g., LED, OLED, LCD, etc.). In some embodiments, the display 166 is incorporated into a touchscreen having both display and manual input functions. Furthermore, in some embodiments where the client device 104 is a wearable device, the display 166 is a transparent visibility component (e.g., the lens of smart glasses) that includes integrated electronic components. For example, the display 166 may include micro-LED or OLED electronic equipment embedded in the lens of smart glasses.
[0028] The network interface 160 includes hardware, firmware, and / or software configured to enable a client device 104 to exchange electronic data with a computing system 102 via the network 110. For example, the network interface 160 may include a cellular communication transceiver, a WiFi transceiver, and / or transceivers for one or more other wired and / or wireless communication technologies.
[0029] Figure 1 shows the client device 104 as a single component that communicates directly with the computing system 102 (i.e., via the network 110), but in some embodiments, the subcomponents of the client device 104 shown in Figure 1 are instead divided into two or more user-side devices. For example, smart glasses may include a processor 162, memory 164, and a display 166, while a smartphone may include other processing units, other memory, other displays, and a network interface 160. The smart glasses may then communicate with a smartphone (e.g., via Bluetooth) as needed to enable the operations described herein.
[0030] Returning to the computing system 102, the semantic copy generator 130 generally operates by acquiring a source image (for example, by accessing the database 180 or receiving it directly from the content provider 106) and generating both semantic (text) information and visual information based on the same source image. The semantic information is generated by the captioner 140, and possibly by the prompt variant 142, while the visual information is generated by the visual embedder 146, and possibly by the image processor 144, as will be described in more detail below. Next, the image generator 148 generates a new image (semantic copy) based on both the semantic and visual information, as will also be described in more detail below. The captioner 140 generates semantic information (captions) using a first generative AI model 150, and the image generator 148 generates a new image (semantic copy) using a second generative AI model 152. In embodiments that support prompt mutation, the prompt mutation generator 142 uses a third generative AI model 154 to mutate the caption (or other text prompt derived from the caption). In some embodiments, the first generative AI model 150 is a multimodal large language model (LLM), the third generative AI model 154 is a fine-tuned LLM, and the second generative AI model 152 is a pixel spread model. In other embodiments, the second generative AI model 152 is a latent spread model, a normal (non-latest) spread model, or another suitable type of image generation model.
[0031] In some embodiments, as will be described in more detail below, the semantic copy generator 130 utilizes reinforcement learning and / or other feedback mechanisms to improve / fine-tune the operation of the first generator AI model 150, the second generator AI model 152, and / or the third generator AI model 154. In such embodiments, the semantic copy generator 130 acquires quality data. The computing system 102 may generate the quality data or acquire (e.g., receive) quality data from another system or device, depending on the embodiment. In the exemplary system 100 of Figure 1, the quality data is stored in the quality database 184. The quality data may be in any format or type suitable for indicating performance in a desired context. In a digital advertising context, for example, the quality data / metrics may include manually generated scores (e.g., based on human review of images), scores generated by the computing system 102 or other systems or devices (e.g., based on predictive machine learning models), or measured or predicted performance metrics such as click-through rate (CTR) or conversion rate (CVR). As a more specific example, quality data / metrics may include a set of scores (manually or computer-generated) for each image, including an aesthetic score (e.g., how "professional" the image looks), a performance score (e.g., how effective the image is in the desired context), and a relevance score (e.g., how relevant the image is to the information the advertiser wants to promote).
[0032] If the prompt mutation generator 142 is present in the semantic copy generator 130, a third generative AI model 154 is used to automatically mutate the text prompt (i.e., the caption generated by the captioner 140, or other text derived from or based on the caption). While there is some randomness in how the third generative AI model 154 mutates any text prompt, the feedback technique described in relation to Figure 3A below can significantly improve the ability of the third generative AI model 154 to modify the text prompt in a useful or helpful way (i.e., in a way that is likely to generate a new image with good / high-quality metrics).
[0033] If the image processor 144 is present in the semantic copy generator 130, it automatically modifies the source image in some way before the visual embedder 146 operates on the source image to create a vectorized visual embedding. For example, the image processor 144 may use a machine learning model to identify one or more protruding regions of the source image and crop all other parts of the source image. For example, the image processor 144 may remove background objects (to enable a second generative AI model 152 to generate an image with better type(s) or variety of background objects), remove unwanted overlays such as text or buttons / controls, and / or crop areas of the source image to focus more on the central subject of the source image. The visual embedder 146 may comprise one or more embedding layers of a machine learning model (e.g., layers(s) of the second generative AI model 152 used to generate a new image).
[0034] The image generator 148 operates by applying both visual embeddings and (potentially mutated) text prompts as input to a second generative AI model 152, which outputs a new image (i.e., a semantic copy of the source image). In some embodiments, the computing system 102 (or other system) then provides the new image to a client device (e.g., client device 104) and presents it on the display via a user interface (e.g., presented in the user interface of application 170 on display 166). For example, the new image may be provided and presented within a digital advertising context as described above.
[0035] Figure 2 shows an exemplary process 200 for performing a semantic-based copy, which creates a new image 202 based on a source image 204 (for example, from a database 180). Process 200 may be implemented by the computing system 102 in Figure 1 (for example, by software instructions for a semantic copy generator 130 executed by processor 122) or by other suitable applications and / or computing systems. For ease of explanation, process 200 is described below with reference to elements of the exemplary system 100 in Figure 1.
[0036] In stage 210, in one path of process 200, the captioner 140 generates descriptive text (i.e., a caption) for the source image 204 by applying the source image 204 as input to a first generative AI model 150 (e.g., a multimodal LLM that outputs a caption for the source image 204). In some embodiments, the caption output by the first generative AI model 150 itself is used as a text prompt 212 for subsequent processing, while in other embodiments, the text prompt 212 is the result of processing the caption in some way (e.g., annotating the caption in a predetermined language or removing restricted languages).
[0037] In stage 214, the prompt mutationr 142 modifies / mutates the text prompt 212 by applying it as input to a third generative AI model 154 (e.g., an LLM finely tuned for prompt mutation in a desired context such as digital advertising), and outputs the mutated text prompt. In an alternative embodiment, process 200 omits stage 214.
[0038] In stage 220, in another path of process 200, the image processor 144 processes the source image 204 to produce the processed image 222. Stage 220 may include identifying any protruding region(s) of the source image 204, and cropping the protruding or non-protruding regions of the source image 204. In stage 224, the visual embedder 146 uses one or more embedding layers to produce a visual embedding of the processed image 222. In other embodiments, process 200 omits stage 220, and instead the visual embedder 146 produces a visual embedding of the original source image 204.
[0039] In stage 230, the image generator 148 generates a new image 202 by applying the text prompt 212 (including or not including the variation in stage 214, depending on the embodiment) and the visual embedding from stage 224 as input to a second generative AI model 152 that generates / outputs a new image 202. In some embodiments, the second generative AI model 152 performs both stages 224 and 230. For example, the semantic copy generator 130 may apply the source image 204 (after any post-processing in stage 220) as input to the embedding layer(s) of the second generative AI model 152, and then apply both (1) the output of the embedding layer(s) and (2) the text prompt 212 (or a variation of the text prompt 212) as input to a subsequent layer of the second generative AI model 152.
[0040] Figures 3A and 3B show examples of processes 300 and 320, respectively, which are similar to process 200 in Figure 2, but also use feedback to improve performance over time. Like process 200 in Figure 2, processes 300 and 320 may be implemented by the computing system 102 in Figure 1 (for example, by software instructions of the semantic copy generator 130 executed by processor 122) or by other suitable applications and / or computing systems. Again, for ease of explanation, processes 300 and 320 are described below with reference to elements of the exemplary system 100 in Figure 1.
[0041] In Figures 3A and 3B, elements with the same labels as those in Figure 2 (e.g., images 202 and / or 204, and / or stages 210, 220, etc.) may be similar to or identical to similarly labeled elements in Figure 2. However, in Figure 3A, a prompt variation in stage 314 incorporates feedback based on quality metrics of the new image (semantic copy) generated by previous iterations of process 300. More specifically, in process 300, the semantic copy generator 130 generates or obtains quality metrics for a given new image 202 in stage 316. The “quality metrics” may be a single metric or value (e.g., CTR, or performance score, etc.) or a set of multiple metrics or values (e.g., CTR and CVR, or relevance, aesthetic, and performance scores, etc.).
[0042] Stage 316 may include, for example, obtaining the quality metric from another server that measures or calculates one or more components of the quality metric. Alternatively or additionally, Stage 316 may include the computing system 102 itself measuring or calculating one or more components of the quality metric. In some embodiments, the quality metric includes one or more predicted scores and / or other values. For example, the semantic copy generator 130 (or software of another computing system or device) may predict the performance score of the new image 202 when it is used in digital advertising, using a machine learning model that is independent of and different from models 150, 152, and 154 shown in Figure 1.
[0043] In any case, the semantic copy generator 130 provides feedback to stage 314 based on quality metrics (specifically, to fine-tune the third generative AI model 154). In some embodiments, the semantic copy generator 130 (or other components of the computing system) applies the feedback as part of a reinforcement learning technique with rewards and / or penalties. For example, the quality metrics may include an aesthetic score, a performance score, and a relevance score, each applied as a reward component. In other embodiments, the semantic copy generator 130 (or other components of the computing system) applies the feedback as an additional sample to fine-tune the third generative AI model 154. Other types of feedback are also possible to improve the performance of the third generative AI model 154. By providing feedback as shown in Figure 3A, process 300 can improve the quality or usefulness of the prompt variations in stage 314 in future iterations, thereby improving the quality of new images output by stage 230 in those future iterations.
[0044] In process 320 of Figure 3B, the captioning in stage 310 incorporates feedback based on a quality metric of the new image (semantic copy) generated by previous iterations of process 320. Specifically, in process 320, the semantic copy generator 130 generates or obtains a quality metric of a given new image 202 in stage 316. The quality metric, and the manner in which the quality metric is generated or obtained, may be as described above, for example, with reference to Figure 3A. In process 320, the semantic copy generator 130 provides feedback to stage 310 (specifically, to the first generative AI model 150) based on the quality metric. In some embodiments, the semantic copy generator 130 (or other components of the computing system) applies the feedback as part of a reinforcement learning technique with rewards and / or penalties. For example, the quality metric may include an aesthetic score, a performance score, and a relevance score, each applied as a reward component. In other embodiments, the semantic copy generator 130 (or other computing system component) applies feedback as additional samples to fine-tune the first generative AI model 150. Other types of feedback are also possible to improve the performance of the first generative AI model 150. By providing feedback as shown in Figure 3B, the process 300 can improve the quality or usefulness of the captioning at stage 310 in future iterations, thereby improving the quality of the new images output by stage 230 in those future iterations.
[0045] In some embodiments, the semantic copy generator 130 employs both feedback techniques of process 300 and process 320 simultaneously, and / or employs one or more other feedback techniques. For example, the semantic copy generator 130 can improve the performance of the second generative AI model 152 used in stage 230 for image generation, either alternatively or by using similar feedback techniques.
[0046] In some embodiments, the image generation stage 230 (e.g., during processes 200, 300, or 320) includes providing a display of the new image 202 in a desired aspect ratio as input to a second generative AI model 152. In some of these embodiments, the desired aspect ratio is input to the second generative AI model 152, for example, as a condition, independently of (potentially mutated) text prompts. In other embodiments, the semantic copy generator 130 automatically annotates or otherwise modifies the text prompts with the desired aspect ratio. In any case, by using the disclosed techniques (and in particular by jointly using both the semantic / text path and the visual embedding path), the second generative AI model 152 can more seamlessly change the aspect ratio of the source image (e.g., source image 204). For example, the aspect ratio can be seamlessly changed from portrait to landscape, or vice versa (for example, without the format change causing objects to be positioned in an aesthetically displeasing manner, and / or without emphasizing or minimizing the features of the new image in a way that impairs the functionality of the new image).
[0047] Figure 4 shows an example of a semantic copy that a computing system 102 (e.g., by the semantic copy generator 130) might create using, for example, process 200 in Figure 2, process 300 in Figure 3A, or process 320 in Figure 3B (e.g., corresponding to a new image 202). In Figure 4, the source image is shown on the left, and the corresponding new image is shown on the right. As seen in Figure 4, each new image significantly modifies the corresponding source image while closely adhering to the specific visual qualities of the source image. Therefore, for example, a semantic copy of a source image that is a digital advertisement for a company may maintain the visual qualities (style, brand colors, etc.) associated with that company or its advertisement.
[0048] Figure 5 shows examples of other semantic copies that may be created by computing system 102 (e.g., semantic copy generator 130) in an alternative embodiment where a new image is generated using text prompts, but without visual embedding (e.g., the upper paths of processes 200, 300, or 320 are used, but their respective lower paths are not). However, as seen in Figure 5, such approaches may be inferior to the approaches of processes 200, 300, or 320, and the new images may not maintain significant visual similarity to the source images. For example, the upper right semantic copy in Figure 5 is thematically similar to the upper left source image (both showing a house with a garden and sky), but the upper right image has a different overall feel and represents a completely different style of house. Similarly, the lower right semantic copy in Figure 5 is thematically similar to the lower left source image (both showing elements of a bathroom), but the lower right image has a different overall feel and represents completely different elements within a bathroom.
[0049] Figure 6 shows a flowchart of an exemplary method 600 for semantic-based copying of a source image. Method 600 can be implemented, for example, by the computing system 102 (e.g., semantic copy generator 130) shown in Figure 1.
[0050] In block 602, a text prompt is generated. Block 602 includes applying the source image to a first generative AI model (e.g., first generative AI model 150) to generate a descriptive caption for the source image. Block 602 may be, for example, stage 210 of process 200 or 300, or stage 310 of process 320.
[0051] In block 604, a visual embedding is generated based on the source image. Block 604 may be similar to, for example, stage 224 of process 200, 300, or 320.
[0052] In block 606, a new image is generated based on a text prompt and a visual embedding using a second generative AI model (e.g., second generative AI model 152). Block 606 may be similar to, for example, stage 230 of process 200, 300, or 320. In some embodiments, block 606 includes generating a mutated text prompt by applying a text prompt to a third generative AI model (e.g., third generative AI model 154) and then applying the mutated text prompt and visual embedding to the second generative AI model. For example, block 606 may be similar to the combination of stages 214 and 230 in Figure 2 or Figure 3B, or the combination of stages 314 and 230 in Figure 3A.
[0053] Method 600 may include iterations on multiple new images and may include a feedback mechanism. For example, Method 600 may include generating multiple new images based on the same text prompt and visual embedding using a second generative AI model, wherein generating these new images may include generating mutated text prompts by applying the text prompt to a third generative AI model, and generating each of the new images by applying one of each of the mutated text prompts to the second generative AI model.
[0054] Method 600 may include one or more additional blocks not shown in Figure 6. For example, Method 600 may include a first additional block in which quality metrics are generated or acquired, each quality metric corresponding to one of the new images described above, and a second additional block in which a third generative AI model is fine-tuned based on the quality metrics (e.g., using reinforcement learning or other techniques).
[0055] It should be understood that the blocks in Figure 6 do not need to be carried out in the order shown. For example, block 604 may occur before block 602 or concurrently with block 602.
[0056] As is evident from the above description, the technology disclosed herein uses artificial intelligence to generate high-performance images. Artificial intelligence (AI) is a segment of computer science that focuses on creating models that can perform tasks with little or no human intervention. Systems of artificial intelligence can utilize, for example, machine learning, natural language processing, and computer vision. Machine learning and its subsets, such as deep learning, focus on developing models that can infer outputs from data. Outputs may include, for example, prediction and / or classification. Natural language processing focuses on analyzing and generating human language. Computer vision focuses on analyzing and interpreting images and videos. Systems of artificial intelligence can include generative models that generate new content, such as images, videos, text, audio, and / or other content, in response to input prompts and / or based on other information.
[0057] Exemplary machine learning models include neural networks or other multi-layered nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some exemplary machine learning models can leverage attention mechanisms such as self-attention. For example, some machine learning models can include multi-head self-attention models (e.g., transformer models).
[0058] Models can be trained using a variety of training or learning techniques. Training can include supervised learning, unsupervised learning, and reinforcement learning. Techniques such as backpropagation can be used during training. For example, a loss function can be backpropagated through the model(s) to update one or more of the model(s) parameters (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent can be used to iteratively update parameters over several training iterations. Several generalization techniques (e.g., weight decay, dropout) can be used to improve the generalization ability of the trained model.
[0059] Prior to domain-specific alignment, the model(s) can be pre-trained. For example, a model can be pre-trained on a general corpus of training data and then fine-tuned to a more targeted corpus of training data. The model can be aligned using prompts designed to elicit domain-specific outputs. The prompts can be designed to include learned prompt values (e.g., soft prompts). The trained model(s) may be validated before use using input data other than the training data and may be further updated or improved based on additional feedback / input during use.
[0060] In some embodiments, the computing system 102 can use one or more of the machine learning models or techniques described above to perform one or more of the operations discussed herein in relation to machine learning. For example, the computing system 102 can use one or more such machine learning techniques to pre-train and / or fine-tune a first generative AI model 150, a second generative AI model 152, and / or a third generative AI model 154, and, in some cases, pre-train and / or fine-tune a model that predicts the performance of an image (e.g., generating the quality metrics described above).
[0061] While the preceding text describes in detail numerous different aspects and embodiments of the present invention, it should be understood that the scope of the patent is defined by the claims language set out at the end of this patent. The embodiments for carrying out the invention should be interpreted as illustrative only and do not describe all possible embodiments, for describing all possible embodiments would be impractical, if not impossible. Numerous alternative embodiments can be implemented using the current art or art developed after the filing date of this patent, and these remain within the scope of the claims. The disclosure herein assumes at least the following examples: [Examples]
[0062] Example 1. A semantic-based image copying method comprising: generating a text prompt by one or more processors, the generation of the text prompts including applying the source image to a first generative artificial intelligence (AI) model to generate a descriptive caption for the source image; generating a visual embedding based on the source image by one or more processors; and generating a new image based on the text prompt and the visual embedding by one or more processors using a second generative AI model.
[0063] Example 2. The method according to Example 1, wherein generating the new image comprises generating a mutated text prompt by applying the text prompt to a third generative AI model, and applying the mutated text prompt and the visual embedding to the second generative AI model.
[0064] Example 3. The method of Example 2, wherein one or more processors generate a plurality of new images based on the text prompt and the visual embedding using the second generative AI model, comprising at least in part (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative AI model, and (ii) generating each of the plurality of new images by applying each of the plurality of mutated text prompts to the second generative AI model.
[0065] Example 4. The method according to Example 3, comprising: generating or acquiring a plurality of quality metrics corresponding to each of the plurality of new images by one or more processors; and fine-tuning the third generative AI model based on the plurality of quality metrics by one or more processors.
[0066] Example 5. The method according to Example 1, wherein generating the new image comprises applying the text prompt and the visual embedding to the second generative AI model.
[0067] Example 6. The method according to any one of Examples 1 to 5, wherein generating the new image based on the text prompt and the visual embedding includes generating the new image to have a different aspect ratio from the source image.
[0068] Example 7. The method according to any one of Examples 1 to 6, wherein generating the visual embedding based on the source image comprises processing the source image and generating the visual embedding based on the processed source image.
[0069] Example 8. The method according to Example 7, wherein processing the source image includes cropping a portion of the source image.
[0070] Example 9. The method according to any one of Examples 1 to 8, wherein the text prompt is the descriptive caption.
[0071] Example 10. The method according to any one of Examples 1 to 9, wherein the second generative AI model is a pixel spread model.
[0072] Example 11. A system comprising one or more processors and one or more memories for storing instructions which, when executed by the one or more processors, cause the one or more processors to (i) generate a text prompt, the generation of the text prompt includes applying the source image to a first generative artificial intelligence (AI) model to generate a descriptive caption for the source image, (ii) generate a visual embedding based on the source image, and (iii) use a second generative AI model to generate a new image based on the text prompt and the visual embedding.
[0073] Example 12. The system according to Example 11, wherein generating the new image includes generating a mutated text prompt by applying the text prompt to a third generative AI model, and applying the mutated text prompt and the visual embedding to the second generative AI model.
[0074] Example 13. The system according to Example 12, wherein the instruction causes one or more processors to generate a plurality of new images based on the text prompts and the visual embeddings, by at least partially (i) generating a plurality of mutated text prompts by applying the text prompt to the third generative AI model, and (ii) generating each of the plurality of new images by applying one of the plurality of mutated text prompts to the second generative AI model.
[0075] Example 14. The system according to Example 13, wherein the instruction causes one or more processors to generate or acquire a plurality of quality metrics corresponding to each of the plurality of new images, and to fine-tune the third generative AI model based on the plurality of quality metrics.
[0076] Example 15. The system according to Example 11, wherein generating the new image includes applying the text prompt and the visual embedding to the second generative AI model.
[0077] Example 17. The system according to any one of Examples 11 to 16, wherein generating the new image based on the text prompt and the visual embedding includes generating the new image to have a different aspect ratio from the source image.
[0078] Example 18. The system according to any one of Examples 11 to 17, wherein generating the visual embedding based on the source image comprises processing the source image and generating the visual embedding based on the processed source image.
[0079] Example 19. The system according to Example 18, wherein processing the source image includes cropping a portion of the source image.
[0080] Example 20. The system according to any one of Examples 11 to 19, wherein the text prompt is the descriptive caption.
[0081] Example 21. The system according to any one of Examples 11 to 20, wherein the second generative AI model is a pixel spread model.
[0082] Example 22. One or more non-temporary computer-readable media storing instructions, wherein, when executed by one or more processors, the one or more processors are instructed to generate text prompts, the generation of the text prompts comprising applying the source image to a first generative artificial intelligence (AI) model to generate a descriptive caption for the source image, generating a visual embedding based on the source image, and using a second generative AI model to generate a new image based on the text prompts and the visual embedding.
[0083] Example 23. One or more non-temporary computer-readable media according to Example 22, wherein generating the new image comprises generating a mutated text prompt by applying the text prompt to a third generative AI model, and applying the mutated text prompt and the visual embedding to the second generative AI model.
[0084] Example 24. One or more non-temporary computer-readable media according to Example 23, wherein the instruction causes one or more processors to generate a plurality of new images based on the text prompts and the visual embeddings, by at least partially (i) generating a plurality of mutated text prompts by applying the text prompts to the third generative AI model, and (ii) generating each of the plurality of new images by applying one of the plurality of mutated text prompts to the second generative AI model.
[0085] Example 25. One or more non-temporary computer-readable media according to Example 24, wherein the instruction causes one or more processors to generate or acquire a plurality of quality metrics corresponding to each of the plurality of new images, and to fine-tune the third generative AI model based on the plurality of quality metrics.
[0086] Example 26. One or more non-temporary computer-readable media according to Example 22, wherein generating the new image includes applying the text prompt and the visual embedding to the second generative AI model.
[0087] Example 27. One or more non-temporary computer-readable media according to any one of Examples 22 to 26, wherein generating the new image based on the text prompt and the visual embedding includes generating the new image to have a different aspect ratio from the source image.
[0088] Example 28. One or more non-temporary computer-readable media according to any one of Examples 22 to 27, wherein generating the visual embedding based on the source image comprises processing the source image and generating the visual embedding based on the processed source image.
[0089] Example 29. One or more non-temporary computer-readable media according to Example 28, wherein processing the source image includes cropping a portion of the source image.
[0090] Example 30. The second generative AI model is a pixel spread model, in one or more non-temporary computer-readable media as described in any one of Examples 22 to 29.
[0091] Example 31. One or more non-temporary computer-readable media according to any one of Examples 22 to 30, wherein the text prompt is the descriptive caption.
[0092] The following additional considerations apply to the preceding discussion and the appended claims. Throughout this specification, multiple examples may implement a component, operation or structure described as a single example. While individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed simultaneously, and nothing requires the operations to be performed in the order illustrated. Structures and functionalities represented as separate components in the examples may be implemented as a combined structure or component. Similarly, structures and functionalities represented as single components may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this disclosure.
[0093] Unless otherwise evident from the context of use, references in this disclosure to the same set of “one or more processors” (or the same “multiple processors,” etc.) performing multiple operations may include embodiments in which the performance of the operations is divided among the processors in any suitable manner. For example, “the production of X by one or more processors and the production of Y by one or more processors” may include (1) embodiments in which a first set of one or more processors (e.g., in a first computing device) produces X and a second set of another one or more processors (e.g., in a different second computing device) independently produces Y; (2) embodiments in which all processors within the set of one or more processors (e.g., all in the same device or distributed across multiple devices) contribute to the production of both X and Y; and (3) other variations.
[0094] Unless otherwise expressed, any discussion in this disclosure using terms such as “process,” “calculate,” “determine,” “represent,” “display,” and similar terms may refer to the actions or processing of a machine (e.g., a computer) that operates or transforms data that is represented as a physical (e.g., electronic, magnetic, or optical) quantity within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0095] Where used in this disclosure, any reference to “one implementation” or “an implementation” means that any particular element, feature, structure, or characteristic described in relation to an implementation is included in at least one implementation. The phrase “in one implementation” appearing in various places in this specification does not necessarily refer to the same implementation in all instances.
[0096] Where used in this disclosure, “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variation thereof are intended to encompass non-exclusive inclusion. For example, a process, method, item, or apparatus that includes a list of elements is not necessarily limited to these elements alone, and may include other elements not expressly enumerated or inherent in such process, method, item, or apparatus. Furthermore, unless expressly stated otherwise, “or” refers to an inclusive or not an exclusive or. For example, condition A or B is satisfied by any one of the following: A is true (or exists) and B is false (or does not exist), A is false (does not exist) and B is true (or exists), and both A and B are true (or exist).
[0097] Those skilled in the art will understand, by reading this disclosure, further additional alternative structural and functional designs through the principles described herein. Therefore, while specific embodiments and applications have been described and explained, it should be understood that the disclosed embodiments are not limited to the exact structures and components disclosed herein. Various modifications, changes, and variations obvious to those skilled in the art may be made in the arrangement, operation, and details of the methods and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
Claims
1. A semantic-based method of copying images, The process involves generating a text prompt using one or more processors, wherein the generation of the text prompt includes applying the source image to a first generative artificial intelligence (AI) model in order to generate a descriptive caption for the source image. The one or more processors generate a visual embedding based on the source image, A method comprising: using one or more processors to generate a new image based on the text prompt and the visual embedding using a second generative AI model.
2. The generation of the aforementioned new image is By applying the aforementioned text prompt to a third generative AI model, a mutated text prompt is generated. The method according to claim 1, comprising applying the mutated text prompt and the visual embedding to the second generative AI model.
3. The process involves generating a plurality of new images based on the text prompt and the visual embedding using the second generative AI model with the one or more processors, at least in part. (i) To generate a plurality of mutated text prompts by applying the text prompt to the third generative AI model, (ii) The method of claim 2, comprising generating each of the plurality of new images by applying each of the plurality of mutated text prompts to the second generative AI model.
4. The one or more processors generate or acquire a plurality of quality metrics corresponding to each of the plurality of new images, The method according to claim 3, further comprising fine-tuning the third generative AI model based on the plurality of quality indicators using one or more processors.
5. The generation of the aforementioned new image is The method according to claim 1, comprising applying the text prompt and the visual embedding to the second generative AI model.
6. Generating the new image based on the text prompt and the visual embedding is, The method according to any one of claims 1 to 5, comprising generating the new image having a different aspect ratio from the source image.
7. Generating the visual embedding based on the source image is Processing the aforementioned source image, The method according to any one of claims 1 to 6, comprising generating the visual embedding based on the processed source image.
8. Processing the aforementioned source image is The method according to claim 7, comprising cropping a portion of the source image.
9. The method according to any one of claims 1 to 8, wherein the text prompt is the descriptive caption.
10. The method according to any one of claims 1 to 9, wherein the second generative AI model is a pixel spread model.
11. It is a system, One or more processors, A system comprising: one or more memories that, when executed by the one or more processors, store instructions causing the one or more processors to perform the method according to any one of claims 1 to 10.
12. One or more non-temporary computer-readable media, which, when executed by one or more processors, store instructions causing the one or more processors to perform the method according to any one of claims 1 to 10.