Image library recommendation service based on controllable diffusion model
By introducing a controllable diffusion model and dynamic image frames into the image gallery recommendation system, users can collaborate with the system through a variety of interactive methods to dynamically adjust the generated images, solving the problem of dynamic limitations in capturing user preferences in existing technologies and achieving more personalized and efficient image search.
Patent Information
- Application Number
- CN202480012123.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-27
- Filing Date
- 2024-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
Existing image gallery recommendation systems have inherent limitations in capturing the dynamics of user preferences, and the ways in which users can interact with the system are limited, making it difficult to achieve highly personalized and dynamic image search.
It adopts a controllable diffusion model, combined with dynamic image frames and interactive components, allowing users to interact with the model through a variety of input methods, including text, sketches, images, etc., and dynamically adjust the generated image to meet personalized needs.
It provides a more natural and dynamic image search experience. Users can gradually fine-tune the generated images through a simple interactive process to meet precise requirements, thereby improving the relevance of image recommendations and user satisfaction.
Smart Images

Figure CN120677473A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image search and recommendation systems, and more particularly to dynamic image search using a controllable diffusion model in an image gallery recommendation service. Background Art
[0002] Image gallery recommendation systems (also known as vision-based or image-based discovery systems) play an increasingly important role in modern applications across many different domains, including e-commerce, social media, and entertainment. The main goal of image gallery recommendation systems is to predict and recommend one or more relevant images that may be of interest or appeal to a user. To achieve this, image gallery recommendation systems utilize a variety of techniques, such as collaborative filtering, content-based filtering, and deep learning, to provide personalized recommendations based on the user's characteristics, preferences, and behavior.
[0003] Collaborative filtering involves analyzing the previous behavior and preferences of similar users to make recommendations. By examining historical data of users with similar tastes and preferences as a given user, image gallery recommendations can more accurately identify images that the user may be interested in. Collaborative filtering can be item-based, user-based, or a combination of both. The former focuses on similarities between items (images), while the latter focuses on similarities between users.
[0004] Content-based filtering refers to analyzing the content and features of the image(s) themselves to make recommendations. Image gallery recommendation systems can extract relevant information such as color, texture, shape, and other visual attributes from images and use this extracted information to find similar images (using feature similarity, distance metrics, etc.). By recommending images that are visually similar to images that the user has already expressed interest in, content-based filtering aims to capture user preferences based on image characteristics.
[0005] Deep learning techniques such as convolutional neural networks (CNNs), variational autoencoders (VAEs), and transformer networks have revolutionized image recommendation systems. CNNs process images through multiple layers of interconnected neurons, learning complex patterns and features (hierarchical representations) from images. By training on large datasets, these networks can capture complex relationships (local and global image features) and make accurate predictions about user preferences based on image content.
[0006] VAEs are generative models that learn a compact representation (latent space) of input data. In the context of image recommendation, VAEs can learn low-dimensional representations of images that capture the underlying structure and variation in the dataset. By leveraging this latent space, VAEs can generate new and diverse images that match user preferences, thereby enhancing the recommendation capabilities of image gallery recommendation services.
[0007] Transformer networks were originally designed for natural language processing tasks but have since proven successful in a range of other applications, including image recommendation (e.g., in computer vision). Transformers can model long-range dependencies and capture contextual information in the data. In image gallery recommendation systems, transformer networks can be used to learn complex contextual relationships between images and generate more accurate recommendations based on this contextual information.
[0008] The image gallery recommendation system can also rely on user behavior data (user interaction) to improve user satisfaction, engagement, and the overall user experience. In terms of user interaction, the image gallery recommendation system can provide users with several ways to engage with the system. For example, in an implementation where (multiple) users interact with the image gallery recommendation system through a user interface (such as a mobile app or website), an initial set of selected images can be presented to the user. The user can then interact with the system by viewing images (e.g., scrolling through a collection of recommended images), liking / disliking images, saving images, sharing images (e.g., via a coupled social media platform), and / or otherwise interacting actively or negatively with one or more images in the image gallery. These user interactions can be used as feedback to the system to better understand user preferences and refine future recommendations.
[0009] It's worth noting that while user interactions play an important role in training and refining image gallery recommendation systems, these interactions are somewhat limited—notably, users have no direct control over the underlying algorithms and model parameters of image gallery recommendation systems. While systems can learn from aggregated user data to improve recommendations for both individual users and the user base as a whole, individual users typically only engage with image galleries through a limited number of pathways (viewing a static image, liking / disliking it, saving it, commenting on it, etc.). Unfortunately, these techniques have inherent limitations in capturing the full dynamics of user preferences. Summary of the Invention
[0010] Embodiments of the present invention are directed to a method for performing dynamic image search in an image gallery recommendation service using a controllable diffusion model. A non-limiting example method includes displaying an image gallery having a plurality of gallery images and a dynamic image frame. The dynamic image frame may include a generated image and an interactive component. The method may include receiving user input in the interactive component and, in response to receiving the user input, generating an updated generated image by inputting the user input into the controllable diffusion model. The method may include replacing the generated image in the dynamic image frame with the updated generated image.
[0011] In some embodiments, the method includes receiving an image query in an information column of an image gallery. In some embodiments, the generated image is generated by inputting the image query into a controlled diffusion model. In some embodiments, the plurality of gallery images and the generated image are selected based on a degree of matching with one or more features in the image query.
[0012] In some embodiments, the method includes determining one or more constraints in the image query. In some embodiments, the generated image is generated by inputting the one or more constraints into a controlled diffusion model.
[0013] In some embodiments, the one or more constraints in the image query include at least one of a pose skeleton and an object boundary. In some embodiments, determining the one or more constraints includes extracting the object boundary when the feature in the image query includes one of a structure and a geological feature. In some embodiments, determining the one or more constraints includes extracting the pose skeleton when the feature in the image query includes one of a person and an animal.
[0014] In some embodiments, the plurality of gallery images are from an image database.
[0015] In some embodiments, the interactive component includes a text information field. In some embodiments, receiving user input in the interactive component includes receiving a text string input into the text information field.
[0016] In some embodiments, the interactive component includes one or more of a drop-down menu, a checkbox, a slider, a color picker, a canvas interface for drawing or sketching, and a rating button.
[0017] In some embodiments, the interactive component includes a canvas for magic wand input. In some embodiments, receiving user input in the interactive component includes receiving a magic wand input that graphically selects one of a specific feature and a specific area in the generated image. In some embodiments, the interactive component also includes a text information field. In some embodiments, receiving user input in the interactive component also includes receiving a text string input that includes contextual information specific to the magic wand input.
[0018] Embodiments of the present invention relate to a system for dynamic image search using a controllable diffusion model in an image gallery recommendation service. A non-limiting example system includes a memory having computer-readable instructions and one or more processors for executing the computer-readable instructions. The computer-readable instructions control the one or more processors to perform various operations. The operations include receiving an image query from a client device communicatively coupled to the system. The operations include providing a plurality of gallery images and generated images to the client device based on a degree of match with one or more features in the image query. The operations include receiving user input from the client device and, in response to the user input, generating an updated generated image by inputting the user input into the controllable diffusion model. The operations include providing the updated generated image to the client device.
[0019] Embodiments of the present invention relate to a system for performing dynamic image search in an image gallery recommendation service using a controllable diffusion model. A non-limiting example system includes a memory having computer-readable instructions and one or more processors for executing the computer-readable instructions. The computer-readable instructions control the one or more processors to perform various operations. These operations include receiving a plurality of gallery images and a generated image from an image gallery recommendation service communicatively coupled to the system. These operations include displaying an image gallery including a plurality of gallery images and a dynamic image frame, the dynamic image frame including a generated image and an interactive component. These operations include receiving user input in the interactive component, transmitting the user input to the image gallery recommendation service, and receiving an updated generated image from the image gallery recommendation service. These operations include replacing the generated image in the dynamic image frame with the updated generated image.
[0020] The above features and advantages and other features and advantages of the present disclosure will become apparent from the following detailed description when taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The specific content of the exclusive rights claimed herein is particularly pointed out and expressly stated in the claims at the end of the specification. The foregoing and other features and advantages of the embodiments of the present invention will be apparent from the following detailed description taken in conjunction with the accompanying drawings:
[0022] Figure 1 depicts a block diagram for using a controlled diffusion model in accordance with one or more embodiments;
[0023] Figure 2 depicts a block diagram for dynamic image search using a controllable diffusion model according to one or more embodiments;
[0024] Figure 3 depicts an example image gallery according to one or more embodiments;
[0025] Figure 4 Depicts the following user interaction according to one or more embodiments. Figure 3 A gallery of example images;
[0026] Figure 5 Depicts the following user interaction according to one or more embodiments. Figure 4 A gallery of example images;
[0027] Figure 6 Depicts the following user interaction according to one or more embodiments. Figure 5 A gallery of example images;
[0028] Figure 7 Depicts the following user interaction according to one or more embodiments. Figure 6 A gallery of example images;
[0029] Figure 8 depicts an example image gallery according to one or more embodiments;
[0030] Figure 9 Depicts the following user interaction according to one or more embodiments. Figure 8 A gallery of example images;
[0031] Figure 10 depicts a block diagram of a computer system according to one or more embodiments; and
[0032] Figure 11 A flow chart is depicted of a method for performing dynamic image search within an image gallery recommendation service using a controllable diffusion model, according to one or more embodiments.
[0033] The schematic diagrams depicted herein are for illustrative purposes only. Many modifications may be made to the schematic diagrams or operations described therein without departing from the spirit of the present invention. For example, the operations may be performed in a different order, or operations may be added, removed, or modified.
[0034] In the accompanying drawings and the subsequent detailed description of the embodiments of the present invention, the various elements illustrated in the accompanying drawings are provided with double or triple digit reference numbers. With a few exceptions, the leftmost digit(s) of each reference number corresponds to the drawing in which the element first appears. DETAILED DESCRIPTION
[0035] Image gallery recommendation systems are widely used in various fields, including e-commerce, social media, and entertainment, to provide personalized recommendations to users. These systems use various techniques, such as collaborative filtering, content-based filtering, and deep learning, to recommend images to users. However, these techniques have inherent limitations in capturing the dynamics of user preferences. Typically, users engage with these systems solely by viewing static images, clicking "like / dislike," saving them, and / or sharing and commenting on them.
[0036] The present disclosure introduces the application of a so-called controllable diffusion model for dynamic image search in an image gallery recommendation service. A diffusion model refers to a class of generative models that utilize a diffusion process to generate high-quality synthetic data. Diffusion refers to the gradual spreading or diffusion of information or noise in a data space (e.g., an image), and the diffusion process in the diffusion model involves iteratively transforming an initial noise vector into a sample by applying a series of diffusion steps. Each diffusion step adds controlled noise to the image while gradually reducing the noise level (noise vector), thereby gradually refining the generated image. By carefully controlling the noise process, the diffusion model can generate high-quality images that exhibit convincing details. In the context of image recommendation, the diffusion model can be used to generate realistic, abstract, synthetic, re-textured, artistic, and other types of images from user prompts that are visually similar to real images in the training dataset. For example, a diffusion model can create a novel watercolor painting of a sailboat on a river from the prompt "Paint a sailboat and a river with watercolor."
[0037] A "controllable" diffusion model refers to a diffusion model that can be dynamically guided and fine-tuned through additional user interaction after the image is generated. For example, in some embodiments, the controllable diffusion model can be accessed via a dynamic image frame, which includes an image (itself an output of the diffusion model) and an interactive component that can be selected, edited, and / or otherwise manipulated by the user. In some embodiments, the interactive component can receive input from the user, such as text input, sketches, images, etc. In some embodiments, the controllable diffusion model can generate new images and / or change previously generated images using the user input received via the interactive component of the dynamic image frame as guidance. Continuing with the previous example, in response to the user entering additional text "Change sailboat to motorboat" into the interactive component, the controllable diffusion model can modify the generated watercolor painting of a sailboat on a river by replacing the sailboat with a motorboat.
[0038] In some embodiments, as part of an image gallery recommendation service, a dynamic image frame and its corresponding image are positioned within a collection of other images within an entire image gallery (referred to herein as gallery images). The gallery images may include images retrieved from an image database (non-generated images). In some embodiments, the retrieved images match an image query received from a user and / or are images that match one or more characteristics of the user. In this way, a user can quickly browse through an image collection (including generated and retrieved images) to find an image(s) of interest.
[0039] Advantageously, according to one or more embodiments, utilizing a controllable diffusion model for dynamic image search in an image gallery recommendation service enables providing users with a more natural and dynamic image search and image gallery experience. Unlike traditional diffusion models (which can frustrate users due to their inherent limitations), the controllable diffusion model and dynamic image frame described herein allow users to easily guide the output of the diffusion model within the image recommendation framework to obtain result images that better meet user needs. In short, the controllable diffusion model and dynamic image frame allow the user and the diffusion model to collaborate and iteratively interact in a simple process to gradually fine-tune the generated (multiple) images to meet the user's precise requirements. Ultimately, the image gallery recommendation service is able to efficiently generate highly relevant images in a collaborative and engaging manner with the user.
[0040] Figure 1 A block diagram 100 is depicted for using a controlled diffusion model 102 according to one or more embodiments of the present invention. Figure 1 As shown in FIG, an image query 104 is received through a topic identification and constraint graph 106. The image query 104 may be received by an external system (e.g., a client device, see FIG. Figure 2 ) generated and / or otherwise obtained. The image query 104 (also referred to as a prompt) can include text and other types of input modalities as described above. For example, the image query 104 may include the text string "show me paintings". In another example, the image query 104 may include a user-provided sketch, stick figure drawing, etc. of a person or animal. In yet another example, the image query 104 may include the text string "retexture this in natural wood" and an image of a metal chair.
[0041] In some embodiments, the topic identification and constraint graph 106 receives and processes the image query 104. In some embodiments, the topic identification and constraint graph 106 includes a module configured to identify subject(s) and / or subject(s) (collectively, "topics") conveyed in the image query 104. Topic identification helps the controlled diffusion model 102 better understand the content and / or context of the prompt, which can ensure more relevant and coherent output. For example, by identifying that the topic in the prompt is a birthday celebration, the controlled diffusion model 102 can adjust its response to be consistent with the expected subject matter (e.g., displaying a birthday cake, candles, gifts, etc.), thereby producing a more accurate and meaningful output.
[0042] For text prompts, topic identification can involve analyzing the text to extract key information representing the subject and / or theme of interest. In some embodiments, the topic identification and constraint diagram 106 includes a natural language processing (NLP) module configured for NLP topic extraction (such as keyword extraction, named entity recognition, and / or topic modeling). These methods help identify important keywords, entities, or topics within the prompt.
[0043] For image cues, topic identification can involve analyzing visual content to understand the objects, scenes, entities, and / or concepts depicted in the image. In some embodiments, the topic identification and constraint graph 106 includes (multiple) visual processing (VP) modules configured for object detection, scene recognition, image captioning, and / or other techniques for extracting relevant information from visual input.
[0044] For audio prompts (including voice / speech data), topic identification can involve applying automatic speech recognition (ASR) techniques (whether or not subsequent NLP methods are used) to extract relevant information from the prompts. In some embodiments, the topic identification and constraint diagram 106 includes (multiple) ASR modules that are configured to: transcribe the audio input (e.g., spoken words) into text, pre-processing steps to clean and normalize the text data (e.g., this may involve removing punctuation, converting the text to lowercase, and handling any specific text challenges associated with the audio transcription process), and topic modeling (such as latent Dirichlet allocation (LDA) and non-negative matrix factorization (NMF)) to discover potential themes and / or topics within the transcribed text.
[0045] Similar techniques can be used for topic identification for other input modalities, and the techniques provided herein merely illustrate the breadth of available techniques.Once a topic is identified, it can be fed as input to the controlled diffusion model 102 .
[0046] In some embodiments, topic identification and constraint graph 106 includes a module configured to identify one or more constraints within image query 104. Constraints provide additional information and / or restrictions that must be adhered to when generating output and help ensure that the generated results will meet specific requirements or exhibit desired characteristics. For example, constraints can include two-dimensional (2D) or three-dimensional (3D) object boundaries and pose skeletons. Other constraint types are also possible.
[0047] Object boundary lines (two-dimensional and three-dimensional) can be used as constraints to guide the generation of images with specific shapes or outlines. For example, if a user wants to generate an image of a car, the user can provide a rough outline or boundary of the car as a constraint. In some embodiments, the controlled diffusion model 102 can extract this constraint and then use it to generate an image of the vehicle that aligns with the specified boundary lines.
[0048] A pose skeleton represents the underlying structure or spatial arrangement of body parts in an image (e.g., a human body). The pose skeleton can define the relative positions and orientations of various body parts (e.g., joints and limbs) of the human body(ies) in the image. By extracting the pose skeleton as a constraint, the controlled diffusion model 102 can generate images that conform to a specified pose.
[0049] Other types of constraints include, for example, style constraints, environmental constraints, spatial constraints, and color constraints. Style constraints can specify a particular style, aesthetic, etc. that the user wants the generated output to exhibit. Environmental constraints capture the desired environment to influence the generation process. Environmental constraints can, for example, include factors such as the time of day, weather conditions, the presence of (multiple) specific objects in (multiple) specific areas of the output image. Spatial constraints are related to the spatial relationships between objects in the image. For example, a spatial constraint might specify the relative position of two objects in a scene (e.g., a chair is placed to the right of a person). Color constraints define the color(s) to be used in the generated output. This can include extracting specific colors from the prompt as well as a color palette, dominant colors, and color distribution.
[0050] Constraint identification involves identifying and extracting relevant constraints from the prompts and may vary depending on the type of constraints used. For example, constraint identification may involve analyzing explicit annotations, understanding natural language descriptions, processing visual elements within the prompts, etc. Once constraints are identified using the topic identification and constraint graph 106, they can be incorporated into the generation of the controlled diffusion model 102.
[0051] In some embodiments, the controlled diffusion model 102 receives the identified topic(s) and any constraints from the topic identification and constraint graph 106. In some embodiments, the controlled diffusion model 102 is configured to create generated image(s) 108 from the identified topic(s) and constraints in the image query 104.
[0052] In some embodiments, the controllable diffusion model 102 also receives user input 110. For example, the controllable diffusion model 102 may receive user input 110 via an interactive component of a dynamic image frame (see Figure 2 ). For example, the user input 110 may include text and / or other input modalities for expressing and / or defining additional constraints on the generated image(s) 108. For example, the user input 110 may include the text "add more trees," which, when used in conjunction with the image query 104 including a scene with several trees near a lake, may cause the controllable diffusion model 102 to create a generated image 108 with more trees near a lake. Figure 2 The implementation of user input 110 and interactive components is discussed in more detail in .
[0053] In some embodiments, user input 110 is passed directly to controlled diffusion model 102. In some embodiments, user input 110 is passed to controlled diffusion model 102 via topic identification and constraint graph 106. For example, topic identification and constraint graph 106 may identify one or more constraints within user input 110 and pass these constraints to controlled diffusion model 102. In some embodiments, user input 110 is passed both directly to controlled diffusion model 102 and as additional constraints extracted by topic identification and constraint graph 106.
[0054] Figure 2 A block diagram 200 is depicted for performing dynamic image search using the controlled diffusion model 102 according to one or more embodiments of the present invention. Figure 2 As shown in FIG, the controllable diffusion model 102 (see Figure 1 For more internal details) can be incorporated into the image gallery recommendation service 202 or incorporated as part of the image gallery recommendation service 202. The implementation of the image gallery recommendation service 202 is not meant to be particularly limited, but may, for example, include a remote or local server (or a service running on or in conjunction with the server), an application accessible through a browser and / or mobile device app (e.g., a web-based application), (multiple) content management systems, etc. In some embodiments, the image gallery recommendation service 202 is an application and / or service integrated / embedded within another service or platform (e.g., within a social media platform, within a browser search page, etc.).
[0055] In some embodiments, the image gallery recommendation service 202 is accessed by a client device 204. The client device 204 is not meant to be particularly limiting, but may include, for example, a personal computer (desktop, laptop, e-reader, etc.), a smartphone, a tablet, a smart home device, a wearable device (smart watch, fitness tracker, etc.), a smart TV, a streaming device, a gaming console, a headset (virtual reality, augmented reality, etc.), and / or any other type of device used by a consumer to access information streams.
[0056] In some embodiments, a client device 204 submits an image query 104 to an image gallery recommendation service 202 and receives one or more gallery images 206 in response. In some embodiments, the gallery images 206 are obtained from a gallery module 208. The gallery module 208 may be incorporated into or collaborate with the image gallery recommendation service 202. In some embodiments, the gallery module 208 retrieves one or more gallery images 206 from an image database 210. In some embodiments, the image database 210 includes a collection of images, and the gallery module 208 is configured to select a subset of the collection of images in response to the image query 104. In some embodiments, the images in the image database 210 are tagged or otherwise associated with metadata for retrieval. For example, an image of a horse may be tagged as "animal," "horse," or the like. In this way, the image gallery recommendation service 202 can provide gallery images 206 relevant to the image query 104. For example, in response to the image query 104 "show me paintings," the image gallery recommendation service 202 can retrieve images of various paintings from the image database 210.
[0057] In some embodiments, the client device 204 includes a user interface 212 configured to display an image gallery 214. In some embodiments, the client device 204 and / or the image gallery recommendation service 202 configures the image gallery 214 to graphically display the gallery images 206 within the user interface 212. In this manner, the image gallery recommendation service 202 can provide one or more recommended images to a user in the context of an image search.
[0058] In some embodiments, the client device 204 submits an image query 104 to the image gallery recommendation service 202 and receives, in response, a generated image 108 and one or more gallery images 206. In some embodiments, the client device 204 and / or the image gallery recommendation service 202 configures the image gallery 214 to graphically display the generated image 108 and the gallery images 206. In some embodiments, the image gallery 214 includes a dynamic image frame 216 that displays the generated image 108. In some embodiments, the user interface 212 and / or the dynamic image frame 216 includes an interactive component 218. In some embodiments, the interactive component 218 includes a user-interactive field and / or button in which the user can provide user input 110. In this way, the image gallery recommendation service 202 can provide a more dynamic image search experience, as described in more detail herein.
[0059] The user-interactive fields and / or buttons of the interactive component 218 are not meant to be particularly limited. In some embodiments, the interactive component 218 includes a text input field (see Figure 3). The text input field allows the user to enter user input 110 (e.g., text prompts or instructions) directly into the dynamic image frame 216. The user can enter keywords, descriptions, and / or specific requests to guide the image generation process of the controlled diffusion model 102. In some embodiments, the interactive component 218 includes one or more drop-down menus. The drop-down menus can provide a list of predefined options for the user to select from to fine-tune their image. For example, these options can include predetermined categories, styles, and / or other attributes that the user can select to fine-tune their prompt (image query 104). For example, the drop-down menu may include options such as "Desaturate" and "Restyle." In some embodiments, the interactive component 218 includes one or more checkboxes. The checkboxes allow the user to select one or more options from a predefined list and can be used to specify certain preferences, constraints, and / or features that should be included in the generated image 108. For example, the checkboxes may include an option to retain specific features / elements discovered by the topic identification and constraint graph 106 (e.g., checking this option to retain a person in the image background, etc.). In some embodiments, the interactive component 218 includes one or more sliders. A slider enables a user to adjust (multiple) values within a certain range by dragging the slider handle. A slider can be used to capture continuous preferences or numerical constraints, such as controlling the level of detail, intensity, and / or size of any feature in the generated image 108. For example, a slider can allow a user to dynamically scale (increase or decrease) an object (e.g., the moon, a building, etc.) in the generated image 108. In some embodiments, the interactive component 218 includes a color picker. The color picker allows a user to select a specific color by selecting a color from a palette and / or by selecting a color code, and can be useful for specifying color preferences or constraints for the generated image 108. In some embodiments, the interactive component 218 includes a drawing or sketching interface or canvas. The drawing or sketching interface enables a user to create or modify visual input within the dynamic image frame 216 using a pen, mouse, touch input, and / or stylus. For example, a user can draw object outlines, depict gestures, and provide other visual cues within the interactive component 218. In some embodiments, the interactive component 218 includes one or more rating buttons to allow the user to express their preferences or opinions numerically or on a relative scale. The rating buttons allow the user to rate specific attributes, such as the quality, style, and relevance of the generated output. For example, the interactive component 218 can include rating buttons for "I like this" and "I don't like this" to further fine-tune the generated image 108.
[0060] It is worth noting that the interactive component 218 can be configured to receive a variety of input types. In some embodiments, the interactive component 218 can receive one or more (even mixed) multimodal inputs, including but not limited to text data (e.g., natural language sentences, documents, transcribed text, etc.), image data (visual representations, drawings, sketches, photos, etc.), video data (e.g., continuous image frames with and without accompanying audio), audio data (e.g., recordings, music, speech, and other forms of audio signals, etc.), sensor data (e.g., data collected from sensors such as temperature sensors, accelerometers, GPS devices, environmental sensors, etc.), gestures (e.g., physical movements and gestures captured by devices such as motion sensors and depth cameras), metadata (e.g., , descriptive and / or contextual information associated with other modalities, such as timestamps, user demographics, location data, etc.), structured data (e.g., tabular data or other structured data formats, including numerical data, categorical variables, relational databases, etc.), emotional data (e.g., information related to emotional state, expression, and mood, which can be expressed and inferred through text, audio, tone of voice, facial expressions, etc.), biometric data (e.g., an individual's physical and physiological data, such as fingerprints, iris scans, heart rate, brainwave patterns, etc.), and social data (e.g., data related to social interactions, social networks, and social graphs, capturing connections, relationships, and communication patterns between individuals, etc.).
[0061] In some embodiments, the interactive component 218 can receive so-called "magic wand" input, which refers to a user interface technique that allows a user to graphically select or otherwise indicate the desirability of specific features and / or regions in an image. The "magic wand"-style user input can then be used to guide the fine-tuning or updating process of the controllable diffusion model 102 to emphasize and / or remove certain features in the generated image 108. In some embodiments, a user can use a brush tool or other type of selection tool within the dynamic image frame 216 to mark regions or features of an image (e.g., the initially generated image 108) that the user wants to enhance, remove, or otherwise alter. In effect, the marked regions or features act as guidance signals to the controllable diffusion model 102, indicating regions that should be emphasized or suppressed during the update process. Utilizing the "magic wand"-style technique in this manner, a user can more precisely control the output of the controllable diffusion model 102, thereby better customizing or personalizing the generated image 108 based on their unique preferences.
[0062] In some embodiments, interactive component 218 can receive a combination of multimodal inputs. For example, in some embodiments, interactive component 218 can receive textual data as well as "magic wand"-style selections. To illustrate, consider a generated drawing of a sailboat (or motorboat, etc.) on a river. In some embodiments, a user can use interactive component 218 to provide a magic wand to select a bank area near the river and textual input such as "add trees" or "add buildings." In some embodiments, controllable diffusion model 102 can use the input received via interactive component 218 to update generated image 108. Continuing with this example, the drawing of the sailboat can be modified to add (or remove) trees and buildings alongside the river. In another example, a user can use a magic wand to circle a feature (such as clouds in the drawing) and enter the text "zoom in" (or manipulate a size slider) to cause controllable diffusion model 102 to alter (zoom in) the clouds. Other combinations of multimodal inputs are possible (e.g., gesture data combined with audio data, textual data combined with biometric data, textual data, magic wand input combined with gestures, etc.), and all such configurations are within the contemplated scope of this disclosure.
[0063] For the sake of clarity, an exhaustive list of all possible interactions and configurations of the interactive component 218 is omitted. However, it should be understood that the examples provided are generally illustrative of the dynamic, generative image process between a user and a gallery of images provided by the interactive component 218 and the controllable diffusion model 102. Other configurations (user prompt type, selection of multimodal inputs, input ranges, number of nested fine-tuning inputs, types of interactive buttons, sliders, dialog boxes, etc.) can be implemented using the interactive component(s) 218, and all such configurations are within the contemplated scope of the present disclosure.
[0064] Figure 3 An example image gallery 214 is depicted according to one or more embodiments of the present invention. The image gallery 214 may be displayed in a user interface (e.g., Figure 2 is presented to the user in the user interface 212). Figure 3 As shown in FIG, the image gallery 214 may include the image query 104 (here, the string "Paintings"), one or more gallery images 206 (here, a collection of paintings), and a dynamic image frame 216. The image gallery 214 is shown to have a single dynamic image frame 216 and a specific number (here, 11) and arrangement of gallery images 206 for ease of discussion, but this is not meant to be particularly limiting. The image gallery 214 may include any number of dynamic image frames 216 and any desired arrangement of gallery images 206. In some embodiments, the image gallery recommendation service 202 (see FIG. 2 ) is used to select the gallery images 206 in response to the image query 104. Figure 2) retrieves and / or creates the gallery images 206 and the generated images 108. In some embodiments, the gallery images 206 are populated from a database (e.g., image database 210). In some embodiments, the generated images 108 are generated using the controlled diffusion model 102 (see Figure 1 and Figure 2 ) is created dynamically.
[0065] In some embodiments, the dynamic image frame 216 includes the generated image 108 (here, a sketch of a painting of a sailboat near a building) and an interactive component 218 (here, a selectable button with pre-generated text "Make me a painting"). The configuration of the dynamic image frame 216 and the interactive component 218 is shown for illustrative purposes only and may include any number of additional aspects or features described herein (e.g., sliders, canvas areas, check boxes, drop-down menus, etc.).
[0066] In some embodiments, the pre-generated text of the interactive component 218 can be generated from the image query 104. In some embodiments, the image query 104 can be provided to the image gallery recommendation service 202 having a topic identification and constraint graph 106, which is configured to identify topics of interest within the image query 104 (e.g., using the NLP described above, etc.). In some embodiments, the pre-generated text includes the identified topics. For example, for an image query 104 for "Paintings", the interactive component 218 may include the pre-generated text "Make me a painting" (as shown).
[0067] Figure 4 Depicts the interactive component 218 having pre-generated text "Draw Me a Picture" after a user clicks or otherwise selects it, according to one or more embodiments of the present invention. Figure 3 The example image gallery 214. Figure 4 As shown in FIG, the generated image 108 in dynamic image frame 216 has been changed to a new, dynamically generated painting. In some embodiments, the new, dynamically generated painting includes a fuller image based on the earlier painting sketch. Note that the painting still includes the sailboat near the building, but additional details, textures, and elements have been added to provide a more complete appearance, making it more similar to the appearance of an actual painting.
[0068] In some embodiments, the pre-generated text of the interactive component 218 (see Figure 3) can be overwritten by user input 110. For example, a user may enter the string "Style like Sistine Chapel by Michelangelo" into the interactive component 218 (e.g., via a text message field).
[0069] Figure 5 Depicts the following after a user enters user input 110 into interactive component 218 according to one or more embodiments of the present invention. Figure 4 The example image gallery 214. Figure 5 As shown in FIG, the generated image 108 in the dynamic image frame 216 has been changed again and is now a new version of the painting recreated in the style of Michelangelo. In some embodiments, the input 110 is consistent with the image query 104 (see Figure 1 and Figure 2 ) is provided in conjunction with the controllable diffusion model 102. In some embodiments, in response to receiving the input 110 via the interactive component 218, the controllable diffusion model 102 creates an updated generated image 108. Figure 5 As further shown in FIG, user input 110 has been changed (now to the string "Change to Van Gogh style").
[0070] Figure 6 Depicts the following after a user enters new user input 110 into interactive component 218 according to one or more embodiments of the present invention. Figure 5 The example image gallery 214. Figure 6 As shown in FIG, the generated image 108 in the dynamic image frame 216 has been changed again and is now a new version of the painting recreated in the style of Van Gogh (here borrowing the style of the oil on canvas painting "The Starry Night").
[0071] like Figure 6As further shown in FIG. 2 , user input 110 has been altered (now to the string "Make it bigger"). Furthermore, dynamic image frame 216 now includes a magic wand input 602 and / or other types of graphical tools for selecting or otherwise indicating specific features and / or areas within the image, as previously described herein. In some embodiments, magic wand input 602 and user input 110 (which may include magic wand input 602) work together to fully define the user's intent. For example, the phrase "make it bigger" is ambiguous by itself, making it unclear which feature(s) of generated image 108 is being referred to. However, combined with the graphical selection of the moon feature in the upper right corner of generated image 108 via magic wand input 602, the intended feature is clearly visible. The collaboration between the multimodal inputs within interactive component 218 is not meant to be particularly limited, and other combinations are possible. For example, interactive component 218 may include a slider that, when combined with the selection of the moon via magic wand input 602, can gradually and continuously increase or decrease the size of the moon (not separately shown).
[0072] Figure 7 Depicts the interactive component 218 after the user enters new user input 110 and magic wand input 602 according to one or more embodiments of the present invention. Figure 6 The example image gallery 214. Figure 7 As shown in , the generated image 108 in the dynamic image frame 216 has been slightly adjusted. Note that although the painting remains essentially the same (still in the style of Van Gogh's "Starry Night"), the moon in the upper right corner has become larger.
[0073] like Figure 7 As further shown in FIG, the user input 110 has been replaced with a new pre-generated text (here, "More like this"). In some embodiments, due to the number of consecutive inputs and selections by the user, the image gallery recommendation service 202 can infer that the corresponding user is particularly interested in Figure 7 The resulting fine-tuned generated image 108 is of interest. In this way, the image gallery recommendation service 202 can adjust the pre-generated text due to ongoing user-system interaction. Although the "more similar content" interaction component 218 is omitted for clarity, selecting the "more similar content" interaction component 218 causes the dynamic image frame 216 and / or any gallery image 206 to change to additional paintings created in the fine-tuned style of the generated image 108. For example, the new images may include various versions of the generated image 108 with different moon sizes.
[0074] Figure 8An example image gallery 214 is depicted according to one or more embodiments of the present invention. The image gallery 214 may be in the form of Figures 3 to 7 The discussion is similar to the way in which the user interface (e.g. Figure 2 is presented to the user in the user interface 212 in FIG. Figure 3 Unlike the image gallery 214 shown in , the image query 104 is empty. This situation may occur, for example, when a user first accesses the image gallery recommendation service 202 (e.g., via the “IMAGES” icon below the image query 104). This situation may also occur during interaction with the image gallery recommendation service 202.
[0075] In any case, please note that the image gallery 214 can still include one or more gallery images 206 (here referring to a collection of various images such as birds, cities, works of art) and one or more dynamic image frames 216. For example, one or more dynamic image frames 216 include a sketch of a cat and a sketch of a room. In the event that the image gallery recommendation service 202 is unable to utilize the image query 104, it can still use the available information related to the user to make image recommendations. In some embodiments, the user can be identified via a user session identifier (ID), a device ID, an account ID, etc. Once the user is identified, the gallery images 206 are populated from a database (e.g., the image database 210) based on known and / or inferred information related to the user.
[0076] Note that the term "identified user" as used herein does not necessarily mean a strict identification of an individual, but rather can mean a relative identification of various characteristics of a user (i.e., non-personally identifiable information) that can indicate the types of images that the user may be interested in. For example, in some embodiments, based on the user's available information, identifying information, and / or non-identifying information, the image gallery recommendation service 202 can recommend one or more gallery images 206 and one or more generated images 108. For example, the user information can include the user's search history, the user's previously indicated preferences (i.e., prior image selections and prompts), the user's location / country / region (e.g., inferred via metrics related to the client device 204 and / or the network used to access the image gallery recommendation service 202), the user's preferred language, and / or various other usage metrics (types of images that the user typically saves, likes, shares, etc.).
[0077] In some embodiments, the image gallery recommendation service 202 can score each image from the pool of available images of the initial group of the image gallery 214. In some embodiments, the image gallery recommendation service 202 can select any number of the highest-scoring images for the gallery images 206. Images can be scored according to any predetermined criteria, such as by a matching metric (e.g., distance measurement) with one or more characteristics of the user. It is known to use a distance metric to define the match between object features, and any suitable method (e.g., Euclidean distance, Tanimoto distance, Jaccard similarity coefficient, etc.) can be used. For example, for a user known to prefer animals, the image gallery recommendation service 202 may give a high score to bird images and a low score to construction site images. Other predetermined criteria are also possible, such as scoring images in whole or in part according to business indicators. For example, the image gallery recommendation service 202 may give a high score to images associated with advertising partners and / or images with relatively high display revenue, and give a low score to images associated with market competitors and / or images with relatively low display revenue.
[0078] Figure 9 Depicts the following after a user selects the "Draw me a modern room" interactive component 218 according to one or more embodiments of the present invention. Figure 8 The example image gallery 214. Figure 9 , the generated image 108 in the dynamic image box 216 in the upper right corner has been changed to a new, dynamically generated drawing of a modern room. In some embodiments, the new, dynamically generated drawing of the modern room includes a more fleshed-out image based on the earlier drawing sketch. Note that the new drawing of the modern room still includes the bookshelf, bed, desk, and chair, but additional details, textures, and elements have been added to provide a more complete appearance that more closely resembles the appearance of the actual room.
[0079] Additionally, note that the generated image 108 in the dynamic image box 216 in the lower left corner has also changed. This generated image 108 now shows an island scene with a sailboat and a sunset (the sketch of the cat has been replaced). In some embodiments, the image gallery recommendation service 202 can infer the user's preference for generated animal images ( Figure 8 ) is not interested in the other option presented in the example, and instead decides to interact with the interactive component 218 having a sketch of a modern room.
[0080] In some embodiments, the image gallery recommendation service 202 may change one or more generated images 108 in one or more dynamic image frames 216 to a new sketch with a new prompt (here, "draw me a beach with a palm tree"). In some embodiments, the new prompt(s) / image(s) may include the second-best guess (e.g., the second-highest score) of the image(s) that the user may be interested in. In some embodiments, the image gallery recommendation service 202 may update any estimate / score based on user interaction. For example, the image gallery recommendation service 202 may lower the score of an image in which an animal is in the user's field of view and ignore that option in a previous interaction. Note that in some cases, the size (height, width) of the new generated image 108 of the dynamic image frame 216 may be different from the previous image. In some embodiments, the image gallery recommendation service 202 may dynamically adjust the size of the dynamic image frame 216 to accommodate the size change of the corresponding generated image 108 (as shown).
[0081] Figure 10 Various aspects of an embodiment of a computer system 1000 are illustrated, which can perform various aspects of the embodiments described herein. In some embodiments, (multiple) computer systems 1000 can implement any of the workflows, processes, systems, and services (e.g., image gallery recommendation service 202) previously described herein, and / or be otherwise incorporated into or combined with any of the items. In some embodiments, computer system 1000 can be implemented on the client side. For example, computer system 1000 can be configured to display and execute the functionality of user interface 212. In some embodiments, computer system 1000 can be implemented on the server side. For example, computer system 1000 can be configured to receive image queries 104 and / or user input 110, and provide generated images 108 and / or gallery images 206 in response.
[0082] The computer system 1000 includes at least one processing device 1002, which typically includes one or more processors or processing units for performing a variety of functions, such as for example, for completing the previously described Figures 1 to 9Any part of the dynamic image search workflow described. Components of the computer system 1000 also include a system memory 1004 and a bus 1006 that couples various system components, including the system memory 1004, to the processing device 1002. The system memory 1004 can include a variety of computer system readable media. Such media can be any available media accessible to the processing device 1002, including volatile and non-volatile media, removable and non-removable media. For example, the system memory 1004 includes non-volatile memory 1008 (such as a hard drive) and may also include volatile memory 1010 (such as random access memory (RAM) and / or cache). The computer system 1000 may also include other removable / non-removable, volatile / non-volatile computer system storage media.
[0083] The system memory 1004 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the embodiments described herein. For example, the system memory 1004 stores various program modules that typically perform the functions and / or methods of the embodiments described herein. One or more modules 1012, 1014 may be included to perform functions related to the dynamic image search previously described herein. The computer system 1000 is not limited in this regard, as other modules may be included depending on the functionality required for the computer system 1000. The term "module" as used herein refers to a processing circuit system that may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group), and memory that executes one or more software or firmware programs, combinational logic circuits, and / or other suitable components that provide the functionality described.
[0084] The processing device 1002 may also be configured to communicate with one or more external devices 1016, such as, for example, a keyboard, a pointing device, and / or any device (e.g., a network card, a modem, etc.) that enables the processing device 1002 to communicate with one or more other computing devices (e.g., the client device 204 and / or the image gallery recommendation service 202). Communication with various devices may occur via input / output (I / O) interfaces 1018 and 1020.
[0085] The processing device 1002 may also communicate with one or more networks 1022, such as a local area network (LAN), a wide area network (WAN), a bus network, and / or a public network (e.g., the Internet), via a network adapter 1024. In some embodiments, the network adapter 1024 is or includes a fiber optic network adapter for communicating over a fiber optic network. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with the computer system 1000. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, and data archiving storage systems.
[0086] Now see Figure 11 , according to an embodiment, a flowchart 1100 is generally shown for performing dynamic image search using a controllable diffusion model in an image gallery recommendation service. Figures 1 to 10 A flowchart 1100 is described and may be included in Figure 11 Additional steps not described in . Although Figure 11 The blocks depicted in the drawings are shown in a particular order, but in some embodiments, the blocks may be rearranged, subdivided, and / or combined.
[0087] At block 1102, the method includes displaying an image gallery having a plurality of gallery images. The image gallery also includes a dynamic image frame having generated images and an interactive component.
[0088] At block 1104 , the method includes receiving user input in an interactive component.
[0089] At block 1106 , the method includes, in response to receiving the user input, generating an updated generated image by inputting the user input into the controllable diffusion model.
[0090] At block 1108 , the method includes replacing the generated image in the dynamic image frame with the updated generated image.
[0091] In some embodiments, the method includes receiving an image query in an information column of an image gallery. In some embodiments, the generated image is generated by inputting the image query into a controlled diffusion model. In some embodiments, the plurality of gallery images and the generated image are selected based on a degree of matching with one or more features in the image query.
[0092] In some embodiments, the method includes determining one or more constraints in the image query. In some embodiments, the generated image is generated by inputting the one or more constraints into a controlled diffusion model.
[0093] In some embodiments, the one or more constraints in the image query include at least one of a pose skeleton and an object boundary. In some embodiments, determining the one or more constraints includes extracting the object boundary when the feature in the image query is one of a structure and a geological feature. In some embodiments, determining the one or more constraints includes extracting the pose skeleton when the feature in the image query is one of a person and an animal.
[0094] In some embodiments, the plurality of gallery images are obtained from an image database. In some embodiments, one or more of the plurality of gallery images are not generated images (i.e., are obtained images rather than images dynamically generated from a diffusion model). In some embodiments, one or more of the plurality of gallery images are previously generated images.
[0095] In some embodiments, the interactive component includes a text information field. In some embodiments, receiving user input in the interactive component includes receiving a text string input into the text information field.
[0096] In some embodiments, the interactive component includes one or more of a drop-down menu, a checkbox, a slider, a color picker, a canvas interface for drawing or sketching, and a rating button.
[0097] In some embodiments, the interactive component includes a canvas for magic wand input. In some embodiments, receiving user input in the interactive component includes receiving a magic wand input that graphically selects one of a specific feature or a specific area in the generated image. In some embodiments, the interactive component also includes a text information field. In some embodiments, receiving user input in the interactive component also includes receiving a text string input having contextual information for the magic wand input.
[0098] Although the present disclosure has been described with reference to various embodiments, it will be understood by those skilled in the art that changes may be made thereto and that elements thereof may be replaced with equivalents without departing from the scope thereof. The various tasks and process steps described herein may be incorporated into more comprehensive programs or processes having additional steps or functionality not described in detail herein. Additionally, various modifications may be made to adapt specific circumstances or materials to the teachings of the present disclosure without departing from the basic scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed, but encompasses all embodiments within its scope.
[0099] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0100] Various embodiments of the present invention are described herein with reference to the accompanying drawings. The drawings described herein are for illustrative purposes only. Various modifications may be made to the schematic diagrams and / or steps (or operations) described therein without departing from the spirit of the present disclosure. For example, the actions may be performed in a different order, or operations may be added, deleted, or modified. All such modifications are considered part of this disclosure.
[0101] The terms used herein are used only to describe specific embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" used herein also include the plural forms. It should also be understood that the terms "include" and / or "comprising" used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. Unless the context clearly indicates otherwise, the term "or" means "and / or".
[0102] The terms "received from," "received from," "delivered to," and "delivered to" describe a communication path between two elements and, unless otherwise specified, do not imply a direct connection between the elements without intervening elements / connections. The corresponding communication path may be a direct or indirect communication path.
[0103] The corresponding structures, materials, acts, and their equivalents of all means or step plus function elements in the claims are intended to include any structure, material, or act for performing the function in combination with other claim elements in the particular claim.
[0104] For the sake of brevity, this document may or may not describe in detail conventional techniques related to making or using various aspects of the present invention. In particular, various aspects of computing systems and specific computer programs for implementing the various technical features described herein are well known. Therefore, for the sake of brevity, many conventional implementation details are only briefly mentioned or omitted entirely, and well-known system and / or process details are not provided.
[0105] The present invention may be a system, method, and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or medium) having computer-readable program instructions thereon for causing a processor to perform various aspects of the present invention.
[0106] Various embodiments are described herein with reference to the flowcharts and / or block diagrams of methods, devices (systems) and computer program products. It should be understood that each block in the flowcharts and / or block diagrams and combinations of blocks in the flowcharts and / or block diagrams can be implemented by computer-readable program instructions.
[0107] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create components for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can instruct the computer, programmable data processing apparatus, and / or other device to operate in a specific manner, such that the computer-readable storage medium having the instructions stored therein comprises an article of manufacture including instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0108] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device, thereby producing a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.
[0109] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functionality and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, an instruction segment or an instruction portion, which includes one or more executable instructions for realizing the specified (multiple) logical functions. In some alternative implementations, the functions noted in the box may not be performed in the order shown in the accompanying drawings. For example, the two boxes shown in succession can actually be executed substantially at the same time, or can sometimes be executed in the opposite order, which specifically depends on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, as well as the box combination in the block diagram and / or flow chart, can be realized by a dedicated hardware system that performs the specified function or action or performs a combination of dedicated hardware and computer instructions.
[0110] The description of the various embodiments described herein is for illustrative purposes only and is not intended to be exhaustive or limited to the disclosed (multiple) forms. These embodiments are selected and described to best explain the principles of the present disclosure. Many modifications and variations are apparent to those of ordinary skill in the art without departing from the scope and spirit of the embodiments. The terms used herein are intended to best explain the principles of the various embodiments, practical applications, or technical improvements to existing technologies on the market, or to enable other persons of ordinary skill in the art to understand the embodiments described herein.
Claims
1. A method comprising: displaying an image gallery 214 comprising a plurality of gallery images 206 , the image gallery 214 also comprising a dynamic image frame 216 comprising the generated image 108 and an interactive component 218 ; receiving user input 110 in the interactive component 218; In response to receiving the user input 110 , generating an updated generated image 108 by inputting the user input 110 into the controllable diffusion model 102 ; as well as The generated image 108 in the dynamic image frame 216 is replaced with the updated generated image 108 .
2. The method according to claim 1, further comprising: An image query is received in an information column of the image gallery. 3 . The method of claim 2 , wherein the generated image is generated by inputting the image query into the controlled diffusion model. 4 . The method of claim 2 , wherein the plurality of gallery images and the generated image are selected based on a degree of matching with one or more features in the image query.
5. The method according to claim 2, further comprising: One or more constraints in the image query are determined. 6 . The method of claim 5 , wherein the generated image is generated by inputting the one or more constraints into the controlled diffusion model. The method of claim 5 , wherein the one or more constraints in the image query include at least one of a pose skeleton and an object boundary.
8. The method of claim 7, wherein determining the one or more constraints comprises: When the feature in the image query includes one of a structural feature and a geological feature, the object boundary is extracted.
9. The method of claim 7, wherein determining the one or more constraints comprises: When the feature in the image query includes one of a person and an animal, the posture skeleton is extracted.
10. The method of claim 1, wherein the plurality of gallery images are from an image database.
11. The method of claim 1 , wherein the interactive component comprises a text information field, and receiving the user input in the interactive component comprises: A text string input into the text information field is received.
12. The method of claim 1, wherein the interactive component comprises one or more of: a drop-down menu, a checkbox, a slider, a color picker, a canvas interface for drawing or sketching, and a rating button.
13. The method of claim 1 , wherein the interactive component comprises a canvas for magic wand input, and receiving the user input in the interactive component comprises receiving a magic wand input that graphically selects one of a specific feature or a specific area in the generated image.
14. The method of claim 13, wherein the interactive component further comprises a text information field, and receiving the user input from the interactive component further comprises: A text string input is received, the text string input having context information for the magic wand input.
15. A system 202 comprising a memory 1104, computer-readable instructions, and one or more processors 1002 configured to execute the computer-readable instructions, wherein the computer-readable instructions control the one or more processors 1002 to perform operations, the operations comprising: receiving an image query 104 from a client device 204 communicatively coupled to the system 202; providing a plurality of gallery images 206 and a generated image 108 to the client device 204 based on a degree of matching with one or more features in the image query 104; receiving user input 110 from the client device 204; generating an updated generated image 108 in response to the user input 110 by inputting the user input 110 into the controllable diffusion model 102 ; as well as The updated generated image 108 is provided to the client device 204 .
16. The system of claim 15, wherein the generated image is generated by inputting the image query into the controlled diffusion model.
17. The system of claim 15, further comprising: One or more constraints in the image query are determined.
18. The system of claim 17, wherein the generated image is generated by inputting the one or more constraints into the controlled diffusion model.
19. A system 204 comprising a memory 1104, computer-readable instructions, and one or more processors 1002 configured to execute the computer-readable instructions, wherein the computer-readable instructions control the one or more processors 1002 to perform operations comprising: receiving a plurality of gallery images 206 and the generated image 108 from an image gallery recommendation service 202 communicatively coupled to the system 204 ; displaying an image gallery 214 including the plurality of gallery images 206 and a dynamic image frame 216 including the generated image 108 and an interactive component 218; receiving user input 110 in the interactive component 218; Transmitting the user input 110 to the image gallery recommendation service 202; receiving an updated generated image 108 from the image gallery recommendation service 202; as well as The generated image 108 in the dynamic image frame 216 is replaced with the updated generated image 108 .
20. The system of claim 19, wherein the interactive component comprises a text information field, and receiving the user input in the interactive component comprises: A text string input into the text information field is received.