Method and system for item recommendation and searching for image generation using generative ai classification

By combining image generation based on classification system with image-based search, photo-level realistic images are generated as the basis for search queries, the problem of insufficient quality and personalization of search results for visual or context-based complex projects in the prior art is solved, and higher quality and personalized search results are achieved.

CN120179838APending Publication Date: 2025-06-20EBAY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411894043.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing search engines have difficulty delivering high-quality and personalized search results when dealing with visual or contextually complex project searches.

Method used

By combining image generation based on classification system with image-based search, photo-level realistic images are generated as the basis for search queries, thereby improving the quality and personalization of search results.

Benefits of technology

Enhanced search results quality and personalization, and capture common attributes of categories by generating more robust images, improving search engines' understanding and response to user intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179838A_ABST
    Figure CN120179838A_ABST
Patent Text Reader

Abstract

And image generation based on the classification system is used for item search, so that the quality and individuation of search results are improved. Previously interacted items are classified into a category classification system. The generative AI model may be used to classify previously interacted items by generating categories or assigning to existing categories. A set of previously interacted items is selected from one of the categories and provided to an image model that generates a photo-level realistic image in response. A photo-level realistic image includes a rendering of an item (i.e., a representation of the item being rendered). An item search using a search engine may be performed based on the generated photo-level realistic image. For example, a photo-level realistic image or portion thereof may be provided as a search query that uses an image-based search or is described to perform a text-based search. Search results are identified for a search query and provided as a response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of search. Background Art

[0002] A web search engine is a software system designed to search for information on the web. Search engines use algorithms to determine the relevance and ranking of items in response to a search query. The main function of these engines is to identify and retrieve items as search results based on keywords, phrases, or images provided by the search query, and the items can include web pages, images, videos, goods and services for sale, and other types of digital content. Over time, the technology behind search engines has evolved, incorporating advanced features such as natural language processing, machine learning, and personalized search capabilities to improve the accuracy and relevance of the search results provided to users. Summary of the Invention

[0003] At a high level, the present technology relates to classification-system-based image generation for item search. Combining classification-system-based image generation with image-based search can improve the quality and personalization of search results that can be identified and returned in response to a search query.

[0004] An example method classifies previously interacted items into categories in a category classification system. The previously interacted items can be items that have received a certain interaction by the user. The previously interacted items can be classified using a generative AI (artificial intelligence) model.

[0005] After classifying the previously interacted items, a set of the previously interacted items can be selected from one of the categories. The set of the previously interacted items from the category is provided to an image model, which generates a photo-realistic image as a response. The photo-realistic image includes item renderings (i.e., illustrations of the items). For example, if the image model generates a lady wearing a dress and holding a handbag, the dress and the handbag can be considered renderings of the items with the photo-realistic image. In one aspect, the photo-realistic image can be provided as a recommendation for items similar to the item renderings.

[0006] In one aspect, an item search using a search engine can be performed based on the generated photo-realistic image. For example, the photo-realistic image or a portion thereof can be provided as a search query for image-based search. Search results are identified for the search query and provided as a response, thus enhancing the way in which database items can be explored through item search.

[0007] In addition, compared with other methods, by using a method based on a classification system, the images generated by an image model can have enhanced searchable quality because the inputs used to generate the images are classified into the same categories.

[0008] This Summary is intended to introduce a selection of concepts in a simplified form that will be further described in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. Additional objectives, advantages, and novel features of the technology will be set forth in part in the following description, and in part will become apparent to those skilled in the art upon examination of the present disclosure, or may be learned by practice of the technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The present technology will be described in detail below with reference to the drawings, in which:

[0010] Figure 1 illustrates an example operating environment in which aspects of the present technology can be used;

[0011] Figure 2 illustrates an example classification of previously interacted items according to aspects described herein;

[0012] Figure 3 illustrates an example ranking of categories for a category classification system according to aspects described herein;

[0013] Figure 4 illustrates a selection of a set of previously interacted items according to aspects described herein;

[0014] Figure 5 illustrates an example generation of a photo-realistic image according to aspects described herein;

[0015] Figure 6 illustrates an example segmentation of a photo-realistic image of Figure 5 according to aspects described herein;

[0016] Figure 7 illustrates an example search query for a photo-realistic image based on Figure 5 according to aspects described herein;

[0017] Figure 8 illustrates an example item search using the search query according to aspects described herein;

[0018] Figure 9 illustrates a flowchart of an example item search method according to aspects described herein; and

[0019] Figure 10An example computing device suitable for implementing aspects of the present technology in accordance with aspects described herein is shown. Detailed Description

[0020] Text-based Internet search refers to the process of using a search engine to find information on a network (e.g., the Internet) by entering text. In this method, a user enters a series of words or phrases into a search engine, and the search engine returns a list of items that may include web pages, documents, images, items for sale or services, videos, music, or other types of digital files that are considered relevant to the query.

[0021] Traditionally, search engines process text-based queries to understand their intent. This can involve: parsing the query, correcting spelling mistakes, and sometimes using synonyms or related terms to expand the query (i.e., a process known as query expansion). Many search engines maintain an extensive index of web pages and other online content. The processed query is used to search the index to find matching results.

[0022] In many cases, search engines use ranking algorithms to rank the results based on various factors (such as relevance, page quality, and the number of inbound links, etc.). A list of ranked results is presented to the user, typically with a title, a summary of the content, and the URL (Uniform Resource Locator) of the landing page related to the item. Then, the user can click on the item or link to access the web page and access the information they are searching for.

[0023] Text-based Internet search has changed significantly since its inception. Early search engines mainly used keyword matching and had a less sophisticated understanding of the context or semantics of queries. Modern search engines use complex algorithms that combine machine learning, natural language processing, and other advanced technologies to provide more accurate and contextually relevant results.

[0024] Image-based search (sometimes referred to as "reverse image search") is a search that uses an image as the search query. In this method, a user uploads an image to a search engine, and the search engine then analyzes the image and returns a list of items that may include similar images or related information. The search engine can use various techniques (such as feature extraction, color histograms, and machine learning algorithms) to identify patterns, shapes, and other characteristics within the image to further optimize the search results.

[0025] For example, some conventional search engines process images to extract key features such as color distribution, texture, and shape. The extracted features are used to search a database of indexed images to find matches. Similar to text-based search, the search engine ranks the results based on a similarity metric and returns the ranked results to the user's computing device.

[0026] Image-based search is particularly useful when users cannot adequately describe in text what they are looking for. For example, it is often easier to use an image to identify a landmark or a work of art than to use a text description. Image-based search can be beneficial for users who are looking for a product but do not know the exact name or brand. They can simply upload a picture of the item to find similar products.

[0027] In addition, image-based search provides a rich contextual framework for queries. A single image can encompass multiple elements (such as color, texture, lighting, etc.) that would otherwise require a large amount of text description from the user. This visual complexity allows for more precise and nuanced search results.

[0028] Additionally, images can resolve the ambiguity that often exists in text queries provided by users. For example, an image of an apple immediately clarifies whether the user is referring to the fruit or the technology company. The multi-faceted information contained in an image (which can include multiple objects or concepts related in a specific way) can provide a comprehensive understanding of the user's intent. Moreover, images capture non-verbal elements (such as emotion, style, and atmosphere) that are difficult to convey through text or would at least require the user to enter a large and lengthy text query to attempt to capture these elements. Finally, the universality of images transcends language barriers, making them particularly useful in global contexts where text-based keywords may not translate effectively.

[0029] One limitation of image-based search is the need for relevant images to act on the query. This can be a problem in various scenarios (such as when you encounter a product you like in a physical store but cannot take a picture to search online). Similarly, if you remember an image or a scene but do not have a copy, text-based search may be the only option. Additionally, technical limitations may hinder the practicality of image-based search; for example, limited device capabilities or poor internet connectivity may prevent users from uploading images, thus restricting them to text-based queries.

[0030] Search technologies, including image-based search, are an integral part of the technological functionality of the Internet, mainly because of the vast amount and variety of content available online. The Internet hosts billions of web pages, images, videos, and other forms of data. Without sophisticated search and ranking algorithms, it would be almost impossible for users to find relevant information in this vast ocean of content. As an example, at the time of filing this application, a search query for "City of Fountains" returned over 92.6 million results. This demonstrates how important search and ranking technologies are for the operation of the Internet, as users cannot sift through each of these results and must rely on search engines and their ability to identify and rank search results.

[0031] As the Internet continues to evolve, the complexity and diversity of queries are also increasing. Simple text-based search queries may not be sufficient to meet the search needs for all types of items, especially those with complex visual or contextual elements. This is where specialized search technologies, such as image-based search, come into play, providing alternative ways to navigate the digital landscape and allowing search engines to provide useful results from among trillions of possibilities.

[0032] In essence, search technologies are not just a useful feature but a prerequisite for the Internet to function as a useful resource. They serve as organizing principles that make the Internet accessible and navigable, transforming vast amounts of data into a structured and user-accessible environment. Improvements in search technologies will necessarily enhance Internet functionality.

[0033] The solution provided by the technology disclosed herein improves web search technology by using a taxonomy for image generation of images that can serve as the basis for search queries.

[0034] An example method that can capture some of these benefits uses categorized previously interacted-with items to generate photo-realistic images that can be used as the basis for search queries. Previously interacted-with items include items that a user has interacted with. For example, a user can interact with an item by clicking on the item, hovering an input indicator over the item for a specific duration, watching a video or a portion of a video for a specific duration, providing an emotional indication (e.g., liking or disliking the item), or other means of indicating the user's attention to a specific item.

[0035] These interacted-with items can be classified into a category classification system that includes an organizational classification system for grouping items based on some shared attributes. Examples include category classification systems based on clothing styles. For example, various clothing items can be classified based on the style they embody.

[0036] After classifying the previously interacted items, some or all of the previously interacted items can be selected from a specific category. In other words, the set of previously interacted items includes a selection of one or more interacted items among the interacted items in one of the categories.

[0037] The selected previously interacted items in the set of the previously interacted items can be provided to an image model. For example, the title, image, text, or other parts of each previously interacted item can be provided to the image model, and the image model generates an image as a response. Since the set of previously interacted items provided to the model is all included in one category, the image represents that specific category and generally captures the common attributes of that category. A diffusion model can be used to generate the image, resulting in a photo-realistic image including the rendering of the item. The item rendering includes an image rendered to have features corresponding to an actual item recognizable by a search engine. The photo-realistic image can be provided for display at a user computing device as a recommendation for an actual item similar to the rendered item. In some cases, the photo-realistic image is used as a basis for item search to identify items similar to the item rendering. In some aspects, the photo-realistic image is segmented to identify and separate the item rendering. Then, the user can select the item rendering to perform a search for an item visually similar to the separated item rendering.

[0038] To perform the search, a search query is determined from the photo-realistic image. The search query can include the photo-realistic image or a portion thereof. In one aspect, the search query includes one or more separated item renderings. The search engine can use any of these image-based search queries to perform an image-based search. In another aspect, a text description of the photo-realistic image or the item rendering therein is generated and used for the search.

[0039] Advantageously, this technique, as well as other techniques to be further described, helps to address many problems inherent in conventional text-based search methods and image-based search methods. For example, by using categories to generate images, more robust images can be generated that are more likely to capture the attributes of the category. In this way, the various attributes common to different categories can be expressed in the various photo-realistic images generated by the image model and used for search. Thus, the photo-realistic images used as the basis for search queries will be more likely to capture the attributes that the user wants to find in the actual items returned as search results. This, in turn, enables the search engine to more accurately identify and return search results. Additionally, the generated photo-realistic images can include enhanced recommendations, which can be a mix of items listed within different groups by the platform. This may be because the photo-realistic images are generated from categories determined according to a generative AI method that can place items from previous interactions in different groupings into the same category. In this way, searches using the generated photo-realistic images can identify actual items with similar category characteristics within the various item groupings in which the items defined by the platform are included.

[0040] The techniques described herein also help to provide visual recommendations that may not otherwise be apparent if a non-taxonomic image generation scheme were used. In other words, aspects of the present technique generate photo-realistic images based on a classification system that can be provided for display at a computing device to recommend or identify items similar to a project rendering. In addition to using the generated photo-realistic images as the basis for item search, as a supplement or alternative, the generated photo-realistic images can also be provided as recommendations.

[0041] Additionally, an image generation taxonomy for search engine searches helps to explore areas of a database that might otherwise be inaccessible. In other words, given that some databases contain billions of potential return values, these methods can help identify and return items that might be missed using conventional searches. For example, by using an AI model to generate photo-realistic images, the model attempts to predict objects in the image that naturally conform to the prompts it is provided. In this case, the model can generate items that naturally conform to the categories in which items from previous interactions were classified. As such, it can predict items relevant to the input that might otherwise be missed using traditional query expansion techniques, and thus return search results related to the predicted items that might otherwise be missed using conventional text-based search queries based on user input. Even in cases where the search query determined from the photo-realistic image is a textual description of the image or item rendering, the textual description generated by describing the rendered image is more robust and describes predicted items that would not otherwise be entered or predicted by a standard search query based on user input.

[0042] Furthermore, the techniques provided herein can help improve aspects of the computing device itself. For example, search queries generated from photo-realistic images are generally more comprehensive and robust in nature, describing in more detail what the user is looking to search for than what is provided in traditional search query inputs. This allows for a more in-depth and targeted exploration of the search index when retrieving items. In essence, more accurate search results can be identified and returned, thus reducing the number of search results transmitted over the network in response to a query. Additionally, since fewer search results are transmitted, the response time experienced is shorter, which allows for a faster processing speed of the query queue and thus can reduce overall system latency. Moreover, as previously mentioned, there are cases where the computing device has low network connectivity, which reduces the likelihood that a user can upload an image for searching. However, in aspects of the present technology, search images can simply be generated from previous item interactions. This can allow for backend image generation that produces images suitable for item searching without the user having to upload an image from their own device.

[0043] Moreover, many of the techniques described are not well-known, routine, or conventional in the relevant art. For example, the process of using a taxonomy for image generation to generate photo-realistic images, where items from previous interactions are classified and used to generate images corresponding to the categories, and the images can be used as a basis for search queries, is not thought to be a process readily used by traditional search engines.

[0044] It will be recognized that the methods described previously are merely examples that can be practiced in accordance with the following description and are provided to more easily understand the present technology and appreciate its benefits. Additional examples are now described with reference to the accompanying drawings.

[0045] Now referring Figure 1 , an example operating environment 100 is provided in which aspects of the present technology can be used. The operating environment 100 includes a server 102, a computing device 104, and a database 106 that communicate with a search engine 110 via a network 108, as well as other components or engines not shown.

[0046] The database 106 generally stores information, including data, computer instructions (e.g., software program instructions, routines, or services), or models used in embodiments of the described technology. For example, the database 106 can store computer instructions for implementing functional aspects of the search engine 110. Although described as a single database component, the database 106 can be embodied as one or more databases or can be embodied in the cloud.

[0047] The network 108 can include one or more networks (e.g., a public network or a virtual private network [VPN]), as shown by the network 108. The network 108 can include, but is not limited to, one or more local area networks (LANs), wide area networks (WANs), or any other communication network or method.

[0048] Generally, the server 102 is a computing device that implements functional aspects of the operating environment 100, such as one or more functions of the search engine 110 for facilitating item searches from photo-realistic images generated according to a category classification system. A suitable example of a computing device that can serve as the server 102 is described as the computing device 1000 with respect to Figure 10 . In an implementation, the server 102 represents a backend or server-side device.

[0049] The computing device 104 is generally a computing device that can be used to initiate item searches. For example, a user can use functional aspects of the search engine 110 initiated at the computing device 104 to receive and display search results for item searches. For example, the computing device 104 can receive input corresponding to an interaction with an item from an input component. In other words, the input at the computing device can indicate an interaction with an item, thus identifying which items are stored or otherwise indicated as items corresponding to previous interactions with the user (e.g., a user account, the computing device the user is using, or an address, etc.).

[0050] Interactions with an item include measurable engagement with the item. For example, this can include direct selection of varying degrees of passive engagement actions, such as clicking to access item details, hovering over an item for a threshold duration, watching a video or listening to music for a specified time, opening a document, or other similar interactions with the item based on the type of the item. In one aspect, previously interacted-with items are items that have received an interaction during a defined threshold time period (e.g., the past 90 days, 180 days, etc.). Previously interacted-with items can include items that have received an interaction during an indefinite time period or during any defined threshold period.

[0051] Like Figure 1 the other components of, computing device 104 is intended to represent one or more computing devices. A suitable example of a computing device that can function as computing device 104 is described with respect to Figure 10 computing device 1000. In an implementation, computing device 104 is a client-side or front-end device.

[0052] In addition to server 102, computing device 104 can also implement functional aspects of operating environment 100, such as one or more functions of search engine 110. It will be understood that some implementations of the present technology will include client-side or front-end computing devices, back-end or server-side computing devices, or both, that perform any combination of the functions in search engine 110 and other functions or combinations of functions (including those not shown).

[0053] Search engine 110 generally performs an item search using a search query determined from a photo-realistic image generated from a category-based classification system and provides search results for the item as a response. The search results can be included within a search engine results page (SERP). An item can include any one of a web page, an image, a video, an infographic, an article, a research paper, a document, a good or service for sale, and other types of files, and can include an associated description or hyperlink. Search engine 110 can be configured for general Internet or web searches, or can be configured to search a specific database or website, e.g., as an example, a search engine configured to return a list of items on an e-commerce platform. The search results returned in response to a search query can include one or more items.

[0054] Broadly, search engine 110 can be used, either alone or in coordination with other components or systems, to perform item searches and provide a SERP of items as a response. In doing so, search engine 110 determines a search query from a photorealistic image generated from items based on previous interactions with a category in a category classification system. As an example, search engine 110 can classify previously interacted-with items into categories in a category classification system. A set consisting of one or more previously interacted-with items can be selected from one of the categories. The set of previously interacted-with items is provided to an image model, which generates a photorealistic image. The photorealistic image includes an item rendering, which is an object in the image that can correspond to a visually similar item recognizable on the network. A search query can be determined from the photo-realistic image, and the search query is executed to identify and return search results.

[0055] To this end, the illustrated example search engine 110 uses an item classification engine 112, a category ranker 114, an item selector 116, a photorealistic image generator 118, a context determiner 120, a segmentation engine 122, and a searcher 124. Again, note that search engine 110 is intended as an example suitable for implementing the present technology. However, other arrangements and architectures of components for generating a photorealistic image for determining a search query to be executed, as well as functions, are intended to be included within the scope of the present disclosure and are understood by those practicing the present technology.

[0056] Generally, the item classification engine 112 can be used to classify previously interacted-with items. The previously interacted-with items can be classified according to a category classification system. In some aspects, the category classification system is a predefined classification system that can define categories and describe the attributes common to each category, through which the previously interacted-with items can be classified. In another aspect, the category classification system is determined based on descriptions of items (including previously interacted-with items or other items). For example, a model (e.g., a generative model) can be used to identify the attributes common to items and classify the items based on the determined common attributes, thereby generating a category classification system based on the items themselves.

[0057] In the illustrated example, the item classification engine 112 uses the item classifier 126 to classify previously interacted items into categories. Generally, the item classifier 126 can be a machine learning model that receives a previously interacted item as input and determines a category as a response. The item classifier 126 can be trained to receive text-based or image-based input or both and predict or generate a category from the input. Thus, by receiving a previously interacted item, the item classifier 126 can receive all or part of a description of the item (including a text description or an image corresponding to the previously interacted item). For example, some items (e.g., documents, item lists of items for sale, web pages, etc.) can include a title that can be used as input to the item classifier 126. Other text portions of the item description can be provided as input. An image identifying or representing the item can be provided to the item classifier 126 to determine into which category the item should be classified.

[0058] In one aspect, the item classifier 126 is a generative AI model. The generative AI model can contextualize the input and output the category into which the previously interacted item is classified. In doing so, the generative AI model can generate a category for the previously interacted item, identify a previously generated category for the previously interacted item, or can identify a category from a predefined category classification system. This can be done based on the type of the generative AI model and its training.

[0059] As an example, the item classifier 126 can be trained to classify previously interacted items into categories in a category classification system. The training data can include a general database of text, visual, and general descriptive information. For example, the training data can include a wide variety of sources capable of generating a broad understanding of human language, such as websites, books, Wikipedia, scientific articles, news media, technical manuals, movie scripts, programming code, educational materials, and other text. The item classifier 126 can be trained on a data set including such text.

[0060] For example, for a model including a transformer architecture with multiple layers and millions of parameters, the training objective is to minimize the difference between the predicted value and the actual value of the next word under a given word sequence of the training data (e.g., general text materials). This can be achieved using a loss function (e.g., cross entropy). The parameters of the item classifier 126 can be optimized using, for example, a gradient-based optimization algorithm (e.g., Adam). This results in a pre-trained base model, which in some cases can be used as the item classifier 126.

[0061] In some aspects, the item corpus 132 can be used, at least in part, to train the item classifier 126. The item corpus includes items and item descriptions corresponding to the items. In some aspects, this can include items sold on an e-commerce platform and item descriptions corresponding to the items.

[0062] In one aspect, the item classifier 126 is fine-tuned on a specific document set, which can be based on the use case of the item classifier 126. The fine-tuning can be performed using an algorithm similar to the algorithm described when training the initial base model. For example, in a broad use case (e.g., general Internet search), the item classifier 126 can be trained on the general training database described. In cases where the item classifier 126 is used for a specific task or specific context, a corpus of documents related to the task or concept can be used to train or fine-tune the item classifier 126. As an example, for use on an e-commerce website, the item classifier 126 can be trained based on items sold on the e-commerce website and item descriptions of the items (e.g., items sold on the e-commerce website and item descriptions of the items included in the item corpus 132). By doing so, the item classifier 126 can better contextualize the input provided in the context of searching the item list on the e-commerce website.

[0063] In one aspect, the resulting trained item classifier 126 is a generative AI model, such that the generative AI model generates new content in response to a prompt based on the training. In such aspects, the item classifier 126 can be used to classify items by receiving, as a prompt, instructions for classifying an item and one or more items. Based on the item description of the item or based on the item description relative to other item descriptions, the item classifier 126 outputs a classification for one or more items. In the context of previously interacted items, the previously interacted items and a prompt for identifying the item classification can be provided to the item classifier 126, and based on the prompt, the item classifier 126 generates an item classification for the previously interacted items.

[0064] In some aspects, a predetermined category classification system is used. In such cases, an example method for classifying an item (e.g., a previously interacted item) includes: providing the categories in the category classification system to the item classifier 126. This can include: providing a description of the categories within the category classification system. After receiving this information, the item classifier 126 can be prompted to classify the item (including the previously interacted item) according to the provided classification system.

[0065] In some aspects, the item classifier 126 can determine a classification based on an image of an item (e.g., a previously interacted-with item). Various models can be used for object detection and classification (e.g., CNN (Convolutional Neural Network); regression models (e.g., YOLO (You Only Look Once)); deep neural networks (e.g., SSD (Single Shot Detector)); or other similar models).

[0066] As an example, such models can be trained based on an image dataset. The images can include labeled images. These labels can correspond to categories in a category classification system. In some aspects, the image dataset can be a collection of images of objects that have been manually labeled with the category that the object is intended to belong to. In some cases, an image model that generates images based on a prompt (e.g., a diffusion model) can be used to generate images.

[0067] Various loss functions can be used during training. One example includes cross-entropy loss. This can be minimized during training by an optimization function (e.g., Adam, SGD (Stochastic Gradient Descent), RMSprop, or other similar functions). During training, the item classifier 126 learns to identify an object in the image (e.g., an item rendering) and outputs a classification for that object according to the category classification system.

[0068] Aspects of the present technology use a multimodal model as the item classifier 126. For example, a multimodal model can use text from an item description as well as an image of the item, leveraging one or more models to classify an item (e.g., a previously interacted-with item). The multimodal model can include any one or more of the previously described models, or other model architectures that are trained to classify items based on text descriptions and images.

[0069] It should be noted that the foregoing training methods involving pre-training and fine-tuning are provided as illustrative examples and are not intended to limit the scope of potential training methods that can be used. Other scenarios can include: reinforcement learning with human feedback (RLHF), transfer learning based on related tasks, where the model is initially trained on a task that is similar but not identical to the target task and then fine-tuned on the specific task of interest; multi-task learning, where the model is trained to perform multiple tasks simultaneously, sharing representations between them to improve overall performance; or other training methods. These training methods can be standalone scenarios or integrated with other techniques to create more robust multi-purpose models, and can also be incorporated as new methods are developed.

[0070] In addition to generative AI models and the other models discussed previously, as a supplement or alternative, other classification models can also be used as the item classifier 126. In some aspects, some discriminative models can be trained and used to classify items (e.g., previously interacted items). Some example models that can be used include: SVM (Support Vector Machine); logistic regression model; decision tree-based models (e.g., random forest, GBM (Gradient Boosting Machine), etc.); KNN (K-Nearest Neighbor); non-generative neural networks (e.g., the neural networks described previously, or other neural networks (e.g., CNN and computer vision techniques)); or other similar models can be used. These models are typically trained on a labeled dataset, and in some aspects, the item corpus 132 can be used to train or fine-tune these models.

[0071] Based on its training, the item classifier 126 can be suitable for classifying items into categories in a category classification system. Thus, the item classifier 126 can classify previously interacted items into the category classification system. As a result, the item classification engine 112 can use the item classifier 126 to classify previously interacted items into categories in the category classification system. This can be done by receiving all or part of the description of the previously interacted item as input. The description of the previously interacted item includes one or more of the following: text description of the item, item image, video of the item, audio describing the item, or other similar modes of conveying item information. In response to the input, the item classifier 126 outputs one or more categories, thus identifying or generating the one or more categories into which the item classification engine 112 classifies the item. In one example implementation, the item classification engine 112 uses the title in the description of the previously interacted item to classify the corresponding previously interacted item into a category.

[0072] In some aspects, the video of the item can be reduced to individual frames, and the individual frames can be provided to the item classification engine 112 for classifying the previously interacted item into a category. In some aspects, the description of the previously interacted item includes the audio of the item, and a speech-to-text model can be used to reduce the audio to text. There are various speech-to-text models applicable to the present technology, and these models will be well-known to those of ordinary skill in the art.

[0073] Figure 2 An illustration of an example classification of a previously interacted item 202 using the item classifier 126 is provided. In this example, the previously interacted item 202 includes items 1-9 respectively labeled as 204-220. All or part of the corresponding description of the previously interacted item is input into the item classifier 126.

[0074] In response, item classifier 126 classifies the previously interacted items into a category classification system 222. In this example, the category classification system 222 includes category A 224, category B 226, and category C 228. As previously described, each item can be classified into one or more categories in the category classification system 222 based on the characteristics of the item, where each item in a category in the category classification system 222 shares common characteristics. In a specific implementation of this technology, each category represents a clothing style, and each item in the category shares a common style. Thus, the term "characteristic" can include physical characteristics or intangible characteristics (e.g., style). In the illustrated example, item 2 206, item 3 208, and item 7 216 are classified into category A 224; item 1 204, item 3 208, item 4 210, item 8 218, and item 9 220 are all classified into category B 226; and, item 4 210 and item 5 212 are classified into category C 228. As previously described, since some items have characteristics common to more than one category, each previously interacted item may be classified into one or more categories. In the illustrated example, item 3 208 is classified into category A 224 and category B 226. Item 4 210 is classified into category B 226 and category C 228. The remaining items are classified into only one category.

[0075] Category ranker 114 generally ranks the categories in the category classification system. The categories can be ranked based on the previously interacted items classified into each category. For example, the categories can be ranked based on the number of previously interacted items classified into each category. For example, a category with a larger number of previously interacted items can be ranked higher than a category with a relatively smaller number of previously interacted items classified into it.

[0076] Figure 3 A diagram illustrating an example ranking using category ranker 114 is provided. Here, category ranker 114 ranks the categories (including category A 224, category B 226, and category C 228) in the category classification system 222 shown in Figure 2 . In the illustrated example, category ranker 114 ranks the categories based on the number of items (in this case, previously interacted items), where the category with a larger number of items is ranked relatively higher. As shown, the ranked categories 302 include category B 226, which is the highest ranked category because it has five items. Category A 224 is ranked second because it has three items, fewer than the number of items in category B 226. Category C 228 is ranked third because it has two items, fewer than the number of items in category B 226 and category A 224.

[0077] As described above, an image for searching can be generated according to a category classification system. In other words, items from previous interactions of a specific category can be used to generate an image for searching for items.

[0078] The item selector 116 generally selects previously interacted items from one or more categories. The item selector 116 can select previously interacted items from one or more ranked categories according to the ranking. For example, one of the categories from which items are selected can be a top-ranked category. Items can be selected for any number of categories.

[0079] When selecting items from a category, the item selector 116 can select any number of items. In one aspect, the item selector 116 selects all items classified into a specific category (e.g., a top-ranked category). In another aspect, the item selector 116 selects a subset of the previously interacted items from a specific category. The selected items are included in the set of previously interacted items.

[0080] In one aspect, the item selector 116 can select additional supplementary items to include in the set of previously interacted items. In other words, in some implementations, the item selector 116 can identify the previously interacted items and select items that are supplementary to the previously interacted items. The supplementary items can be items selected from the same category as the previously interacted items. In some cases, the supplementary items are selected because they include the same or similar features (e.g., identified by common tags, descriptions, or visual features of the items). By doing so, the item selector 116 can establish a robust set of previously interacted items from which to generate photo-realistic images. In doing so, it can contribute to the rendering of more diverse photo-realistic images of items for recommendation or search.

[0081] The item selector 116 can generate one or more sets of previously interacted items from a single category. For example, the item selector 116 can generate a single set of previously interacted items that includes the previously interacted items selected from that category. All or a selected portion of the previously interacted items classified into that category can be selected and used to generate an image for searching for items based on that category, as will be further described.

[0082] In another aspect, multiple sets of previously interacted items are selected. Thus, in one aspect, the item selector 116 can generate a first set of previously interacted items and at least a second set of previously interacted items. For each set of previously interacted items, different previously interacted items from that category can be selected. As a result, the search engine 110 can be configured to generate or access and provide multiple different images that can be used for item search.

[0083] Figure 4 An example of previously interacted items that can be selected using the item selector 116 is shown. In the example shown, category B 226 is provided as the specific category from which items are selected for image generation. This can be done based on category B 226 being the highest ranked category. Although two sets of previously interacted items are generated, it will be appreciated that any one or more sets can be generated using the item selector 116. In the example shown, the set 402 of previously interacted items is generated from a portion of the previously interacted items classified into category B 226. In this example, the item selector 116 has selected item 1 204, item 3 208, and item 4 210 to be included in the set 402 of previously interacted items. The item selector 116 has generated a second set 404 of previously interacted items that includes item 4 210, item 8 218, and item 9 220. As shown, previously interacted items can be selected to be included in more than one set of previously interacted items. As will be described, by generating multiple sets of previously interacted items with different previously interacted items, different images for the same category can be generated for item search from the database.

[0084] Although Figure 4 not shown, the item selector 116 can also select sets of previously interacted items from additional categories based on ranking. For example, the item selector 116 can select sets of previously interacted items from any number of categories according to a ranking based on a threshold ranking value. In other words, the threshold ranking value can define the number of categories from which the item selector 116 selects sets of previously interacted items. For example, if the threshold ranking value is five, the item selector 116 can select one or more sets of previously interacted items from each of the top five ranked categories. The search engine 110 can be configured to provide images generated from previously interacted items from any number of categories.

[0085] Generally, the photo-realistic image generator 118 can be used to generate photo-realistic images from sets of previously interacted items using the image model 128. The generated photo-realistic images can include one or more item renderings and can be used to determine search queries for item search.

[0086] Typically, a photo-realistic image is an output generated by an image model 128 in response to an input. The input can be based on a collection of previously interacted items and can be modified with prompts for the image model 128 to generate a photo-realistic image. A photo-realistic image can mimic the appearance and quality of a photograph. A photo-realistic image can include a synthetic image generated by the image model 128 using generative techniques and simulating a real-world scene with objects. The objects in these images can represent the generated synthetic items, and the synthetic items can be similar to real-world items that can be searched via indexing.

[0087] These synthetic items in the photo-realistic image are generally referred to as item renders. A render is a digital representation of an item that is potentially searchable via a search engine. They are constructed based on the input parameters provided to the image model 128 and integrated into the overall photo-realistic image. Depending on the type of model used as the image model 128 and the specific input prompts using the collection of previously interacted items, the item renders can vary widely in complexity and detail.

[0088] To generate a photo-realistic image from an input based on a collection of previously interacted items, the image model 128 can include a machine learning model. In one example, the image model 128 is a generative AI model that receives text, images, or both and outputs a photo-realistic image in response. In a specific example, the image model 128 is a diffusion model. While generally referring to diffusion models (or more broadly, generative AI models), it will be understood that other image generation models can be adopted or developed, and these models are intended to be within the scope of the present disclosure. Some non-limiting examples can include generative adversarial networks (GANs), variational autoencoders (VAEs), transformer models, etc. The image model 128 can be a single AI model or can be a combination of various models for generating images.

[0089] In the context of a diffusion model, the training process for generating an image can include a two-stage mechanism that first corrupts the original image by iteratively adding noise and then reverses the process to generate new image samples based on the input. Generally, there are numerous image datasets that can be used. Some examples include Flickr 30k, IMBD-Wiki, Berkeley Deep Drive, etc. The input is encoded into the latent space to condition the generation process. During the diffusion process, the image model 128 learns to map this latent representation to a series of noisy image states, thus effectively learning the transition dynamics between the input and the corresponding image.

[0090] The image model 128 is trained to minimize the difference between the generated image and the actual image corresponding to the input. As an example, mean squared error, cross-entropy, or other similar functions can be used as the loss function for training. Optimization is typically performed using gradient-based algorithms (e.g., Stochastic Gradient Descent (SGD), Adam, etc.). Once trained, the image model 128 can take the input and iteratively refine the noisy image until it generates a new image that closely matches the input, thereby generating a photo-realistic image based on the input. Other training methods can be used as they are developed or can be used based on the use of a specific model. Based on its training, the image model 128 receives an input that can include text or an image of a collection of items from a previous interaction. In response, the image model 128 generates a photo-realistic image representing the collection of items of the previous interaction of the input, the photo-realistic image potentially including item renderings.

[0091] Figure 5 An example of the photo-realistic image generator 118 using the image model 128 to generate a photo-realistic image 502 is shown. Continuing Figures 1 to 4 with the example shown, the collection of items 402 from the previous interaction includes item 1 204, item 3 208, and item 4 210, all of which are used as inputs to the image model 128, as shown. In some aspects, the captions from each of item 1 204, item 3 208, and item 4 210 are provided to the image model 128. Generally, the input to the image model 128 can use all or part of the text descriptions of each of the items from the collection of items 402 from the previous interaction. The text description can include text describing the characteristics of the items of the previous interaction. As described above, the input to the image model 128 can include an image, and thus, one or more images of the items in item 1 204, item 3 208, and item 4 210 can be included in the input. In multimodal aspects, both text and image are provided as inputs. In some aspects, the image model 128 can receive a prompt (e.g., "generate an image based on the following description or image") as an input and include the text description or image of the items from the previous interaction. Depending on the model used as the image model 128 and its training, other prompts and prompt manipulation techniques can be used. Based on these inputs, the image model 128 generates a photo-realistic image 502, which, as will be further discussed, can include item renderings that can form an item search for the items in the database.

[0092] In one aspect, an input can be provided to the photo-realistic image generator 118 that causes the image model 128 to generate a photo-realistic image with a project rendering and a contextual background. The contextual background can be generated to provide context for the generated object or context related to a user performing a project search. For example, if the project rendering is related to skiing equipment, the contextual background can include a mountain range or other snow-related scenes. If the project rendering is related to swimwear, the contextual background can be related to a beach or a pool. The contextual background can also be user-related. For example, if the user lives in a coastal area (e.g., determined by the user account or IP address), the contextual background can include coastal scenes. The contextual background can be based on the time of year. For example, if the project search performed using the generated photo-realistic image is close to the Christmas holiday, the contextual background can be Christmas holiday-themed. These are just some examples of how a contextual background can be generated, and it will be understood that other examples are also applicable to this technology.

[0093] To generate a contextual background for the photo-realistic image, the search engine 110 can use the context determiner 120 to determine the context. When generating the photo-realistic image, context information (e.g., context information related to location, date, project information, etc.) can be identified and passed to the image model 128. In some implementations, the context determiner 120 uses the context model 130 to determine the context information for the contextual background that can be used to generate the photo-realistic image.

[0094] To generate the contextual background, in addition to the context information, the prompt that is also provided as an input to the image model 128 can include information about the items from previous interactions (as described above). The prompt can be manipulated to indicate that the context information should be generated as a background for the photo-realistic image. For example, the prompt can include "generate an image with a beach background", or any other prompt that indicates that the context information should be generated as a contextual background.

[0095] In the example shown by Figure 5 the context information is determined by the context determiner 120 and passed to the image model 128 for generating the photo-realistic image 502. In this example, the context determiner 120 uses the context model 130 to determine the context information from the set 402 of previous interaction items that includes the item 1 204, the item 3 208, and the item 4 210. The context information determined by the context model 130 is passed to the image model 128, which generates a contextual background corresponding to the context information.

[0096] In one example, the context model 130 is a generative AI model trained to understand text or images. For example, a large language model trained on a general database of textual, visual, and general descriptive information learns a broad understanding of text or visual information related to human language and comprehension, and based on this, can determine context. Some example models that can identify context from project information, including a textual description from a project or a portion of an image, include GPT-4, Mistral 7B, LLaMa 2, etc. As a result, in some implementations, the context determiner 120 can provide a prompt to the generative AI model to determine the context of project information from previously interacted-with projects.

[0097] Although Figure 5 a single photo-realistic image is shown (i.e., photo-realistic image 502), any number of photo-realistic images can be generated. In one implementation, the photo-realistic image generator 118 can generate one or more photo-realistic images for a single category. In such cases, the set of previously interacted-with projects used to generate each photo-realistic image can be different, as Figure 4 shown. This provides different photo-realistic images with potentially different project renderings, where the user can select from the different project renderings for project search. As a result, the photo-realistic image generator 118 can generate a plurality of photo-realistic images including a first photo-realistic image and at least a second photo-realistic image, the first photo-realistic image being generated from a first set of previously interacted-with projects, the second photo-realistic image being generated from a second set of previously interacted-with projects of the same category as the first set of previously interacted-with projects, where the second set of previously interacted-with projects has a different combination of previously interacted-with projects than the first set of previously interacted-with projects. Each generated image can be presented, and an image or a project rendering from the image can be selected for project search.

[0098] In addition, in one aspect, the photorealistic image generator 118 generates one or more images from each set of a collection of multiple previously interacted items (e.g., from the top three or five categories). This provides different photorealistic images corresponding to different categories, which include different item renderings, where the user can select from the different item renderings to initiate an item search. Here, the photorealistic image generator 118 can generate multiple photorealistic images including a first photorealistic image and at least a second photorealistic image, where the first photorealistic image is generated from a first set of previously interacted items, and the second photorealistic image is generated from a second set of previously interacted items from a category different from the first set of previously interacted items. Each generated image can be presented, and an image or an item rendering from the image can be selected for an item search.

[0099] As will be described below, the photorealistic images generated by the photorealistic image generator 118 can be used to determine a search query for performing an item search. The search query can include a photorealistic image, or a portion of a photorealistic image (e.g., a separated item rendering) or a text description thereof.

[0100] To identify and separate item renderings from photorealistic images, the search engine 110 can use a segmentation engine 122. Generally, the segmentation engine 122 identifies and separates item renderings. The separated item renderings can be used as a basis for a search query for an item search.

[0101] In some aspects, the segmentation engine 122 can be used to identify an item rendering in a photorealistic image and separate it from the rest of the photorealistic image. The segmentation engine 122 can apply a segmentation mask over the region in the photorealistic image corresponding to the item rendering.

[0102] For example, for image segmentation, a CNN can be used. The CNN can be trained using a labeled image dataset, such as ImageNet, COCO (Common Objects in Context), PASCAL VOC (Visual Object Classes), etc. These datasets include labeled images with a rich set of annotations, which include thousands of object categories. As a supplement or alternative to the datasets described above, more specific datasets can also be used for item rendering recognition. A more specific dataset includes millions of items sold or offered for sale on eBay. These items and item titles or other labels can be used to train or fine-tune a model to accurately identify item renderings in photorealistic images corresponding to items similar to those sold on an e-commerce platform.

[0103] There are various methods for training a CNN for image segmentation. Among them, the Fully Convolutional Network (FCN) and U-Net are examples. The FCN is designed to process inputs of various sizes, making them suitable for use with this technology. The U-Net can extend the capabilities of the FCN by incorporating skip connections, which helps to preserve details of the photo-realistic image from the input across the entire network architecture. In an example training method, a labeled image dataset (where each pixel in the image is assigned to a specific class or category) is used to train the CNN. During training, the CNN network learns to recognize patterns and features through backpropagation, with the goal of minimizing the difference between its predicted segmentation and the ground truth labels.

[0104] Image classification can be done by the same or different networks to identify item renderings in an image. The CNN can also be suitable for classification tasks. During training, each pixel can be labeled with a specific classification in the training image dataset. These classifications can be related to many different objects in the image. For example, pixel classification can be related to various items. To minimize these differences during training, mean squared error, cross-entropy loss, or other similar algorithms can be used. This is just one example, and other models and other training methods can be used for image segmentation.

[0105] Models other than the CNN (such as RNN (Recurrent Neural Network), autoencoder, GAN, Transformer, etc.) can also be used to identify and separate item renderings from photo-realistic images. Various training methods can be used (such as supervised training methods using the previously provided training dataset). The discussion in this article aims to provide some examples for this technology. These networks can be used alone or in combination with other networks, and other models suitable for use can be developed.

[0106] The segmentation engine 122 can use the trained image recognition model to segment a photo-realistic image for item renderings. For example, the segmentation engine 122 identifies pixels in the input photo-realistic image and assigns a classification to each pixel (including classifications corresponding to items learned from training). In doing so, the segmentation engine 122 identifies the edges of item renderings in the photo-realistic image, as the edges can correspond to pixels of an item rendering adjacent to pixels assigned to another classification or another item rendering. The segmentation engine 122 can separate the item rendering by removing or rendering as transparent the pixels in the photo-realistic image that are not assigned to the specific item rendering to be separated. The remaining pixels in the photo-realistic image that are assigned to the item rendering provide the separated item rendering.

[0107] Figure 6Shows an example of how the segmentation engine 122 segments the photo-realistic image 502 to identify and separate the item renders 602A - 602D. One or more of these item renders can be identified and separated by the segmentation engine 122, as shown by the separated item renders 604A - 604D. Any one or more of the item renders can be used for item search.

[0108] In some aspects, the segmented photo-realistic image can be presented to the user. A selection of the segmented photo-realistic image can be received. The selection can correspond to an area in the segmented photo-realistic image that is identified as an item render, and thus indicates that the selected specific item render will be used as the basis for a search query for item search. This is an example method where a portion of the photo-realistic image corresponding to a specific item render can be identified and used for item search.

[0109] To perform an item search, the search engine 110 can use the searcher 124. The searcher 124 can use text-based or image-based search techniques to identify items using the search query and return search results. Generally, the searcher 124 uses a search query determined from the photo-realistic image to perform the item search. The search query can be all or part of the photo-realistic image (e.g., the item render).

[0110] Generally, the search query used to perform an item search can be determined from the photo-realistic image, which means that the search query can be determined from all or part of the photo-realistic image (e.g., one or more item renders). The search query can include the photo-realistic image or a portion thereof (e.g., one or more separated item renders). In some aspects, the search query is a text-based query, where the text of the text-based query is determined from the photo-realistic image or a portion thereof (e.g., one or more separated item renders). Various models can be used to generate text from an image, such as: a CNN-RNN architecture that can utilize a CNN to extract image features and an RNN to generate descriptive text from the extracted features; a generative AI model trained with both image and text descriptions; or other model types and training. An example model that can be suitable for use is the NIC (Neural Image Caption) generator.

[0111] Figure 7An example of a search query determined from a photorealistic image 502 is shown. As shown, the search query 704 is a text-based search query. Here, the search query 704 is a text description 702 generated from the photorealistic image 502. The photorealistic image 502 can be used as an input to an image-to-text generation model that outputs the text description 702 to be used as the search query 704, as described. In one aspect, a search query 706 is generated as an image-based search query, where the photorealistic image 502 is used as the image.

[0112] The search query 710 is generated from a portion of the photorealistic image 502, which in this instance is the isolated item rendering 604A, as Figure 6 shown. Although shown as using one isolated item rendering 604A, any one or more isolated item renderings can be used to determine the search query. In one aspect, the user selects an isolated item rendering (e.g., the isolated item rendering 604A) for generating the search query. Continuing Figure 7 ,the search query 710 is a text-based search query. Here, the search query 710 is a text description 708 generated from the isolated item rendering 604A. The isolated item rendering 604A can be used as an input to an image-to-text generation model that outputs the text description 708 to be used as the search query 710, as described. In one aspect, a search query 712 is generated as an image-based search query, where the isolated item rendering 604A is used as the image.

[0113] The searcher 124 uses a search query (e.g., any one of the search queries 704, 706, 710, or 712) to perform an item search and returns search results. The search performed can return one or more items. In one aspect, the searcher 124 uses a text-based search query or an image-based search query to perform the item search. In one aspect, the searcher 124 performs the item search by combining a text-based search query search with an image-based search query search. This can be done for search queries originating from the same photorealistic image or a portion thereof. In short, using Figure 7 as an example, the searcher 124 can use both the search query 704 and the search query 706 to perform an item search for the photorealistic image 502. Similarly, the searcher 124 can use both the search query 710 and the search query 712 to perform an item search for the isolated item rendering 604A. One or more items identified during the search can be provided as search results (e.g., recommendations or a search results page).

[0114] Now turning toFigure 8 , an example implementation of a searcher 124 for performing item searches is provided. As shown, the searcher 124 performs an item search using a search query 802, which can be a text-based search query, an image-based search query, or a combination of both determined from a photo-realistic image or a portion thereof. The item search returns search results 804, which are shown to include search result 1 806, search result 2 808, search result 3 810, and search result 4 812. These are items that were identified during the search that may be similar to the photo-realistic image or the isolated item rendering.

[0115] Now refer to Figure 9 , a block diagram of an example method for performing item searches is provided. Each block of the method can include a computational process performed using any combination of hardware, firmware, or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method can also be embodied as computer-usable instructions stored on a computer storage medium. The method can be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service), or a plug-in for another product, just to list a few possibilities. The method can be implemented, in whole or in part, by components of the operating environment 100.

[0116] In block 902, method 900 classifies previously interacted items. This can be done using an item classification engine 112. In one aspect, the item classification engine 112 uses an item classifier 126 to perform the classification of the previously interacted items. In one aspect, a generative AI model is used to perform the classification. Another type of model (e.g., any of the models described previously with respect to the classification engine 112) can be used for the classification. The generative AI model can classify the previously interacted items according to a predefined category classification system, or can generate a classification based on the input previously interacted items and classify the previously interacted items according to the generated category classification system. When classifying the previously interacted items, the item classification engine 112 can receive the title of the previously interacted items, the text description of the previous items included in the previously interacted items, the image of the previously interacted items, or any combination thereof.

[0117] In block 904, method 900 selects a set of previously interacted items from a category in a category classification system. This can be done by item selector 116. Item selector 116 can select items from other categories to form other sets of previously interacted items. In one aspect, the categories in the category classification system are ranked according to the number of previously interacted items classified into each category. In one aspect, item selector 116 selects a set of previously interacted items from the top-ranked categories. Item selector 116 can select multiple sets of previously interacted items from an item category (e.g., the top-ranked categories) in such a way that each set of previously interacted items selected from the category includes different previously interacted items.

[0118] In block 906, method 900 accesses a photorealistic image including item renderings. The accessed photorealistic image can be generated from the set of previously interacted items. This can be done using an image model (e.g., image model 128). In one aspect, accessing the photorealistic image includes providing a cue generated from the set of previously interacted items to the image model and receiving the photorealistic image as a response. In some aspects, accessing the photorealistic image includes generating the photorealistic image using photorealistic image generator 118.

[0119] In one aspect, multiple photorealistic images can be generated from the same category or different categories. In some cases, the generated photorealistic images are presented to the user. As will be described below, when performing an item search, a photorealistic image can be selected from the presented photorealistic images and used, or an item render can be selected from the presented photorealistic images and used.

[0120] In some cases, a photorealistic image with a contextual background can be generated. This can be done by determining the context from the previously interacted items in the set of previously interacted items, from the user or computer location, or from some other user data from the user account. The context information can be passed to photorealistic image generator 118 along with instructions to use the context information as a background when generating the photorealistic image.

[0121] In block 908, method 900 performs an item search for an item. The performed item search can include: searching for items corresponding to the item renderings by using a search query determined from the photorealistic image. The performed search can return one or more items as search results (e.g., item recommendations).

[0122] In one aspect, item rendering is used to determine a search query. The item rendering can be a detached item rendering. For example, a photo-realistic image can be provided to the segmentation engine 122 to identify and isolate the item rendering. One or more item renderings can be selected and used to determine a search query for an item search being performed.

[0123] An overview of some embodiments of the technology has been described. The following describes example computing environments in which embodiments of the technology can be implemented to provide a general context for various aspects of the technology. Specifically, now referring to Figure 10 , an example operating environment for implementing embodiments of the technology is shown and is generally labeled as computing device 1000. Computing device 1000 is only one example of a suitable computing environment and is not intended to imply any limitation as to the scope of use or functionality of the technology. Computing device 1000 should not be construed as having any dependency or requirement related to any one or combination of the components shown.

[0124] The technology can be described in the general context of computer code or machine usable instructions, including computer executable instructions (such as program modules) executed by a computer or other machine (such as a cellular phone, personal data assistant, or other handheld device). Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The technology can be practiced in a variety of system configurations, including handheld devices, consumer electronics, general purpose computers, more specialized computing devices, etc. The technology can also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0125] Referring to Figure 10 , computing device 1000 includes a bus 1010 that directly or indirectly couples the following devices: a memory 1012, one or more processors 1014, one or more presentation components 1016, input / output (I / O) ports 1018, input / output components 1020, and an example illustrative power supply 1022. Bus 1010 represents one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although, for clarity, Figure 10 the various boxes of Figure 10The figures are only illustrative of example computing devices that can be used in conjunction with one or more embodiments of the present technology. No distinction is made among categories such as "workstation", "server", "laptop computer", "handheld device", etc., because all such categories are within the scope of Figure 10 and are referred to as "computing devices".

[0126] Computing device 1000 generally includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1000 and includes volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media (also referred to as communication components) includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to: RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 1000. Computer storage media does not itself include signals.

[0127] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" refers to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example and not limitation, communication media includes wired media (e.g., a wired network or direct wired connection) and wireless media (e.g., acoustic wireless media, radio frequency (RF) wireless media, infrared wireless media, and other wireless media). Combinations of any of the above should also be included within the scope of computer-readable media.

[0128] Memory 1012 includes computer storage media in the form of volatile or non-volatile memory. The memory can be removable, non-removable, or a combination thereof. Example hardware devices include solid state memory, hard disk drives, optical disk drives, etc. Computing device 1000 includes one or more processors that read data from various entities, such as memory 1012 or I / O component 1020. Presentation component 1016 presents data indications to a user or other device. Examples of presentation components include display devices, speakers, printing components, vibration components, etc.

[0129] The I / O port 1018 allows the computing device 1000 to be logically coupled to other devices including the I / O component 1020, some of which may be built-in. The components schematically shown include a microphone, a joystick, a gamepad, a satellite antenna, a scanner, a printer, a wireless device, etc. The I / O component 1020 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network unit for further processing. The NUI can implement any combination of the following: speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, and air gestures, head and eye tracking, or touch recognition associated with the display of the computing device 1000. The computing device 1000 can be equipped with a depth camera for gesture detection and recognition, such as a stereo camera system, an infrared camera system, an RGB (red-green-blue) camera system, touchscreen technology, other similar systems, or a combination of these. Additionally, the computing device 1000 can be equipped with an accelerometer or a gyroscope capable of detecting motion. The output of the accelerometer or gyroscope can be provided to the display of the computing device 1000 to render immersive augmented reality or virtual reality.

[0130] At a low level, the hardware processor executes instructions selected from the machine language (also known as machine code or native) instruction set of a given processor. The processor recognizes the native instructions and performs corresponding low-level functions related to, for example, logical, control, and memory operations. Low-level software written in machine code can provide more complex functions for higher-level software. As used herein, computer-executable instructions include any software, including low-level software written in machine code; higher-level software such as application software; and any combination thereof. In this regard, a component that searches for items using images generated using a category classification system can manage resources and provide the described functions. Any other variants and their combinations are contemplated within the embodiments of the present technology.

[0131] Briefly return to the reference Figure 1 , note and emphasize again that any additional or fewer components in any arrangement can be used to achieve the desired functions within the scope of the present disclosure. Although, for clarity, Figure 1 the various components are shown by lines, in reality, the boundaries between the various components are not so clear, and metaphorically, the lines could more accurately be gray or blurred. Although Figure 1Some of the components are depicted as a single component, but these depictions are intended as examples in nature and number and should not be construed as limiting for all embodiments of the present disclosure. The functionality of the operating environment 100 can be further described based on the functionality and characteristics of its components. Other arrangements and elements (e.g., machines, interfaces, functions, commands, and groupings of functions) can be used to supplement or replace the shown arrangements and elements, and some elements can be omitted entirely.

[0132] In addition, with respect to Figure 1 Some of the units described (e.g., the units described with respect to the search engine 110) can be implemented as discrete components or distributed components or functional entities combined with other components and can be implemented in any suitable combination and location. The various functions described herein are performed by one or more entities and can be performed by hardware, firmware, or software. For example, the various functions can be performed by a processor executing computer-executable instructions stored in a memory (e.g., the database 106). In addition, the functionality of the search engine 110 and other functions can be performed by the server 102, the computing device 104, or any other component in any combination. For example, in a particular aspect, the search engine 110 can identify and select a set of previously interacted items, provide them to another server for generating an image, and then receive the image back from the server. This is merely an example, and those of ordinary skill in the art will understand other example combinations and configurations within the context of the content of the present disclosure.

[0133] Generally, after identifying the various components in the present disclosure by reference to the drawings and the description, it should be understood that any number of components and arrangements can be employed to achieve the desired functionality within the scope of the present disclosure. For example, for clarity of concept, the components in the embodiments depicted in the drawings are shown with lines. Other arrangements of these and other components can also be implemented. For example, although some components are depicted as a single component, many of the units described herein can be implemented as discrete or distributed components or combined with other components and implemented in any suitable combination and location. Some units can be omitted entirely. In addition, the various functions described herein that are performed by one or more entities can be performed by hardware, firmware, or software. For example, the various functions can be performed by a processor executing instructions stored in a memory. Thus, other arrangements and units (e.g., machines, interfaces, functions, commands, and groupings of functions, etc.) can be used to supplement or replace the shown arrangements and units.

[0134] The above embodiments can be combined with one or more of the specifically described alternatives. Specifically, the claimed embodiments can include references to more than one other embodiment in the alternatives. The claimed embodiments can specify other limitations of the claimed subject matter.

[0135] This document specifically describes the subject matter of the present technology to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Instead, the inventors have contemplated that the claimed or disclosed subject matter may also be embodied in other ways, to incorporate other existing technologies or future technologies to include different steps or combinations of steps similar to those described in this document. Additionally, although the terms "step" or "block" may be used herein to denote different elements of the methods employed, such terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and except when the order of the individual steps is explicitly described.

[0136] For the purposes of this disclosure, the words "comprising," "having," and other similar words and their derivatives have the same broad meaning as the word "including," and the word "access" includes "receive," "reference," or "retrieve" or their derivatives. Additionally, the word "communicate" has the same broad meaning as the words "receive" or "send," such words "receive" or "send" as facilitated by a software- or hardware-based bus, receiver, or transmitter using the communication media described herein.

[0137] Furthermore, unless otherwise stated, words such as "a," "an," etc. include the plural as well as the singular. Thus, for example, in the presence of one or more features, the constraint regarding "a feature" is satisfied. Additionally, the term "or" includes conjunction, disjunction, and both (thus a or b includes: a, or b, and a and b).

[0138] For the purposes of the detailed discussion above, embodiments of the present technology are described with reference to a distributed computing environment. However, the distributed computing environment described herein is merely an example. Components may be configured to perform novel aspects of the embodiments, where the term "configured to" or "configured as" may mean "programmed to" perform a particular task or implement a particular abstract data type using code. Additionally, although embodiments of the present technology may generally be referred to a distributed data object management system and the diagrams described herein, it should be understood that the described technology may be extended to other implementation contexts.

[0139] From the foregoing, it can be seen that the present technology is well-suited to achieve all of the above objectives and purposes, including other advantages that are apparent or inherent in the structure. It will be understood that some features and subcombinations are useful and may be used without reference to other features and subcombinations. This is contemplated by the claims and within the scope of the claims. Since many possible embodiments of the described technology may be made without departing from this scope, it should be understood that all matter described herein or shown in the accompanying drawings should be interpreted as illustrative and not restrictive.

[0140] Some example aspects that can be practiced in accordance with the foregoing description include, but are not limited to, the following examples:

[0141] Aspect 1: A method (system or medium) executed by one or more processors, the method comprising: classifying items of a previous interaction into a category classification system using a generative artificial intelligence (AI) model; selecting a set of items of the previous interaction from a category in the category classification system; generating a photo-realistic image including a rendering of the item using an image model, the photo-realistic image being generated based on the set of items of the previous interaction; and performing an item search for an item corresponding to the item rendering using a search query determined from the photo-realistic image.

[0142] Aspect 2: A system (method or medium) comprising: at least one processor; and one or more computer storage media having computer-readable instructions stored thereon, the computer-readable instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising: ranking the categories in the category classification system by the number of items of the previous interaction classified into each category; selecting a set of items of the previous interaction from the top-ranked categories; generating a photo-realistic image including a rendering of the item using an image model, the photo-realistic image being generated from the set of items of the previous interaction; and performing an item search for an item corresponding to the item rendering using a search query determined from the photo-realistic image.

[0143] Aspect 3: One or more computer storage media (method or system) having computer-readable instructions stored thereon, the computer-readable instructions, when executed by a processor, cause the processor to perform a method, the method comprising: classifying items of a previous interaction into a category classification system using a generative artificial intelligence (AI) model; selecting a set of items of the previous interaction from a category of the category classification system; accessing a photo-realistic image including a rendering of the item, the photo-realistic image being generated from the set of items of the previous interaction by an image model; and performing an item search for an item corresponding to the item rendering using a search query determined from the photo-realistic image.

[0144] Aspect 4: The solution according to any one of Aspect 1 or 3, further comprising: ranking the categories in the category classification system by the number of items of the previous interaction classified into each category, wherein a category in the category classification system from which the set of items of the previous interaction is selected corresponds to a top-ranked category.

[0145] Aspect 5: The solution according to Aspect 2, further comprising: classifying items of the previous interaction into categories in the category classification system using a generative artificial intelligence (AI) model.

[0146] Aspect 6: The solution according to any one of Aspects 1-5 further includes: using an image model to generate a plurality of photo-realistic images, the plurality of photo-realistic images including the photo-realistic image and a second photo-realistic image generated from a second set of items of previous interactions from the one category, the second set of items of previous interactions including a combination of items of previous interactions different from the set of items of previous interactions, wherein the search query is determined based on the received selection of the photo-realistic image.

[0147] Aspect 7: The solution according to any one of Aspects 1-6 further includes: identifying an item rendering from a plurality of item renderings in the photo-realistic image; and separating the item rendering from the plurality of item renderings, wherein a search query is determined and an item search is performed based on the separated item rendering.

[0148] Aspect 8: The solution according to any one of Aspects 1-7 further includes: determining the context of the set of items of previous interactions, wherein a context background for generating the photo-realistic image is generated to correspond to the context.

[0149] Aspect 9: The solution according to any one of Aspects 1-8, wherein the title of the description of the item of previous interaction is provided to a generative AI model for classifying the corresponding item of previous interaction into a category in a category classification system.

[0150] Aspect 10: The solution according to any one of Aspects 1-9, wherein the generative AI model is a multi-modal model, and the classification of the item of previous interaction is performed using at least a part of the text description of the description of the item of previous interaction and the item image of the description of the item of previous interaction respectively.

Claims

1. A method performed by one or more processors, the method comprising: Use generative artificial intelligence (AI) models to categorize previously interacted items into a category classification system; selecting a collection of previously interacted items from a category in the category classification system; generating, using the image model, a photorealistic image including a rendering of an item, the photorealistic image generated from the collection of previously interacted items; as well as An item search is performed for an item corresponding to the item rendering using a search query determined from the photorealistic image.

2. The method according to claim 1, wherein: The titles of the previously interacted item descriptions are provided to the generative AI model for classifying the corresponding previously interacted items into categories in the category classification system.

3. The method according to claim 1, wherein: The generative AI model is a multimodal model, and the classification of previously interacted items is performed using a text description of the previously interacted item description and at least a portion of the item image of the previously interacted item description, respectively.

4. The method according to claim 1, further comprising: The categories in the category classification system are ranked by the number of previously interacted items classified into each category, wherein a category in the category classification system from which the set of previously interacted items is selected corresponds to a top-ranked category.

5. The method according to claim 4, further comprising: Generating a plurality of photorealistic images using the image model, the plurality of photorealistic images comprising the photorealistic image and a second photorealistic image generated from a second set of previously interacted items from the one category, the second set of previously interacted items comprising a combination of previously interacted items that is different from the set of previously interacted items, wherein the search query is determined based on a received selection for the photorealistic image.

6. The method according to claim 1, further comprising: identifying an item rendering from a plurality of item renderings in the photorealistic image; as well as The item rendering is separated from the plurality of item renderings, wherein the search query is determined and the item search is performed based on the separated item rendering.

7. The method according to claim 1, further comprising: A context of the set of previously interacted items is determined, wherein a contextual background of the photo-realistic image is generated to correspond to the context.

8. A system comprising: at least one processor; as well as one or more computer storage media having computer readable instructions stored thereon, which when executed by the at least one processor cause the at least one processor to perform operations comprising: ranking the categories in the category classification system by the number of previously interacted items that were classified into each category; Select a collection of previously interacted items from the top-ranked categories; generating, using the image model, a photorealistic image comprising a rendering of an item, the photorealistic image generated from the collection of previously interacted items; and An item search is performed for an item corresponding to the item rendering using a search query determined from the photorealistic image.

9. The system according to claim 8, wherein: The operations also include: using a generative artificial intelligence (AI) model to classify previously interacted items into categories in the category classification system.

10. The system according to claim 8, wherein: The classification into each category is based on the title of the previously interacted item description of the corresponding previously interacted item.

11. The system according to claim 8, wherein: The classification into each category is based on a text description of a previously interacted item description and at least a portion of an item image of the previously interacted item description.

12. The system according to claim 8, wherein: The operations also include generating a plurality of photo-realistic images using the image model, the plurality of photo-realistic images including the photo-realistic image and a second photo-realistic image generated from a second set of previously interacted items from the top-ranked category, the second set of previously interacted items including a combination of previously interacted items that is different from the set of previously interacted items, wherein the search query is determined based on a received selection for the photo-realistic image.

13. The system according to claim 8, wherein: The operations also include: identifying an item rendering from a plurality of item renderings in the photorealistic image; and The item rendering is separated from the plurality of item renderings, wherein the search query is determined and the item search is performed based on the separated item rendering.

14. The system according to claim 8, wherein: The operations also include determining a context for the set of previously interacted items, wherein a contextual background for the photo-realistic image is generated to correspond to the context.

15. One or more computer storage media having computer readable instructions stored thereon, which when executed by a processor cause the processor to perform a method comprising: Categorize previously interacted items into a category classification system; selecting a collection of previously interacted items from a category in the category classification system; accessing a photorealistic image including a rendering of an item, the photorealistic image generated by an image model from the collection of previously interacted items; as well as An item search is performed for an item corresponding to the item rendering using a search query determined from the photorealistic image.

16. The medium according to claim 15, wherein The titles of the previously interacted item descriptions are provided to a generative artificial intelligence (AI) model for classifying the corresponding previously interacted items into categories in the category classification system.

17. The medium according to claim 15, wherein The classification of previously interacted items is performed using a generative artificial intelligence (AI) model, which is a multimodal model, and the classification of previously interacted items is performed using a text description of the previously interacted item description and at least a portion of the item image of the previously interacted item description.

18. The medium according to claim 15, wherein The method further includes ranking categories in the category classification system by the number of previously interacted items classified into each category, wherein a category in the category classification system from which the set of previously interacted items was selected corresponds to a top-ranked category.

19. The medium according to claim 15, wherein The method also includes accessing a plurality of photorealistic images generated by the image model, the plurality of photorealistic images including the photorealistic image and a second photorealistic image generated from a second set of previously interacted items from the one category, the second set of previously interacted items including a combination of previously interacted items that is different from the set of previously interacted items, wherein the search query is determined based on a received selection for the photorealistic image.

20. The medium according to claim 15, wherein The method further comprises: identifying an item rendering from a plurality of item renderings in the photorealistic image; and The item rendering is separated from the plurality of item renderings, wherein the search query is determined and the item search is performed based on the separated item rendering.