Method and system for enhanced product analysis through aesthetic, emotional, and material metadata matching
Patent Information
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- SHAW IND GROUP INC
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-31
AI Technical Summary
Users face difficulty in narrowing down product choices on digital platforms due to the vast variety of options, particularly in flooring products, as they need to consider numerous pieces of information including appearance, materials, installation time, and environmental factors, making the review process cumbersome and inefficient.
A method and system that utilizes a descriptor generator, leveraging machine learning techniques, to extract aesthetic, texture, and pattern descriptors from user-provided images or phrases, matching them with product catalog descriptors for rapid and accurate product recommendations, simplifying the selection process by considering multiple product attributes beyond traditional keyword searches.
Enables users to quickly find relevant products by matching image or phrase descriptors with product catalog attributes, providing a more thorough and efficient search experience by considering texture, patterns, and aesthetics, reducing the time and effort required to find suitable flooring products.
Abstract
Description
METHOD AND SYSTEM FOR ENHANCED PRODUCT ANALYSIS THROUGH AESTHETIC, EMOTIONAL, AND MATERIAL METADATA MATCHINGCROSS- REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of, and priority to, U.S. provisional application Ser. No. 63 / 625,820, filed January 26, 2024, the contents of which is hereby incorporated by reference herein in its entirety.BACKGROUND
[0002] Digital platforms provide users with potentially thousands of options for products. It can be difficult for a user to narrow down the choices given the variety of options. To determine a product choice, a potential buyer / user of the platform generally has to consider numerous pieces of information.
[0003] In the context of flooring products that may be provided via a digital platform, there are often many variations of designs and products, which can make review and consideration of available products of the platform too onerous. For example, in the case of flooring products, the relevant information includes not only the appearance, shape, and size of the flooring product, but can also include information about the installation time, the history of how the product is made, available inventory, curing time, environmental stresses, types of materials used, and the like.
[0004] For these and other reasons, a need exists for improved methods and systems to identify items or products which match a user’s preferences.SUMMARY
[0005] This disclosure is directed to systems and methods for identifying items from a product catalog which are most relevant to a user based on a selected image or on a phrase.
[0006] A method is described herein for providing, for display on a client device, a set of items. The method generates a plurality of catalog descriptors associated with each of a plurality of items. The method receives an image or a phrase. The method extracts from the image or from the phrase at least one descriptor. Each of the plurality of items is ranked based on a match between the at least one descriptor and the plurality of catalog descriptors. A set ofitems is provided for display on a client device. The set of items provided is based on the ranked matching between the descriptor and the catalog descriptors.
[0007] A system including at least one processor, computer memory and a database is described herein. The at least one processor receives computer program instructions from the computer memory and accesses the database to perform a method. The method provides, for display on a client device, a set of items. The method generates a plurality of catalog descriptors associated with each of a plurality of items. The method receives an image or a phrase. The method extracts from the image or from the phrase at least one descriptor. Each of the plurality of items is ranked based on a match between the at least one descriptor and the plurality of catalog descriptors. A set of items is provided for display on a client device. The set of items provided is based on the ranked matching between the descriptor and the catalog descriptors
[0008] The image may be processed and sent to a descriptor generator to extract an image descriptor. Extracting the at least one descriptor from the phrase may include processing the phrase using natural language processing and sending the processed phrase to the descriptor generator to extract a phrase descriptor. The match between the at least one descriptor and the plurality of catalog descriptors may include an image match or it may include a phrase match. The image match is a match between an image descriptor and the plurality of catalog descriptors. The phrase match is the match between a phrase descriptor and the plurality of catalog descriptors. Each of the plurality of items in the catalog may have at least a catalog color descriptor, a catalog texture descriptor, a catalog pattern descriptor, and a catalog aesthetic descriptor.
[0009] Extracting the at least one descriptor may include: extracting a color descriptor, extracting a texture descriptor; extracting a pattern descriptor; and extracting an aesthetic descriptor. Ranking each of the plurality of items may include any of: calculating a color match score between the color descriptor and the catalog color descriptor; calculating a texture match score between the texture descriptor and the catalog texture descriptor; calculating a pattern match score between the pattern descriptor and the catalog pattern descriptor; or calculating an aesthetic match score between the aesthetic descriptor and the catalog aesthetic descriptor. An overall match score between the received image (or the received phrase) and each of the plurality of items may including fusing the color match score, the texture match score, the pattern match score, and the aesthetic match score.
[0010] Each of the color match score, the texture match score, the pattern match score, and the aesthetic match score may be based on a similarity calculation between an encoding of each of the catalog descriptors and an encoding of the at least one descriptor.
[0011] Calculating any of the texture match score, the pattern match score, or the aesthetic match score may be based on a classification of the image as belonging to one of a set of textures, a set of patterns, or a set of aesthetic criteria and a probability associated with the classification.
[0012] Extracting at least one descriptor from the received image may also include segmenting the received image into segmented images and processing each of the segmented images to extract at least one descriptor from each of the segmented images. The image color descriptor and the catalog color descriptor may be considered to match when an absolute value of the percentage difference between each of the color coordinates of the image color descriptor and the corresponding color coordinates for the catalog color descriptor are within a percentage threshold value of each other. The image color descriptor and the catalog color descriptor may use color coordinates and they may be considered matched if a distance in color coordinate space between the image color descriptor and the catalog color descriptor is less than a threshold value.
[0013] The methods and techniques described in this disclosure simplify selecting a product for a user relative to conventional keyword only searches. The conventional approach for identifying a product, for example, a flooring product, is to use a keyword search based on, for example, a color or a type of product and to be presented with options. To achieve this end, each product in the product catalog must be tagged with metadata or other attribute information related to color or other aspects of the product. A descriptor generator, as described in this document, can leverage machine learning techniques to efficiently add descriptors to each product of the product catalog across a wide variety of attributes, thus enabling more rapid and more accurate searches. In addition, the descriptor generator described herein enables a user to enter an inspirational phrase or an inspirational image and correlate that with the products in the product catalog. In some implementations, this is accomplished by having the descriptor generator parse the inspiration phrase / image and generate therefrom descriptors that are then correlated with the descriptors of the products in the product catalog..
[0014] An advantage of such techniques is the simplification of selecting a product by a user and also more easily presenting a user with similar or related products. In addition, the techniques enable a more thorough and rapid search by taking into account additional, oftensecondary, considerations such as matching also based on texture, patterns, aesthetics, etc. in addition to matching based on color. A user can be presented with an array of similar products and can more easily filter them than by using a conventional techniques
[0015] Other systems, devices, methods, features and advantages of the subject matter described herein will be or will become apparent to one with skill in the art upon examination of the following figures and detailed description. It is intended that all such additional systems, methods, features and advantages be included within this description, be within the scope of the subject matter described herein, and be protected by the accompanying claims. In no way should the features of the example embodiments be construed as limiting the appended claims, absent express recitation of those features in the claims.BRIEF DESCRIPTION OF FIGURES
[0016] The details of the subject matter set forth herein, both as to its structure and operation, may be apparent by study of the accompanying figures, in which like reference numerals refer to like parts. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the subject matter. Moreover, all illustrations are intended to convey concepts, where relative sizes, shapes and other detailed attributes may be illustrated schematically rather than literally or precisely.
[0017] FIG. 1 illustrates the components of the system for matching inspirations to products.
[0018] FIG. 2 illustrates an example flow chart of part of the method.
[0019] FIG. 3 illustrates an example flow chart of part of the method.
[0020] FIG. 4 illustrates a flow chart describing the method.
[0021] FIG. 5 illustrates a screen shot of an example of images used for inspiration.
[0022] FIG. 6. illustrates a screen shot of an example output of the system.
[0023] FIG. 7 illustrates a flow chart of an alternative example of the method.
[0024] FIG. 8 illustrates a computing device.DETAILED DESCRIPTION
[0025] Before the present subject matter is described in detail, it is to be understood that this disclosure is not limited to the particular embodiments described, as such may, of course, vary. The terminology used herein is for the purpose of describing particular embodimentsonly, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.
[0026] When making a decision to purchase an item (e.g., a flooring product), a buyer generally considers and evaluates numerous (e.g., hundreds or thousands) of pieces of information regarding various items. Wading through listing of hundreds / thousands of items can be time consuming, laborious, and may still not surface items that a buyer is actually looking for. In some instances, buyers may be looking for certain types of items that align with their preferences, but without knowing exactly the type of item that the buyer wants or needs.
[0027] The present disclosure describes a method and a system for enabling a user to provide an image(s) or a natural language phrase or sentence (e.g., an inspirational phrase) and then analyzing the chosen image(s) or the natural language phrase or sentence to determine and present (for display on an end-user / client device), the most relevant product choices available in the product catalog — taking into account various aspects or features of the chosen image or phrase.
[0028] FIG. 1 illustrates the components of the system for matching inspirations to products. The system includes a digital platform 310 which is connected over a network 390 to a client device 350.
[0029] The client device 350 can run an application 350-A. There can be multiple client devices 350, but for simplicity of illustration a single client device is illustrated. The user application 350-A can be an application provided by the digital platform 310 to the user device 350-A. The network 390 can be a wired network or a wireless (i.e., radio-frequency communications) network. The digital platform 310 can include at least a natural language processor 320, an image processor 322, a category assigning engine 324, a descriptor generator 106, a product catalog 104, one or more individual matching models 326-1, 326-2, etc., an overall match model 328, inspiration descriptors 110, and product descriptors 120. The client device 350 can be used to access the digital platform 310 by the application 350-A over the network 390. The client device 350 can transmit to the digital platform 310 an image or a phrase or both. The image can be an inspirational image. The phrase can be a search query or a prompt or an inspirational phrase or a combination. The image is processed by the image processor 322 to determine features of the image. The features of the image are submitted to the descriptor generator 106 which generates descriptors of the image based on the output of the image processor 322. The category assignment engine 324 assigns, to each descriptor, one or more categories. The natural language processor 320 translates the phrase to English, if the phrasewas not submitted in English. The natural language processor 320 parses the phrase and submits the results to the descriptor generator 106. The descriptor generator 106 generates the inspired descriptors 110, which are inspired from the phrase or the image. The descriptor generator 106 generates descriptors for each item of the product catalog 104. The category assigning engine 324 assigns a category to each descriptor of the parsed phrase or assigns a category to each descriptor of the image. In an example, the image contains a main color which is identified by the image processor 322 with a the example descriptor “deep red” and the category assigning engine 324 assigns the descriptor “deep red” to the color category. In an example, the descriptor generator 106 generates the descriptor “bamboo-like texture” based on the input phrase. The category assigning engine 324 assigns the descriptor “bamboo-like texture” to the category “texture.” Each of these engines and components are described in further detail with reference to FIGS. 2-7 below.
[0030] FIG. 2 illustrates an example system described in this disclosure. Although the following description is provided in the context of flooring products, one skilled in the art will appreciate that the below-described systems and techniques can be applied in the context of any type of item and need not be limited to flooring products. The system 100 includes a product catalog 104 which includes details about products provided on the digital platform 310. In an example, the product catalog 104 can be a database of flooring products. In an example, the details of the products can include colors, yarn types, pile types, patterns, textures, installation requirements, tile sizes, tile shapes, etc.
[0031] The digital platform 310 includes a descriptor generator 106 which can be a machine learning (ML) model trained to receive an image or features of an image identified by an image processor 322 and to generate descriptors of the image. The descriptor generator 106 can also receive an inspirational phrase or text supplemented by input from a natural language processor 320 and output one or more inspirational descriptors 110 related to the phrase. In an example, a user submits, from a user device 350 an inspirational image 102 or text 101 (the text may be an inspirational phrase or simply a search query) to the descriptor generator 106. The text may first be processed by a natural language processor (NLP) 320. The descriptor generator 106 produces a set of inspiration descriptors 110 for the input image 102 and / or text 101. The descriptor generator 106 can be a machine learning (ML) model, trained to identify words, phrases, or features identified in an image or a section of text.
[0032] In FIG. 2, five descriptors are shown - color, texture, pattern, aesthetic, and other - although any number of additional or other descriptors may be used. The descriptor generator106 is also used to parse the product catalog 104 to produce a set of product descriptors 120. In an example, prior to the product catalog 106 being opened to searches by user devices 350, one or more models (e.g., the descriptor generator) can be applied to the product catalog 106 including, for example, a database of images of products and product descriptions, to generate corresponding product descriptors 120. Generating product descriptors 120 for the product catalog 106 can occur before receiving the inspiration descriptors 110, the inspirational phrase, or the inspirational image. Each of the inspiration descriptors 110 and the product descriptors 120 can be assigned to one or more categories by, for example, the category assigning engine 324.
[0033] The category assigning engine 324 receives a descriptor and assigns to the descriptor one or more categories which can be used subsequently for a matching operation. In an example, an inspirational phrase “bamboo-like” can be assigned to multiple categories such as color, texture, and pattern. In an example the phrase “red” may be assigned to a single category such a color only. Each of the inspiration descriptors 110 is matched with the equivalent product descriptor 120 based on the category or type of descriptor. For instance, the inspiration descriptor 110 for color is compared with the product descriptor for color 120 to produce a color match 130. Similarly, the inspiration descriptor for texture is compared with the product descriptor for texture to produce a texture match 132. The comparison of the inspiration and product descriptors thus produces a set of matches: color match 130, texture match 132, pattern match 134, aesthetic match 136, and other match 138. The match need not be an exact match. For example, a match score can be a score between 0 and 1, based on a degree of match. An example color match score is the inverse of a Euclidian distance between color coordinates. An example texture match is a cosine similarity between the inspiration texture descriptors and the product texture descriptors.
[0034] Each of these matches (e.g., each of the match scores) is then taken into account to produce an overall match score 140 between the image 102 and each product of the product catalog 104 or between the phrase 101 and the each of the products. In an example, the overall match score is an average of the normalized individual match scores. In an example, the overall match score is a weighted sum of the normalized individual match scores by the weights based, for example, on the user’s inputs or on the previous search history, or on other factors. In an example, the overall match score is a geometric mean of the normalized individual match scores. Based on the overall match score 140, each flooring product of the product catalog 104 can be ranked. A selected number of the topmost ranked products can be presented to the userfor evaluation. In an alternative, any ranked product above a predetermined threshold value for the overall match score 140 may be presented to the user for evaluation.
[0035] In certain circumstances, each of the inspiration descriptors 110 and each of the product descriptors 120 may include sub-descriptors. For example, the inspiration descriptor 110 for color may comprise multiple segments of an image and the colors used for each of the segments may be taken into account when performing the matching with the product descriptor 120 for color. As an example, an image can be divided into multiple segments and the dominant colors in each of the segments. In another example, an image can be segmented by color and the first dominant color may be the color whose segment is the largest in area. When generating the product catalog 104, the various color schemes may already be included so that it may be possible to skip the step of submitting the product catalog images to the descriptor generator 106 and have the product catalog 104 simply share the product descriptors 120 already stored rather than have them generated anew.
[0036] Similarly, the product catalog or product database may include the relevant criteria needed to perform the matching with the inspiration descriptors. The relevant product information may include color schemes, textures, patterns, and aesthetic criteria in addition to data and measurements related to the flooring product. Examples of the latter include location and inventory data, testing data, sizing, measured data, and the like.
[0037] Descriptor generator model
[0038] A descriptor generator model 106, which can be a machine learning model, is trained to create a set of descriptors when a phrase 101 (e.g., an inspirational phrase or a simple search query) or an image 102 (e.g., an inspirational image) is submitted to it. The model is trained using training data of many images and includes detailed sub-models that may segment the image into segments and can be used to determine the most relevant colors in each of the segments. Analogously, training a language model calls for, in general, a large amount of text for training. In addition, the descriptor generator may identify, for each segment of an image, color pairs, color triplets, color quadruplets, etc., or may determine that one particular color is the primary (most important) color of the image segment and may also determine which other colors are of secondary importance and also the ratio of colors within the image segment. An image may be segmented based on any aspect of the image such as, for example, by location within the image, pattern, texture, color, by objects within the image, by object location, or by other aspects.
[0039] The descriptor generator model 106 may also determine other descriptors of the phrase 101 or of the image 102 (or of a product from the product catalog 104) such as, for example, pattern, texture, and aesthetic descriptors. Training data for the descriptor generator model 106 may include many images outside of the product catalog and may call for fine tuning of the model so that it produces accurate descriptors that a human can understand and which also align with human-assigned labels or descriptors. For example, an aesthetic descriptor may be a word (e.g., “bold”, “striking”) or words which people may use to describe the pattern or the flooring product. When using such a trained model, images of the flooring product can be fed into the model to produce a set of such descriptors for every flooring product in a catalog. It is also possible to include additional information related to each of the flooring products in a database and this information does not necessarily need to be generated by the trained model.
[0040] In an example, a descriptor generator model may be trained, tested, and validated by using many images available publicly or by a combination of proprietary images along with publicly available images. Training a model may involve dividing the available data into two, three or more groups. The model may be trained on a first fraction of the data (e.g., 80% of the images) called the training data and then maybe tested or validated on the remaining data - the withheld fraction (e.g., 20%). An alternative method for training the model involves k-fold cross validation in which all the data is divided into k groups and each group is withheld for testing after the model is trained on the remaining k - 1 groups. A final model involves comparing the results of each of the k trained models and integrating them together. In an example, the parameters from each of the k trained models may be averaged together. In another example, the parameters from the best performing model are retained and the other k- 1 models perform within a predesignated range. Other methods of training the descriptor generator model may also be employed.
[0041] FIG. 3 illustrates an example flow chart of part of the method. This example includes a method for sequential texture, pattern, and color matching. In this example, unlike the example from FIG. 2 above, only texture, pattern, and color matching are performed. In this example, the first match is a texture match 132, followed by a pattern match 134, and then by a color match 130. These three matches are then used to determine the overall match score 140. In an example, the texture matching 132 is performed first to produce a first set of matched products. The pattern matching 134 is performed second and the matched products are limited to those which were successfully matched in the texture matching 132. Then the color matching 130 is performed and the matched products are limited to those which were successfullymatched in the previous texture and color matching steps. In an example, a product descriptor having a match score above a certain threshold value with an inspiration descriptor can be successfully matched. In an example, a number of products can be displayed up to a cutoff number, ranked by, for example, an individual match score or by the overall match score.
[0042] FIG. 4 illustrates a flow chart of the method. The method 400 starts with collating a collection of product images, descriptors, and metadata at step 402. The product catalog can include images of product, physical data of the product (e.g., shape, dimensions, color, yam type, pile type, etc.), meta-data (e.g., current inventory, delivery time, inventory locations, installation methods, etc.) The descriptor generator 106 may be applied to each product of the product catalog 104 to generate product descriptors 120. For example, the descriptor generator may be applied to each of the images and the descriptions in the product catalog at step 404 to produce the product descriptors 120. In addition, the descriptor generator may be trained using training images and training descriptors prior to being used to generate the product descriptors of the product catalog.
[0043] A user may submit a phrase (e.g., an inspirational phrase or a search query) at step 406a. The natural language processor (NLP) 320 may be applied to the user-provided phrase 101 at step 408a to extract context and some meaning. In an example, the NLP may look for a quick match with product descriptors from the product catalog. The descriptor generator model 106 may be applied to the parsed phrase from step 408a at step 410a and may generate inspiration descriptors 110 based on the user-supplied phrase. An alternative path is illustrated in that a user may enter an image or images at step 406b . An example set of inspirational images is provided in FIG. 5. The example set of inspirational images shows various designs and various color schemes which the user liked. Image processing may occur on the image at step 408b, such as, for example, segmenting the image into segments based on location, color, pattern, etc. or based on identified objects in the image. At step 410b, the descriptor generator model 106 is applied to the processed image to produce inspirational descriptors 110. Both the phrase pathway (e.g., steps 406a, 408a, and 410a) and the image pathway (e.g., steps 406b, 408b, and 410b) can be performed independently but also together. If a user enters only a phrase, the only the phrase pathway may be taken. If the user enters only images, then only the image pathway may be taken. If the user enters both a phrase and an image then both pathways may be taken. In addition, the method may be iterated so that first the phrase pathway is taken, and then an image is added so that the image pathway is taken. When the descriptor generator produces the descriptors — either the product descriptors or the inspirational descriptors, at leastone category is assigned to each of the descriptors. In an example, the inspirational image may have red as a prominent color so that the descriptor “red” is assigned the category color. In an example, the descriptor may have a square pattern so that the descriptor “square” may be assigned the category pattern.
[0044] At step 412, each of image / phrase / inspiration descriptors 110 are matched with each of the product descriptors 120. The matching can include matching based on a category of the descriptor. In an example, product has several descriptors each with several categories. For example, the product may have the descriptor “red” assigned to the category color and “black” also assigned to the category color. The product may have the descriptor “bricks” to the category pattern and may have the descriptor “wood-like” to the category texture.
[0045] At step 414, the results of the matchings of each of the image descriptors 110 with each of the product descriptors 120 based on category are fused into an overall match (e.g., calculating an overall match score) for each of the products in the product catalog 104. Then the user may be presented with a list of products at step 416, ranked by the overall match score between each of the image / phrase / inspiration descriptors 110 with each of the product descriptors 120. An example output of the method is shown in FIG. 6. The example output 600 includes a sample product 602 based on the overall match score of the inspirational descriptors and the product descriptors. The example output includes not just a proposed sample product 602 based on the overall match score but also some variations of shapes and sizes 604 of the product, variation of color schemes 606, and variations on yam types 608. Other variations of the product can also be provided such as installation criteria, delivery date, additional sample products with lower overall match scores, etc..
[0046] FIG. 7 illustrates an example method 700 for analyzing a phrase to return a list of the information relevant to a product. A phrase is entered and translated into English, if needed, at step 702. Natural language processing (NLP) can be used to translate from a first language to a target language. In the examples provided, the example product catalog has been assembled in English, so that the translating from another language to English is most efficient, but other options also exist. In an example, a trained, large language model, is used to identify the first language, to parse the meaning of the entered phrase, and to generate an English equivalent phrase. The translation can take place by a trained ML model or by a translation engine.
[0047] At step 704, text corrections and keyword replacements may be applied to the phrase. In an example, frequently mis-typed or mis-spelled words may be corrected. In an example, a phrase such as “a primary color of red with to secondary colors black and white”would be corrected as “a primary color of red with two secondary colors black and white.” In an example, a keyword “six-sided” may be replaced with “hexagonal” since the second term is more standard for a given product catalog. At step 706, regional specific references may be applied (e.g. “colour” or “grey” for British English may be converted to “color” or “gray” for American English) or a specific language model for that region may be selected and used. For example “with a bat theme” may produce a result related to the flying mammal in, say, Singapore, may produce a cricket bat in a British location, and may produce a baseball bat in an American location. Based on these results, a cognitive search query is created at step 708. The cognitive search query can include identifying some of the parsed phrase to a search query relevant for the product catalog. The cognitive search query may include certain phrases as being assigned to certain categories. For example, the inspirational phrase may include both “vibrant color” and “to install before the end of summer”. The “vibrant color” phrase may be assigned to the category of color and also to the category of aesthetic. The “install before the end of summer” coupled with the location may be under the category of metadata for installation and also under the inventory near that location.
[0048] At step 710, documents in a database (e.g., a product catalog) are searched and the most relevant text or the most relevant documents are selected. Again, the search can occur by matching the various aspects based on category and then determining an overall match score based on the match scores of the individual categories. At optional step 714, the documents may be translated, if necessary, and the content filtered based on various constraints. For example, if the user is interested in only thick carpet, the filter may constrain the results to be only those products with a carpet pile greater than 1 cm in height. In another example, if the user specifies requiring installation before a certain date, then only existing products already in inventory may be provided to the user. At step 716 a full response is generated and sent to the user device. The full response may include a product sheet with variations as shown in an example of FIG. 6. At step 718, the response and additional information or links to additional resources are provided to the user device. For example, the nearby inventory of already manufactured product may be provided to the user device including likely installation deadlines. In an example in which a user prefers to have the product manufactured, estimated delivery dates could be included. In an example in which a user prefers to submit an additional or slightly modified inspirational phrase or an additional phrase with the same images or an additional images, a follow-on or iterative execution of the method may result.
[0049] Images
[0050] A customer interested in purchasing an item (e.g., a flooring product) can select an image or images or a phrase or set of words or some combination of these several inspirations. For instance, a customer / user may input as an inspiration Japanese woodblock prints of iris flowers in bloom and further enter the phrase “relaxing, calm, pale purple and ivory with overtones of green as in these pictures”. The image(s) and the phrase(s) are fed into the descriptor generator model 106 which then may produce a series of descriptors. In this example, the inspiration 102 includes both an image and a phrase; however, a user may submit a single image, multiple images, a single phrase, or a mixture of images and phrases. Thus natural language processing may be used to parse the entered phrase and image processing may be used to analyze the image or images. In an example in which the phrase refers to the images (e.g., “like this picture of the mountains but lighter green”). Thus, the image analysis module and the natural language processing module may integrate information referred to in the other submission, such as when the phrase refers to an image. In the example above, the images may be analyzed to determine which of the colors are pale purple, ivory, and green, and then color descriptors produced which are associated with those color terms.
[0051] The descriptors produced by the model given the input of an image and a phrase can be matched with the catalog descriptors already produced for the catalog of flooring product. The catalog descriptors may be generated by feeding images of the catalog of product into the descriptor generator model, but the catalog descriptors may also be generated by other methods and may also include other data. The descriptors from the image and the descriptors from the catalog may be matched. This matching will result in a series of classifications of the inspiration with different flooring products on different axes or aspects. For example, an image of pine tree bark may match a hexagonal pattern to a certain percentage but may match a particular type of wooden flooring to a higher percentage. In addition the models may produce a match score based on texture and a match or matches based on color with an associated probability or percentage score. In an example, the associated probability or percentage can be the match score. By taking into account the various individual match scores generated, an overall match score may be produced for each flooring product in the catalog with the user’s chosen inspiration.
[0052] As a first step of analyzing the existing products, an image or a set of images of the material can be taken and submitted for analysis. This analysis may extract from the image of the product such things as a color element or multiple color elements, a texture, a pattern, andthe like. It is also possible to have already determined this information without an explicit image of the product in question. Next, the model builder associates with each item in the product catalog a description, a pattern type, a texture type, a classification according to aesthetic qualities, and other information. In this disclosure, aesthetic qualities may refer to descriptive terms used by humans to describe human feelings when seeing an image or a piece of artwork or feeling a texture or the like. For instance, an image may be viewed as “bold and brazen” and another image may be viewed as “calm and comforting” yet another image may be viewed as “thought provoking and uplifting”. A customer or user may want to have some of these emotional / aesthetic qualities associated with their product choices. Of course, truly aesthetic choices are individualistic, but by incorporating a large enough number of human- assigned attributes to an image, it is possible to say with some reasonable probability whether any general user would associate certain feelings or certain aesthetic qualities with a particular image.
[0053] The model can output a variety of classifications along with a probability associated with each of the classifications and a classification type, if there is more than one classification type (e.g., a unified model for both pattern and texture). For instance, the model receives an image and determines that the image belongs to the category “pastel floral” with probability 12% and category “Chinese landscape painting” with probability 8%, and category “calming” with probability 7.5%. In a subsequent matching step, these classifications and probabilities would be used to present to a user those items which are most closely associated with any of these three categories. If other items from the product catalog are either in categories the user is not interested in or match to the specified categories below a certain probability threshold, say below 5% probability, then those items may not be presented to a user. This probability threshold may be changed or tuned as needed and may also depend on which descriptor is being classified. For instance, the probability threshold for a pattern descriptor may be 35%, but the probability threshold for a texture or an aesthetic quality may be 22%.
[0054] Natural Language Processing
[0055] A phrase may be entered as an inspiration or a request. This phrase may be parsed by a natural language processor (NLP).
[0056] The user may type in a phrase, a search query, or an inspiration, which may be independent of any images or it may relate to an image or images also submitted. Some preprocessing may be performed on the phrase such as: translation to English from another language, check and correct for common typographical errors, check and replace commonbusiness terms, adjust for region specific references. The resulting pre-processed phrase may be submitted as a search query to an Al-assisted search tool which uses natural language processing. Using natural language processing enables a better and more precise search to be conducted and also may better engage users.
[0057] Such a search may first pre-process the received phrase, as noted above, and then perform some contextual analysis on the phrase. Such a contextual analysis may include identifying an emotional or aesthetic terms with the phrase (or with the general search query). The customer database of products may include product data such as specification data, product number, product name, color data, texture, pattern, and the like.
[0058] In an example, a system or method using natural language processing (NLP) may include the following steps. The system may receive the phrase and modify the phrase to enhance its grammatical structure. The system parses the modified phrase using code logic designed to identify patterns containing information relevant to a product in the catalog / database such as, for example, a selling style number or a selling style name. The system may then search a cached list of the product information (e.g., selling style numbers, selling style names).
[0059] The next step can involve one of two paths: knowledge retrieval or a direct match. In the direct match path, if a selling style number or a selling style name is detected in the query phrase, an additional search is performed on data specific to the identified product.
[0060] In the knowledge retrieval path, the query phrase may be incorporated in a search which may be performed over the product catalog and other related documents. This approach is also called the retrieval augmented generation (RAG) approach. A next step for the RAG approach may include searching documents, evaluating the document contents, and extracting relevant data. This data may then be appended with the query and also with custom instructional prompts.
[0061] In a subsequent step, the refined data may be sent to another model for further processing and analysis (e.g., the OpenAI GPT Model). The response to the query may then be shaped and defined via additional logic steps.
[0062] In the Natural Language Processing (NLP) pipeline, Retrieval-Augmented Generation (RAG) may be integrated to enhance the capabilities of Large Language Models (LLMs). This integration involves augmenting the LLM with an additional layer of data retrieval, utilizing special software (e.g., Azure Al Search) as the primary information retrieval system. In such an example, a database of business documents is systematically processed andindexed prior to the query being received, making the information accessible quickly and easily through a specialized search engine (e.g., Azure Al Search). This setup allows for precise control over the grounding data employed by the LLM during response generation, ensuring that the output (e.g., the answer to the query) is anchored to enterprise-specific content derived from the indexed documents.
[0063] Upon receiving a query, the NLP service may extract relevant information by querying such a search service (e.g., the Azure Al). This content is then concatenated with the initial query and any guiding prompts, forming an enriched input for the LLM. The LLM, thus, generates responses that are not only contextually relevant but also grounded in the enterprisespecific data provided.
[0064] This process of retrieval and information augmentation may occur faster than if the product data had not been prepared beforehand. Concurrently, the system ensures scalability and timeliness in the indexing process, continuously updating the information retrieval system with newly added content. In an example, a remote system (e.g., Azure, AWS, software as a service) may provide the image processing, the language processing, or the descriptor generator for the product catalog. In an example, the descriptor generator may be a pre-trained ML model applied to the product catalog. In an example, the descriptor generator may be trained with some data from the product catalog and also trained with data outside of the product catalog.
[0065] Color
[0066] An image or images selected by a user may be analyzed to determine the main colors of the image or to determine the main colors of segments (sub-sections) of the image. The images may be segmented by color as well as by location or by object. If, for instance, the majority of the image is black or dark grey, then the image may be assigned “black” as the color descriptor for the primary color. If the image is evenly divided between black and yellow then both black and yellow may be returned as primary color descriptors.
[0067] A color descriptor may be quantified in a variety of fashions. For instance, color can be determined from the RGB framework or from one of the CIE color coordinate spaces, or from some other framework (e.g. Hex, Pantone’s PMS, Benjamin Moore’s paint swatches, yam poms, etc.). Once the color descriptor or descriptors have been determined, a range of colors around the individual color descriptors may also be included when searching the product catalog for a match. Color may also be determined in various segments of the image and then combined in various combinations.
[0068] Based on the output of the first model a product or multiple products may be returned. The returned products may also need to match the desired products on color (in addition to pattern or texture) to within the threshold amount for each. For instance if the model determines that the image includes black, red, and white with a very glossy texture, then products with a high probability of being assigned to a glossy texture class and including the colors black, red, and white may be returned.
[0069] In an example, the RGB color values from the segmented image may match the RGB values of the products in the vision analysis table to within a particular variance or range. This variance may be ±15%, ±10%, ±5%, or even ±1%. In an example, the absolute value of the differences in the image RGB values with the color values of a flooring product may determine whether the image color descriptor matches a product color descriptor. Thus if the absolute value of the difference in the RGB values is less than this threshold value in variance, then the product is considered a match (e.g., the match score exceeds the threshold value). Such a method may be applied to all three of the RBG color values. For instance if the variance is 5%, then |(Rimage - Rcataiog)| / Rimage < 5%, and similarly for Blue and Green color coordinates: |(Gimg - Gcat)| / Gimg < 5% and |(Bimg—Bcat)| / Bimg < 5%).
[0070] In another example, each of the RGB color values for the image (Ro, Go, Bo) could have independent variances (AR, AG, AB). In such an instance, the products returned would have RGB values with R = Ro ± AR, G = Go ± AG, and B = Bo ± AB. In yet another example, the color coordinates in the CIE color space could be used and a radius around those color coordinates could determine the product space for products to be presented to a user based on a color match. In this example, a product matches the image color if (R,G,B) is within (Ro, Go ,Bo) ± AColor, where AColor is the radius around the image color in color space.
[0071] The color descriptors may also include not just a primary color but also include additional colors from the image such as a secondary color, a tertiary color, and so on. The color descriptor may also include information about the ratio of the colors, such as if a fraction of the image contains red and another fraction of the image contains yellow, then the ratio of these two fractions may be returned as an additional color descriptor and may be matched to existing product from the product catalog.
[0072] Combining the NLP with the color match is also possible. The input text is parsed to determine if the user wishes to alter an originally submitted color or to select a particular color from an accompanying image. This method can be applied to one or more of the color coordinates (e.g., RGB values) or applied to one or more of the colors selected as part of thecolor matching. The base color coordinates for the match can be changed or modified accordingly and the color match may proceed with a modified set of base color coordinates or a modified range of color coordinates for the matching.
[0073] Pattern + Texture
[0074] Pattern and texture classification may operate in a similar manner. An image is analyzed, for instance, by applying various filters over the image to determine whether there are any significant features such as many straight up and down lines, many straight left and right lines, diagonal lines, or diagonal lines in two orthogonal directions, or one big wavy line, or many wavy lines, or many circles, etc. Pattern recognition and determining the most important features of an image are well known in the art of image analysis. In the specific instance of matching to a particular product, product knowledge may be included or incorporated by limiting the number of descriptors or by preferring particular features over other features in images.
[0075] Once the features of the image have been identified, a classification model may be applied to the identified image features. Such a model, which may be hosted at a remote site or by a third party (e.g. Azure), may use, for example, a trained convolutional neural network to classify the pattern or texture as belonging to a particular category and also assign a confidence value (e.g., a probability) that the image feature belongs with that category. The training process for the classification model may involve highly curated imagery with meticulously chosen attributes, ensuring precision in identifying specific patterns (e.g., checkerboard or herringbone) and textures (e.g., wire brushed, brick, rustic) from images to meet specific requirements.
[0076] Aesthetics
[0077] Products can also be matched based on data derived from a design aesthetic classification of products by designers. In an implementation, the system may maintain the following statistics for a variety of design aesthetic classes: Design Aesthetic; Pattern; Average, Min, Max, and Standard Deviation of Probability of the Pattern for the Design Aesthetic.
[0078] In a particular example aesthetics criteria for products may be determined by building a database of human-assigned aesthetics for particular products. In an example, designers may be polled for a variety of products and the aesthetic or emotional words or descriptors they associate with a particular product. By polling enough designers, aesthetic or emotional qualities associated with a distribution of people may result in an image being assigned a descriptor approximately of what an average person may feel or sense when seeinga particular product. Such a classification may take into account all of the above-described dimensions: color, texture, and pattern, as well as other information.
[0079] Fusion of model s / Overall matching
[0080] In an example, the system may comprise four specialized models: pattern assignment model, texture assignment model, color model, and design aesthetics model. (Of course, additional models may also be included, but in this example, only four are considered.) Once the images or inspirational phrases have been parsed and the inspirations assigned to descriptors based on these four models, the descriptors are compared with the descriptors of the existing products in a product catalog. This procedure is sometimes called fusing models or blending models or determining an overall match score. The goal is to determine a comprehensive or overall match score between a user-selected image or phrase and a list of products to be displayed to the user which are best matched to the image or to the phrase, or to both.
[0081] In an example, the various image descriptors may be considered to match a particular product if the probability of assignment to the same class of the image is within a variance or standard deviation of the probability that the product is also assigned to that class. In short, if the product from the product catalog is classified to a particular class with an average probability, Pavg, then the image is a good match with the product if the image’s probability, P, of being classified to the same class falls within a standard deviation (e.g., a threshold value) of the product probability. Mathematically if this statement (Pavg - StdDev) < P < (Pavg + StdDev) is true, then the product is considered to match the image and the product may be displayed to the user.
[0082] In another example, the image / inspiration descriptors could be encoded in a vector representation including all the information described above. Similarly, the products in the product catalog could also have vector representations created for each of the products. A match could occur by performing a cosine similarity calculation using these two vectors and returning as a match those products which have the highest cosine similarity a list of selected number products ranked by the highest cosine similarity (e.g., the highest match scores). Other similarity metrics or similarity calculations such as an inverse Euclidian distance or an inverse Minkowski distance might also be employed to quantify how well a product matches an image or a search phrase (i.e., using a similarity metric as a match score).
[0083] Implementation Examples and Details
[0084] In an example, a system and method may be implemented through software as a service. A remote server may already have the trained models available and a user may access the server through a local interface, connecting to the remote server over a network. In a specific example, an outside service provider such as the Azure Cognitive Search Service API may implement the method and the Azure OpenAI GPT model may impose limits on, for example, the tokens used or the number of tokens transmitted. One example of a method for connecting to a remote server uses Javascript Object Notation (JSON) web tokens (JWTs) to exchange information. The user’s data, image, search phrase or inspirational phrase may be encoded in a JSON web token (JWT) which may include a header, a payload / body, and a cryptographic signature for validating the security of the web token. Other methods for securing communication between a user’s electronic device and a remote server may also be employed.
[0085] In one implementation, an image may be received by the system and analyzed, as detailed elsewhere in this disclosure. A dominant color (and / or additional secondary or tertiary colors) of the image may be determined. The determined color(s) may be compared with the determined color(s) (dominant and / or secondary and tertiary colors) for every product in the catalog. Each product may have already been classified by a pattern and texture model or by a pattern model and also by a texture model. Depending on the model used, such a classification may include not just an assigned class but also an associated probability that the assigned class is correct.
[0086] In an example, a user may submit an unusual brick pattern as their image inspiration. The pattern and texture model may detect the diagonal lines of the brick pattern and return as the first class: “herring bone”, probability 35%, classification type: pattern. The model may also output a second class: “rough brick”, probability 25%, classification type: texture. Thus the model returns a set of classes associated with the image along with a probability associated with each class and a classification type. In an implementation which uses a single pattern detection model, the classification type may be left out. In an alternative implementation which uses, e.g., two models - a pattern detection model and a texture detection model, each of the models would return a list of classes and a probabilities associated with each class assignment for each of the two models. In a subsequent step the results of these two models may be integrated or fused together.
[0087] In another example implementation, the image may be matched to a product in the product catalog based on a probability being greater than a certain percent for a certain classification. For example, a product may be matched to the image when the classification is texture and the probability of the texture classification is greater than or equal to 35%. A subsequent limiting of the number of matches could occur based on, e.g., a color matching of those products which met the first criteria, namely the texture classification probability above a set threshold. For this second matching step for, e.g., color, the RGB color values for a product from the product catalog would be evaluated against a color threshold (e.g., 5.8% ) of the RGB color values from the image, or from a segment of the image, to be considered a match. Thus one method for providing a list of products which match the user- selected image or the user-entered phrase is to perform a sequence of comparisons or matches with each step in the sequence further limiting the number products from the product catalog to be presented to the user.
[0088] In yet another example implementation, the image may be segmented and a primary and a secondary color may be determined which would then be matched with a primary and a secondary color for each product in the product catalog to within a color threshold. Then a second matching step might occur for pattern matching and a third matching step might occur for aesthetic matching.
[0089] It is also possible to include particularly relevant data by merging or including information from both images and from a phrase. For example, if a user enters the phrase “fire- resistant tile for a camp office” then the product list provided to the user would be limited to those types of product above a certain fire resistance level.
[0090] Achieving a high-quality match (i.e., a high overall match score or high individual match scores) involves a unique process that varies depending on the specific requirements of the consuming user or application. For example, one user might need a more practical result as they are designing primary or secondary school cafeterias, whereas another user may be designing for a glamorous hotel ballroom. This distinction is categorized into the following groups: Standard, Probability, First 3 Colors, and Design Aesthetic.
[0091] Example 1 Sequential Texture, Pattern, and Color Matching
[0092] In this example, each image submitted by a user may be segmented by an object detection process and then each segment may be classified by the custom pattern and texture models. The result is a list of classification types for pattern and a list of classification typesfor texture, such as "Striped," "Marbled," "Wood Grain," etc. presented in conjunction with unique probability / confidence level for each classification.
[0093] Receiving the image initiates a primary search for a match in the top classification. In an example, an image segment may be classified as having a "Distressed" texture with high probability. The objective is then to identify a congruent match within the catalog which are also in the "Distressed" category. In the absence of a direct match at this level, the algorithm may shift to the second strongest classification, and may require a minimum matching probability (e.g., 35%) or it returns no results.
[0094] This sequential process is replicated for pattern results. Subsequently, a query for color matching is conducted as described elsewhere in this disclosure.
[0095] In an example, the color matching may occur using any of the identified most dominant colors in any order. For example the color matching may occur using the first, the third, and the fifth most dominant colors extracted from the image. In another example, only the first and second most dominant colors of the image would be used. In yet another example, the two most dominant colors from the two largest image segments could be used, for four colors in total.
[0096] Example 2: Pattern and Texture Matching Only
[0097] Image segments undergo the same image processing to identify features and to classify the features as belonging to a particular category. The results are used to query the database of products for matching products. The algorithm may begin by seeking a match in the highest ranked classification (i.e., the classification with the highest confidence level). In an example, for a successful match, the product's probability / confidence level must be within a ±5% variance of the image’s classification probability for the top classification. This criterion may be applied to both the texture and the pattern classifications. In other examples, different values may be applied to the variance as well as to the number of top classifications used.
[0098] Importantly, color matching may not be performed in this query, as the primary emphasis may be on achieving a match based on the top classification with a strict probability threshold without consideration of color. Such a search query may occur when a designer is looking for a very close match to a particular pattern but does not care at all about the color.
[0099] Example 3: Top 3 Colors Only
[0100] Image segments may be processed through a color detection algorithm. The color detection algorithm identifies the main colors in each of the image segments. The algorithm may rank each of the colors in each of the segments. For example the color with the most number of pixels could be ranked first. Alternatively a color with the largest continuous andcompact feature could be assigned the top rank. The results from this process may then be utilized to conduct a search within the product database, which also contains color data for each product. This search may apply the color matching criteria detailed elsewhere, focusing specifically on the colors identified by the algorithm. Naturally the search may be limited further such as only using the first, second, and third most dominant colors in the image or only the two most dominant colors or the first, third, and fifth most dominant colors may be used.
[0101] The probability score associated with each query result may be computed based on the absolute difference in the RGB (Red, Green, Blue) values of the product and the received image. This probability score may be calculated by subtracting the individual RGB values of the image from the individual RGB values of the product and taking the absolute value of each of these differences. This method provides a quantitative measure of color similarity between the image segments and the products in the database. In an example, a permitted tolerance range of ±15 RGB units may be used which can be written mathematically as: { |RProduct - Rimagel ± |Gproduct>Gimage| ± |Bproduct - Bimage| } < 15.
[0102] Example 4: Design Aesthetic
[0103] For Design Aesthetic output, the catalog products may be matched to the image or the parsed phrase based on statistical data derived from the Design Aesthetic curated classification of sample of products. In the instance of flooring products, commercial flooring designers may rank all the designs in the product catalog. In another example, an automated system may be applied to all the products in the catalog.
[0104] The system may maintain the following stats for each design aesthetic classification of the image: design aesthetic category; pattern category; average, min, max, and standard deviation of probability (confidence level) of the pattern for the design aesthetic. The matching criteria are discussed elsewhere in this disclosure.
[0105] The product catalog may also include much information not necessarily required by most users but of great use to a small minority of users. For instance, the product catalog may include a database of product details including, for the example of flooring, a thickness, a firerating, a carpet pile size, a thread density, inventory data, a closest location, a time to custommake a product, testing data, and the like.
[0106] Computing Device or System
[0107] FIG. 8 is block diagram of an example computer system 800 that can be used to perform operations described above. The system 800 includes a processor 810, a memory 820, a storage device 830, and an input / output device 840. Each of the components 810, 820, 830,and 840 can be interconnected, for example, using a system bus 850. The processor 810 is capable of processing instructions for execution within the system 800. In some implementations, the processor 810 is a single-threaded processor. In another implementation, the processor 810 is a multi -threaded processor. The processor 810 is capable of processing instructions stored in the memory 820 or on the storage device 830.
[0108] The memory 820 stores information within the system 800. In one implementation, the memory 820 is a computer-readable medium. In some implementations, the memory 820 is a volatile memory unit. In another implementation, the memory 820 is a non-volatile memory unit.
[0109] The storage device 830 is capable of providing mass storage for the system 800. In some implementations, the storage device 830 is a computer-readable medium. In various different implementations, the storage device 830 can include, for example, a hard disk device, an optical disk device, a storage device that is shared over a network by multiple computing devices (e.g., a cloud storage device), or some other large capacity storage device.
[0110] The input / output device 840 provides input / output operations for the system 800. In some implementations, the input / output device 840 can include one or more of a network interface devices (e.g., an Ethernet card), a serial communication device (e.g., a USB port), and / or a wireless interface device (e.g., an 802.11 card). In another implementation, the input / output device can include driver devices configured to receive input data and send output data to peripheral devices 860, e.g., keyboard, printer and display devices. Other implementations, however, can also be used, such as mobile computing devices, mobile communication devices, set-top box television client devices, etc.
[0111] Although an example processing system has been described in FIG. 8, implementations of the subject matter and the functional operations described in this specification can be implemented in other types of digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them.
[0112] Embodiments of the subj ect matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions, encoded on computer storage media (or medium) for execution by, or to control the operation of, data processing apparatus.Alternatively, or in addition, the program instructions can be encoded on an artificially- generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium can be, or be included in, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be a source or destination of computer program instructions encoded in an artificially-generated propagated signal. The computer storage medium can also be, or be included in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0113] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0114] The term “data processing apparatus” encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, a system on a chip, or multiple ones, or combinations, of the foregoing. The apparatus can include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can also include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The apparatus and execution environment can realize various different computing model infrastructures, such as web services, distributed computing and grid computing infrastructures.
[0115] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub-programs, or portions of code). A computer program can bedeployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0116] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0117] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing actions in accordance with instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few. Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and optical disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0118] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., an LCD (liquid crystal display) monitor or a light emitting diode (LED) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documentsfrom a device that is used by the user; for example, by sending web pages to a web browser on a user’s client device in response to requests received from the web browser.
[0119] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a frontend component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), an internetwork (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0120] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e.g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[0121] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0122] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirableresults. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0123] It should be noted that all features, elements, components, functions, and steps described with respect to any embodiment provided herein are intended to be freely combinable and substitutable with those from any other embodiment. If a certain feature, element, component, function, or step is described with respect to only one embodiment, then it should be understood that that feature, element, component, function, or step can be used with every other embodiment described herein unless explicitly stated otherwise. This paragraph therefore serves as antecedent basis and written support for the introduction of claims, at any time, that combine features, elements, components, functions, and steps from different embodiments, or that substitute features, elements, components, functions, and steps from one embodiment with those of another, even if the following description does not explicitly state, in a particular instance, that such combinations or substitutions are possible. It is explicitly acknowledged that express recitation of every possible combination and substitution is overly burdensome, especially given that the permissibility of each and every such combination and substitution will be readily recognized by those of ordinary skill in the art.
[0124] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise.
[0125] While the embodiments are susceptible to various modifications and alternative forms, specific examples thereof have been shown in the drawings and are herein described in detail. It should be understood, however, that these embodiments are not to be limited to the particular form disclosed, but to the contrary, these embodiments are to cover all modifications, equivalents, and alternatives falling within the spirit of the disclosure. Furthermore, any features, functions, steps, or elements of the embodiments may be recited in or added to the claims, as well as negative limitations that define the inventive scope of the claims by features, functions, steps, or elements that are not within that scope.
Claims
CLAIMSWHAT IS CLAIMED IS1. A method comprising: generating a plurality of catalog descriptors associated with each of a plurality of items; receiving an image or a phrase; extracting, from the image or the phrase, at least one descriptor; ranking each of the plurality of items based on a match score between the at least one descriptor and the plurality of catalog descriptors; and providing, for display on a client device, a set of items from among the plurality of items based on the ranked matching.
2. The method of claim 1, wherein extracting the at least one descriptor from the image comprises processing the image and sending the processed image to a descriptor generator to extract an image descriptor and where extracting the at least one descriptor from the phrase comprises processing the phrase using natural language processing, and sending the processed phrase to the descriptor generator to extract a phrase descriptor.
3. The method of claims 1 or 2, wherein the match score between the at least one descriptor and the plurality of catalog descriptors comprises an image match score between an image descriptor and the plurality of catalog descriptors or a phrase match score between the phrase descriptor and the plurality of catalog descriptors.
4. The method of any of claims 1-3, wherein the plurality of catalog descriptors comprises, for each of the plurality of items, at least a catalog color descriptor, a catalog texture descriptor, a catalog pattern descriptor, and a catalog aesthetic descriptor.
5. The method of any of claims 1-4, wherein extracting the at least one descriptor comprises: extracting a color descriptor; extracting a texture descriptor; extracting a pattern descriptor; extracting an aesthetic descriptor; andranking each of the plurality of items comprises: calculating a color match score between the color descriptor and the catalog color descriptor; calculating a texture match score between the texture descriptor and the catalog texture descriptor; calculating a pattern match score between the pattern descriptor and the catalog pattern descriptor; calculating an aesthetic match score between the aesthetic descriptor and the catalog aesthetic descriptor; and fusing the color match score, the texture match score, the pattern match score, and the aesthetic match score to determine an overall match score between the received image and each of the plurality of items.
6. The method of any of claims 1-5, wherein each of the color match score, the texture match score, the pattern match score, and the aesthetic match score is based on a similarity calculation between an encoding of each of the catalog descriptors and an encoding of the at least one descriptor.
7. The method of claim any of claims 1-5, wherein calculating any of the texture match score, the pattern match score, or the aesthetic match score is based on a classification of the image as belonging to one of a set of textures, a set of patterns, or a set of aesthetic criteria and a probability associated with the classification.
8. The method of any of claims 1-7, wherein extracting at least one descriptor from the received image further comprises, segmenting the received image into segmented images and processing each of the segmented images to extract at least one descriptor from each of the segmented images.
9. The method of any of claims 1-8, wherein the image color descriptor and the catalog color descriptor match when an absolute value of the percentage difference between each of the color coordinates of the image color descriptor and the corresponding color coordinates for the catalog color descriptor are within a percentage threshold value of each other.
10. The method of any of claims 1-9, wherein the image color descriptor and the catalog color descriptor use color coordinates, and are considered well matched if a distancein color coordinate space between the image color descriptor and the catalog color descriptor is less than a threshold value.
11. A system comprising: at least one processor; a computer memory; and a database, wherein the at least one processor receives computer program instructions from the computer memory and accesses the database to perform a method comprising: generating a plurality of catalog descriptors associated with each of a plurality of items and stored in the database; receiving an image or a phrase; extracting, from the image or the phrase , at least one descriptor; ranking each of the plurality of items based on a match score between the at least one descriptor and the plurality of catalog descriptors; and providing, for display on a client device, a set of items from among the plurality of items based on the ranked matching.
12. The system of claim 11, wherein extracting the at least one descriptor from the image comprises processing the image and sending the processed image to a descriptor generator to extract an image descriptor and where extracting the at least one descriptor from the phrase comprises processing the phrase using natural language processing, and sending the processed phrase to the descriptor generator to extract a phrase descriptor.
13. The system of any of claims 11-12, wherein the match score between the at least one descriptor and the plurality of catalog descriptors comprises an image match score between an image descriptor and the plurality of catalog descriptors or a phrase match score between the phrase descriptor and the plurality of catalog descriptors.
14. The system of any of claims 11-13, wherein the plurality of catalog descriptors comprises, for each of the plurality of items, at least a catalog color descriptor, a catalog texture descriptor, a catalog pattern descriptor, and a catalog aesthetic descriptor.
15. The system of any of claims 11-14, wherein extracting the at least one descriptor comprises: extracting a color descriptor; extracting a texture descriptor; extracting a pattern descriptor;extracting an aesthetic descriptor; and ranking each of the plurality of items comprises: calculating a color match score between the color descriptor and the catalog color descriptor; calculating a texture match score between the texture descriptor and the catalog texture descriptor; calculating a pattern match score between the pattern descriptor and the catalog pattern descriptor; calculating an aesthetic match score between the aesthetic descriptor and the catalog aesthetic descriptor; and fusing the color match score, the texture match score, the pattern match score, and the aesthetic match score to determine an overall match score between the received image and each of the plurality of items.
16. The system of any of claims 11-15, wherein each of the color match score, the texture match score, the pattern match score, and the aesthetic match score is based on a similarity calculation between an encoding of each of the catalog descriptors and an encoding of the at least one descriptor.
17. The system of any of claims 11-15, wherein calculating any of the texture match score, the pattern match score, or the aesthetic match score is based on a classification of the image as belonging to one of a set of textures, a set of patterns, or a set of aesthetic criteria and a probability associated with the classification.
18. The system of any of claims 11-17, wherein extracting at least one descriptor from the received image further comprises, segmenting the received image into segmented images and processing each of the segmented images to extract at least one descriptor from each of the segmented images.
19. The system of any of claims 11-18, wherein the image color descriptor and the catalog color descriptor match when an absolute value of the percentage difference between each of the color coordinates of the image color descriptor and the corresponding color coordinates for the catalog color descriptor are within a percentage threshold value of each other.
20. The system of any of claims 11-19, wherein the image color descriptor and the catalog color descriptor use color coordinates, and are considered well matched if a distance in color coordinate space between the image color descriptor and the catalog color descriptor is less than a threshold value.