Finding Complementary Digital Images Using Conditional Generative Adversarial Networks
By using the Conditional Generative Adversarial Network (CGAN) generator component, the problem that existing recommendation engines are difficult to identify complementary images is solved, and efficient and accurate complementary image recommendation is achieved, improving the user experience.
Patent Information
- Application Number
- CN201980089261.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-16
- Filing Date
- 2019-12-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2039-12-10
AI Technical Summary
Existing recommendation engines have difficulty accurately identifying occasions where users are trying to find complementary digital items, such as complementary item images of clothing, and existing click count information is not enough to subtly locate user needs.
Generator components trained by machine generate complementary images through a conditional generative adversarial network (CGAN) and retrieve output images from the data repository based on the generated image information, providing user interface rendering.
It improves the accuracy of the recommendation system, can effectively recommend complementary images, reduce user search time, and improve user experience.
Smart Images

Figure CN113330455B_ABST
Abstract
Description
Background Art
[0001] Given a user's initial selection of digital items, a computing system often employs a recommendation engine to suggest one or more digital items. In a common strategy, the recommendation engine may consider click count information to deem digital items as related. For example, if many users who click (or otherwise interact with) a first digital item also click a second digital item, the recommendation engine may determine that the two digital items are related. However, this measure of relatedness is noisy because it encompasses various reasons for a user's interaction with a digital item. Therefore, this measure of relatedness is insufficient to subtly identify those occasions where a user is attempting to find two complementary digital items. For example, when a user is attempting to find images of complementary items of clothing. Summary of the Invention
[0002] The computer-implemented techniques described herein are used to retrieve at least one recommended output image. In one implementation, the techniques use a machine-trained generator component to transform a first portion of image information associated with an image of a first portion selected by a user into one or more instances of a second portion of generated image information. Each instance of the second portion of generated image information complements the first portion of image information. The generator component is trained by a computer-implemented training system using a conditional generative adversarial network (CGAN). The techniques further include: retrieving one or more second portion output images from a data repository based on the instance(s) of the second portion of generated image information; generating a user interface presentation that presents the image of the first portion and the one or more second portion output images; and displaying the user interface presentation on a display device.
[0003] In one implementation, the image of the first portion depicts a first clothing item (such as a pair of pants) selected by the user, and each second portion output image depicts a second clothing item that complements the first clothing item (such as a shirt, a jacket, a blouse, etc.). More generally, the image of the first portion and each second portion output image may depict any complementary image content, such as complementary furniture items, decorative items, parts of a car, etc.
[0004] In one implementation, the training system generates a generator component by: identifying a set of synthetic images, each synthetic image consisting of at least an image of a first portion and a complementary second portion of an image; segmenting the synthetic images to generate a plurality of first portion extraction images and a plurality of second portion extraction images; and training the generator component using cGAN based on the plurality of first portion extraction images and the plurality of second portion extraction images.
[0005] The techniques summarized above can be embodied in various types of systems, devices, components, methods, computer-readable storage media, data structures, graphical user interface presentations, articles of manufacture, etc.
[0006] The present disclosure is intended to introduce some concepts in a simplified form; these concepts are further described in the detailed description below. The present disclosure is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 An illustrative training system for generating a trained generator component is shown. The training system performs this task using a conditional generative adversarial network (cGAN).
[0008] Figure 2 A recommendation generation system that uses the generator component to recommend output images is shown, where the generator component is generated using the Figure 1 training system.
[0009] Figure 3 A computing device that can be used to implement the Figure 1 and Figure 2 training system and recommendation-generation system is shown.
[0010] Figure 4 An illustrative aspect of the Figure 1 training system is shown.
[0011] Figure 5 An illustrative aspect of the Figure 2 recommendation-generation system is shown.
[0012] Figure 6 An illustrative user interface presentation generated by the Figure 2 recommendation-generation system is shown.
[0013] Figure 7 An illustrative convolutional neural network (CNN) component that can be used to implement various aspects of the Figure 1 and Figure 2 training system and recommendation-generation system is shown.
[0014] Figure 8 Is a flowchart showing an illustrative way of Figure 2 operating the recommendation-generation system.
[0015] Figure 9 Is a flowchart showing an illustrative way of Figure 1 operating the training system.
[0016] Figure 10 An illustrative type of computing device that can be used to implement any aspect of the features shown in the above figures is shown.
[0017] Like reference numerals are used throughout the disclosure and the figures to denote like components and features. Series 100 numbers refer to features initially found in Figure 1 , series 200 numbers refer to features initially found in Figure 2 , series 300 numbers refer to features initially found in Figure 3 , and so on. DETAILED DESCRIPTION
[0018] The present disclosure is organized as follows. Section A describes a computing system for training a generator component and using the generator component to recommend an image given a selected input image. Section B sets forth illustrative methods that explain the operation of the computing system of Section A. And Section C describes illustrative computing functions that can be used to implement any aspect of the features described in Sections A and B.
[0019] As a citation, the term "hardware logic circuitry" corresponds to one or more hardware processors (e.g., CPU, GPU, etc.) that execute machine-readable instructions stored in a memory, and / or one or more other hardware logic components (e.g., FPGA) that perform operations using a fixed and / or programmable set of logic gates. Section C provides additional information regarding one implementation of the hardware logic circuitry. Each of the terms "component" and "engine" refers to a portion of the hardware logic circuitry that performs a particular function.
[0020] In one case, separating the various parts in the figures into different units may reflect the use of corresponding different physical and tangible parts in an actual implementation. Alternatively or additionally, any single component illustrated in the figures may be implemented by multiple actual physical parts. Alternatively or additionally, the depiction of any two or more separated parts in the figures may reflect different functions performed by a single actual physical part.
[0021] Other figures depict concepts in flowchart form. In this form, certain operations are described as constituting different boxes that are executed in a certain order. This implementation is illustrative and non-limiting. The specific boxes described herein may be combined together and executed in a single operation, certain boxes may be broken down into multiple component boxes, and certain boxes may be executed in an order different from that illustrated herein (including executing the boxes in parallel). In one implementation, the boxes shown in the flowcharts regarding process-related functions may be implemented by the hardware logic circuitry described in Section C, which in turn is implemented by one or more hardware processors and / or other logic components including a set of task-specific logic gates.
[0022] Regarding terminology, the phrase "configured to" encompasses various physical and tangible mechanisms for performing the identified operations. The mechanisms can be configured to perform the operations using the hardware logic circuits of Section C. The term "logic" likewise encompasses various physical and tangible mechanisms for performing tasks. For example, each process-related operation illustrated in a flowchart corresponds to a logic component for performing that operation. The logic component can use the hardware logic circuits of Section C to perform its operation. When implemented by a computing device (however implemented), the logic component represents an electrical component that is a physical part of the computing system.
[0023] Any one of the storage resources described herein, or any combination of these storage resources, can be considered a computer-readable medium. In many cases, a computer-readable medium represents some form of physical and tangible entity. The term computer-readable medium also encompasses propagated signals, such as those transmitted or received via physical ducts and / or air or other wireless media. However, the specific term "computer-readable storage medium" expressly excludes the propagated signals themselves while including all other forms of computer-readable media.
[0024] The following explanations may identify one or more features as "optional". This type of statement is not to be construed as an exhaustive listing of features that may be considered optional; that is, other features may be considered optional even though not expressly identified in the text. Additionally, the description of a single entity is not intended to exclude the use of multiple such entities; similarly, the description of multiple entities is not intended to exclude the use of a single entity. Further, while the description may interpret certain features as alternative ways of performing the identified functions or implementing the identified mechanisms, these features may also be combined in any combination. Finally, the terms "exemplary" or "illustrative" refer to one implementation among potentially many implementations.
[0025] A. Illustrative Computing System
[0026] Figure 1 An illustrative training system for generating a trained generator component 104 is shown. The training system 102 performs this task using a conditional generative adversarial network (cGAN) 106. In Figure 2 The recommendation-generation system shown and described below uses the trained generator component 104 to map a received input image into one or more other images of a complementary input image. That is, in the terminology used herein, the recommendation-generation system maps a selected first portion of an image into one or more second portions of an image, where each second portion of the image complements the first portion of the image.
[0027] In most of the illustrative examples presented herein, the image of the first portion depicts a selected clothing item (such as a pair of pants), and the image of the complementary second portion depicts complementary clothing items (such as various shirts, jackets, blouses, etc. determined to match the selected pants). In another case, the image of the selected first portion depicts a selected furniture item (such as a sofa), and the image of the complementary second portion depicts complementary furniture items (such as a side table, coffee table, etc.). In another case, the image of the selected first portion depicts a home decor item (such as a vase), and the image of the complementary second portion depicts complementary decor items (such as a wall hanging, etc.). In another case, the image of the selected first portion depicts a selected automotive part (such as a selected seat style), and the image of the complementary second portion depicts complementary automotive parts (such as a dashboard trim, etc.), and so on. The above cases are provided by way of example and not limitation.
[0028] Further note that Figure 1 it is shown that the training system 102 produces only a single generator component 104. However, in practice, the training system 102 can produce multiple generator components. Each generator maps between different image content category pairs. For example, given a selected pair of pants, a first generator component can suggest a shirt, while given a selected shirt, a second generator component can suggest pants. Given a selected pair of shoes, a third generator component can suggest pants, and so on. The recommendation-generation engine can utilize any number of such generator components to generate its recommendations, depending on how it is configured.
[0029] The training system 102 includes a preprocessing function 108 for generating a training image corpus. The cGAN 106 then uses the training images to generate a generator component 104. As a first stage, a matching component 110 retrieves an initial set of matching images, where each matching image of the matching images includes an image of a first part and an image of a complementary second part. For example, the matching component 110 may retrieve an initial set of images of human objects, where each image of the image set includes an image of a first part showing the pants worn by the human object and an image of a second part showing the shirt worn by the human object. In one case, the matching component 110 performs its function by submitting one or more search queries to a search engine, such as the BING search engine provided by MICROSOFT CORPORATION of Redmond, Washington. The search engine then retrieves the initial set of matching images from a data repository associated with a wide area network (e.g., the Internet). The matching component 110 may submit, but is not limited to, the following queries: (1) "Men's autumn fashion"; (2) "Men's autumn fashion 2018"; (3) "Women's autumn fashion"; (4) "Women's latest autumn fashion", etc. As indicated by these query examples, the model developer can pick an initial matching target set by formulating the input query in an appropriate manner. That is, the model developer can select queries that identify the theme of the images being sought, the date associated with the images being sought, etc. The matching component 110 stores the initial set of matching images in a data repository 112.
[0030] A training set generation component (TSGC) 114 generates a training corpus based on the initial set of matching images. Figure 4 The operation of this component will be described in detail. As an overview, the TSGC 114 first extracts a plurality of synthetic images from the initial set of matching images. Each synthetic image includes an image of a first part and an image of a second part. For example, each synthetic image corresponds to a portion of an initial matching image showing the complete body of a human object, excluding other portions of the initial matching image that do not belong to the object's body. The TSGC 114 may then extract a plurality of images of the first part (such as images showing pants) and a plurality of images of the second part (such as images showing shirts) from the synthetic images. These extracted images are referred to herein as the extracted images of the first part 116 and the extracted images of the second part 118, respectively. The TSGC 114 may optionally perform other processing on the extracted images, such as centering the images, normalizing the images, compressing the images, etc.
[0031] The data repository 120 stores the extracted image 116 of the first part and the extracted image 118 of the second part. Alternatively, the TSGC 114 can convert the image into feature information, such as using a deep neural network (DNN). The data repository 120 can then store the feature information associated with each extracted image, instead of the extracted image itself, or in addition to the extracted image. The data repository 120 can represent each instance of the feature information as a feature vector.
[0032] More generally, the term "image information of the first part" as used herein refers to the pixels associated with the image of the first part, or some kind of feature - space representation of the image of the first part, such as a vector describing the features associated with the image of the first part. Similarly, the term "image information of the second part" as used herein refers to the pixels associated with the image of the second part, or some kind of feature - space representation of the image of the second part.
[0033] Using the above terms, the purpose of the trained generator component 104 is to map the image information of the first part to the generated image information of the second part. The image information of the first part can correspond to the original pixels associated with the image of the first part, or the feature - space representation of the image of the first part. Similarly, the generated image information of the second part can correspond to the original pixels associated with the generated image of the second part, or the feature - space representation of the generated image of the second part. The modifier "generated" as used herein indicates that the image information produced by the generator component 104 reflects the insights obtained from multiple real pre - existing images, but does not necessarily match any of those pre - existing images. That is, the generated image is created, rather than retrieved from the empirical distribution of pre - existing images.
[0034] The cGAN 106 includes a generator component 104’ (G) that maps the image information y of the first part, together with an instance of random information z obtained from the distribution p z into the generated image information of the second part. For example, the generator component 104’ can first concatenate the vector representing the random information z with the vector representing the image information y of the first part, and then map the concatenated information into the generated information of the second part. Note that the generator component 104’ represents the training - phase counterpart of the fully trained generator component 104. That is, once the generator component 104’ becomes fully trained, it constitutes the generator component 104. At any stage of training, the behavior of the generator component 104’ is given by a set of weighted values θ g given.
[0035] The cGAN also includes a discriminator component 124 (D) for generating a score that reflects the confidence level that an instance of the image information fed into its second part is real, as opposed to being generated. The real instances of the image information of the second part correspond to the images that extract the actual second part from the synthetic images. The generated instances of the image information of the second part are produced by the generator component 104'. At any stage during training, the behavior of the discriminator component 124 is represented by a set of weighting values θ d is represented.
[0036] The training component 126 feeds instances of the image information into the generator component 104' and the discriminator component 124. The training component 126 also updates the weighting values for the generator component 104' and the discriminator component 124 based on the currently observed performance of these two components. In practice, the training component 126 can update the weighting values of the discriminator component 124 at a rate r and the weighting values of the generator component 104' at a rate s; in some implementations, r > s.
[0037] More specifically, the training component 126 updates the generator component 104' and the discriminator component 124 according to the following objective function:
[0038]
[0039] In this equation, y represents an instance of the image information of the first part (e.g., an extracted image depicting a pair of pants), and x represents an instance of the image information of the second part (e.g., an extracted image depicting a complementary shirt). p d represents the distribution of the discriminator component 124 over x. Given an instance of the random information z and an instance of the image information y of the first part, G(z|y) represents an instance of the generated information of the second part generated by the generator component 104'. Given y, D(x|y) represents the probability that x represents a real instance of the image information of the second part.
[0040] The training component 126 attempts to train the discriminator component 124 to maximize the first part of Equation (1), i.e., (log D(x|y)). By doing so, the training component 126 will maximize the probability that the discriminator component will correctly identify the information of the second part as being real and generated instances. At the same time, the training component 126 attempts to train the generator component 104' to minimize the second part of Equation (1), i.e., (log(1 - D(G(z|y)))). By doing so, the training component 126 increases the ability of the generator component 104' to "fool" the discriminator component 124 into classifying the generated image as real. Overall, cGAN is considered adversarial because the generator component 104' and the discriminator component 124 compete with each other; as each component improves its performance, it makes the function of the other component more difficult to execute. cGAN 106 is limited to being "conditional" because the generator component 104' generates the image information of the second part based on the image information y of the first part rather than on z alone.
[0041] Different implementations can formulate the above objective function in different environment-specific ways. In one implementation, the training component 126 can attempt to train the generator component 104' to solve:
[0042]
[0043] N represents a plurality of pairs of the extracted images of the first part and the second part in the training corpus, and each such pair is represented by (y k , x k ). Lcf(·) represents the loss function of the generator component 104'. represents the instance of the generated image information of the second part generated by the generator component 104' with respect to the image information y of the first part k .
[0044] In one implementation, the generator loss function L cf (·) can be defined as a weighted combination of multiple individual loss components that model the corresponding desired characteristics of the generated images of the second part. For example, in one case, the generator loss function is given by:
[0045] L cf = α1L MSE + α2L percept + α3Ladv (3).
[0046] L MSE 、L percept 、and L adv represent individual loss components, while α1, α2, and α3 represent weighted factor values.
[0047] More specifically, LMSE Represents the mean squared error loss between the image information x of the first part and the generated information of the second part. The training component 126 can calculate this loss as follows:
[0048]
[0049] W and H respectively correspond to the width and height of the image of the first part and the image of the second part. ||·|| F Corresponds to the Frobenius norm.
[0050] Lpercept represents the perceptual loss between the image information x of the first part and the generated information of the second part. In one implementation, the training component 126 can calculate this loss by performing a two-dimensional discrete cosine transform (DCT) on the image information x of the first part and the generated information of the second part, as follows:
[0051]
[0052] W and H again represent the width and height of the image information. Ω k Represents a two-dimensional DCT matrix of size k×k, applied to the image information channel-wise.
[0053] Finally, L adv Represents the adversarial loss of the generator component 104'. The training component 126 can calculate this loss as follows:
[0054]
[0055] Given y, the inner term of this equation reflects the probability that the generator component 104' will produce an effective reconstruction of the image information of the second part. The training component 126 sums this probability over all N training samples in the training corpus.
[0056] In one implementation, the discriminator component 124 performs its function by receiving the actual image information x of the second part. In another implementation, instead of receiving the actual image information x of the second part, the discriminator component 124 receives the transform L(x) of the actual image information x of the second part. In one implementation, L(x) represents the transform of a linear combination of x and the generated copy of x, such as:
[0057]
[0058] The symbol β represents a random value between 0 and 1. If β is greater than a specified threshold T, the transformed image information L(x) is considered to be real. If β ≤ T, L(x) is considered to be generated (i.e., "fake"). The training component 126 increases as training progresses. This has the net effect of providing smoother training of the cGAN 106. In one implementation, the training system 102 can implement the transformation L(x) as a deep neural network that is co-trained with the cGAN 106.
[0059] In one example, the training component 126 can feed image pairs to the cGAN 106 in any order. This corresponds to the random batch approach (RBA). In another example, the training component 126 can first form clusters of pairs and then feed the image pairs to the cGAN 106 on a cluster-by-cluster basis. This constitutes the clustered batch approach (CBA). For example, a cluster of pairs can show pants arranged in a particular pose.
[0060] Now proceed to Figure 2 , which shows the recommendation-generation system 202 that uses the trained generator component 104 to recommend images, and the trained generator component 104 is produced by the training system 102 using Figure 1 As an initial step, the user can interact with the search engine 204 via a browser application to select an image of a first part, such as a pair of pants. The user can perform this task in various ways, such as by searching for, finding, and clicking on the image of the first part. Here, "clicking" broadly refers to any mechanism for selecting an image, such as selecting it via a mouse device, touching on a touch-sensitive surface, etc. Or the user can click on a page that includes the image of the first part as part of other information. In other cases, the user can click on a specific part of the retrieved image, such as clicking on the lower garment shown in the image; the selected part constitutes the image of the first part. Generally, the user selects the image of the first part by submitting one or more input signals to the browser application and / or some other application and by interacting with an input device (mouse device, touch-sensitive surface, keyboard, etc.). In some cases, if information about the image of the first part does not already exist, the optional preprocessing component 206 transforms the selected image of the first part into a feature space representation of the image of the first part.
[0061] The trained generator component 104 then maps the first part y to the generated image information of the second part In an example. In other cases, the trained generator component can map the image information y of the first part into multiple instances of the generated image information of the second part. The trained generator component 104 can perform this task by executing it for different instances of the input noise information z for each iteration. In some cases, each instance of the generated image information of the second part represents a representation of the pixel levels of the generated image of the second part. In other cases, the instance of the generated information of the second part represents a feature-space representation of the generated image of the second part.
[0062] As shown, the recommendation generation system 202 can alternatively provide multiple trained generator components 208, and each component of the multiple trained generator components 208 performs different pairs of mapping transformations. The recommendation-generation system 202 can select a suitable generator component based on knowledge of the content type included in the image of the first part that the user has selected, and knowledge of the types of complementary images that should be presented to the user. The recommendation-generation system 202 can use various techniques to detect the content of the image of the first part, such as by analyzing the metadata associated with the image of the first part, and / or using a machine-trained classification model to identify the content associated with the image of the first part, etc. The recommendation-generation system 202 can determine the type of complementary image that should be presented to the user based on, for example, preconfigured settings defined by an administrator or the user himself or herself. Alternatively, the recommendation-generation system 202 can allow the user to input a real-time command that conveys the type of complementary image that the user is interested in, such as by allowing the user to click on a graphical command such as "Show me matching shirts".
[0063] The recommendation-generation component 210 retrieves one or more output images of the second part based on the generated image information of the second part generated by the trained generator component 104. The recommendation-generation component 210 will be described in detail below with reference to Figure 5 As an overview, the recommendation-generation component 210 can select one or more output images of the second part from the data repository 212, and these images are determined to match the generated image information of the second part.
[0064] Finally, the user interface (UI)-generation component 214 generates a user interface presentation that displays the output images of the second part generated by the recommendation-generation component 210. In some cases, the UI-generation component 214 can display the image of the first part selected by the user; this will allow the user to visually judge the compatibility between the image of the first part and each output image of the second part.
[0065] Generally, given a user specification of an image of a first portion, the recommendation-generation system 202 provides relevant recommendations for an image of a second portion. The recommendation-generation system 202 can provide good recommendations because it uses a recent corpus of synthetic images, a generator component 104 trained using cGAN 106, where each image in the synthetic images shows an image of the first portion and an image of the second portion. The recommendation-generation system 202 can also be said to effectively utilize computing resources because it quickly provides good recommendations to the user, reducing the need for the user to engage in a lengthy search session.
[0066] Figure 3 illustrates computing equipment 302 for a training system 102 and a recommendation-generation system 202 that can be used to implement Figure 1 and 2 In one implementation, one or more computing devices (such as, one or more servers 304) implement the training system 102. One or more computing devices (such as, one or more servers 306) implement the recommendation-generation system 202. Multiple data repositories 308 store images for processing by the training system 102 and the recommendation-generation system 200
[0067] An end user can interact with the recommendation-generation system 202 via multiple user computing devices 310, via a computer network 312. The user computing device 310 can correspond to any combination of: a desktop personal computing device, a laptop device, a handheld device (such as, a tablet device, a smart phone, etc.), a game console device, a mixed reality device, a wearable computing device, etc. The computer network 312 can correspond to a wide area network (such as, the Internet), a local area network, one or more point-to-point links, or any combination thereof.
[0068] Figure 3 Indicates that all functions of the recommendation-generation system 202 are assigned to the (multiple) servers 306. The user computing device can interact with this functionality via a browser program running on the user computing device. Alternatively, any portion of the functionality of the recommendation-generation system 202 can be implemented by the local user computing device.
[0069] Figure 4 illustrates illustrative aspects of a training set generation component (TSGC) 114, which is Figure 1 part of the training system 102 introduced in. As the name indicates, the purpose of the TSGC 114 is to generate a corpus of training images. The TSGC 114 will be described in the context of extracting images of lower body clothing (e.g., pants) and upper body clothing (e.g., shirts, jackets, blouses, etc.). However, the principles described below can be extended to the extraction of any paired complementary images.
[0070] The first segmentation component 402 receives a matching image from the matching component 110. The illustrative matching image 404 includes a synthetic image 406 showing a complete human body, optionally together with other content (such as background content, in this case, a building). The first segmentation component 402 extracts the synthetic image 406 from the remainder of the matching image 404. The first segmentation component 402 can perform this task using any computer automated segmentation technique. Alternatively or additionally, the first segmentation component 402 can perform this task with the manual assistance of a human reviewer 408 via a crowdsourcing platform.
[0071] The second segmentation component 410 receives the synthetic image 406 from the first segmentation component 402. The second segmentation component 410 then extracts an extracted image 412 of a first portion and an extracted image 414 of a second portion from the synthetic image 406. Here, the extracted image 412 of the first portion shows the pants worn by the human object, while the extracted image 414 of the second portion shows the jacket and shirt worn by the human object. Again, the second segmentation component 410 can perform this task using any computer automated segmentation technique. Alternatively or additionally, the second segmentation component 410 can perform this task with the manual assistance of a human reviewer 416 via a crowdsourcing platform.
[0072] In one implementation, the automated segmentation technique can first use machine-trained components, such as a convolutional neural network (CNN) component, to classify the pixels of the image. The classification of each pixel indicates the type of content to which it most likely belongs. For example, the first segmentation component 402 can use a machine-trained component that differentiates between pixels showing a human object and pixels showing non-human content. The second segmentation component 410 can use a machine-trained component to differentiate between lower-body clothing pixels and upper-body clothing pixels, etc. The automated segmentation technique can then draw regions of interest (ROIs) around each group of pixels sharing the same classification. The automated segmentation technique can rely on different techniques to perform this function. In one method, the automated segmentation technique draws a rectangle that encompasses all pixels having the same classification, optionally ignoring peripheral pixels that do not have adjacent pixels having the same classification. In another method, the automated segmentation component iteratively merges image regions sharing the same classification in the image. In another method, the automated segmentation technique can use a region proposal network (RPN) to generate ROIs. In the overall context, background information on the subject of region proposal networks is described in Rend et al., "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks," arXiv:1506.01497v3 [cs.CV], January 6, 2016, 9 pages.
[0073] The post - processing component 418 performs post - processing operations on the extracted images (412, 414). For example, the post - processing component 418 can center the main content of each image, normalize each image, adjust the contrast of each image, etc. The post - processing component 418 can also reject images that do not meet one or more environment - specific tests. For example, if the resolution of an image is below a specified threshold, or its main content is occluded by one or more objects, etc., the post - processing component 418 can reject the image.
[0074] In some implementations, the post - processing component 418 can also convert each image into a feature - space representation. As noted above, the post - processing component 418 can use machine - trained components to perform this task, such as a CNN component. In the general context, background information on the subject of CNN components for converting images into feature - space representations includes: Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision”, arXiv:1704.04861v1 [cs.CV], April 17, 2017, 9 pages; He et al., “Deep Residual Learning for Image Recognition”, arXiv:1512.03385v1 [cs.CV], December 10, 2015, 12 pages; and Szegedy et al., “Going Deeper with Convolutions”, arXiv:1409.4842v1 [cs.CV], September 17, 2014, 12 pages.
[0075] In some implementations, the post - processing component 418 can also compress the original image itself or its feature - space representation. For example, the post - processing component can apply principal component analysis (PCA) to reduce the size of each feature - space representation of the image. The processing operations described above are described in an illustrative rather than limiting spirit; other implementations can apply other types of post - processing operations.
[0076] Figure 5 Shows Figure 2 Illustrative aspects of the recommendation - generation system 202. Assume that the recommendation - generation system 202 receives multiple instances of the second part of the generated image information (s1, s2,..., s n ) from the trained generator component 104. For example, for a specified instance of the first part of the image information y (e.g., it shows a pair of pants), the trained generator component 104 can generate multiple such instances by successively inputting different instances of random information z. Each such instance of the generated image information can be formally expressed as Suppose each instance of the generated image information in the second part corresponds to an image of a top with a specified style.
[0077] In one implementation, the recommendation-generation component 202 can display to the user instances (s1, s2,..., s n ) of the generated image information in the second part. However, note that each instance of the generated information in the second part does not necessarily have a real counterpart image and thus does not necessarily correspond to a real product that the user may purchase. However, the user may find the generated information in the second part helpful because it provides general advice on products that can complement the input image information y in the first part.
[0078] In another implementation, the recommendation-generation component 202 identifies one or more pre-existing images that are similar to any of the instances (s1, s2,..., s n ) in the generated image information in the second part. It then presents these pre-existing images to the user, instead of or in addition to the generated image information. In some implementations, each pre-existing image is associated with a product that the user can purchase or otherwise select.
[0079] For example, in one implementation, the adjacent retrieval component 502 first identifies an adjacent set N ij of the extracted images in the second part in the data repository 120, i which matches each instance of the generated image information s i . Recall that the training system 102 generates these extracted images in the second part based on the selected set of synthetic images. The adjacent retrieval component 502 can perform its function by looking for those extracted images in the second part that are within a specified threshold distance of any instance s
[0080] The output item retrieval component 504 can then map the extracted image of each second part to one or more output images of the second part provided in the data repository 506. For example, the output image of each second part can show a product page or advertisement associated with a product that can be actually purchased by the user or otherwise selected by the user. The output item retrieval component 504 can match the extracted image of each second part with one or more output images of the second part in the same way as the adjacent retrieval component 502, for example, by using any distance metric (e.g., cosine similarity, Euclidean distance, etc.) to find a set of output images within a specified threshold distance of the extracted image of the second part. The output item retrieval component 504 can rank the identified output images of the second part by their distance scores (e.g., their cosine similarity scores), and then select the top k ranked output image sets. Finally, the UI-generation component displays the selected image of the first part and the top-ranked output images of the second part together.
[0081] In other implementations, the recommendation-generation component 202 can omit the adjacent retrieval component 502. In that case, the output item retrieval component 504 directly compares each instance of the generated image information si with each image of the output images of the candidate second part.
[0082] Figure 6 shown by Figure 2 An illustrative user interface presentation 602 generated by the recommendation-generation system 202 is shown. In this non-limiting scenario, the presentation 602 shows an image 604 of the first part and four output images 606 of the second part. The user interface presentation 602 can optionally show graphical cues (608, 610, 612, 614) that allow the user to select any of the output images 606 of the second part. For example, the user can click on the graphical cue 608 to view a product page or advertisement associated with the image 616 of the second part. Alternatively or additionally, the recommendation-generation component 202 can show the original instance of the generated image information si for the user to review, which may not correspond to the actual product the user may purchase.
[0083] In other scenarios, the recommendation-generation component 202 can provide multiple recommendation sets for user consideration within one or more user interface presentations. For example, assume the user starts again by selecting the image 604 of the first part. The recommendation-generation component 202 can use the first trained generator component to find and display a complementary set of tops (as Figure 6 shown in). It can also use the second trained generator component to find and display a complementary set of shoes (not shown in Figure 6 ). Each set of output images complements a pair of pants initially selected by the user.
[0084] In other cases, the recommendation-generation component 202 can treat the output images of one or more second parts as input examples and look for one or more sets of output images that complement these input examples. For example, the recommendation-generation system 202 can treat the shirt shown in image 616 of the second part as a new input example. It can then look for and display one or more necklaces, etc. that complement the shirt. The recommendation-generation system 202 can extend this recommendation chain any number of steps.
[0085] The user interface rendering 602 can also include a real-time configuration option 618 that allows the user to control the (if any) recommended sets being displayed. Alternatively or additionally, the user interface rendering 602 includes a link 620 to a separate configuration page (not shown) that allows the user to more generally specify the recommendation behavior of the UI-generation component 214. Although not shown, the UI-generation component 214 can include a learning component that learns the user's recommendation-set preferences and automatically invokes these settings for that user.
[0086] Figure 7 An illustrative convolutional neural network (CNN) component 702 is shown, which can be used to implement Figure 1 and Figure 2 aspects of the training system 102 and the recommendation-generation system 202. For example, the generator components (104, 104') and the discriminator component 124 can be implemented using the CNN component 702, respectively. Figure 7 The general characteristics of the CNN function are described.
[0087] The CNN component 702 performs analysis in a stage pipeline. One of the multiple convolutional components 704 performs a convolution operation on the input image 706. One or more pooling components 708 perform downsampling operations. One or more feedforward components 710 each provide one or more fully connected neural networks, each neural network including any number of layers. More specifically, the CNN component 702 can intersperse the above three components in any order. For example, the CNN component 702 can include more than two convolutional components interleaved with pooling components. In some implementations, any type of post-processing component can operate on the output of the previous layer. For example, a classifier (Softmax) component can operate on the image information provided by the previous layer using the normalized exponential function to generate the final output information.
[0088] In each convolution operation, the convolution component moves an n×m kernel across the input image (where "input image" in this overall context refers to any image fed into the convolution component). In one case, at each position of the kernel, the convolution component generates the dot product of the kernel values and the underlying pixel values of the image. The convolution component stores this dot product as an output value in the output image at the position corresponding to the current position of the kernel.
[0089] More specifically, the convolution component can perform the above operations for different sets of kernels with different machine learning kernel values. Each kernel corresponds to a different pattern. In earlier processing layers, the convolution component may apply kernels that serve to identify relatively primitive patterns in the image (such as edges, corners, etc.). In later layers, the convolution component may apply kernels that look for more complex shapes (such as shapes similar to a human leg, arm, etc.).
[0090] In each pooling operation, the pooling component moves a window of a predetermined size across the input image (where the input image corresponds to any image fed into the pooling component). The pooling component then performs some aggregation / summarization operation on the values of the input image enclosed by the window, such as by identifying and storing the maximum value in the window, generating and storing the average value in the window, etc.
[0091] The feedforward component can start its operation by forming a single input vector. It can perform this task by concatenating the rows or columns of the input image (or images) fed to it to form a single input vector z1. The feedforward component then processes the input vector z1 using a feedforward neural network. Typically, the feedforward neural network can include N layers of neurons that map the input vector z1 to an output vector q. The values in any layer j can be given by the formula z j =f(W j z j-1 +bj) for j = 2,...N. The symbol W j represents the jth machine learning weight matrix, and the symbol b j represents the optional jth machine learning bias vector. The function f(·) (referred to as the activation function) can be represented in different ways, such as the tanh function, the ReLU function, etc.
[0092] The training component 126 iteratively generates weighted values that control the operation of at least the (multiple) convolution components 704 and the (multiple) feedforward components 710, and optionally the (multiple) pooling components 708. These values together constitute the machine training model. The training component 126 can perform its learning by iteratively operating on the set of images in the data repository 120. The training component 126 can use any technique to generate the machine training model, such as gradient descent techniques.
[0093] Other implementations of the functionality described herein employ other machine learning models (in addition to, or in place of, the CNN model) including, but not limited to: recurrent neural network (RNN) models, support vector machine (SVC) models, classification tree models, Bayesian models, linear regression models, etc. RNN models may use long short-term memory (LSTM) cells, gated recurrent units (GRUs), etc.
[0094] B. Illustrative Processes
[0095] Figure 8 and 9 Processes (802, 902) are shown in flowchart form that explain the operation of the computing system of Section A. Since the basic principles of the computing system operation have been described in Section A, certain operations will be addressed in this section in a summary fashion. As described in the preamble of the detailed description, each flowchart represents a series of operations to be performed in a particular order. However, the order of these operations is merely illustrative and can be changed in any way.
[0096] Figure 8 Shows a Figure 2 Process 802 that provides an overview of one way of operating the recommendation-generation system 202. In block 804, in response to at least one input signal provided by a user using an input device, the recommendation-generation system 202 identifies a first portion of an image. In block 806, the recommendation-generation system 202 receives first portion image information that describes the first portion of the image. In block 808, the recommendation-generation system uses the generator component 104 to transform the first portion image information into one or more instances of second portion generated image information, where each instance of the second portion generated image information complements the first portion image information. In block 810, the recommendation-generation system 202 retrieves one or more second portion output images 606 from the data repository 506 based on the instance(s) of the second portion generated image information. In block 812, the recommendation-generation system 202 generates a user interface presentation 602 that presents the first portion of the image and the second portion output image(s) 606. The user interface presentation 602 includes one or more graphical cues (608, 610, 612, 614) that allow the user to select any of the second portion output images. In block 814, the recommendation-generation system 202 displays the user interface presentation 602 on a display device for presentation to the user. The generator component 104 is trained by the computer-implemented training system 102 using a conditional generative adversarial network (cGAN) 106.
[0097] Figure 9 Shows a Figure 1Process 902 of an illustrative mode of operation of training system 102. In block 904, training system 102 uses matching component 110 to identify a set of synthetic images, each synthetic image being composed of at least an image of a first portion and an image of a second portion. In block 906, training system 102 segments the synthetic images to generate a plurality of extracted images of the first portion and a plurality of extracted images of the second portion. In block 908, training system 102 uses a conditional generative adversarial network (cGAN) 106 to train generator component 104' based on the plurality of extracted images of the first portion and the plurality of extracted images of the second portion. Operations using cGAN 106 for multiple instances of image information of the first portion include: in block 910, using generator component 104' to map an instance of image information of the first portion, together with an instance of random information obtained from a distribution, into an instance of generated image information of the second portion; in block 912, using discriminator component 124 to provide a score reflecting a confidence level that an instance of generated image information of the second portion is a representation of a true pre-existing image of the second portion; and in block 914, adjusting weight values of generator component 104' and discriminator component 124 based at least on the score.
[0098] C. Representative Computing Functions
[0099] Figure 10 FIG. shows a computing device 1002 that can be used to implement any aspect of the mechanisms set forth in the above figures. For example, referring to Figure 3 , Figure 10 the type of computing device 1002 shown therein can be used to implement any computing device associated with any of training system 102 or recommendation-generation system 202 or user device 310 etc. In all cases, computing device 1002 represents a physical and tangible processing mechanism.
[0100] Computing device 1002 can include one or more hardware processors 1004. The (multiple) hardware processors can include, but are not limited to, one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs), and / or one or more application specific integrated circuits (ASICs), etc. More generally, any hardware processor can correspond to a general-purpose processing unit or a special-purpose processor unit.
[0101] The computing device 1002 may also include a computer-readable storage medium 1006 corresponding to one or more computer-readable media hardware units. The computer-readable storage medium 1006 holds any kind of information 1008, such as machine-readable instructions, settings, data, etc. For example, the computer-readable storage medium 1006 may include, but is not limited to, one or more solid-state devices, one or more magnetic hard disks, one or more optical discs, magnetic tapes, etc. Any instance of the computer-readable storage medium 1006 may use any technology for storing and retrieving information. In addition, any instance of the computer-readable storage medium 1006 may represent a fixed or removable component of the computing device 1002. Further, any instance of the computer-readable storage medium 1006 may provide volatile or non-volatile retention of information.
[0102] The computing device 1002 may utilize any instance of the computer-readable storage medium 1006 in different ways. For example, any instance of the computer-readable storage medium 1006 may represent a hardware memory unit (such as random access memory (RAM)) for storing transient information during program execution by the computing device 1002, and / or a hardware storage unit (such as a hard disk) for retaining / archiving information on a more permanent basis. In the latter case, the computing device 1002 also includes one or more drive mechanisms 1010 (such as a hard disk drive mechanism) for storing and retrieving information from an instance of the computer-readable storage medium 1006.
[0103] When the (one or more) hardware processors 1004 execute computer-readable instructions stored in any instance of the computer-readable storage medium 1006, the computing device 1002 may perform any of the above functions. For example, the computing device 1002 may execute computer-readable instructions to perform each process block described in Section B.
[0104] Alternatively or additionally, the computing device 1002 may rely on one or more other hardware logic components 1012 to perform operations using a set of task-specific logic gates. For example, the (one or more) hardware logic components 1012 may include a fixed configuration of hardware logic gates, e.g., the hardware logic gates are created and set at the time of manufacture and are thereafter unchangeable. Alternatively or additionally, the (one or more) other hardware logic components 1012 may include a set of programmable hardware logic gates that can be set to perform different application-specific tasks. The latter classification of devices includes, but is not limited to, programmable array logic devices (PALs), generic array logic devices (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), etc.
[0105] Figure 10Generally, the hardware logic circuitry 1014 includes any combination of (one or more) hardware processors 1004, computer-readable storage media 1006, and / or (one or more) other hardware logic components 1012. That is, the computing device 1002 may employ any combination of (one or more) hardware processors 1004 that execute machine-readable instructions provided in the computer-readable storage media 1006, and / or one or more other (one or more) hardware logic components 1012 that perform operations using a fixed set of hardware logic gates and / or a programmable set of hardware logic gates. More generally, the hardware logic circuitry 1014 corresponds to one or more hardware logic components of any type that perform operations based on logic stored in and / or otherwise embodied in the (one or more) hardware logic components.
[0106] In some cases (e.g., where the computing device 1002 represents a user computing device), the computing device 1002 may also include an input / output interface 1016 for receiving various inputs (via the input device 1018) and for providing various outputs (via the output device 1020). Illustrative input devices include keyboard devices, mouse input devices, touchscreen input devices, digitizers, one or more still image cameras, one or more video cameras, one or more depth camera systems, one or more microphones, speech recognition mechanisms, any motion detection mechanisms (e.g., accelerometers, gyroscopes, etc.), and the like. A particular output mechanism may include a display device 1022 and an associated graphical user interface rendering (GUI) 1024. The display device 1022 may correspond to a liquid crystal display device, a light emitting diode display (LED) device, a cathode ray tube device, a projection mechanism, and the like. Other output devices include printers, one or more speakers, haptic output mechanisms, archival mechanisms (for storing output information), and the like. The computing device 1002 may also include one or more network interfaces 1026 for exchanging data with other devices via one or more communication conduits 1028. One or more communication buses 1030 communicatively couple the above components together.
[0107] (One or more) communication conduits 1028 may be implemented in any manner, e.g., via a local computer network, a wide area computer network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. (One or more) communication conduits 1028 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. managed by any protocol or combination of protocols.
[0108] Figure 10 A computing device composed of a discrete set of discrete units is shown. In some cases, the set of units may correspond to discrete hardware units provided in a computing device chassis having any form factor.Figure 10 An illustrative form factor is shown in its bottom portion. In other cases, computing device 1002 may include a hardware logic component that integrates the functions of more than two of the units shown in Figure 1 For example, computing device 1002 may include a system-on-chip (SoC or SOC), which corresponds to an integrated circuit that combines the functions of more than two of the units shown in Figure 10 the units shown.
[0109] The following overview provides a non-exhaustive set of illustrative aspects of the techniques set forth herein.
[0110] According to a first aspect, one or more computing devices for retrieving one or more output images are described. The (multiple) computing devices include hardware logic circuitry that corresponds to: (a) one or more hardware processors that perform operations by executing machine-readable instructions stored in a memory, and / or (b) one or more hardware logic components that perform operations using a set of task-specific logic gates. The operations include: in response to at least one input signal provided by a user using an input device, identifying a first portion of an image; receiving first image information describing the first portion of the image; using a generator component to transform the first image information into one or more instances of second generated image information, each instance of the second generated image information supplementing the first image information, the generator component having been trained by a computer-implemented training system using a conditional generative adversarial network (CGAN); retrieving one or more output images of a second portion from a data repository based on the (multiple) instances of the second generated image information; generating a user interface presentation that presents the first portion of the image and the (multiple) output images of the second portion, the user interface presentation including one or more graphical cues that allow the user to select any of the output images of the second portion of the (multiple) output images of the second portion; and displaying the user interface presentation on a display device for presentation to the user.
[0111] According to a second aspect, the first portion of the image shows a first clothing item selected by the user, and each output image of the second portion shows a second clothing item that complements the first clothing item.
[0112] According to a third aspect, each instance of the second generated image information describes features associated with the generated second portion of the image.
[0113] According to a fourth aspect, each instance of the second generated image information describes pixels associated with the generated second portion of the image.
[0114] According to a fifth aspect, the operations further include selecting a generator component from a set of generator components, each generator component being associated with a pair of complementary item types.
[0115] According to a sixth aspect, the operation further includes applying two or more generator components, each generator component being associated with a pair of complementary item types.
[0116] According to a seventh aspect, the retrieval operation further includes identifying one or more extracted images of a second part that match any instance of the (multiple) instances of the generated image information of the second part, each extracted image of the second part being derived from a composite image that includes at least an image of a first part and an image of a second part as parts of the composite image; and identifying (multiple) output images of the second part that match any of the extracted images of the (multiple) second parts, each output image of the second part showing an image of a selectable product.
[0117] According to an eighth aspect, the training system generates a generator component by: identifying a set of composite images, each composite image being composed of at least an image of a first part and an image of a second part; segmenting the composite images to generate a plurality of extracted images of the first part and a plurality of extracted images of the second part; and training the generator component using a cGAN based on the plurality of extracted images of the first part and the plurality of extracted images of the second part.
[0118] According to a ninth aspect, relying on the eighth aspect, the operation of using a cGAN includes, for multiple instances of the image information of the first part: using a generator component to map an instance of the image information of the first part together with an instance of random information taken from a distribution into an instance of the generated image information of the second part; using a discriminator component to provide a score that reflects the confidence level that the instance of the generated image information of the second part represents a real pre-existing image of the second part; and adjusting the weight values of the generator component and the discriminator component based at least on the score.
[0119] According to a tenth aspect, a method implemented by one or more computing devices for retrieving one or more output images is described. The method includes: in response to at least one input signal provided by a user using an input device, identifying a first portion of an image; receiving first portion image information describing the first portion of the image; using a generator component to transform the first portion image information into one or more instances of second portion generated image information, each instance of the second portion generated image information complementing the first portion image information, retrieving one or more second portion output images from a data repository based on the instance(s) of the second portion generated image information; generating a user interface rendering presenting the first portion of the image and the second portion output image(s), the user interface rendering including one or more graphical cues that allow the user to select any of the second portion output images; and displaying the user interface rendering on a display device for presentation to the user. The generator component has been trained by a computer-implemented training system using a conditional generative adversarial network (CGAN) by: identifying a set of synthetic images, each synthetic image consisting of at least a first portion of an image and a second portion of an image; segmenting the synthetic images to generate a plurality of first portion extracted images and a plurality of second portion extracted images; and training the generator component using the cGAN based on the plurality of first portion extracted images and the plurality of second portion extracted images.
[0120] According to an eleventh aspect, dependent on the tenth aspect, the first portion of the image shows a first clothing item selected by the user, and each second portion output image shows a second clothing item that complements the first clothing item.
[0121] According to a twelfth aspect, dependent on the eleventh aspect, the first clothing item is an upper body clothing item, and the second clothing item is a lower body clothing item.
[0122] According to a thirteenth aspect, dependent on the tenth aspect, the first portion of the image shows a first selectable product item selected by the user, and each second portion output image shows a second selectable product item that complements the first selectable product item.
[0123] According to a fourteenth aspect, dependent on the tenth aspect, each instance of the second portion generated image information describes features associated with the generated second portion of the image.
[0124] According to a fifteenth aspect, dependent on the tenth aspect, each instance of the second portion generated image information describes pixels associated with the generated second portion of the image.
[0125] According to a sixteenth aspect, dependent on the tenth aspect, the method further includes selecting a generator component from a set of generator components, each generator component being associated with a pair of complementary item types.
[0126] According to the seventeenth aspect, relying on the tenth aspect, the operation further includes applying more than two generator components, each generator component being associated with a pair of complementary item types.
[0127] According to the eighteenth aspect, relying on the tenth aspect, the retrieval operation further includes: identifying one or more extracted images of the second part that match any instance of the (multiple) instances of the generated image information of the second part; and identifying the (multiple) output images of the second part that match any image of the (multiple) extracted images of the second part, each output image of the second part showing an image of a selectable product.
[0128] According to the nineteenth aspect, relying on the tenth aspect, the operation of using cGAN includes, for multiple instances of the image information of the first part: using a generator component to map an instance of the image information of the first part together with an instance of random information taken from a distribution into an instance of the generated image information of the second part; using a discriminator component to provide a score that reflects the confidence level that the instance of the generated image information of the second part represents a real pre-existing image of the second part; and adjusting the weight values of the generator component and the discriminator component based at least on the score.
[0129] According to the twentieth aspect, a computer-readable storage medium for storing computer-readable instructions is described. The computer-readable instructions, when executed by one or more hardware processors, perform a method that includes: using a matching component to identify a set of synthetic images, each synthetic image being composed of at least an image of the first part and an image of the second part; segmenting the synthetic images to generate multiple extracted images of the first part and multiple extracted images of the second part; and training a generator component based on the multiple extracted images of the first part and the multiple extracted images of the second part using a conditional generative adversarial network (cGAN). The operation of using cGAN includes, for multiple instances of the image information of the first part: using a generator component to map an instance of the image information of the first part together with an instance of random information taken from a distribution into an instance of the generated image information of the second part; using a discriminator component to provide a score that reflects the confidence level that the instance of the generated image information of the second part represents a real pre-existing image of the second part; and adjusting the weight values of the generator component and the discriminator component based at least on the score.
[0130] The twenty-first aspect corresponds to any combination of the first to twentieth aspects above (e.g., any logically consistent permutation or subset).
[0131] The twenty-second aspect corresponds to any method counterpart, device counterpart, system counterpart, means-plus-function counterpart, computer-readable storage medium counterpart, data structure counterpart, article of manufacture counterpart, image user interface presentation counterpart, etc. associated with the first to twenty-first aspects.
[0132] Finally, the functions described herein can employ various mechanisms to ensure that any user data is processed in a manner that complies with applicable laws, social norms, and the expectations and preferences of individual users. For example, the function can allow users to explicitly opt in (and then explicitly opt out) of the provisions of the function. The function can also provide appropriate security mechanisms to ensure the privacy of user data (such as, data sanitization mechanisms, encryption mechanisms, password protection mechanisms, etc.).
[0133] In addition, the description may have set forth various concepts in the context of illustrative challenges or problems. This manner of explanation is not intended to suggest that others have understood and / or elucidated the challenges or problems in the manner specified herein. Further, this manner of explanation is not intended to suggest that the subject matter recited in the claims is limited to addressing the identified challenges or problems; that is, the subject matter in the claims can be applied in the context of challenges or problems other than those described herein.
[0134] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the above specific features and acts are disclosed as example forms of implementing the claims.
Claims
1. A computing device for retrieving one or more output images, comprising: Hardware logic circuitry corresponding to: (a) one or more hardware processors that perform operations by executing machine-readable instructions stored in a memory, and / or (b) one or more other hardware logic components that perform the operations using a set of task-specific logic gates, the operations including: Identifying a first portion of an image in response to at least one input signal provided by a user using an input device; Receiving image information of a first portion of the image that describes the first portion of the image; Using a generator component implemented by the hardware logic circuitry to transform the image information of the first portion into one or more instances of second portion generated image information using a neural network, each instance of the second portion generated image information supplementing the image information of the first portion and being fake image information created by the generator component rather than retrieved from an empirical distribution of pre-existing images; The generator component having been trained by a computer-implemented training system using a conditional generative adversarial network (cGAN); Retrieving one or more second portion output images from a data repository based on the one or more instances of the second portion generated image information; Generating a user interface presentation that presents the first portion of the image and the one or more second portion output images, the user interface presentation including one or more graphical cues that allow the user to select any of the one or more second portion output images; and Displaying the user interface presentation on a display device for presentation to the user.
2. The computing device according to claim 1, wherein the first portion of the image shows a first clothing item selected by the user, and each second portion output image shows a second clothing item that supplements the first clothing item.
3. The computing device according to claim 1, wherein each instance of the second portion generated image information describes features associated with the generated second portion of the image.
4. The computing device according to claim 1, wherein each instance of the second portion generated image information describes pixels associated with the generated second portion of the image.
5. A computing device according to claim 1, wherein the operation further comprises: Selecting the generator component from a set of generator components implemented by the hardware logic circuitry, each generator component being associated with a pair of complementary item types.
6. A computing device according to claim 1, wherein the operation further comprises: Applying two or more generator components implemented by the hardware logic circuitry, each generator component being associated with a pair of complementary item types.
7. The computing device according to claim 1, wherein the retrieval includes: Identifying one or more second portion extracted images that match any of the one or more instances of the second portion generated image information; And Identify the output images of the one or more second parts as one or more corresponding images that match the extraction images of any of the second parts in the extraction images of the one or more second parts, with the output image of each second part showing an image of a selectable product.
8. A computing device according to claim 1, wherein the training system generates the generator component through the following process: Identify a set of synthetic images, each synthetic image consisting of at least an image of a specific first part and an image of a specific second part; Segment the synthetic images to generate a plurality of extraction images of the first parts and a plurality of extraction images of the second parts; And Use the cGAN to train the generator component based on the plurality of extraction images of the first parts and the plurality of extraction images of the second parts.
9. A computing device according to claim 8, wherein using the cGAN includes, for a plurality of instances of the image information of the first part: Use the generator component to map a specific instance of the image information of the first part together with a specific instance of random information taken from a distribution into an instance of the generated image information of the second part; Use a discriminator component implemented by the hardware logic circuit to provide a score that reflects the confidence level that the specific instance of the generated image information of the second part represents a real pre-existing image of the second part; And Adjust the weight values of the generator component and the discriminator component based at least on the score.
10. A computing device according to claim 8, wherein using the cGAN includes, for a specific instance of the image information of the first part: Use the generator component to map the specific instance of the image information of the first part into a specific instance of the generated image information of the second part; Use a discriminator component implemented by the hardware logic circuit to provide a score that reflects the confidence level that the specific instance of the generated image information of the second part represents a real pre-existing image of the second part; And The discriminator component operates on the combined image information, which depends on a mixture of the specific instance of the generated information of the second part produced by the generator component and the actual instance of the image information of the second part, and the combined image information is considered real or generated based on the value controlling the mixture.
11. A computing device according to claim 1, wherein the one or more instances of the generated image information of the second part generated by the generator component include a plurality of instances of the generated image information of the second part, and Among them, the generator component generates the plurality of instances of the image information of the second part based on the plurality of corresponding instances of the image information of the first part and the random information.
12. A computing device according to claim 1, wherein the user interface presentation also displays the one or more instances of the generated image information of the second part.
13. A method for retrieving one or more output images implemented by one or more computing devices, including: In response to at least one input signal provided by a user using an input device, identify an image of a first portion; Receive image information of a first portion of the image of the first portion; Use a generator component implemented by the one or more computing devices to transform the image information of the first portion into one or more instances of generated image information of a second portion using a neural network, each instance of the generated image information of the second portion supplementing the image information of the first portion and being fake image information created by the generator component, rather than retrieved from an empirical distribution of pre-existing images; Based on the one or more instances of the generated image information of the second portion, retrieve one or more output images of the second portion from a data repository; Generate a user interface presentation presenting the image of the first portion and the one or more output images of the second portion, the user interface presentation including one or more graphical cues allowing the user to select any of the one or more output images of the second portion; And Display the user interface presentation on a display device for presentation to the user; The generator component has been trained by a computer-implemented training system using a conditional generative adversarial network (cGAN) through the following process: Identify a set of synthetic images, each synthetic image consisting of at least a specific image of a first portion and a specific image of a second portion; Segment the synthetic images to generate a plurality of extracted images of the first portion and a plurality of extracted images of the second portion; And Use the cGAN to train the generator component based on the plurality of extracted images of the first portion and the plurality of extracted images of the second portion.
14. The method according to claim 13, wherein the identified image of the first portion shows a first clothing item selected by the user, and each retrieved output image of the second portion shows a second clothing item supplementing the first clothing item.
15. The method according to claim 14, wherein the first clothing item is an upper body clothing item, and the second clothing item is a lower body clothing item.
16. The method according to claim 13, wherein the identified image of the first portion shows a first selectable product item selected by the user, and each retrieved output image of the second portion shows a second selectable product item supplementing the first selectable product item.
17. The method according to claim 13, wherein each instance of the generated image information of the second portion describes features associated with the generated image of the second portion.
18. The method according to claim 13, wherein each instance of the generated image information of the second portion describes pixels associated with the generated image of the second portion.
19. The method according to claim 13, wherein the method comprises: Select the generator component from a set of generator components implemented by the one or more computing devices, each generator component being associated with a pair of complementary item types.
20. The method according to claim 13, wherein the method further comprises: Apply two or more generator components implemented by the one or more computing devices, each generator component being associated with a pair of complementary item types.
21. The method according to claim 13, wherein the retrieval includes: Identify one or more extracted images of a second portion that match any instance of the one or more instances of the generated image information of the second portion; And Identify the output images of the one or more second portions as one or more corresponding images that match the extracted image of any second portion among the extracted images of the one or more second portions, and each output image of the second portion shows an image of a selectable product.
22. The method according to claim 13, wherein using a cGAN includes, for multiple instances of the image information of the first portion: Using the generator component to map a particular instance of the image information of the first portion together with an instance of random information taken from a distribution into a particular instance of the generated image information of the second portion; Using a discriminator component implemented by the one or more computing devices to provide a score that reflects the confidence level that the particular instance of the generated image information of the second portion represents a real pre-existing image of the second portion; And Adjusting the weight values of the generator component and the discriminator component at least based on the score.
23. A computer-readable storage medium for storing computer-readable instructions that, when executed by one or more hardware processors, perform a method that includes: Using a matching component to identify a set of synthetic images, each synthetic image being composed of at least an image of a first portion and an image of a second portion; Segmenting the synthetic images to generate multiple extracted images of the first portion and multiple extracted images of the second portion; And Using a conditional generative adversarial network (cGAN) to train a generator component based on the multiple extracted images of the first portion and the multiple extracted images of the second portion, wherein using the cGAN includes, for multiple instances of the image information of the first portion: Using the generator component to map an instance of the image information of the first portion together with an instance of random information taken from a distribution into an instance of the generated image information of the second portion; Using a discriminator component to provide a score that reflects the confidence level that the instance of the generated image information of the second portion represents a real pre-existing image of the second portion; And Adjusting the weight values of the generator component and the discriminator component at least based on the score, wherein the matching component, the generator component, and the discriminator component are implemented by the computer-readable instructions.
Citation Information
Patent Citations
Garment classification and collocation recommending method and garment classification and collocation recommending system based on deep convolution neural network
CN106504064A
Method, system and medium for garment wearing recommendation based on conditional generative adversarial nets
CN108829855A