Information processing device, information processing method, and information processing program
The use of a diffusion model with UNets and Set Transformer layers in an information processing device generates harmonious coordinated outfits by training on harmonious item interactions, addressing the lack of holistic harmony in conventional techniques.
Patent Information
- Application Number
- JP2024179863
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-10-15
AI Technical Summary
Conventional techniques fail to generate coordinated outfits that are harmonious as a whole, lacking the ability to provide information on highly harmonious fashion coordinates.
An information processing device and method using a diffusion model with multiple neural networks, specifically UNets and Set Transformer layers, to generate item images from noise images, considering the harmony between items, and incorporating interactions between items in coordinated outfits.
Enables the generation of highly harmonious coordinated outfits by training the model on harmonious item combinations, allowing for interactive image generation and editing, and suggesting ideal outfit combinations.
Smart Images

Figure 0007746503000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] 2. Description of the Related Art Conventionally, there are known techniques for generating fashion coordinates, such as techniques for generating fashion coordinates based on predetermined rules. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-235528 Summary of the Invention [Problem to be solved by the invention]
[0004] However, conventional techniques are unable to generate coordinates based on the harmony of all items, and therefore are unable to provide information on highly harmonious coordinates as a whole.
[0005] The present application has been made in view of the above, and aims to provide information that is highly harmonious as a whole in coordination. [Means for solving the problem]
[0006] The information processing device according to the present application is Remove the noise Generate item images By using an image generation model with multiple neural networks realized by a diffusion model trained as above, we can generate individual item images from individual noise images by incorporating the interactions between the items that make up the coordinates. Generate multiple item images from multiple noise images, taking into consideration the harmony with each other. Equipped with an image generation unit It is characterized by: [Effects of the Invention]
[0007] According to one aspect of the embodiment, it is possible to provide information that is highly harmonious as a whole coordinated outfit. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is an explanatory diagram showing an overview of an information processing system according to an embodiment. [Figure 2] FIG. 2 is a diagram showing an example of the configuration of an image generation model UN used by the server device when generating an image in consideration of harmony. [Figure 3] FIG. 3 is a diagram showing a first example of a UI for specifying image generation conditions. [Figure 4] FIG. 4 is a diagram showing a second example of a UI for specifying image generation conditions. [Figure 5] Figure 5 is an explanatory diagram outlining the design concept of the Outfit Diffusion architecture. [Figure 6] FIG. 6 is a diagram showing an example of the result of coordination. [Figure 7] FIG. 7 is a diagram illustrating an example of the configuration of a terminal device according to the embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of the configuration of a server device according to the embodiment. [Figure 9] FIG. 9 is a flowchart showing a processing procedure according to the embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an information processing device, an information processing method, and an information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to these embodiments. Furthermore, the same components in the following embodiments will be denoted by the same reference numerals, and duplicated descriptions will be omitted.
[0010] [1. Overview of the information processing system] First, an overview of an information processing system according to an embodiment will be described with reference to Fig. 1. Fig. 1 is an explanatory diagram showing an overview of an information processing system according to an embodiment. As shown in Fig. 1, an information processing system 1 according to an embodiment includes a terminal device 10 and a server device 100. These various devices are connected to each other via a network N in a wired or wireless manner so as to be able to communicate with each other. This enables the terminal device 10 to cooperate with the server device 100. The network N is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network) such as the Internet.
[0011] Terminal device 10 is an information processing device used by a user U. For example, terminal device 10 may be a smart device such as a smartphone or tablet terminal, a mobile phone such as a feature phone (Gala-ke or Gala-ho), a personal computer (PC), a personal digital assistant (PDA), a game console or AV device with communication functions, an information appliance or digital appliance, a car navigation system, a wearable device such as a smart watch, a head-mounted display, or smart glasses. Terminal device 10 may also be a house or building, a car, a home appliance, an electronic device, or the like that is compatible with the Internet of Things (IOT).
[0012] In this embodiment, the terminal device 10 is a smart device such as a smartphone or tablet terminal used by a user U, and is a mobile terminal device capable of communicating with any server device via a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation: fifth generation mobile communication system). The terminal device 10 has a screen such as a liquid crystal display with a touch panel function, and accepts various operations on displayed data such as content, such as tapping, sliding, and scrolling, performed by the user U with a finger or a stylus. Note that an operation performed on an area of the screen where content is displayed may be considered an operation on the content. The terminal device 10 may be not only a smart device, but also an information processing device such as a desktop PC (Personal Computer) or a notebook PC.
[0013] In addition, the terminal device 10 can connect to the network N via a wireless communication network such as LTE, 4G, or 5G, or via short-range wireless communication such as Bluetooth (registered trademark) or wireless LAN (Local Area Network), and communicate with the server device 100.
[0014] The server device 100 is, for example, a computer such as a PC or a blade server, or a mainframe or a workstation, etc. The server device 100 may be realized by cloud computing.
[0015] In this embodiment, the server device 100 is an information processing device that works in conjunction with the terminal device 10 of each user U and provides API (Application Programming Interface) services for various applications (hereinafter referred to as apps) and various data to the terminal device 10 of each user U, and is realized by a computer, a cloud system, etc.
[0016] The server device 100 may also be an information processing device that provides some kind of online web service to the terminal device 10 of each user U. For example, the server device 100 may provide the following web services: internet connection, search service, social networking service (SNS), electronic commerce (EC), electronic payment, online games, online banking, online trading, hotel and ticket reservations, video and music distribution, news, maps, route search, route guidance, line information, operation information, and weather forecast. In practice, the server device 100 may cooperate with various servers that provide the above-mentioned web services and act as an intermediary for the web services or may be responsible for processing the web services.
[0017] The server device 100 can acquire user information about the user U. For example, the server device 100 acquires, as the user information, information (attribute information) about the attributes of the user U, such as the gender, age, and residential area of the user U. The server device 100 can also acquire information about the attributes of the user U, such as demographic attributes, psychographic attributes, geographic attributes, and behavioral attributes. The server device 100 may also acquire, as the user information, a segment to which the user U belongs in the marketing field or a persona (personality). The server device 100 then stores and manages the information (attribute information) about the attributes of the user U together with identification information (such as a user ID) that identifies the user U.
[0018] The server device 100 also acquires various types of history information (log data) indicating the behavior of the user U from the terminal device 10 of the user U or from various servers based on the user ID, etc. For example, the server device 100 acquires a location history, which is a history of the user U's location and date and time, from the terminal device 10. The server device 100 also acquires a search history, which is a history of search queries entered by the user U, from an e-commerce server, a content server, a search server (search engine), etc. The server device 100 also acquires a browsing history, which is a history of content (fashion items, fashion coordination, etc.) viewed by the user U, from an e-commerce server, a content server, etc. The server device 100 also acquires a purchase history (payment history), which is a history of the user U's product purchases and payment processes, from an e-commerce server, a payment processing server, etc. The server device 100 may also acquire a listing history, which is a history of the user U's listings on the marketplace, and a sales history from an e-commerce server, a payment processing server, etc. Furthermore, the server device 100 acquires a posting history, which is a history of posts made by the user U, from a posting server that provides a word-of-mouth posting service, an SNS server, an e-commerce server, or the like. Note that the above-mentioned various servers may be the server device 100 itself. In other words, the server device 100 may function as the above-mentioned various servers.
[0019] Furthermore, the number of devices included in the information processing system 1 shown in Fig. 1 is not limited to that shown in the figure. For example, in Fig. 1, for the sake of simplicity, only one terminal device 10 is shown, but this is merely an example and is not limiting, and two or more devices may be included.
[0020] [2. Coordinate Image Generation Using Diffusion Model] A diffusion model is a type of generative AI (artificial intelligence) that is trained to learn the process of gradually adding noise to a target image to degrade it (the diffusion process), and the process of gradually removing noise (denoising) by reversing the degradation process and reconstructing the image (the de-diffusion process).
[0021] The server device 100 generates item images (product images) from noise images using a diffusion model. For example, the server device 100 generates multiple item images from multiple noise images, taking into consideration the harmony between each other. For example, the server device 100 generates individual item images by incorporating the interactions between items that are components of a coordinated outfit. A coordinated outfit is a set of items.
[0022] Here, the details of the process by which the server device 100 generates multiple item images from multiple noise images while taking into consideration the harmony with each other will be described. While one UNet is shown in FIG. 1, in reality, there are as many UNets as there are images to be generated. For example, if three images are generated simultaneously, there are three UNets. For each UNet, a Set Transformer layer ("Set Transformer" layer) is incorporated in an arbitrary layer of the UNet. The Set Transformer layer is connected to Set Transformer layers arranged in the same layer of other UNets, and is configured to output values based on the values of the Set Transformer layer of each UNet.
[0023] Fig. 2 is a diagram showing an example of the configuration of an image generation model UN used by the server device 100 when generating an image in consideration of harmony. As shown in Fig. 2, the image generation model UN has multiple networks UN1 and UN2. Note that in reality, the image generation model UN has the same number of networks as the number of images to be generated simultaneously, but the example shown in Fig. 2 shows two networks UN1 and UN2.
[0024] Network UN1 is a neural network that generates image P1 from noise image N1 and is realized by a so-called diffusion model. More specifically, network UN1 is a neural network that has been trained so that, when input with noise image N1 and category information C1 indicating an object for which an image is to be generated (for example, a fashion item category such as tops (e.g., shirts), jackets / outerwear (e.g., coats), pants, skirts, dresses, hats, bags, shoes, legwear (e.g., socks), and accessories (e.g., necklaces, rings, earrings, and the like)), network UN1 generates, from noise image N1, an image of the object indicated by the category information C1 (e.g., an image of the object, an image including the object, etc.). Note that such network UN1 can be realized by a learning method similar to the learning method of a so-called diffusion model, a model known in the art.
[0025] Network UN2 is Network UN 1 Similarly, it is a neural network that generates a target image P2 indicated by category information C2 from a noise image N2, and is realized by a so-called diffusion model learning method.
[0026] Here, networks UN1 and UN2 (hereinafter sometimes collectively referred to as network UN) have a structure called UNet. For example, network UN1 has layers UN1-1, ST1-1, UN1-2, ST1-2, UN1-3, ST1-3, UN1-4, ST1-4, UN1-5, and ST1-5, which are intermediate layers. Layers UN1-1, UN1-2, UN1-3, UN1-4, and UN1-5 (hereinafter collectively referred to as layer UN1) are layers that constitute a so-called UNet. As shown in FIG. 2, network UN1 has a configuration in which layers ST1-1, ST1-2, ST1-3, ST1-4, and ST1-5 (hereinafter collectively referred to as layer ST1), which are Set Transformer layers, are located after each layer of UNet constituted by layers UN1-1, UN1-2, UN1-3, UN1-4, and UN1-5.
[0027] Also, for example, network UN2 has layers UN2-1, ST2-1, UN2-2, ST2-2, UN2-3, ST2-3, UN2-4, ST2-4, UN2-5, and ST2-5, which are intermediate layers. Layers UN2-1, UN2-2, UN2-3, UN2-4, and UN2-5 (hereinafter collectively referred to as layer UN2) are layers that make up what is called a UNet. As shown in Figure 2, network UN2 has a configuration in which layers ST2-1, ST2-2, ST2-3, ST2-4, and ST2-5 (hereinafter collectively referred to as layer ST5), which are Set Transformer layers, are located after each layer of the UNet made up of layers UN2-1, UN2-2, UN2-3, UN2-4, and UN2-5.
[0028] Here, layers ST1 and ST2 are interconnected and are attention layers that learn to output values calculated based on values output from the previous layer (i.e., layers UN1 and UN2) and values received by other connected layers. For example, layer ST1-1 receives values output from layer UL1-1. Also, layer ST2-1 receives values output from layer UL2-1. In such a case, layer ST1-1 calculates a value to be output taking into account not only the value received from layer UL1-1 but also the value received by layer ST2-1 from layer UL2-1. Similarly, layer ST2-1 calculates a value to be output taking into account not only the value received from layer UL2-1 but also the value received by layer ST1-1 from layer UL1-1.
[0029] Similarly, layers ST1-2 to ST1-5 and layers ST2-2 to ST2-5 are similarly connected, and calculate output values taking into account not only the output from the previous layer but also the values received from the previous layer by other connected layers. If there are networks other than networks UN1 and UN2, each network has the same structure, and layers ST located at the same positions are interconnected. For example, if the image generation model UN includes network UN3, and layer ST3-1 is located at the same position as layers ST1-1 and ST2-1 in network UN3, layer ST1-1 will calculate an output value from the values received from the previous layer by layer ST1-1, taking into account all the values received by layers ST1-1, ST1-2, and ST1-3.
[0030] The image generation model UN, which is made up of such multiple networks, undergoes training similar to that of a diffusion model so as to generate (restore) different images from multiple noise images and multiple pieces of category information. In this case, by preparing a group of images of multiple harmonious objects and multiple pieces of category information indicating each object as training data, the image generation model UN, upon receiving multiple pieces of category information, can generate a group of images of multiple objects that are harmonious as a whole, which are a set of images including the objects indicated by each piece of category information.
[0031] For example, the server device 100 acquires, as training data, coordinated outfits published in magazines or e-commerce sites (i.e., coordinated outfits evaluated as harmonious by the publisher), coordinated outfits posted on social media, etc. (i.e., coordinated outfits evaluated as harmonious by the poster), and a group of images of multiple products that make up these coordinated outfits that have been highly rated by viewers (i.e., coordinated outfits that viewers also evaluate as harmonious), along with a plurality of category information indicating each product. The server device 100 may also acquire, as training data, a group of images of multiple products manually selected by a stylist or the like, along with a plurality of category information indicating each product, as a group of images of products that make up coordinated outfits evaluated as harmonious, along with a plurality of category information indicating each product. In addition to these methods, any combination of images and a plurality of category information collected by any method or any evaluation axis may be acquired as the correct answer data, as long as it acquires images of each product and category information indicating each product for multiple products that may be evaluated as harmonious. Note that "harmonious" may mean harmonious from an overall perspective, or harmonious from a specific perspective (for example, color, silhouette, size, design, etc., but not limited to the examples). Furthermore, the server device 100 may acquire, as learning data, information regarding the brand, color, silhouette, size, design, material, comfort, item image, genre (adult casual, etc.), season (winter outfit, etc.), occasion (wedding, etc.), era (Y2K, which incorporates trends from around the year 2000), wearer (physical information including body type, skin color, etc.), number of likes, desired outfit image, etc., and train the diffusion model.
[0032] Next, for a certain set of correct answer data, the server device 100 generates images by gradually adding noise to each image included in the correct answer data. Then, the server device 100 corrects the connection coefficients of the network included in the image generation model UN using various known learning methods similar to the diffusion model so that, for example, when multiple images with n levels of noise added and multiple pieces of category information are input to the image generation model UN, an image with n-1 levels of noise added is generated. By repeating this learning process, the image generation model UN does not simply generate multiple images, but generates images in which multiple objects included in each generated image are harmonious (for example, have a harmonious appearance).
[0033] The server device 100 then inputs a plurality of noise images and category information indicating each category of a plurality of objects for which images are desired to be generated into the image generation model UN that has undergone such learning, and by repeatedly inputting the images generated by the image generation model UN together with the category information into the image generation model UN, it is possible to generate images of each object indicated by the category information that have an overall harmonious appearance (i.e., images that are likely to be well-received on social media, or images that could be manually coordinated by a stylist, etc.).
[0034] Note that the server device 100 does not need to input category information, for example, when networks UN1 and UN2 are to generate target images of fixed categories. For example, if network UN1 corresponds to clothing and network UN2 corresponds to shoes, the server device 100 collects, as training data, image pairs of clothing and shoes that are evaluated as harmonious. The server device 100 then performs training so that, when noise images are input to networks UN1 and UN2, network UN1 reconstructs a clothing image that serves as ground truth data, and network UN2 reconstructs a shoe image that serves as ground truth data (a shoe image that is evaluated as harmonious with the clothing image reconstructed by network UN1). The image generation model UN that has undergone such training can output clothing images and shoe images that are estimated to be harmonious, for example, simply by inputting a plurality of noise images.
[0035] At this time, the server device 100 may set image generation conditions. Then, the server device 100 generates item images according to the image generation conditions. Note that the image generation conditions may include image generation conditions for each item and image generation conditions for the entire item. For example, the image generation conditions for each item may be image generation conditions that specify the category, brand, color, silhouette, size, design, material, comfort, desired image of the item, etc., and are not limited to the examples shown. The image generation conditions may also be image generation conditions for a specific part. For example, they may be image generation conditions that specify the silhouette of the collar. Furthermore, for example, the image generation conditions for the entire item may be image generation conditions that specify the genre (such as adult casual), season (such as winter outfits), occasion (such as weddings), era (such as Y2K, which incorporates trends from around the year 2000), wearer (such as physical information including body type, skin color, etc.), number of likes, desired coordinated image, etc., and are not limited to the examples shown. Note that the image generation conditions exemplified for each item may be used as the image generation conditions for the whole, and vice versa.
[0036] The server device 100 may also acquire one or more fixed item images. The server device 100 then generates item images taking into consideration harmony with the fixed item images. The fixed item images are images of items that the user owns, items that the user has registered as favorites, items that the user has searched for or viewed, and other items that the user wishes to include in a coordination. The server device 100 may acquire, as the fixed item images, images that are registered or posted on an e-commerce site, a coordination site, or the like, or may acquire images that the user has uploaded.
[0037] The server device 100 also acquires one or more pre-edited item images, sets image editing conditions, and generates an item image from the pre-edited item image via a noise image according to the image editing conditions. Specifically, the server device 100 adds noise to all or a portion of the pre-edited item image (also referred to as the portion to be edited; for example, the collar portion if the image editing conditions specify the silhouette of the collar of a dress) according to the image editing conditions, and then removes the noise to generate an item image. The pre-edited item image may be an image of an item owned by the user, an image of an item registered as a favorite by the user, an image of an item searched or viewed by the user, or an image of an item that is close to the image of an item desired to be included in an outfit. The server device 100 may acquire, as the pre-edited item image, an image registered or posted on an e-commerce site or a coordination site, or an image uploaded by the user. The image editing conditions may include image editing conditions for each pre-edited item and image editing conditions for the entire item. For example, the image editing conditions for each pre-edited item may be image editing conditions that specify the category, brand, color, silhouette, size, design, material, comfort, desired image of the item, etc., and are not limited to the examples. The image editing conditions may also be image editing conditions related to a specific part. For example, they may be image editing conditions that specify the silhouette of the collar. Furthermore, for example, the image editing conditions for the entire item may be image editing conditions that specify the genre (e.g., adult casual), season (e.g., winter outfit), occasion (e.g., wedding), era (e.g., Y2K, which incorporates trends from around the year 2000), wearer (e.g., physical information including body type, skin color, etc.), number of likes, desired coordinated image, etc., and are not limited to the examples. The example image editing conditions for each pre-edited item may be the image editing conditions for the entire item, or the example image editing conditions for the entire item may be the image generation conditions for each pre-edited item.
[0038] The server device 100 also constructs a diffusion model that learns a diffusion process that adds noise to each of a plurality of harmonious item images to generate a plurality of noise images, and a de-diffusion process that removes noise from each of the plurality of noise images to generate a plurality of original item images.The server device 100 then generates the item images using the diffusion model.
[0039] Furthermore, the terminal device 10 of the user U cooperates with the server device 100 to display an item image generated from a noise image. For example, the terminal device 10 of the user U displays a plurality of item images generated from a plurality of noise images, taking into consideration harmony with each other.
[0040] Furthermore, the terminal device 10 of the user U cooperates with the server device 100 to display the generated item image as a query image. Thereafter, the terminal device 10 of the user U searches for content (items, coordination) based on the displayed query image. Then, the terminal device 10 of the user U displays the content that is the search result.
[0041] Furthermore, the terminal device 10 of the user U accepts image generation conditions from the user U. The terminal device 10 of the user U may also accept a fixed item image from the user U. Then, the terminal device 10 of the user U cooperates with the server device 100 to display the generated item image in accordance with the image generation conditions.
[0042] Furthermore, the terminal device 10 of the user U receives one or more pre-edited item images and image editing conditions from the user U. Then, the terminal device 10 of the user U cooperates with the server device 100 and displays an item image generated from the pre-edited item image via a noise image in accordance with the image editing conditions.
[0043] Furthermore, the terminal device 10 of the user U searches for items and coordination based on each of the multiple query images in cooperation with the server device 100. Furthermore, the terminal device 10 of the user U displays the search results.
[0044] Furthermore, the terminal device 10 of the user U cooperates with the server device 100 to search for an item based on a query image designated by the user U from among a plurality of query images. The terminal device 10 of the user U displays the search results.
[0045] 1, the server device 100 receives image generation conditions from the terminal device 10 of the user U via the network N and sets the image generation conditions (step S1). At this time, the terminal device 10 of the user U displays a UI (User Interface) for specifying the image generation conditions.
[0046] An image of a UI for specifying image generation conditions will be described with reference to Figures 3 and 4. Figure 3 is a diagram showing a first example of a UI for specifying image generation conditions. Figure 4 is a diagram showing a second example of a UI for specifying image generation conditions.
[0047] As shown in Figures 3 and 4, the UI for specifying image generation conditions and image editing conditions includes a text input field TX, a condition setting button B1 (also called the image generation button), an item image display field IMG, a category etc. selection field CS, an item fixation checkbox CB for setting an item to be fixed, an item search button B2, and an item search result display field SR.
[0048] The text input field TX is a field for inputting image generation conditions and image editing conditions as text. Note that a prompt (instruction) may be input into the text input field TX as an image generation condition. When the condition setting button B1 is pressed by the user, the input text, the selected category, etc. are set as the image generation conditions and image editing conditions. Then, an item image is generated according to the image generation conditions and image editing conditions.
[0049] The item image display field IMG is a field for registering item images. It is also a field for displaying generated item images. The registered item images are item images of items desired to be included in the coordination (fixed item images) or item images of items before editing that are close to the image of the items desired to be included in the coordination (pre-edited item images). In this embodiment, multiple item images are registered. At this time, one item image is displayed in one item image display field IMG. Furthermore, when the user presses the condition setting button B1, a generated item image is displayed in the item image display field IMG, either newly or in place of the pre-edited item image. At this time, one item image is displayed in one item image display field IMG. The item image displayed in the item image display field IMG becomes the query image.
[0050] There are as many category selection fields CS and item fixing check boxes CB as there are item images. That is, there are as many item image display fields IMG. The number of item image display fields IMG can be increased by the user U.
[0051] The category etc. selection field CS is a field for selecting the category etc. of the item to be generated (image generation conditions and image editing conditions). If an item category etc. is specified, an item image will be generated according to the specified category etc. when the user presses the condition setting button B1. Conversely, if an item category etc. is not specified, an item image will be generated randomly regardless of the category etc. Note that even if an item category etc. is not specified, if text is entered in the text input field TX, an item image will be generated according to the text input.
[0052] The item fixation checkbox CB is a configuration for setting an item to be fixed. When the item fixation checkbox CB is checked to fix an item, the corresponding item image becomes a fixed item image, and the item is fixed and unchanging within the coordinate. In other words, when editing an item within the coordinate, it will not be changed to the image of another item. Note that other configurations such as radio buttons may be used instead of checkboxes.
[0053] When the user presses the item search button B2, a search for an item is performed using the query image displayed in the item image display field IMG. The search results are displayed in the item search result display field SR. In the example of FIG. 3, an item search result display field SR exists for each item image (the same number as the number of item images), but in the example of FIG. 4, the item search result display field SR is grouped together for all item images (coordinate units).
[0054] By registering the item image displayed in the item search result display field SR in the item image display field IMG, the item image can be used as the next fixed item image or the pre-edited item image. At this time, it is also possible to register all item images in a coordination in the item image display field IMG on a coordination basis.
[0055] 3 and 4, one set of the text input field TX, the condition setting button B1, the item image display field IMG, the category etc. selection field CS, and the check box CB is provided, but multiple sets may be provided. Furthermore, one or more of the text input field TX, the condition setting button B1, the category etc. selection field CS, and the check box CB may be provided as common elements for multiple sets. That is, item images may be generated by changing the image generation conditions and image editing conditions for each set, or item images may be generated by setting the image generation conditions and image editing conditions to be the same for multiple sets. Then, the user may select a set of item images that they like from the multiple sets of item images and use them as a query image to perform an item search.
[0056] Next, the server device 100 constructs a diffusion model that has learned a diffusion process that adds noise to each of the plurality of item images to generate a plurality of noise images, and a de-diffusion process that removes noise from each of the plurality of noise images to generate the original plurality of item images (step S2). Note that step S2 may be executed before step S1.
[0057] Next, the server device 100 generates an item image from the noise image using a diffusion model (step S3). At this time, the server device 100 acquires one or more fixed item images, generates an item image in consideration of harmony with the fixed item image according to the image generation conditions and image editing conditions, and displays the generated item image on the terminal device 10 of the user U. Alternatively, the server device 100 acquires one or more pre-edited item images, generates an item image from the pre-edited item image via a noise image according to the image editing conditions, and displays the generated item image on the terminal device 10 of the user U.
[0058] Next, the server device 100 displays the generated item images and fixed item images as query image candidates on the terminal device 10 of the user U (step S4). If the generated item image is not to the user's liking, the user may specify the image generation conditions and image editing conditions again to generate an item image, or may generate an item image again using the same image generation conditions and image editing conditions. Furthermore, if the user likes one of the generated item images, the user may fix that item image and generate another item image.
[0059] Next, the server device 100 searches for items based on each of the multiple query images, or based on a query image specified by the user U (user) from among the multiple query images, and displays the search results on the terminal device 10 of the user U (step S5).
[0060] [2-1. Item image generation taking coordination into consideration] Conventionally, coordination suggestions have been made using a set matching model that outputs a set matching score (SetMatchingscore) that indicates the compatibility between a group of multiple input items. However, the feature quantities of the items input into the set matching model (which are also the feature quantities used to extract actual items) are incomprehensible to humans. Therefore, if a user does not like the proposed coordination, it is not possible to determine whether the feature quantities of the items are not suitable for the user (whether the feature quantities of the items should be changed) or whether the actual items extracted based on the feature quantities of the items are not suitable for the user (whether the actual items to be extracted should be changed without changing the feature quantities of the items), which could prevent the system from proposing a truly ideal coordination.
[0061] In this embodiment, the server device 100 generates an item image (query image) to be used in the search before searching for the actual item, thereby making it possible to visualize the feature quantities of the item and to propose a truly ideal coordination. Furthermore, the server device 100 can interactively generate and edit the item image while visualizing it.
[0062] To achieve this, an architecture that satisfies the following two properties is required. · Swap items in your outfit for symmetry Variability of item images to patches
[0063] Therefore, as shown in Figure 1, we add a Set Transformer layer to the diffusion model UNet to incorporate interactions between items in the coordinates. UNet is a type of FCN (fully convolution network) and is a network for estimating image segmentation (where objects are located). The Set Transformer layer mainly consists of the following two parts: (A) and (B).
[0064] (A)SA+FF(SelfAttention + FeedForward) The interaction between coordinated items is calculated by taking the sum of height and width and calculating SelfAttention.
[0065] (B)CA+FF(CrossAttention + FeedForward) By calculating the cross-attention between the output of SA+FF and the original image, the item image is transformed to reflect the interactions between the coordinated items.
[0066] As an application, the server device 100 generates a coordinated outfit by inputting conditions such as a category. Possible conditions include, for example, category, brand, color, time, material, usage scenario on ZOZOTOWN (registered trademark), description of the coordinated outfit, description of each item, location, travel destination, scenario, user information (user cluster, browsing history), number of likes, etc. However, in reality, the conditions are not limited to these examples. In this case, the server device 100 can perform the following information processing.
[0067] (1) Item search The server device 100 generates images of items to be searched for based on images of items that the user actually owns. To obtain images of items that the user actually owns, the server device 100 may accept registration of images of items owned by the user, or may extract images of items included in past purchase history (items with purchase records). In other words, the server device 100 can generate one or more item images (query images) based on fixed item images and pre-edited item images. This enables the server device 100 to suggest ways to mix and match items that the user actually owns and search for highly harmonious items. Another advantage of image generation is that the query image used in item search can be visualized.
[0068] After generating an image of the item to be searched, the server device 100, as a user interaction, presents the image to be used as a query image for the search to the user and asks a confirmation or inquiry such as "Is this query okay?". Specifically, the server device 100 displays the generated item image in the item image display field IMG and asks a confirmation or inquiry such as "Is this query okay?". Alternatively, the server device 100 may accept a designation of an attribute (such as a color image or a V-neck) from the user (accept image generation conditions or image editing conditions from the user) and generate an image to be used as a query image in accordance with the designation. Alternatively, the server device 100 may present each of the generated images to the user and allow the user to select an image to be used as a query image from among them.
[0069] (2) Virtual try-on of fictitious items After generating the item images, the server device 100 generates snapshots using a different method, such as virtual try-on. Specifically, the server device 100 generates an image of the user wearing one or more generated items based on the one or more generated item images and the user image. If a fixed item image is set, the server device 100 generates an image of the user also wearing the fixed item. This allows the user to determine whether the generated item image is appropriate as a query image while keeping in mind the image of the item to be searched.
[0070] (3) Item search, coordination search The server device 100 searches for actual items and coordinations by combining similar item search and similar coordination search techniques. Specifically, the server device 100 embeds each generated item image and generated coordination image (including a coordination image generated by combining each generated item image) into a distributed representation space, and searches for item images of actual items and coordination images using actual items based on the similarity of vectors in the distributed representation space.
[0071] (4) Increase in likes The server device 100 generates item images of some or all of the multiple items that make up an outfit and suggests to the user, "If you wear this, you'll get more likes." Specifically, the server device 100 accepts text, such as the user's desired number of likes or an outfit that is likely to get more likes, as image generation conditions or image editing conditions, and generates item images of some or all of the multiple items that make up an outfit that is likely to get the user's desired number of likes or an outfit that is likely to get more likes, and suggests to the user, "If you wear this, you'll get more likes." The user can then determine whether the query image is likely to find an item or outfit that is likely to get more likes before searching for an item. Note that the server device 100 acquires, as training data, a group of images of the multiple items that make up an outfit for each number of likes and multiple category information indicating each item as correct answer data, and trains a diffusion model using these acquired images.
[0072] (5) Forecasting and Trends As a trend analysis, the server device 100 generates something like an average image of generated images in response to prompts (prompts: instructions) such as "What's popular now?" or "Coordination that was popular in XX year." For example, the server device 100 performs interpolation processing using GAN (Generative Adversarial Networks). This makes it possible to know, for example, the time evolution of the average image of adult casual wear.
[0073] The server device 100 also generates an image of the next desired item. For example, the server device 100 generates an image of the item by predicting future news based on information sources such as what images are frequently generated and what items sell well when certain news is released.
[0074] The server device 100 also receives image generation conditions and image editing conditions, such as coordination that looks like it's from the year 2000 or coordination that will likely be popular in 2025 (future), and generates images of each item. Fashion trends are generally said to occur in 20-year cycles, and by learning from past trends, it is possible to propose coordination that reflects future trends. The server device 100 acquires, as training data, a group of images of multiple items that make up coordination for each era and multiple pieces of category information indicating each item as correct answer data, and trains the diffusion model using these acquired data.
[0075] (6) Personalization The server device 100 learns negative prompts for each individual user.
[0076] (7) Coordinate generation The server device 100 generates a coordination based on items that the user actually owns. Then, the server device 100 performs an image search within ZOZOTOWN (registered trademark) based on the generated results. This differs from a collective search in that it can be performed interactively. Specifically, the server device 100 acquires images of two or more items that the user actually owns (pre-edited item images), and generates a query image from each of the acquired images of the two or more items. Then, the server device 100 uses the generated two or more query images to search for and display images of the actual two or more items.
[0077] (8) Editing items in your outfit The server device 100 edits the items in the outfit. At this time, the server device 100 performs interactive editing. The editing direction may be, for example, "What would happen if we made an outfit using item images from the 1990s more modern?" Specifically, the server device 100 acquires images of two or more items from the 1990s that make up the outfit (item images before editing) and an image editing condition such as an outfit that looks like it's from 2024, and generates new item images from each of the acquired images of the two or more items in accordance with the image editing condition, taking into consideration how they harmonize with each other.
[0078] The server device 100 may also edit items in the outfit according to specific tags (style information such as adult casual) acquired as image editing conditions. The server device 100 also displays that the outfit is likely to increase the number of likes, to show that the outfit suits the user.
[0079] Furthermore, the server device 100 generates coordinations (fill in the blanks) from the perspective of, "What items should be added to old items (e.g., 2019) to create a modern (e.g., 2024) look?" Specifically, the server device 100 acquires images of one or more items from 2019 (fixed item images) and an image generation condition of a 2024-like outfit, and generates images of items to be added in accordance with the image generation condition, taking into consideration harmony with the fixed item images. Treating 360-degree images of items as a set enables the generation and editing of items that are consistent across 360 degrees.
[0080] (9) New forms of search This embodiment is expected to create a new form of search. For example, when a user inputs a search word (image generation conditions or image editing conditions) or performs one-click recommendations based on the products they own (fixed item images or pre-edited item images), multiple images (generated query images) are displayed as candidates with the question "Is this image close to what you had in mind?", and a search method can be realized in which the user can select from among them (the user can search for images of actual items similar to the selected query image).
[0081] (10) Assistance in creating your own avatar The server device 100 can also assist in creating a custom avatar. For example, a game or the like has a function for creating clothes for one's avatar, but creating clothes from scratch can be tedious, so it would be convenient to be able to create a coordination (a coordination image consisting of multiple item images generated in consideration of harmony with each other) from text (image generation conditions). If the user does not like the coordination, it may be possible to allow the user to later correct the details themselves (the generated item image may be used as a pre-edit item image, and the item image may be generated again according to the image editing conditions). It is also possible to have the avatar wear clothing that the user (the user) actually wears.
[0082] In this way, the server device 100 generates product images from noise using a diffusion model. In particular, the server device 100 calculates the interactions between items using the Set Transformer and generates images that are likely to be compatible with the items set as conditions. This makes it possible to fill in gaps in coordination. It also makes it possible to generate an unlimited number of coordinations.
[0083] Note that the objects to be generated are not limited to images of items. For example, the server device 100 may generate snapshots (images of outfits, images of people wearing items, images of people wearing outfits, etc.) in addition to images of items. Snapshots posted by a specific poster on an SNS or the like tend to harmonize because they have similar tastes. For example, the server device 100 may generate a plurality of snapshots similar to those posted by a specific poster. Specifically, the server device 100 may use a diffusion model trained on a plurality of snapshots posted by a specific poster as training data to generate a plurality of snapshots from a noise image while taking into consideration the harmony between each of the snapshots.
[0084] Furthermore, the server device 100 may be configured to add a score output module so that the degree of match and suitability can be measured.
[0085] The server device 100 may generate an item set that increases the set matching score by combining it with other models. The server device 100 may also generate an item set by combining it with FashionGPT.
[0086] [2-2.Outfit Diffusion Architecture Design Concept] The three properties that must be satisfied are as follows: (a) The symmetry of item replacement within the outfit (regardless of the order of the product images) (b) Flexibility in the number of items in a coordinated outfit (the number of product images that make up a coordinated outfit can be set arbitrarily) (c) Variability of the item image to the patch (I want to be able to set the height and width arbitrarily, like StableDiffusion)
[0087] If the processing is carried out as follows, the above three conditions (a) to (c) above) can be satisfied.
[0088] (1) Input of product image features The server device 100 receives input of the feature values of product images that make up a coordinated outfit. The feature values of each product have an image-like shape (height, width, channel). For example, image features are expressed as a solid shape of batch size x sequence length x feature dimension.
[0089] (2) Creating features that represent overall characteristics The server device 100 creates a feature quantity representing the overall feature of each product from the patch-like feature quantity of each product image (conversion for each image).
[0090] For example, the following methods can be considered as methods for creating features. How to use classifier tokens like Vision Transformer (ViT) How to use Q-former (query transformer) Simply summing the feature values for each patch against the height and width
[0091] (3) Conversion to features taking into account permutation symmetry The server device 100 converts the feature obtained in (2) above into a feature that takes into account permutation symmetry (conversion for the entire outfit). Typically, this is SA+FF (SelfAttention+FeedForward). This satisfies (a) and (b) above. In addition, the server device 100 converts the feature obtained in (2) above into a feature that takes into account permutation symmetry (conversion for the entire outfit). This is typically SA+FF (SelfAttention+FeedForward). This satisfies (a) and (b) above. In addition, the server device 100 converts the feature obtained by referring to the feature of other products in the outfit, and acquires the overall features of the product image that result in a natural outfit.
[0092] (4) Transformation of input features The server device 100 converts the input feature for each image based on the feature obtained in (3) above (conversion for each image). Typically, this is CA+FF (CrossAttention+FeedForward), which satisfies the property in (c) above.
[0093] (5) Output of product image features The server device 100 outputs the feature quantities of the converted product images.
[0094] The above content can be illustrated in Figure 5. Figure 5 is an explanatory diagram that shows an overview of the design concept of the Outfit Diffusion architecture.
[0095] For example, as shown in Fig. 5, the server device 100 extracts features of each image using a QF (Q-Former) in an ST (Set Transformer module) (step S11). For example, the server device 100 extracts overall features of a product image by inputting the product image into the QF (Q-Former). The QF (Q-Former) consists of two sub-modules, an image converter and a text converter, and in this case, the image converter interacts with an image encoder that accepts the input image to extract visual features.
[0096] Next, the server device 100 performs conversion on the overall features of the extracted product image, taking into account the coordinateability, using SA+FF (SelfAttention+FeedForward) in the ST part (step S12).
[0097] Next, the server device 100 converts the original product image based on the feature amount that takes into account the coordination property by using CA+FF (CrossAttention+FeedForward) of the ST part (step S13).
[0098] FIG. 6 shows examples of coordinated outfits. Examples of results of this embodiment are (1) a coordinated outfit based on the IQQN3000 (2) with the same category; (3) a hat, earrings, coat, and skirt generated from a noise image using Fill in the n blanks; (4) an edited coordinate; and (5) a partially edited coordinate. The original coordinated outfit (1) is a monotone outfit, but in coordinated outfit (2), the categories of each item are maintained (the original categories are set as the image editing conditions for each item) and matching items (shoes, hat, earrings, coat, skirt, gloves, sweater) are generated from noise images. For example, the coat is changed from black to light blue, and the skirt is changed from black to brown. In coordinated outfit (3), the original shoes, gloves, and sweater are fixed, and the categories of the other items (hat, earrings, coat, skirt) are maintained, and matching items are generated from noise images. For example, the coat has been changed to a loose-fitting blue one, and the skirt to a navy blue pattern. Furthermore, the hat and earrings have been changed to brown, adding an accent to the overall color scheme. In the coordinations (4) and (5), the overall coordination and the atmosphere of each item have been maintained, but the texture has been changed.
[0099] [2-3. Supplementary Information] The server device 100 proposes coordination based on the generated multiple product images. The server device 100 has an architecture in which multiple UNets with the same connection coefficients communicate with each other. The feature values between images are aggregated each time. At this time, it is also possible to aggregate only a portion of the feature values between images.
[0100] If the category is given in advance, the server device 100 may insert the Set Transformer only in the last few times. Conversely, if the category is not given in advance, the server device 100 needs to insert the Set Transformer from the beginning. Otherwise, there is a possibility that, for example, two pairs of shoes will be displayed.
[0101] The server device 100 may also generate item images by combining a diffusion model and a set matching model. The set matching model is a model trained to output a set matching score indicating the compatibility between at least a first item group (plurality of item images) and a second item group (plurality of item images) when they are input. The set matching model may also be a model trained to output a set matching score indicating the compatibility between the first item group, the second item group, and the additional information when additional information (e.g., user information of a user wearing each item (e.g., including physical information such as body shape and skin color)) is input in addition to the first item group and the second item group. The set matching model may be trained using images and user information of users wearing outfits posted on e-commerce sites, outfit posting sites, various social networking sites, etc. as training data (correct answer data). Alternatively, the set matching model may be trained using images and user information of users wearing outfits posted on e-commerce sites, outfit posting sites, various social networking sites, etc. that have been evaluated by viewers or evaluators (e.g., evaluated as looking good or harmonious) as training data (correct answer data). In other words, the set matching model may be trained to output a higher set matching score the closer the input data is to the correct answer data. Furthermore, the server device 100 may use a diffusion model to input at least intermediate images during the process of generating (restoring) item images from noise images to the set matching model, output (calculate) a set matching score, and cause the diffusion model to generate (restor) item images that will result in a high set matching score. The network may be a ControlNet instead of a UNet. The ControlNet can create line drawings and generate images based on the line drawings. Examples of line drawings include line drawings corresponding to categories and line drawings desired by users.
[0102] Furthermore, the server device 100 adds noise to an image that it actually has to create a noise image, and then compares the input image with this noise image to remove noise and restore the image without creating restored data.
[0103] The server device 100 restores the items up to the point where it is no longer necessary to add categories. The server device 100 only needs to be able to find items other than those the user owns. There may also be a dedicated model with a function to restore items in a fill-in-the-blank manner.
[0104] [3. Example of terminal device configuration] Next, the configuration of the terminal device 10 will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the configuration of the terminal device 10 according to the embodiment. As shown in Fig. 7, the terminal device 10 includes a communication unit 11, a display unit 12, an input unit 13, a positioning unit 14, a sensor unit 20, a control unit 30 (controller), and a storage unit 40.
[0105] (Communications Department 11) The communication unit 11 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 11 is realized by a NIC (Network Interface Card), an antenna, etc.
[0106] (Display section 12) Display unit 12 is a display device that displays various information such as position information. For example, display unit 12 is a liquid crystal display (LCD) or an organic electro-luminescent display (OLED). Display unit 12 is also a touch panel display, but is not limited to this.
[0107] (Input section 13) The input unit 13 is an input device that accepts various operations from the user U. For example, the input unit 13 has buttons for inputting characters, numbers, etc. The input unit 13 may be an input / output port (I / O port), a USB (Universal Serial Bus) port, etc. If the display unit 12 is a touch panel display, a part of the display unit 12 functions as the input unit 13. The input unit 13 may be a microphone that accepts voice input from the user U. The microphone may be wireless.
[0108] (Positioning unit 14) The positioning unit 14 receives signals (radio waves) transmitted from satellites of a GPS (Global Positioning System), and acquires position information (e.g., latitude and longitude) indicating the current position of the terminal device 10, which is the device itself, based on the received signals. That is, the positioning unit 14 positions the position of the terminal device 10. Note that GPS is merely an example of a GNSS (Global Navigation Satellite System).
[0109] The positioning unit 14 can also measure the position using various methods other than GPS. For example, the positioning unit 14 may measure the position by using various communication functions of the terminal device 10 as an auxiliary positioning means for position correction, etc., as described below.
[0110] (Wi-Fi positioning) For example, the positioning unit 14 uses a Wi-Fi (registered trademark) communication function of the terminal device 10 or a communication network provided by each communication company to measure the position of the terminal device 10. Specifically, the positioning unit 14 performs Wi-Fi communication or the like and measures the distance to a nearby base station or access point, thereby measuring the position of the terminal device 10.
[0111] (Beacon positioning) The positioning unit 14 may also measure the position by using a Bluetooth (registered trademark) function of the terminal device 10. For example, the positioning unit 14 measures the position of the terminal device 10 by connecting to a beacon transmitter connected by the Bluetooth (registered trademark) function.
[0112] (geomagnetic positioning) The positioning unit 14 also measures the position of the terminal device 10 based on a geomagnetic pattern of a structure that has been measured in advance and a geomagnetic sensor that the terminal device 10 has.
[0113] (RFID positioning) Furthermore, for example, if the terminal device 10 has a function of an RFID (Radio Frequency Identification) tag equivalent to a contactless IC card used at station ticket gates, in stores, etc., or has a function of reading an RFID tag, the location where the terminal device 10 was used is recorded together with information on the payment or the like made by the terminal device 10. The positioning unit 14 may obtain such information to determine the location of the terminal device 10. Alternatively, the location may be determined by an optical sensor, an infrared sensor, or the like provided in the terminal device 10.
[0114] The positioning unit 14 may measure the position of the terminal device 10 using one or a combination of the above-mentioned positioning means, as needed.
[0115] (Sensor unit 20) The sensor unit 20 includes various sensors mounted on or connected to the terminal device 10. The connection may be wired or wireless. For example, the sensors may be detection devices other than the terminal device 10, such as wearable devices or wireless devices. In the example shown in FIG. 7 , the sensor unit 20 includes an acceleration sensor 21, a gyro sensor 22, a barometric pressure sensor 23, a temperature sensor 24, a sound sensor 25, a light sensor 26, a magnetic sensor 27, and an image sensor (camera) 28.
[0116] The above-described sensors 21 to 28 are merely examples and are not intended to be limiting. That is, the sensor unit 20 may be configured to include some of the sensors 21 to 28, or may include other sensors such as a humidity sensor in addition to or instead of the sensors 21 to 28.
[0117] The acceleration sensor 21 is, for example, a three-axis acceleration sensor, and detects physical movements of the terminal device 10, such as the direction of movement, speed, and acceleration of the terminal device 10. The gyro sensor 22 detects physical movements of the terminal device 10, such as tilt in three axial directions, based on the angular velocity of the terminal device 10. The air pressure sensor 23 detects, for example, the air pressure around the terminal device 10.
[0118] Since the terminal device 10 includes the acceleration sensor 21, the gyro sensor 22, the atmospheric pressure sensor 23, etc., it is possible to measure the position of the terminal device 10 using a technique such as Pedestrian Dead-Reckoning (PDR) that uses these sensors 21 to 23. This makes it possible to obtain indoor position information that is difficult to obtain using a positioning system such as GPS.
[0119] For example, the number of steps, walking speed, and distance walked can be calculated using a pedometer that uses the acceleration sensor 21. In addition, the direction of travel, line of sight, and body tilt of the user U can be determined using the gyro sensor 22. In addition, the altitude and floor on which the terminal device 10 of the user U is located can be determined from the air pressure detected by the air pressure sensor 23.
[0120] The temperature sensor 24 detects, for example, the temperature around the terminal device 10. The sound sensor 25 detects, for example, the sound around the terminal device 10. The light sensor 26 detects the illuminance around the terminal device 10. The magnetic sensor 27 detects, for example, the geomagnetism around the terminal device 10. The image sensor 28 captures an image around the terminal device 10.
[0121] The above-mentioned air pressure sensor 23, temperature sensor 24, sound sensor 25, light sensor 26, and image sensor 28 can detect the air pressure, temperature, sound, and illuminance, respectively, and capture images of the surroundings, thereby detecting the environment and situation around the terminal device 10. Furthermore, the accuracy of the location information of the terminal device 10 can be improved based on the environment and situation around the terminal device 10.
[0122] (control unit 30) The control unit 30 includes, for example, a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM, input / output ports, etc., and various other circuits. The control unit 30 may also be configured with hardware such as an integrated circuit, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 30 includes a transmitting unit 31, a receiving unit 32, and a processing unit 33.
[0123] (Transmitter 31) The transmission unit 31 can transmit, for example, various information input by the user U using the input unit 13, various information detected by each sensor 21 to 28 mounted on or connected to the terminal device 10, and location information of the terminal device 10 measured by the positioning unit 14 to the server device 100 via the communication unit 11.
[0124] (Receiving unit 32) The receiving unit 32 can receive various types of information provided by the server device 100 and requests for various types of information from the server device 100 via the communication unit 11.
[0125] (Processing unit 33) The processing unit 33 controls the entire terminal device 10, including the display unit 12. For example, the processing unit 33 can output various information transmitted by the transmitting unit 31 and various information received from the server device 100 by the receiving unit 32 to the display unit 12 for display.
[0126] Furthermore, the processing unit 33 may function (operate) as a reception unit 33A, a search unit 33B, and a display control unit 33C as described below by starting an app or the like. That is, the processing unit 33 includes the reception unit 33A, the search unit 33B, and the display control unit 33C.
[0127] (Reception Section 33A) The reception unit 33A receives image generation conditions from the user U via the input unit 13. The reception unit 33A also receives one or more fixed item images. The reception unit 33A also receives one or more pre-edited item images and image editing conditions.
[0128] (Search unit 33B) The search unit 33B cooperates with the server device 100 via the communication unit 11 and searches for content such as items or coordinates based on a query image. For example, the search unit 33B searches for items based on each of a plurality of query images. Alternatively, the search unit 33B searches for items based on a query image designated by the user from among a plurality of query images. Note that the search unit 33B may use a fixed item image or a pre-edited item image received by the receiving unit 33A as the query image.
[0129] (Display control unit 33C) The display control unit 33C cooperates with the server device 100 via the communication unit 11, and displays content that is a search result by the search unit 33B on the display unit 12. For example, the display control unit 33C displays the search results for each of a plurality of query images on the display unit 12. Alternatively, the display control unit 33C displays the search results based on a query image designated by the user on the display unit 12.
[0130] Furthermore, the display control unit 33C displays the generated item image on the display unit 12 in accordance with the image generation conditions received by the reception unit 33A. Furthermore, the display control unit 33C displays the item image generated from the pre-edited item image via a noise image on the display unit 12 in accordance with the image editing conditions received by the reception unit 33A.
[0131] The display control unit 33C cooperates with the server device 100 via the communication unit 11, and displays an item image generated from a noise image on the display unit 12. For example, the display control unit 33C displays a plurality of item images generated from a plurality of noise images, taking into consideration harmony with each other, on the display unit 12.
[0132] (Storage unit 40) The storage unit 40 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), an optical disk, etc. The storage unit 40 stores various programs, various data, etc.
[0133] [4. Server device configuration example] Next, the configuration of the server device 100 according to the embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing an example of the configuration of the server device 100 according to the embodiment. As shown in Fig. 8, the server device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.
[0134] (Communication unit 110) The communication unit 110 is realized by, for example, a network interface card (NIC), etc. The communication unit 110 is connected to a network N by wire or wirelessly.
[0135] (Storage unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as an HDD, an SSD, an optical disk, etc. The storage unit 120 may store attribute information and history information (log data) of the user U together with identification information (such as a user ID) indicating the user U.
[0136] (control unit 130) The control unit 130 is a controller, and is realized by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like, executing various programs (corresponding to an example of an information processing program) stored in a storage device inside the server device 100 using a storage area such as a RAM as a working area. In the example shown in FIG. 8, the control unit 130 has an acquisition unit 131, a setting unit 132, a construction unit 133, an image generation unit 134, and a provision unit 135.
[0137] (Acquisition part 131) The acquisition unit 131 acquires a search query input by a user U. For example, when the user U inputs a search query into a search engine or the like to perform a keyword search, the acquisition unit 131 acquires the search query via the communication unit 110. That is, the acquisition unit 131 acquires, via the communication unit 110, the keywords input by the user U into the search box of a search engine, website, or app.
[0138] Furthermore, the acquisition unit 131 acquires user information about the user U via the communication unit 110. For example, the acquisition unit 131 acquires identification information (such as a user ID) indicating the user U, location information of the user U, attribute information of the user U, etc. from the terminal device 10 of the user U. Furthermore, the acquisition unit 131 may acquire the identification information indicating the user U, attribute information of the user U, etc. when the user U is registered. Then, the acquisition unit 131 stores the user information in the storage unit 120.
[0139] Furthermore, the acquisition unit 131 acquires various types of history information (log data) indicating the behavior of the user U via the communication unit 110. For example, the acquisition unit 131 acquires various types of history information indicating the behavior of the user U from the terminal device 10 of the user U or from various servers based on the user ID or the like. Then, the acquisition unit 131 stores the various types of history information in the storage unit 120.
[0140] The acquisition unit 131 also acquires one or more fixed item images. The acquisition unit 131 also acquires one or more pre-edited item images. For example, the acquisition unit 131 acquires an item image to be a query image to be searched from the terminal device 10 of the user U via the communication unit 110.
[0141] (Setting unit 132) The setting unit 132 sets image generation conditions. For example, the setting unit 132 receives an instruction for the image generation conditions from the terminal device 10 of the user U via the communication unit 110, and sets the image generation conditions. The setting unit 132 also sets image editing conditions. For example, the setting unit 132 receives an instruction for the image editing conditions from the terminal device 10 of the user U via the communication unit 110, and sets the image editing conditions.
[0142] (Construction Section 133) The construction unit 133 constructs a diffusion model that learns a diffusion process that adds noise to each of a plurality of harmonious item images to generate a plurality of noise images, and a dediffusion process that removes noise from each of the plurality of noise images to generate a plurality of original item images. In other words, the construction unit 133 is a learning unit that generates a diffusion model by machine learning.
[0143] (Image generation unit 134) The image generation unit 134 generates an item image using a diffusion model. The image generation unit 134 also generates an item image in accordance with image generation conditions.
[0144] Furthermore, the image generating unit 134 generates an item image from a noise image. For example, the image generating unit 134 generates a plurality of item images from a plurality of noise images, taking into consideration the harmony with each other.
[0145] Furthermore, the image generating unit 134 generates the item image while taking into consideration harmony with the one or more fixed item images acquired by the acquiring unit 131.
[0146] Furthermore, the image generating unit 134 generates an item image from one or more pre-edited item images acquired by the acquiring unit 131 through a noise image in accordance with the image editing conditions, taking into consideration harmony with each of them.
[0147] (Provider 135) The providing unit 135 provides the generated item image to the terminal device 10 of the user U via the communication unit 110. For example, the providing unit 135 provides an image of an item in a coordinated outfit to the terminal device 10 of the user U via the communication unit 110. In addition, the providing unit 135 provides a search result of an item search based on the query image to the terminal device 10 of the user U via the communication unit 110.
[0148] [5. Processing Procedure] Next, a processing procedure by the server device 100 according to the embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the processing procedure according to the embodiment. Note that the processing procedure shown below is repeatedly executed by the control unit 130 of the server device 100.
[0149] For example, as shown in Fig. 9, the acquisition unit 131 of the server device 100 acquires one or more fixed item images and one or more pre-edited item images (step S101). Note that this step may be omitted.
[0150] Next, the setting unit 132 of the server device 100 receives instructions on image generation conditions and image editing conditions from the terminal device 10 of the user U via the communication unit 110, and sets the image generation conditions and image editing conditions (step S102).
[0151] Next, the construction unit 133 of the server device 100 constructs a diffusion model that has learned a diffusion process that adds noise to each of a plurality of harmonious item images to generate a plurality of noise images, and a de-diffusion process that removes noise from each of the plurality of noise images to generate a plurality of original item images (step S103). Note that this step may be performed before step S101 or step S102.
[0152] Next, the image generation unit 134 of the server device 100 uses a diffusion model to generate an item image from the noise image in accordance with the image generation conditions and image editing conditions, while taking into consideration harmony with each other (step S104). When generating an item image in accordance with the pre-edited item image and the image editing conditions, the image generation unit 134 of the server device 100 adds noise to all or part of the pre-edited item image, and then generates an item image from there, while taking into consideration harmony with each other.
[0153] Next, the providing unit of the server device 100 provides the generated item image to the terminal device 10 of the user U via the communication unit 110 (step S105). At this time, the terminal device 10 of the user U displays the generated item image.
[0154] Next, the control unit 130 of the server device 100 searches for an item image of an actual item based on the generated item image (query image) (step S106). Note that the control unit 130 of the server device 100 may search for an item image of an actual item based on an item image (query image) specified by the user from among the generated item images. Note that the number of searched item images of actual items may not be one per query image, but may be multiple.
[0155] Next, the providing unit 136 of the server device 100 provides the search results (item images of actual items) to the terminal device 10 of the user U via the communication unit 110 (step S107). At this time, the terminal device 10 of the user U displays the search results (item images of actual items). Note that, when there are multiple item images of actual items for one query image, the providing unit 136 of the server device 100 may provide item images of actual items that are more similar to the query image with higher priority. Specifically, the terminal device 10 of the user U displays item images of actual items that are more similar to the query image in more prominent or higher positions.
[0156] [6. Modifications] The terminal device 10 and the server device 100 described above may be implemented in various different forms other than the above embodiment. Therefore, modifications of the embodiment will be described below.
[0157] In the above embodiment, some or all of the processing executed by the server device 100 may actually be executed by the terminal device 10 of the user U (or an application running on the terminal). For example, the processing may be completed in a stand-alone manner (by the terminal device 10 alone). In this case, the terminal device 10 is assumed to have the functions of the server device 100 in the above embodiment. Furthermore, in the above embodiment, the terminal device 10 is linked to the server device 100, and therefore, from the perspective of the user U, it appears that the processing of the server device 100 is also being executed by the terminal device 10. In other words, from another perspective, the terminal device 10 can also be said to be equipped with the server device 100.
[0158] In the above embodiment, part or all of the processing executed by the terminal device 10 of the user U may actually be executed by the server device 100.
[0159] In the above embodiment, the terminal device 10 of the user U and the server device 100 may be the same device (one device). That is, the processes executed by the terminal device 10 of the user U and the server device 100 may be executed by the same device (one device).
[0160] In the above embodiment, the image of the item may be a video or a multi-view image, or may be an illustration.
[0161] [7. Effects] As described above, the information processing device (terminal device 10 and server device 100) according to the present application includes an image generation unit 134 that generates an item image from a noise image, and the image generation unit 134 generates a plurality of item images from a plurality of noise images, taking into consideration the harmony with each other.
[0162] The information processing device according to the present application also includes a setting unit 132 that sets image generation conditions, and an image generation unit 134 generates an item image in accordance with the image generation conditions.
[0163] The information processing device according to the present application also includes an acquisition unit 131 that acquires one or more fixed item images, and an image generation unit 134 generates an item image in consideration of harmony with the fixed item image.
[0164] In addition, the information processing device of the present application includes an acquisition unit 131 that acquires one or more pre-edited item images and a setting unit 132 that sets image editing conditions, and an image generation unit 134 generates an item image from the pre-edited item image via a noise image in accordance with the image editing conditions.
[0165] In addition, the information processing device according to the present application includes a construction unit 133 that constructs a diffusion model that learns a diffusion process that adds noise to each of a plurality of harmonious item images to generate a plurality of noise images, and a de-diffusion process that removes noise from each of the plurality of noise images to generate the original plurality of item images, and an image generation unit 134 generates the item images using the diffusion model.
[0166] From another perspective, the information processing device according to the present application The information processing device according to the present application (terminal device 10 and server device 100) includes a display unit 12 that displays an item image generated from a noise image, and the display unit 12 displays a plurality of item images generated from a plurality of noise images, taking into consideration harmony with each other.
[0167] The information processing device according to the present application also includes a search unit 33B that searches for content such as items or coordinates based on a query image, and the display unit 12 displays the content that is the search result.
[0168] The information processing device according to the present application also includes a receiving unit 33A that receives image generation conditions, and the display unit 12 displays the generated item image in accordance with the image generation conditions.
[0169] In addition, the information processing device of the present application includes a reception unit 33A that receives one or more pre-edited item images and image editing conditions, and the display unit 12 displays an item image generated from the pre-edited item image via a noise image in accordance with the image editing conditions.
[0170] Furthermore, the search unit 33B searches for items based on each of the multiple query images, and the display unit 12 displays each search result.
[0171] Furthermore, the search unit 33B searches for an item based on a query image designated by the user from among the plurality of query images, and the display unit 12 displays the search results.
[0172] By using any one or a combination of the above-described processes, the information processing device according to the present application can provide information that is highly harmonious as a whole coordinated item.
[0173] [8. Hardware Configuration] The terminal device 10 and the server device 100 according to the above-described embodiments are realized by a computer 1000 having a configuration as shown in Fig. 10, for example. The following description will be given taking the server device 100 as an example. Fig. 10 is a diagram showing an example of a hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which a calculation device 1030, a primary storage device 1040, a secondary storage device 1050, an output I / F (Interface) 1060, an input I / F 1070, and a network I / F 1080 are connected via a bus 1090.
[0174] The arithmetic device 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, programs read from the input device 1020, and the like, and executes various processes. The arithmetic device 1030 is realized by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like.
[0175] The primary storage device 1040 is a memory device such as a RAM (Random Access Memory) that temporarily stores data used by the arithmetic device 1030 for various calculations. The secondary storage device 1050 is a storage device in which data used by the arithmetic device 1030 for various calculations and various databases are registered, and is realized by a ROM (Read Only Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like. The secondary storage device 1050 may be an internal storage device or an external storage device. The secondary storage device 1050 may also be a removable storage medium such as a USB (Universal Serial Bus) memory or an SD (Secure Digital) memory card. The secondary storage device 1050 may also be cloud storage (online storage), a NAS (Network Attached Storage), a file server, or the like.
[0176] The output I / F 1060 is an interface for transmitting information to be output to an output device 1010 that outputs various types of information, such as a display, a projector, a printer, etc., and is realized by a connector conforming to a standard such as USB (Universal Serial Bus), DVI (Digital Visual Interface), or HDMI (High Definition Multimedia Interface), etc. The input I / F 1070 is an interface for receiving information from various input devices 1020, such as a mouse, a keyboard, a keypad, a button, a scanner, etc., and is realized by a USB, etc.
[0177] Furthermore, the output I / F 1060 and the input I / F 1070 may be wirelessly connected to the output device 1010 and the input device 1020, respectively. That is, the output device 1010 and the input device 1020 may be wireless devices.
[0178] The output device 1010 and the input device 1020 may be integrated into one device, such as a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated into one device as an input / output I / F.
[0179] The input device 1020 may be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), a DVD (Digital Versatile Disc), or a PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0180] The network I / F 1080 receives data from other devices via the network N and sends it to the arithmetic device 1030, and also transmits data generated by the arithmetic device 1030 to other devices via the network N.
[0181] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output I / F 1060 and the input I / F 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.
[0182] For example, when the computer 1000 functions as the server device 100, the arithmetic unit 1030 of the computer 1000 executes a program loaded onto the primary storage device 1040 to realize the functions of the control unit 130. The arithmetic unit 1030 of the computer 1000 may also load a program acquired from another device via the network I / F 1080 onto the primary storage device 1040 and execute the loaded program. The arithmetic unit 1030 of the computer 1000 may also cooperate with the other device via the network I / F 1080 to call and use the functions and data of a program from another program of the other device.
[0183] [9. Other] Although the embodiments of the present application have been described above, the present invention is not limited to the contents of these embodiments. Furthermore, the above-described components include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the scope of so-called equivalents. Furthermore, the above-described components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the above-described embodiments.
[0184] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0185] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0186] For example, the above-mentioned server device 100 may be realized by multiple server computers, and depending on the function, the configuration can be flexibly changed, such as by calling an external platform using an API (Application Programming Interface) or network computing.
[0187] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0188] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]
[0189] 1. Information Processing Systems 10 Terminal Equipment 12 Display section 13 Input section 33A Reception 33B Search Department 33C Display control unit 100 Server device 110 Communications Department 120 Storage section 130 Control Unit 131 Acquisition Department 132 Settings 133 Construction Department 134 Image Generation Unit 135 Provision Department
Claims
1. An image generation unit is provided that generates a plurality of item images from a plurality of noise images, taking into consideration the harmony of each of them, by using an image generation model having a plurality of neural networks realized by a diffusion model that has been trained to remove noise from a noise image containing noise and generate an item image, and by generating individual item images from individual noise images while incorporating interactions between items that are components of a coordinated outfit.
1. An information processing device comprising:
2. a setting unit for setting image generation conditions, the image generation unit generates an item image in accordance with the set image generation conditions.
2. The information processing apparatus according to claim 1, wherein:
3. An acquisition unit that acquires one or more fixed item images, When generating an item image from a noise image, the image generation unit generates the item image while taking into consideration harmony with the acquired fixed item image.
2. The information processing apparatus according to claim 1, wherein:
4. an acquisition unit that acquires one or more pre-edited item images; a setting unit for setting image editing conditions; Equipped with the image generation unit generates an item image from the acquired pre-edited item image via a noise image in accordance with the set image editing conditions; 2. The information processing apparatus according to claim 1, wherein:
5. The method includes a construction unit that constructs a diffusion model that has learned a diffusion process that adds noise to each of a plurality of harmonious item images to generate a plurality of noise images, and a de-diffusion process that removes noise from each of the plurality of noise images to generate the original plurality of item images, The image generation unit generates an item image using the constructed diffusion model.
2. The information processing apparatus according to claim 1, wherein:
6. The present invention includes a providing unit that provides a plurality of item images generated from a plurality of noise images, taking into consideration the harmony with each other, by using an image generation model having a plurality of neural networks realized by a diffusion model that has been trained to remove noise from a noise image containing the noise and generate an item image, and providing individual item images generated from individual noise images by incorporating interactions between items that are components of a coordinated outfit.
1. An information processing device comprising:
7. a search processing unit that searches for content such as items or coordinates using the provided item image as a query image; the providing unit provides the searched content.
7. The information processing apparatus according to claim 6,
8. an acquisition unit for acquiring image generation conditions; the providing unit provides an item image generated in accordance with the acquired image generation conditions.
7. The information processing apparatus according to claim 6,
9. An acquisition unit that acquires one or more pre-edited item images and image editing conditions, the providing unit provides an item image generated from the acquired pre-edited item image via a noise image in accordance with the acquired image editing conditions; 7. The information processing apparatus according to claim 6,
10. the search processing unit searches for items using the provided item images as query images, respectively; the providing unit provides each of the searched items.
8. The information processing apparatus according to claim 7,
11. the search processing unit searches for an item using an item image designated by a user from among the plurality of provided item images as a query image; The providing unit provides the searched item.
8. The information processing apparatus according to claim 7,
12. An information processing method executed by an information processing device, The method includes an image generation process in which an image generation model having a plurality of neural networks realized by a diffusion model trained to remove noise from a noise image containing noise and generate an item image is used to generate individual item images from individual noise images by incorporating interactions between items that are components of a coordinated outfit, thereby generating a plurality of item images from a plurality of noise images while taking into consideration the harmony with each other.
1. An information processing method comprising:
13. An information processing method executed by an information processing device, The method includes a providing step of providing a plurality of item images generated from a plurality of noise images by incorporating interactions between items that are components of a coordinated outfit using an image generation model having a plurality of neural networks realized by a diffusion model trained to remove noise from a noise image and generate an item image, thereby providing a plurality of item images generated from the plurality of noise images while taking into consideration harmony with each other.
1. An information processing method comprising:
14. An image generation procedure for generating a plurality of item images from a plurality of noise images, taking into consideration the harmony of each of them, by using an image generation model having a plurality of neural networks realized by a diffusion model trained to generate item images by removing noise from a noise image containing the noise, and generating individual item images from individual noise images while incorporating interactions between items that are components of a coordinated outfit. An information processing program characterized by causing a computer to execute the above.
15. A procedure for providing a plurality of item images generated from a plurality of noise images, taking into consideration the harmony with each other, by providing individual item images generated from individual noise images by incorporating interactions between items that are components of a coordinated outfit, using an image generation model having a plurality of neural networks realized by a diffusion model trained to generate item images by removing noise from a noise image containing the noise. An information processing program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Attribute generative adversarial network and matched clothes generation method based on attribute generative adversarial network
CN110909754A
Multi-mode-based generative compatible garment matching scheme generation method and system
CN111861672A
Sketch guidance-based paired clothing image generation method
CN113298906A
Fashion coordination generating device, fashion coordination generating system, fashion coordination generating method, program, and recording medium
JP2013235528A
Information processing device, data extraction method, and data extraction program
JP2020098521A