Information processing device, information processing method, and information processing program

The information processing system addresses the challenge of generating harmonious fashion coordination by using a diffusion model with Set Transformer layers to create visually appealing and user-preferred outfit suggestions.

JP2026070013AActive Publication Date: 2026-04-27ZOZO INC
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ZOZO INC
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional fashion coordination techniques fail to generate highly harmonious outfits due to the lack of consideration for the harmony between individual items.

Method used

An information processing system utilizing a diffusion model with Set Transformer layers to generate item images from noise images, taking into account the harmony between multiple items, and incorporating user-specific information and preferences.

Benefits of technology

Enables the creation of highly harmonious fashion coordination by generating visually appealing and user-preferred outfit suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070013000001_ABST
    Figure 2026070013000001_ABST
Patent Text Reader

Abstract

It provides highly harmonious information as a whole. [Solution] The information processing device according to the present invention comprises an image generation unit that generates item images from noise images, characterized in that the image generation unit generates multiple item images from multiple noise images, taking into consideration their harmony with each other. Furthermore, the information processing device according to the present invention comprises a display unit that displays the item images generated from noise images, characterized in that the display unit displays multiple item images generated from multiple noise images, taking into consideration their harmony with each other.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, techniques for generating fashion coordination have been known. For example, techniques for generating fashion coordination based on predetermined rules have been known.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the conventional technology, since coordination cannot be generated based on the harmony of all items, it has not been possible to provide highly harmonious information as a whole coordination.

[0005] The present application has been made in view of the above, and an object thereof is to provide highly harmonious information as a whole coordination.

Means for Solving the Problems

[0006] The information processing apparatus according to the present application includes an image generation unit that generates item images from noise images, and the image generation unit is characterized in that it generates a plurality of item images from a plurality of noise images in consideration of the harmony with each of them.

Effects of the Invention

[0007] According to one aspect of the embodiment, there is an effect that highly harmonious information can be provided as a whole coordination. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is an explanatory diagram showing an overview of the information processing system according to the embodiment. [Figure 2] Figure 2 shows an example of the configuration of the image generation model UN used by the server device when generating images while considering harmony. [Figure 3] Figure 3 shows a first example of a UI for specifying image generation conditions. [Figure 4] Figure 4 shows a second example of a UI for specifying image generation conditions. [Figure 5] Figure 5 is an explanatory diagram illustrating the design philosophy of the Outfit Diffusion architecture. [Figure 6] Figure 6 shows an example of the resulting coordination. [Figure 7] Figure 7 shows an example of the configuration of a terminal device according to this embodiment. [Figure 8] Figure 8 shows an example of the configuration of a server device according to the embodiment. [Figure 9] Figure 9 is a flowchart showing the processing procedure according to the embodiment. [Figure 10] Figure 10 shows an example of a hardware configuration. [Modes for carrying out the invention]

[0009] The following describes in detail, with reference to the drawings, embodiments for implementing the information processing device, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments"). Note that these embodiments do not limit the information processing device, information processing method, and information processing program according to the present application. Furthermore, the same parts are denoted by the same reference numerals in the following embodiments, and redundant descriptions are omitted.

[0010] [1. Overview of the Information Processing System] First, with reference to Figure 1, an overview of the information processing system according to the embodiment will be described. Figure 1 is an explanatory diagram showing an overview of the information processing system according to the embodiment. As shown in Figure 1, the information processing system 1 according to the embodiment includes a terminal device 10 and a server device 100. These various devices are connected to each other via a network N, either by wire or wireless means, enabling communication. As a result, the terminal device 10 can cooperate with the server device 100. The network N is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network) such as the Internet.

[0011] Terminal device 10 is an information processing device used by user U. For example, terminal device 10 may be a smart device such as a smartphone or tablet, a mobile phone such as a feature phone, a PC (Personal Computer), a PDA (Personal Digital Assistant), a game console or AV equipment with communication functions, an information appliance or digital appliance, a car navigation system, a wearable device such as a smartwatch, head-mounted display, or smart glasses. Alternatively, terminal device 10 may be a house or building compatible with the Internet of Things (IoT), a car, a home appliance, or an electronic device.

[0012] In this embodiment, the terminal device 10 is a smart device such as a smartphone or a tablet terminal used by the user U, and is a portable terminal device that can communicate with any server device via a wireless communication network such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation: 5G mobile communication system). Further, the terminal device 10 has a screen such as a liquid crystal display, which has a function of a touch panel, and accepts various operations on display data such as content, such as tap operations, slide operations, and scroll operations, by a finger or a stylus from the user U. Among the screens, an operation performed on the area where the content is displayed may also be regarded as an operation on the content. Further, the terminal device 10 may be not only a smart device but also an information processing device such as a desktop PC (Personal Computer) or a notebook PC.

[0013] Further, such a terminal device 10 can be connected to the network N via a wireless communication network such as LTE, 4G, or 5G, or short-range wireless communication such as Bluetooth (registered trademark) or wireless LAN (Local Area Network), and communicate with the server device 100.

[0014] The server device 100 is, for example, a computer such as a PC or a blade server, or a mainframe or a workstation. The server device 100 may be realized by cloud computing. [[ID=X]]

[0015] In this embodiment, the server device 100 is an information processing device that cooperates with the terminal devices 10 of each user U and provides API (Application Programming Interface) services and various data for the terminal devices 10 of each user U with respect to various applications (hereinafter referred to as apps), etc., and is realized by a computer, a cloud system, or the like.

[0016] Furthermore, the server device 100 may be an information processing device that provides some kind of Web service online to each user U's terminal device 10. For example, the server device 100 may provide services such as internet connection, search services, SNS (Social Networking Service), e-commerce (EC), electronic payment, online games, online banking, online trading, accommodation / ticket reservations, video / music distribution, news, maps, route search, route guidance, route information, service information, and weather forecasts as Web services. In practice, the server device 100 may cooperate with various servers that provide the above-mentioned Web services and act as an intermediary for Web services, or it may be responsible for processing Web services.

[0017] The server device 100 can acquire user information about user U. For example, the server device 100 can acquire information about user U's attributes (attribute information), such as gender, age, and residential area. The server device 100 can also acquire information about user U's demographics, psychographics, geographics, behavioral attributes, etc. The server device 100 may also acquire information about the segment or persona to which user U belongs in the field of marketing, as user information. The server device 100 stores and manages information about user U's attributes (attribute information) along with identification information (user ID, etc.) that identifies user U.

[0018] Further, the server device 100 acquires various types of history information (log data) indicating the actions of the user U from the terminal device 10 of the user U or from various servers or the like based on the user ID or the like. For example, the server device 100 acquires a location history, which is a history of the location and time of the user U, from the terminal device 10. Further, the server device 100 acquires a search history, which is a history of search queries input by the user U, from an e-commerce server, a content server, a search server (search engine), or the like. Further, the server device 100 acquires a browsing history, which is a history of content (such as fashion items and fashion coordinates) browsed by the user U, from an e-commerce server, a content server, or the like. Further, the server device 100 acquires a purchase history (settlement history), which is a history of product purchases and settlement processing of the user U, from an e-commerce server, a settlement processing server, or the like. Further, the server device 100 may acquire a listing history and a sales history, which are histories of the user U's listings on the marketplace, from an e-commerce server, a settlement processing server, or the like. Further, the server device 100 acquires a posting history, which is a history of the user U's posts, from a posting server that provides a word-of-mouth posting service, a SNS server, an e-commerce server, or the like. Note that each of the above-mentioned various servers or the like may be the server device 100 itself. That is, the server device 100 may function as each of the above-mentioned various servers or the like.

[0019] Also, the number of each device included in the information processing system 1 shown in FIG. 1 is not limited to that shown. For example, in FIG. 1, for the sake of simplification of the illustration, only one terminal device 10 is shown, but this is merely an example and is not limited, and two or more may be provided.

[0020] [2. Generation of Coordination Images Using a Diffusion Model] A diffusion model is a type of generative AI that learns a process of gradually adding noise to a target image to degrade it (diffusion process) and a process of reconstructing the image by gradually removing (denoising) the noise so as to reverse the degradation process (reverse diffusion process).

[0021] The server device 100 generates item images (product images) from noise images using a diffusion model. For example, the server device 100 generates multiple item images from multiple noise images, taking into account their harmony with each other. For example, the server device 100 generates individual item images by incorporating the interactions between items that are components of a coordination. A coordination is a set of items.

[0022] Here, we will explain the details of the process by which the server device 100 generates multiple item images from multiple noise images, taking into account their harmony with each other. In Figure 1, there is one UNet, but in reality, there are as many UNets as there are images to be generated. For example, if three images are to be generated simultaneously, there are three UNets. For each UNet, a Set Transformer layer ("Set Transformer" layer) is incorporated into any layer of the UNet. The Set Transformer layer is connected to a Set Transformer layer located in the same layer of another UNet and is configured to output a value based on the value of the Set Transformer layer of each UNet.

[0023] Figure 2 shows an example configuration of the image generation model UN used by the server device 100 when generating images while considering harmony. As shown in Figure 2, the image generation model UN has multiple networks UN1 and UN2. In reality, the image generation model UN has the same number of networks as the number of images generated simultaneously, but in the example shown in Figure 2, only two networks, UN1 and UN2, are shown.

[0024] Network UN1 is a neural network that generates image P1 from noise image N1, and is implemented using a so-called diffusion model. More specifically, Network UN1 is a neural network that, when inputted with noise image N1 and category information C1 indicating the object for which images are to be generated (for example, categories of fashion items such as tops (shirts, etc.), jackets / outerwear (coats, etc.), pants, skirts, dresses, hats, bags, shoes, legwear (socks, etc.), and accessories (necklaces, rings, earrings, etc.)), generates images of the object indicated by category information C1 (images of the object, images containing the object, etc.) from noise image N1. Such Network UN1 can be implemented using a learning method similar to that of publicly known diffusion models.

[0025] Network UN2, like Network UN2, is a neural network that generates the target image P2 indicated by category information C2 from a noise image N2, and is implemented using a learning method known as a diffusion model.

[0026] Here, networks UN1 and UN2 (hereinafter sometimes collectively referred to as networks UN) have a structure called a UNet. For example, network UN1 has intermediate layers UN1-1, ST1-1, UN1-2, ST1-2, UN1-3, ST1-3, UN1-4, ST1-4, UN1-5, and ST1-5. Layers UN1-1, UN1-2, UN1-3, UN1-4, and UN1-5 (hereinafter collectively referred to as layer UN1) are the layers that make up the so-called UNet. As shown in Figure 2, network UN1 has a configuration in which Set Transformer layers ST1-1, ST1-2, ST1-3, ST1-4, and ST1-5 (hereinafter collectively referred to as layer ST1) are located after each layer of the UNet composed of layers UN1-1, UN1-2, UN1-3, UN1-4, and UN1-5.

[0027] Furthermore, for example, network UN2 has intermediate layers UN2-1, ST2-1, UN2-2, ST2-2, UN2-3, ST2-3, UN2-4, ST2-4, UN2-5, and ST2-5. Layers UN2-1, UN2-2, UN2-3, UN2-4, and UN2-5 (hereinafter collectively referred to as layers UN2) are the layers that constitute the so-called UNet. As shown in Figure 2, network UN2 has a configuration in which Set Transformer layers ST2-1, ST2-2, ST2-3, ST2-4, and ST2-5 (hereinafter collectively referred to as layers ST5) are located after each layer of the UNet, which is composed of layers UN2-1, UN2-2, UN2-3, UN2-4, and UN2-5.

[0028] Here, layers ST1 and ST2 are interconnected and are attention layers that are trained to output values ​​calculated based on the values ​​output from the preceding layers (i.e., layers UN1 and UN2) and the values ​​received by the other connected layers. For example, layer ST1-1 receives the values ​​output by layer UL1-1. Similarly, layer ST2-1 receives the values ​​output by layer UL2-1. In such cases, layer ST1-1 calculates its output value considering not only the values ​​received from layer UL1-1 but also the values ​​received by layer ST2-1 from layer UL2-1. Likewise, layer ST2-1 calculates its output value considering not only the values ​​received from layer UL2-1 but also the values ​​received by layer ST1-1 from layer UL1-1.

[0029] Similarly, layers ST1-2 to ST1-5 and layers ST2-2 to ST2-5 are connected in the same way, and the output value is calculated by considering not only the output from the preceding layer, but also the values ​​received by other connected layers from the preceding layer. In addition to networks UN1 and UN2, if there are other networks, each network will have the same structure, and layers ST located in the same position will be interconnected. For example, if network UN3 exists in the image generation model UN, and layer ST3-1 is located in the same position as layers ST1-1 and ST2-1 in network UN3, layer ST1-1 will calculate the output value from the values ​​received by layers ST1-1, ST1-2, and ST1-3.

[0030] The image generation model UN, consisting of multiple such networks, learns in a similar manner to a diffusion model to generate (reconstruct) different images from multiple noise images and multiple categories of information. In this process, by preparing a set of images of multiple harmonized objects and multiple categories of information representing each object as training data, the image generation model UN can, upon receiving multiple categories of information, generate a set of images containing the objects indicated by each category of information, resulting in a set of images of multiple objects that are harmonized as a whole.

[0031] For example, the server device 100 acquires, as training data, a set of images of multiple products that make up a coordinated outfit that has been published in magazines or on e-commerce sites (i.e., evaluated as harmonious by the publisher), a coordinated outfit that has been posted on social media (i.e., evaluated as harmonious by the poster), and a set of category information for each product that makes up a coordinated outfit that has been well-received by viewers (i.e., evaluated as harmonious by viewers), as well as a set of category information for each product. Alternatively, the server device 100 may acquire, for example, a set of images of multiple products manually selected by a stylist, along with a set of category information for each product, as a set of images of products that make up a coordinated outfit that has been evaluated as harmonious, along with a set of category information for each product. In addition to these methods, if the goal is to acquire images of each product and category information for each product for multiple products that may be evaluated as harmonious, then any set of images and category information collected using any method and any evaluation axis may be acquired as ground truth data. Furthermore, "harmonious" can mean harmonious from an overall perspective, or harmonious from a specific perspective (for example, color, silhouette, size, design, etc., not limited to those exemplified). In addition, the server device 100 may acquire information such as brand, color, silhouette, size, design, material, comfort, item image, genre (adult casual, etc.), season (winter outfit, etc.), occasion (wedding, etc.), era (Y2K, incorporating trends from around the year 2000, etc.), wearer (body information including body type, skin tone, etc.), number of likes, and desired outfit image as training data, and use this information to train the diffusion model.

[0032] Next, the server device 100 generates images by progressively adding noise to each image in a given set of ground truth data. Then, the server device 100 modifies the network connection coefficients in the image generation model UN using various known learning methods similar to those used for diffusion models, so that when multiple images with n levels of noise and multiple categories of information are input to the image generation model UN, it generates an image with n-1 levels of noise. By repeatedly performing this learning, the image generation model UN learns to generate not just multiple images, but images in which multiple objects included in each generated image are in harmony (for example, their appearances are in harmony).

[0033] The server device 100 then inputs multiple noise images and category information indicating each category of multiple objects for which image generation is desired to occur to the image generation model UN that has undergone such training. By repeatedly inputting the images generated by the image generation model UN along with the category information, the server device 100 can generate images of each object indicated by the category information that have a harmonious overall appearance (i.e., images that are presumed to be well-received on social media, or images that can be manually coordinated by a stylist, etc.).

[0034] Furthermore, if the server device 100 is to generate images of fixed categories for networks UN1 and UN2, it does not need to input category information. For example, if network UN1 corresponds to clothing and network UN2 corresponds to shoes, the server device 100 collects pairs of clothing and shoe images that are evaluated as harmonious as training data. Then, when a noise image is input to networks UN1 and UN2, the server device 100 can be trained so that network UN1 reconstructs the clothing image that will be the ground truth data, and network UN2 reconstructs the shoe image that will be the ground truth data (an image of a shoe that is evaluated as harmonious with the clothing image reconstructed by network UN1). An image generation model UN that has undergone such training will, for example, be able to output clothing and shoe images that are estimated to be harmonious simply by inputting multiple noise images.

[0035] At this time, the server device 100 may set image generation conditions. The server device 100 then generates item images according to the image generation conditions. The image generation conditions may include conditions for each item and conditions for the overall image. For example, the image generation conditions for each item may specify categories, brands, colors, silhouettes, sizes, designs, materials, comfort levels, and desired image of the item, and are not limited to those exemplified. Furthermore, the image generation conditions may also relate to specific parts. For example, they may specify the silhouette of the collar area. Also, for example, the image generation conditions for the overall image may specify genres (such as adult casual), seasons (such as winter outfits), occasions (such as weddings), eras (such as Y2K incorporating trends from around the year 2000), wearers (such as body shape and skin tone), number of likes, and desired image of the outfit, and are not limited to those exemplified. Furthermore, the image generation conditions exemplified for each item may also be used as image generation conditions for the entire system, and vice versa.

[0036] The server device 100 may also acquire one or more fixed item images. The server device 100 then generates item images, taking into consideration their harmony with the fixed item images. Fixed item images are images of items owned by the user, images of items that the user has favorited, images of items that the user has searched for or viewed, etc., and are images of items that the user wants to include in the outfit. The server device 100 may acquire images registered or posted on e-commerce sites or outfit coordination sites as fixed item images, or it may acquire images uploaded by the user.

[0037] Furthermore, the server device 100 acquires one or more pre-edited item images, sets image editing conditions, and generates an item image from the pre-edited item image via a noise image according to the image editing conditions. Specifically, the server device 100 adds noise to all or part of the pre-edited item image (also called the part to be edited; for example, the collar part if the image editing condition specifies the silhouette of the collar of a dress), removes the noise, and generates an item image. The pre-edited item image is an image of an item owned by the user, an image of an item that the user has favorited, an image of an item that the user has searched for or viewed, and is an image of an item that is close to the image of an item to be included in the outfit. The server device 100 may acquire images registered or posted on e-commerce sites or outfit coordination sites as pre-edited item images, or it may acquire images uploaded by the user. Furthermore, the image editing conditions may include image editing conditions for each pre-edited item and image editing conditions for the overall image. For example, the image editing conditions for each pre-edited item may specify category, brand, color, silhouette, size, design, material, comfort, and desired image of the item, and are not limited to those exemplified. Furthermore, the image editing conditions may also relate to specific parts. For example, they may specify the silhouette of the collar. Also, for example, the image editing conditions for the whole may specify genre (e.g., adult casual), season (e.g., winter outfit), occasion (e.g., wedding), era (e.g., Y2K incorporating trends from around the year 2000), wearer (including body information such as body type and skin tone), number of likes, and desired image of the outfit, and are not limited to those exemplified. Note that the example image editing conditions for each pre-edited item may also be used as the overall image editing conditions, and the example image editing conditions for the whole may also be used as the image generation conditions for each pre-edited item.

[0038] Furthermore, the server device 100 constructs a diffusion model that learns a diffusion process to generate multiple noise images by adding noise to each of the multiple harmonized item images, and a dediffusion process to generate the original multiple item images by removing noise from each of the multiple noise images. Then, the server device 100 generates item images using the diffusion model.

[0039] Furthermore, user U's terminal device 10 works in conjunction with the server device 100 to display item images generated from noise images. For example, user U's terminal device 10 displays multiple item images generated from multiple noise images, taking into consideration their harmony with each other.

[0040] Furthermore, user U's terminal device 10 interacts with the server device 100 and displays the generated item image as a query image. Subsequently, user U's terminal device 10 searches for content (items, outfits) based on the displayed query image. Finally, user U's terminal device 10 displays the search results content.

[0041] Furthermore, user U's terminal device 10 receives image generation conditions from user U. The terminal device 10 may also receive fixed item images from user U. Then, user U's terminal device 10 cooperates with the server device 100 to display the generated item images according to the image generation conditions.

[0042] Furthermore, the user U's terminal device 10 receives one or more pre-edited item images and image editing conditions from the user U. The user U's terminal device 10 then works in conjunction with the server device 100 to display the item image generated from the pre-edited item image via a noise image, according to the image editing conditions.

[0043] Furthermore, user U's terminal device 10 works in conjunction with the server device 100 to search for items and outfits based on each of the multiple query images. User U's terminal device 10 also displays the search results.

[0044] Furthermore, user U's terminal device 10 works in conjunction with the server device 100 to search for items from among multiple query images based on the query image specified by user U. User U's terminal device 10 then displays the search results.

[0045] For example, as shown in Figure 1, the server device 100 receives image generation conditions from the user U's terminal device 10 via the network N and sets the image generation conditions (step S1). At this time, the user U's terminal device 10 displays a UI (User Interface) for specifying the image generation conditions.

[0046] Referring to Figures 3 and 4, the UI for specifying image generation conditions will be explained. Figure 3 shows a first example of the UI for specifying image generation conditions. Figure 4 shows a second example of the UI for specifying image generation conditions.

[0047] As shown in Figures 3 and 4, the UI for specifying image generation and image editing conditions includes a text input field TX, a condition setting button B1 (also called the image generation button), an item image display field IMG, a category selection field CS, an item fixing checkbox CB for fixing items, an item search button B2, and an item search result display field SR.

[0048] The text input field TX is for entering text for image generation conditions and image editing conditions. Alternatively, a prompt (instruction text) may be entered in the text input field TX as an image generation condition. When the user presses the condition setting button B1, the entered text and selected categories are set as image generation and editing conditions. The item image is then generated according to these conditions.

[0049] The item image display field (IMG) is a field for registering item images. It is also a field for displaying the generated item images. The registered item images are either the item images of items to be included in the outfit (fixed item images) or the item images of items that are similar in appearance to the items to be included in the outfit but before editing (pre-edit item images). In this embodiment, multiple item images are registered. At this time, one item image is displayed in one item image display field (IMG). When the user presses the condition setting button B1, a newly generated item image is displayed in the item image display field (IMG), either in place of the pre-edit item image. At this time, one item image is displayed in one item image display field (IMG). The item image displayed in the item image display field (IMG) becomes the query image.

[0050] The category selection field (CS) and the item fixing checkbox (CB) each exist in quantities equal to the number of item images. In other words, they exist in quantities equal to the number of item image display fields (IMG). Note that the number of item image display fields (IMG) can be increased by the user (U).

[0051] The Category Selection field CS is for selecting the category (image generation conditions and image editing conditions) of the item to be generated. If an item category is specified, when the user presses the Condition Setting button B1, the item image will be generated according to the specified category. Conversely, if no item category is specified, an item image will be generated randomly, regardless of the category. Even if no item category is specified, if text is entered in the Text Input field TX, the item image will be generated according to that text input.

[0052] The item-fixing checkbox (CB) is used to set an item to be fixed. When the item is fixed by checking the item-fixing checkbox (CB), the corresponding item image becomes the fixed item image, and that item remains fixed and unchangeable within the outfit. In other words, when editing items within the outfit, the image will not be changed to that of other items. Note that other configurations such as radio buttons can also be used instead of checkboxes.

[0053] When the item search button B2 is pressed by the user, it searches for items using the query image displayed in the item image display field IMG. The search results are displayed in the item search results display field SR. In the example in Figure 3, there is one item search results display field SR for each item image (the same number as the number of item images), but in the example in Figure 4, all item images are grouped together into one (per outfit unit).

[0054] Furthermore, by registering the item image displayed in the item search results display area (SR) to the item image display area (IMG), that item image can be used as the next fixed item image or the item image before editing. At this time, it is also possible to register all the item images within a coordinate to the item image display area (IMG) on a coordinate basis.

[0055] In Figures 3 and 4, one set of text input field TX, condition setting button B1, item image display field IMG, category selection field CS, and checkbox CB is provided, but multiple sets may be provided. Also, one or more of the text input field TX, condition setting button B1, category selection field CS, and checkbox CB may be common to multiple sets. In other words, item images may be generated by changing the image generation conditions and image editing conditions for each set, or item images may be generated by making the image generation conditions and image editing conditions the same for multiple sets. The user may then select a set of item images they like from among the multiple sets of item images and use them as query images for item searches.

[0056] Next, the server device 100 constructs a diffusion model that has learned a diffusion process to generate multiple noise images by adding noise to each of the multiple item images, and a dediffusion process to generate the original multiple item images by removing noise from each of the multiple noise images (step S2). Note that step S2 may be performed before step S1.

[0057] Next, the server device 100 generates item images from the noise image using a diffusion model (step S3). At this time, the server device 100 acquires one or more fixed item images, generates an item image considering harmony with the fixed item image according to the image generation conditions and image editing conditions, and displays it on the user U's terminal device 10. Alternatively, the server device 100 acquires one or more pre-edited item images, generates an item image from the pre-edited item image via the noise image according to the image editing conditions, and displays it on the user U's terminal device 10.

[0058] Next, the server device 100 displays the generated item images and fixed item images as candidates for query images on the user U's terminal device 10 (step S4). If the user does not like the generated item images, they can specify the image generation conditions and image editing conditions again to generate new item images, or they can generate new item images again using the same image generation conditions and image editing conditions. Also, if the user likes any of the generated item images, they can fix that item image and generate other item images.

[0059] Next, the server device 100 searches for items based on each of the multiple query images, or based on the query image specified by user U (user) from among the multiple query images, and displays the search results on user U's terminal device 10 (step S5).

[0060] [2-1. Generating item images with coordination in mind] Traditionally, outfit suggestions have been made using set-matching models that output a set-matching score indicating the compatibility between multiple input item groups. However, the feature quantities of the items input into the set-matching model (which are also the feature quantities used to extract the actual items) are incomprehensible to humans. Therefore, if a user does not like the suggested outfit, it is unclear whether the feature quantities of the items are not suitable for the user (should the feature quantities of the items be changed), or whether the actual items extracted based on the feature quantities of the items are not suitable for the user (should the actual items extracted be changed without changing the feature quantities of the items), which could lead to the inability to propose a truly ideal outfit.

[0061] In this embodiment, the server device 100 can visualize the features of items by generating item images (query images) used for the search before actually searching for items, and can propose truly ideal coordinates. Furthermore, the server device 100 can interactively generate and edit item images while visualizing them.

[0062] To achieve this, an architecture that satisfies the following two properties is required. • Symmetry in swapping items within a coordinated outfit • Variability of item images to patches

[0063] Therefore, as shown in Figure 1, a Set Transformer layer is added to the UNet diffusion model to incorporate the interactions between items in the coordination. UNet is a fully convolutional network (FCN) and is a network for estimating image segmentation (where objects are located). The Set Transformer layer mainly consists of the following two parts: (A) and (B).

[0064] (A)SA+FF(SelfAttention + FeedForward) By summing the height and width values ​​and calculating SelfAttention, the interaction between items in the outfit is calculated.

[0065] (B)CA+FF(CrossAttention + FeedForward) By calculating CrossAttention between the SA+FF output and the original image, the item images are transformed to reflect the interactions between items in the outfit.

[0066] Furthermore, as an application, the server device 100 generates outfits based on conditions such as categories. Possible conditions include, for example, categories, brands, colors, time, materials, usage scenarios on ZOZOTOWN (registered trademark), outfit descriptions, descriptions of each item, locations / travel destinations / scenes, user information (user clusters / browsing history), number of likes, etc. However, in practice, it is not limited to these examples. In this case, the server device 100 can perform the following information processing.

[0067] (1) Item Search The server device 100 generates images of items to be searched based on images of items the user actually owns. For obtaining images of items the user actually owns, the server device 100 may accept registration of images of items the user owns, or it may extract images of items included in past purchase history (items with a purchase record). In other words, the server device 100 can generate one or more item images (query images) based on fixed item images or pre-edited item images. This allows the server device 100 to suggest ways to mix and match items the user actually owns and to search for items with high harmonious combinations. Furthermore, an advantage of image generation is that the query images used for item searches can be visualized.

[0068] The server device 100 then generates an image of the item to be searched, and as a user interaction, presents the user with an image that will be used as the query image for the search and asks for confirmation, such as "Is this query OK?". Specifically, the server device 100 displays the generated item image in the item image display field IMG and asks for confirmation, such as "Is this query OK?". Alternatively, the server device 100 may accept the user's specification of attributes such as color image and V-neck (accepting image generation conditions and image editing conditions from the user) and generate an image that will be used as the query image accordingly. Furthermore, the server device 100 may present each of the generated images to the user and ask the user to select the image that will be used as the query image.

[0069] (2) Virtual try-on of fictional items The server device 100 generates item images and then generates snapshots using a different method, such as virtual try-on. Specifically, the server device 100 generates images of the user wearing one or more of the generated item images, based on the user's image. If fixed item images are set, the server device 100 also generates images of the user wearing the fixed items. This allows the user to have an image of the items being searched for and to determine whether the generated item images are appropriate as query images.

[0070] (3) Item search, outfit search The server device 100 searches for actual items and outfits by combining similar item search and similar outfit search techniques. Specifically, the server device 100 embeds each generated item image and generated outfit image (including outfit images generated by combining each generated item image) into the distributed representation space, and searches for item images of actual items and outfit images using actual items based on the similarity of vectors in the distributed representation space.

[0071] (4) Increase the number of likes The server device 100 generates images of some or all of the items that make up an outfit and suggests to the user, "Wearing this will increase your likes." Specifically, the server device 100 accepts text such as the number of likes the user wants or an outfit that is likely to get more likes as image generation and image editing conditions, and generates images of some or all of the items that make up an outfit that is likely to get the number of likes the user wants or an outfit that is likely to get more likes, and suggests to the user, "Wearing this will increase your likes." The user can then determine whether the query image is likely to be able to search for items or outfits that are likely to get more likes before searching for items. The server device 100 acquires a set of images of multiple items that make up outfits for each number of likes, and multiple category information that represents each item as ground truth data, and trains a diffusion model.

[0072] (5) Prediction and Trend-Related The server device 100 generates an average image of generated images in response to prompts (instructions) such as "What is currently in fashion?" or "What was a popular outfit in XX year?" as part of its trend analysis. For example, the server device 100 performs interpolation-like processing using a GAN (Generative Adversarial Network). This allows us to understand, for example, the temporal changes in the average image of adult casual wear.

[0073] Furthermore, the server device 100 generates images of the next requested items. For example, the server device 100 predicts future news and generates images of items based on information sources such as what kinds of images are frequently generated and what kinds of items sell well when certain news stories are released.

[0074] Furthermore, the server device 100 accepts image generation and image editing conditions, such as outfits that look like they're from 2000 or outfits that are likely to be popular in 2025 (the future), and generates images for each item. Generally, fashion trends are said to follow a 20-year cycle, and by learning past trends, it is possible to propose outfits that reflect future trends. The server device 100 acquires as ground truth data a set of images of multiple items that make up outfits for each decade, and multiple category information that represents each item, and trains a diffusion model.

[0075] (6) Personalization The server device 100 learns negative prompts individually for each user.

[0076] (7) Coordination The server device 100 generates outfits based on items the user actually owns. Then, the server device 100 performs an image search within ZOZOTOWN (registered trademark) based on the generated results. The difference from a set search is that it can be interactive. Specifically, the server device 100 obtains images of two or more items that the user actually owns (item images before editing), and generates a query image from each of the obtained images of the two or more items. Then, the server device 100 uses the two or more generated query images to search for and display images of the two or more actual items.

[0077] (8) Editing items within the outfit The server device 100 edits the items within the outfit. At this time, the server device 100 performs interactive editing. The direction of editing may be, for example, "What would happen if we made an outfit using 1990s item images look modern (contemporary)?" Specifically, the server device 100 obtains images of two or more items from the 1990s that make up the outfit (item images before editing) and an image editing condition of an outfit that looks like it's from 2024. According to the image editing condition, it generates a new item image from each of the two or more obtained item images, taking into consideration their harmony with each other.

[0078] Furthermore, the server device 100 may edit items within the outfit according to specific tags (style information such as "adult casual") obtained as image editing conditions. In addition, the server device 100 may display a message indicating that the outfit is likely to receive more "likes" in order to show that it suits the user.

[0079] Furthermore, the server device 100 generates outfits from the perspective of "what items should be added to older items (e.g., 2019) to make them look current (e.g., 2024)?" (Fill in the n blanks). Specifically, the server device 100 obtains images of one or more items from 2019 (fixed item images) and an image generation condition of a 2024-like outfit, and generates images of items to be added according to the image generation condition, while considering harmony with the fixed item images. By treating the 360-degree images of items as a set, it becomes possible to generate and edit items that are consistent in 360 degrees.

[0080] (9) New forms of search This embodiment is expected to create a new form of search. For example, when a user enters search terms (image generation conditions or image editing conditions), or uses their owned products (fixed item images or pre-edited item images) for one-click recommendations, a search method can be implemented where multiple images (generated query images) are presented as candidates with the question, "Does this image closely match your image?", and the user can select from among them (the system can search for images of actual items similar to the selected query image).

[0081] (10) Assistance in creating custom avatars The server device 100 can also assist in the creation of user-created avatars. For example, games often have a function to create clothes for one's avatar, but creating them from scratch is tedious. It would be convenient if it could create a coordinated outfit (a coordinated outfit image consisting of multiple item images generated considering their harmony) from text (image generation conditions). If the user is not satisfied with the coordinated outfit, they may be allowed to make minor adjustments themselves afterward (the generated item images could be treated as pre-edited item images, and the item images could be regenerated according to image editing conditions). Furthermore, the avatar can be dressed in clothing and other items that the user (themselves) actually wears in real life.

[0082] In this way, the server device 100 generates product images from noise using a diffusion model. Specifically, the server device 100 calculates the interactions between items using a Set Transformer and generates images that seem to be compatible with the items set as conditions. This makes it possible to fill in the gaps in the coordination. It also becomes possible to generate an unlimited number of coordinations without any limitations.

[0083] Furthermore, the generated images are not limited to images of items. For example, the server device 100 may generate not only images of items, but also snapshots (images of outfits, images of people wearing items, images of people wearing outfits, etc.). Snapshots posted by specific users on social media tend to harmonize because they reflect similar tastes. For example, the server device 100 may generate multiple snapshots similar to those posted by a specific user. Specifically, the server device 100 may use a diffusion model trained with multiple snapshots posted by a specific user as training data to generate multiple snapshots from noise images, taking into account their harmony with each other.

[0084] Furthermore, the server device 100 may be modified to include a score output module so that it can measure the degree of match and suitability.

[0085] Furthermore, the server device 100 may generate item sets that increase the set matching score by combining it with other models. The server device 100 may also generate item sets by combining it with FashionGPT.

[0086] [2-2. Design Philosophy of Outfit Diffusion Architecture] The following three properties must be satisfied: (a) Symmetry in swapping items within the outfit (not dependent on the order of the product images) (b) Variability in the number of items in a coordinated outfit (I want to be able to set the number of product images that make up the coordinated outfit arbitrarily). (c) Variableness of the item image patch (I want to be able to set Height and Width arbitrarily, like in StableDiffusion)

[0087] The above three conditions (a) to (c) above) can be satisfied if the process is carried out as follows.

[0088] (1) Input of product image features The server device 100 accepts input of feature quantities from product images that make up a coordinated outfit. The shape of the feature quantities for each product is as shown in the image (Height, Width, Channel). For example, image features are represented as a three-dimensional shape with batch size × sequence length × feature dimension.

[0089] (2) Creating features that represent the overall characteristics The server device 100 creates feature quantities that represent the overall characteristics of each product from the patch-like feature quantities of each product image (image-by-image transformation).

[0090] For example, the following methods can be considered for creating features. How to use classifier tokens like in Vision Transformer (ViT) • How to use Q-former (Query Converter) • A simple method of summing the features for each patch against Height and Width.

[0091] (3) Transformation to features considering permutation symmetry The server device 100 transforms the features obtained in (2) above into features that take symmetry into consideration (transformation for the entire coordination). Typically, this is SA+FF (SelfAttention + FeedForward). This satisfies (a) and (b) above. In addition, the transformation is performed by referring to the features of other products in the coordination to obtain the overall features of the product image that result in a natural coordination.

[0092] (4) Transformation of input features The server device 100 transforms the input features for each image based on the features obtained in (3) above (image-by-image transformation). Typically, this is CA+FF (CrossAttention + FeedForward). This satisfies the property of (c) above.

[0093] (5) Output of product image features The server device 100 outputs the feature quantities of the converted product image.

[0094] The above information can be illustrated as shown in Figure 5. Figure 5 is an explanatory diagram illustrating the design philosophy of the Outfit Diffusion architecture.

[0095] For example, as shown in Figure 5, the server device 100 uses the QF (Q-Former) of the ST unit (Set Transformer module) to extract feature quantities from each image (step S11). For example, the server device 100 inputs a product image to the QF (Q-Former) to extract the overall features of the product image. The QF (Q-Former) consists of two submodules: an image converter and a text converter. In this case, the image converter interacts with the image encoder that receives the input image to extract visual features.

[0096] Next, the server device 100 performs a transformation on the overall characteristics of the extracted product images, taking into account coordination, using the SA+FF (SelfAttention + FeedForward) function of the ST unit (step S12).

[0097] Next, the server device 100 uses the CA+FF (CrossAttention + FeedForward) function of the ST unit to transform the original product image based on feature quantities that take coordination into consideration (step S13).

[0098] Figure 6 shows an example of the coordination results. As a result of implementing this embodiment, (1) based on the IQQN3000 coordination, (2) a coordination with the same category generated, (3) a hat, earrings, coat, and skirt generated from noise images using Fill in the n blanks, (4) an edited coordination, and (5) a partially edited coordination are shown. The original coordination (1) is a monochrome coordination, but in coordination (2), while maintaining the category of each item (the original category is set as the image editing condition for each item), each item that harmonizes with it (shoes, hat, earrings, coat, skirt, gloves, sweater) is generated from noise images. As an example, the coat is changed from black to light blue, and the skirt is changed from black to brown. In coordination (3), the original shoes, gloves, and sweater are fixed, and while maintaining the categories of the other items (hat, earrings, coat, skirt), each item that harmonizes with them is generated from noise images. For example, the coat is blue and loose-fitting, and the skirt is a navy blue patterned one. Furthermore, the hat and earrings are changed to brown, adding an accent to the overall color scheme of the outfit. In outfits (4) and (5), the overall look and feel of the outfit and individual items are maintained, but the textures and other aspects have been changed.

[0099] [2-3. Supplementary Information] The server device 100 makes coordination suggestions based on the multiple product images that have been generated. The server device 100 has an architecture with multiple UNets with the same connection coefficients, which communicate with each other. The features between images are summarized each time. However, it is also possible to summarize only a portion of the features between images.

[0100] Furthermore, if a category is provided beforehand, the server device 100 may only use Set Transformer for the last few times. Conversely, if a category is not provided beforehand, the server device 100 must use Set Transformer from the beginning. Otherwise, for example, two pairs of shoes might appear.

[0101] Furthermore, the server device 100 may generate item images by combining a diffusion model and a set matching model. The set matching model is a model that has been trained to output a set matching score indicating the degree of compatibility between the first item group and the second item group when at least a first item group (multiple item images) and a second item group (multiple item images) are input. In addition to the first and second item groups, the model may also be trained to output a set matching score indicating the degree of compatibility between the first and second item groups and the additional information when additional information (for example, user information of the user wearing each item (including physical information such as body type and skin color)) is input. Furthermore, the set matching model may learn from images and user information of users wearing outfits posted on e-commerce sites, outfit posting sites, and various social networking sites as training data (ground truth data), or it may learn from images and user information of users wearing outfits posted on e-commerce sites, outfit posting sites, and various social networking sites that have been evaluated by viewers or evaluators (for example, evaluated as looking good, evaluated as harmonious) as training data (ground truth data). In other words, the set matching model may be trained to output a higher set matching score the closer the input data is to the ground truth data. The server device 100 may also input at least intermediate images during the process of generating (reconstructing) an item image from a noise image using a diffusion model to the set matching model, causing the set matching model to output (calculate) a set matching score, and causing the diffusion model to generate (reconstruct) an item image that will result in a higher set matching score. The network may be ControlNet in addition to UNet. ControlNet can create line drawings and generate images based on those line drawings. Examples of line drawings include line drawings corresponding to categories or line drawings requested by the user.

[0102] Furthermore, the server device 100 adds noise to the actual image it possesses to create a noised image, and without creating restored data, it removes the noise by comparing the input image with this noised image and then restores the original image.

[0103] The server device 100 restores the data to a point where categories do not need to be entered. The server device 100 only needs to be able to find items other than those the user possesses. Additionally, a dedicated model with a function to fill in gaps in the data may be available.

[0104] [3. Example of terminal device configuration] Next, the configuration of the terminal device 10 will be described using Figure 7. Figure 7 is a diagram showing an example of the configuration of the terminal device 10 according to the embodiment. As shown in Figure 7, the terminal device 10 comprises a communication unit 11, a display unit 12, an input unit 13, a positioning unit 14, a sensor unit 20, a control unit 30 (controller), and a storage unit 40.

[0105] (Communications Section 11) The communication unit 11 is connected to the network N by wire or wireless connection and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 11 can be implemented using a NIC (Network Interface Card) or an antenna.

[0106] (Display section 12) The display unit 12 is a display device that displays various information such as location information. For example, the display unit 12 may be a liquid crystal display (LCD) or an organic electro-luminescent display (OLED). The display unit 12 may also be a touch panel display, but is not limited to this.

[0107] (Input section 13) The input unit 13 is an input device that receives various operations from the user U. For example, the input unit 13 has buttons for inputting characters, numbers, etc. The input unit 13 may also be an input / output port (I / O port) or a USB (Universal Serial Bus) port. If the display unit 12 is a touch panel display, a part of the display unit 12 functions as the input unit 13. The input unit 13 may also be a microphone that receives voice input from the user U. The microphone may be wireless.

[0108] (Positioning unit 14) The positioning unit 14 receives signals (radio waves) transmitted from GPS (Global Positioning System) satellites and, based on the received signals, acquires position information (e.g., latitude and longitude) indicating the current position of the terminal device 10. In other words, the positioning unit 14 determines the position of the terminal device 10. Note that GPS is just one example of a GNSS (Global Navigation Satellite System).

[0109] Furthermore, the positioning unit 14 can determine its position using various methods other than GPS. For example, the positioning unit 14 may use various communication functions of the terminal device 10 to determine its position as an auxiliary positioning means for position correction, etc., as described below.

[0110] (Wi-Fi positioning) For example, the positioning unit 14 determines the location of the terminal device 10 by utilizing the Wi-Fi® communication function of the terminal device 10 and the communication network provided by each telecommunications company. Specifically, the positioning unit 14 determines the location of the terminal device 10 by performing Wi-Fi communication, etc., and determining the distance to nearby base stations and access points.

[0111] (Beacon positioning) Furthermore, the positioning unit 14 may determine the location using the Bluetooth® function of the terminal device 10. For example, the positioning unit 14 determines the location of the terminal device 10 by connecting to a beacon transmitter connected via the Bluetooth® function.

[0112] (Geomagnetic positioning) Furthermore, the positioning unit 14 determines the position of the terminal device 10 based on the geomagnetic pattern of the structure, which has been measured in advance, and the geomagnetic sensor provided by the terminal device 10.

[0113] (RFID positioning) Furthermore, if, for example, the terminal device 10 is equipped with an RFID (Radio Frequency Identification) tag function equivalent to that of a contactless IC card used at a train station ticket gate or in a store, or if it is equipped with a function to read RFID tags, the location where it was used will be recorded along with the information on the payment or other transactions made by the terminal device 10. The positioning unit 14 may determine the location of the terminal device 10 by acquiring such information. Alternatively, the location may be determined by an optical sensor or infrared sensor equipped in the terminal device 10.

[0114] The positioning unit 14 may, if necessary, determine the position of the terminal device 10 using one or a combination of the positioning means described above.

[0115] (Sensor unit 20) The sensor unit 20 includes various sensors mounted on or connected to the terminal device 10. The connection can be wired or wireless. For example, the sensors may be detection devices other than the terminal device 10, such as wearable devices or wireless devices. In the example shown in Figure 7, the sensor unit 20 includes an acceleration sensor 21, a gyro sensor 22, a barometric pressure sensor 23, a temperature sensor 24, a sound sensor 25, a light sensor 26, a magnetic sensor 27, and an image sensor (camera) 28.

[0116] The sensors 21-28 described above are merely examples and not limiting. In other words, the sensor unit 20 may be configured to include some of the sensors 21-28, or it may include other sensors such as humidity sensors in addition to or instead of the sensors 21-28.

[0117] The acceleration sensor 21 is, for example, a 3-axis acceleration sensor and detects the physical movement of the terminal device 10, such as its direction of movement, velocity, and acceleration. The gyro sensor 22 detects the physical movement of the terminal device 10, such as its tilt in the three axes, based on its angular velocity. The barometric pressure sensor 23 detects the atmospheric pressure around the terminal device 10, for example.

[0118] Since the terminal device 10 is equipped with the acceleration sensor 21, gyroscope 22, barometric pressure sensor 23, etc., it becomes possible to determine the position of the terminal device 10 using technologies such as pedestrian dead-reckoning (PDR) that utilize these sensors 21 to 23. This makes it possible to obtain indoor location information that is difficult to obtain with positioning systems such as GPS.

[0119] For example, a pedometer using an accelerometer 21 can calculate the number of steps, walking speed, and distance walked. Additionally, a gyroscope 22 can be used to determine the user U's direction of movement, gaze direction, and body tilt. Furthermore, the barometric pressure detected by the barometric pressure sensor 23 can be used to determine the altitude and floor number of the user U's terminal device 10.

[0120] The temperature sensor 24 detects, for example, the ambient temperature around the terminal device 10. The sound sensor 25 detects, for example, the ambient sound around the terminal device 10. The light sensor 26 detects the ambient illumination around the terminal device 10. The magnetic sensor 27 detects, for example, the Earth's magnetic field around the terminal device 10. The image sensor 28 captures an image of the area around the terminal device 10.

[0121] The aforementioned pressure sensor 23, temperature sensor 24, sound sensor 25, light sensor 26, and image sensor 28 can detect the surrounding environment and conditions of the terminal device 10 by detecting atmospheric pressure, temperature, sound, and illuminance, respectively, and by capturing images of the surroundings. Furthermore, it becomes possible to improve the accuracy of the location information of the terminal device 10 based on the surrounding environment and conditions.

[0122] (Control Unit 30) The control unit 30 includes, for example, a microcomputer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM, input / output ports, and various circuits. Alternatively, the control unit 30 may be composed of hardware such as an integrated circuit (ASIC) or FPGA (Field Programmable Gate Array). The control unit 30 includes a transmission unit 31, a reception unit 32, and a processing unit 33.

[0123] (Transmitter 31) The transmission unit 31 can transmit various information, such as information input by the user U using the input unit 13, various information detected by sensors 21-28 mounted on or connected to the terminal device 10, and location information of the terminal device 10 determined by the positioning unit 14, to the server device 100 via the communication unit 11.

[0124] (Receiver 32) The receiving unit 32 can receive various information provided by the server device 100, as well as requests for various information from the server device 100, via the communication unit 11.

[0125] (Processing 33) The processing unit 33 controls the entire terminal device 10, including the display unit 12. For example, the processing unit 33 can output and display various information transmitted by the transmission unit 31 and various information received from the server device 100 by the reception unit 32 to the display unit 12.

[0126] Furthermore, the processing unit 33 may function (operate) as a reception unit 33A, a search unit 33B, and a display control unit 33C, as shown below, by launching an application or the like. That is, the processing unit 33 includes a reception unit 33A, a search unit 33B, and a display control unit 33C.

[0127] (Reception area 33A) The reception unit 33A receives image generation conditions from the user U via the input unit 13. The reception unit 33A also receives one or more fixed item images. The reception unit 33A also receives one or more pre-edited item images and image editing conditions.

[0128] (Search section 33B) The search unit 33B communicates with the server device 100 via the communication unit 11 to search for content such as items or coordinates based on query images. For example, the search unit 33B searches for items based on each of multiple query images. Alternatively, the search unit 33B searches for items based on a query image specified by the user from among multiple query images. The search unit 33B may also use fixed item images or unedited item images received by the reception unit 33A as query images.

[0129] (Display control unit 33C) The display control unit 33C communicates with the server device 100 via the communication unit 11 and displays the content, which is the search result from the search unit 33B, on the display unit 12. For example, the display control unit 33C displays the search results for each of multiple query images on the display unit 12. Alternatively, the display control unit 33C displays the search results based on the query image specified by the user on the display unit 12.

[0130] Furthermore, the display control unit 33C displays the generated item image on the display unit 12 according to the image generation conditions received by the reception unit 33A. Also, the display control unit 33C displays the item image generated from the pre-edited item image via a noise image on the display unit 12 according to the image editing conditions received by the reception unit 33A.

[0131] The display control unit 33C communicates with the server device 100 via the communication unit 11 and displays item images generated from noise images on the display unit 12. For example, the display control unit 33C displays multiple item images on the display unit 12, which are generated from multiple noise images, taking into consideration their harmony with each other.

[0132] (Storage unit 40) The storage unit 40 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDD (Hard Disk Drive), SSD (Solid State Drive), and optical discs. Various programs and various data are stored in this storage unit 40.

[0133] [4. Example of Server Device Configuration] Next, the configuration of the server device 100 according to the embodiment will be described using Figure 8. Figure 8 is a diagram showing an example of the configuration of the server device 100 according to the embodiment. As shown in Figure 8, the server device 100 includes a communication unit 110, a storage unit 120, and a control unit 130.

[0134] (Communications Department 110) The communication unit 110 is implemented, for example, by a NIC (Network Interface Card). The communication unit 110 is connected to the network N by wire or wireless connection.

[0135] (Storage unit 120) The storage unit 120 is implemented by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as HDDs, SSDs, and optical discs. The storage unit 120 may store identification information (such as a user ID) indicating user U, as well as attribute information and history information (log data) of user U.

[0136] (Control unit 130) The control unit 130 is a controller, and is realized by executing various programs (corresponding to an example of an information processing program) stored in the internal memory of the server device 100 using a memory area such as RAM as a working area, for example, by a CPU (Central Processing Unit), MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), or FPGA (Field Programmable Gate Array). In the example shown in Figure 8, the control unit 130 has an acquisition unit 131, a setting unit 132, a construction unit 133, an image generation unit 134, and a provision unit 135.

[0137] (Acquisition part 131) The acquisition unit 131 acquires the search query entered by the user U. For example, when the user U enters a search query into a search engine or the like and performs a keyword search, the acquisition unit 131 acquires the search query via the communication unit 110. In other words, the acquisition unit 131 acquires the keyword entered by the user U into the search box of a search engine, website, or application via the communication unit 110.

[0138] Furthermore, the acquisition unit 131 acquires user information about user U via the communication unit 110. For example, the acquisition unit 131 acquires identification information (such as user ID), location information, and attribute information of user U from user U's terminal device 10. The acquisition unit 131 may also acquire identification information and attribute information of user U when user U is registered. The acquisition unit 131 then stores the user information in the storage unit 120.

[0139] Furthermore, the acquisition unit 131 acquires various historical information (log data) indicating the user U's actions via the communication unit 110. For example, the acquisition unit 131 acquires various historical information indicating the user U's actions from the user U's terminal device 10, or from various servers based on the user ID, etc. The acquisition unit 131 then stores the various historical information in the storage unit 120.

[0140] Furthermore, the acquisition unit 131 acquires one or more fixed item images. Also, the acquisition unit 131 acquires one or more pre-edited item images. For example, the acquisition unit 131 acquires an item image that will be the query image to be searched from the user U's terminal device 10 via the communication unit 110.

[0141] (Settings section 132) The setting unit 132 sets the image generation conditions. For example, the setting unit 132 receives instructions for image generation conditions from the user U's terminal device 10 via the communication unit 110 and sets the image generation conditions. The setting unit 132 also sets the image editing conditions. For example, the setting unit 132 receives instructions for image editing conditions from the user U's terminal device 10 via the communication unit 110 and sets the image editing conditions.

[0142] (Construction Section 133) The construction unit 133 constructs a diffusion model that has learned a diffusion process to generate multiple noise images by adding noise to each of multiple harmonized item images, and a dediffusion process to generate the original multiple item images by removing noise from each of the multiple noise images. In other words, the construction unit 133 is a learning unit that generates a diffusion model using machine learning.

[0143] (Image generation unit 134) The image generation unit 134 generates item images using a diffusion model. The image generation unit 134 also generates item images according to image generation conditions.

[0144] Furthermore, the image generation unit 134 generates item images from noise images. For example, the image generation unit 134 generates multiple item images from multiple noise images, taking into consideration their harmony with each other.

[0145] Furthermore, the image generation unit 134 generates item images while also considering their harmony with one or more fixed item images acquired by the acquisition unit 131.

[0146] Furthermore, the image generation unit 134 generates item images from one or more pre-edited item images acquired by the acquisition unit 131, taking into account the harmony with each of them, via noise images, according to the image editing conditions.

[0147] (Provider 135) The provisioning unit 135 provides the generated item images to the user U's terminal device 10 via the communication unit 110. For example, the provisioning unit 135 provides images of items within a coordinate to the user U's terminal device 10 via the communication unit 110. The provisioning unit 135 also provides the search results of an item search based on a query image to the user U's terminal device 10 via the communication unit 110.

[0148] [5. Processing Procedure] Next, the processing procedure by the server device 100 according to the embodiment will be described using Figure 9. Figure 9 is a flowchart of the processing procedure according to the embodiment. Note that the processing procedure shown below is repeatedly executed by the control unit 130 of the server device 100.

[0149] For example, as shown in Figure 9, the acquisition unit 131 of the server device 100 acquires one or more fixed item images and one or more pre-edited item images (step S101). This step may be omitted.

[0150] Next, the setting unit 132 of the server device 100 receives instructions for image generation conditions and image editing conditions from the user U's terminal device 10 via the communication unit 110, and sets the image generation conditions and image editing conditions (step S102).

[0151] Next, the construction unit 133 of the server device 100 constructs a diffusion model that has learned a diffusion process to generate multiple noise images by adding noise to each of the multiple harmonized item images, and a dediffusion process to generate the original multiple item images by removing noise from each of the multiple noise images (step S103). Note that this step may be performed before step S101 or before step S102.

[0152] Next, the image generation unit 134 of the server device 100 generates item images from the noise image using a diffusion model, taking into account their harmony with the image, according to the image generation conditions and image editing conditions (step S104). When the image generation unit 134 of the server device 100 generates item images according to the pre-edited item image and image editing conditions, it first adds noise to all or part of the pre-edited item image, and then generates item images from there, taking into account their harmony with the image.

[0153] Next, the server device 100 provides the generated item image to the user U's terminal device 10 via the communication unit 110 (step S105). At this time, the user U's terminal device 10 displays the generated item image.

[0154] Next, the control unit 130 of the server device 100 searches for the item image of the actual item based on the generated item image (query image) (step S106). The control unit 130 of the server device 100 may also search for the item image of the actual item from among the generated item images based on the item image (query image) specified by the user. Note that the item image of the actual item found may not be one for each query image, but may be multiple.

[0155] Next, the server device 100's provision unit 136 provides the search results (actual item images) to the user U's terminal device 10 via the communication unit 110 (step S107). At this time, the user U's terminal device 10 displays the search results (actual item images). If there are multiple item images for a single query image, the server device 100's provision unit 136 may prioritize providing item images of actual items that are similar to the query image. Specifically, the user U's terminal device 10 will display item images of actual items that are similar to the query image in a more prominent or higher position.

[0156] [6. Variant Example] The terminal device 10 and server device 100 described above may be implemented in various other forms besides those of the embodiment described above. Therefore, the following describes modifications of the embodiment.

[0157] In the above embodiment, some or all of the processing performed by the server device 100 may actually be performed by the user U's terminal device 10 (or an application running on the terminal). For example, the processing may be completed in a standalone manner (by the terminal device 10 alone). In this case, the terminal device 10 is assumed to have the same functions as the server device 100 in the above embodiment. Furthermore, in the above embodiment, since the terminal device 10 is in cooperation with the server device 100, from the user U's perspective, it appears as if the processing of the server device 100 is also being performed by the terminal device 10. In other words, from another perspective, the terminal device 10 can be said to have the server device 100.

[0158] Furthermore, in the above embodiment, some or all of the processing performed by the user U's terminal device 10 may actually be performed by the server device 100.

[0159] Furthermore, in the above embodiment, the user U's terminal device 10 and the server device 100 may be the same device (one device). In other words, the processes performed by the user U's terminal device 10 and the server device 100 may be performed by the same device (one device).

[0160] Furthermore, in the above embodiment, the image of the item may be a video or a multi-view image. The image of the item may also be an illustration.

[0161] [7. Effects] As described above, the information processing device (terminal device 10 and server device 100) according to the present invention includes an image generation unit 134 that generates item images from noise images, and the image generation unit 134 generates multiple item images from multiple noise images, taking into consideration their harmony with each other.

[0162] Furthermore, the information processing device according to the present invention includes a setting unit 132 for setting image generation conditions, and an image generation unit 134 generates item images according to the image generation conditions.

[0163] Furthermore, the information processing device according to the present invention includes an acquisition unit 131 that acquires one or more fixed item images, and an image generation unit 134 that generates item images while also considering their harmony with the fixed item images.

[0164] Furthermore, the information processing device according to the present invention includes an acquisition unit 131 that acquires one or more pre-edited item images, and a setting unit 132 that sets image editing conditions, and an image generation unit 134 that generates an item image from the pre-edited item image via a noise image according to the image editing conditions.

[0165] Furthermore, the information processing device according to the present invention includes a construction unit 133 that constructs a diffusion model that has learned a diffusion process to generate multiple noise images by adding noise to each of multiple harmonized item images, and a dediffusion process to generate the original multiple item images by removing noise from each of the multiple noise images, and an image generation unit 134 that generates item images using the diffusion model.

[0166] From another perspective, the information processing device according to the present application (terminal device 10 and server device 100) includes a display unit 12 that displays item images generated from noise images, and the display unit 12 displays a plurality of item images generated from a plurality of noise images, taking into consideration their harmony with each other.

[0167] Furthermore, the information processing device according to the present invention includes a search unit 33B that searches for content such as items or coordinates based on a query image, and a display unit 12 displays the content which is the search result.

[0168] Furthermore, the information processing device according to the present invention includes a receiving unit 33A that receives image generation conditions, and the display unit 12 displays the generated item image according to the image generation conditions.

[0169] Furthermore, the information processing device according to the present invention includes a receiving unit 33A that receives one or more pre-edited item images and image editing conditions, and the display unit 12 displays an item image generated from the pre-edited item image via a noise image according to the image editing conditions.

[0170] Furthermore, the search unit 33B searches for items based on each of the multiple query images, and the display unit 12 displays the respective search results.

[0171] Furthermore, the search unit 33B searches for items based on the query image specified by the user from among multiple query images, and the display unit 12 displays the search results.

[0172] Through any or a combination of the above-described processes, the information processing device according to the present invention can provide highly harmonious information as a whole.

[0173] [8. Hardware Configuration] Furthermore, the terminal device 10 and server device 100 according to the above-described embodiment are realized by a computer 1000 having a configuration such as that shown in Figure 10. The following explanation will use the server device 100 as an example. Figure 10 shows an example of the hardware configuration. The computer 1000 is connected to an output device 1010 and an input device 1020, and has a configuration in which an arithmetic unit 1030, a primary storage device 1040, a secondary storage device 1050, an output interface 1060, an input interface 1070, and a network interface 1080 are connected by a bus 1090.

[0174] The arithmetic unit 1030 operates based on programs stored in the primary storage device 1040 and the secondary storage device 1050, as well as programs read from the input device 1020, and executes various processes. The arithmetic unit 1030 can be implemented using, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array).

[0175] The primary storage device 1040 is a memory device, such as RAM (Random Access Memory), that temporarily stores data used by the arithmetic unit 1030 for various calculations. The secondary storage device 1050 is a storage device where data used by the arithmetic unit 1030 for various calculations and various databases are registered, and can be implemented using ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The secondary storage device 1050 may be internal storage or external storage. The secondary storage device 1050 may also be a removable storage medium such as USB (Universal Serial Bus) memory or SD (Secure Digital) memory card. The secondary storage device 1050 may also be cloud storage (online storage), NAS (Network Attached Storage), file server, etc.

[0176] The output I / F 1060 is an interface for transmitting information to be output to output devices 1010, such as displays, projectors, and printers, and is implemented using connectors of standards such as USB (Universal Serial Bus), DVI (Digital Visual Interface), and HDMI (High Definition Multimedia Interface). The input I / F 1070 is an interface for receiving information from various input devices 1020, such as mice, keyboards, keypads, buttons, and scanners, and is implemented using, for example, USB.

[0177] Furthermore, the output interface 1060 and input interface 1070 may be wirelessly connected to the output device 1010 and input device 1020, respectively. In other words, the output device 1010 and input device 1020 may be wireless devices.

[0178] Furthermore, the output device 1010 and the input device 1020 may be integrated as a touch panel. In this case, the output I / F 1060 and the input I / F 1070 may also be integrated as an input / output I / F.

[0179] The input device 1020 may also be a device that reads information from, for example, an optical recording medium such as a CD (Compact Disc), DVD (Digital Versatile Disc), or PD (Phase Change Rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0180] The network interface 1080 receives data from other devices via network N and sends it to the computing unit 1030, and also transmits data generated by the computing unit 1030 to other devices via network N.

[0181] The arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output interface 1060 and the input interface 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or the secondary storage device 1050 onto the primary storage device 1040 and executes the loaded program.

[0182] For example, when computer 1000 functions as a server device 100, the arithmetic unit 1030 of computer 1000 realizes the functions of the control unit 130 by executing a program loaded onto the primary storage device 1040. Alternatively, the arithmetic unit 1030 of computer 1000 may load a program obtained from another device via the network interface 1080 onto the primary storage device 1040 and execute the loaded program. Furthermore, the arithmetic unit 1030 of computer 1000 may cooperate with other devices via the network interface 1080 and call and use program functions, data, etc., from other programs on other devices.

[0183] [9. Other] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.

[0184] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0185] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0186] For example, the server device 100 described above may be implemented using multiple server computers, and the configuration can be flexibly changed, such as by calling external platforms via APIs (Application Programming Interfaces) or network computing depending on the function.

[0187] Furthermore, the embodiments and modifications described above can be combined as appropriate, provided that the processing content is not inconsistent.

[0188] Furthermore, the terms "section, module, unit" mentioned above can be replaced with "means" or "circuit," etc. For example, the acquisition unit can be replaced with acquisition means or acquisition circuit. [Explanation of Symbols]

[0189] 1. Information Processing System 10 Terminal devices 12 Display section 13 Input section 33A Reception Desk 33B Search Section 33C Display Control Unit 100 Server Devices 110 Communications Department 120 Storage section 130 Control Unit 131 Acquisition Department 132 Settings Section 133 Construction Department 134 Image generation unit 135 Provision Department

Claims

1. It includes an image generation unit that generates item images from noise images, The image generation unit generates multiple item images from multiple noise images, taking into consideration their harmony with each other. An information processing device characterized by the following:

2. It includes a setting unit for setting image generation conditions, The image generation unit generates an item image according to the set image generation conditions. The information processing apparatus according to feature 1.

3. It includes an acquisition unit that acquires one or more fixed item images, The image generation unit generates item images while also considering their harmony with the acquired fixed item images. The information processing apparatus according to feature 1.

4. An acquisition unit that acquires one or more pre-edited item images, A settings section for setting image editing conditions, Equipped with, The image generation unit generates an item image from the acquired pre-edited item image via a noise image, according to the set image editing conditions. The information processing apparatus according to feature 1.

5. The system includes a construction unit that builds a diffusion model that learns a diffusion process that generates multiple noise images by adding noise to each of multiple harmonized item images, and a dediffusion process that generates the original multiple item images by removing noise from each of the multiple noise images. The image generation unit generates item images using the constructed diffusion model. The information processing apparatus according to feature 1.

6. It includes a providing unit that provides item images generated from noise images, The aforementioned providing unit provides multiple item images generated from multiple noise images, taking into consideration their harmony with each other. An information processing device characterized by the following:

7. The system includes a search processing unit that uses the provided item image as a query image to search for content such as items or outfits. The aforementioned provisioning unit provides the retrieved content. The information processing apparatus according to feature 6.

8. It includes an acquisition unit that acquires the conditions for generating an image, The providing unit provides item images generated according to the acquired image generation conditions. The information processing apparatus according to feature 6.

9. It includes an acquisition unit that acquires one or more pre-edited item images and image editing conditions, The providing unit provides an item image generated from the acquired pre-edited item image via a noise image, according to the acquired image editing conditions. The information processing apparatus according to feature 6.

10. The search processing unit uses the provided multiple item images as multiple query images to search for items, The aforementioned provisioning unit provides each of the retrieved items. The information processing apparatus according to feature 7.

11. The search processing unit searches for items using the item image specified by the user as the query image from among the multiple item images provided. The aforementioned providing unit provides the searched items. The information processing apparatus according to feature 7.

12. An information processing method performed by an information processing device, This includes an image generation process that generates item images from noise images, In the image generation process, multiple item images are generated from multiple noise images, taking into consideration their harmony with each other. An information processing method characterized by the following:

13. An information processing method performed by an information processing device, The process includes providing an item image generated from a noise image, In the aforementioned provisioning process, multiple item images are provided, which are generated from multiple noise images, taking into consideration their harmony with each other. An information processing method characterized by the following:

14. The computer is made to perform an image generation procedure that generates item images from noise images. The aforementioned image generation procedure generates multiple item images from multiple noise images, taking into consideration their harmony with each other. An information processing program characterized by causing the computer to execute the following.

15. The computer is made to perform a procedure that provides item images generated from noise images. The aforementioned provision procedure provides multiple item images generated from multiple noise images, taking into consideration their harmony with each other. An information processing program characterized by causing the computer to execute the following.

Citation Information

Patent Citations

  • Attribute generative adversarial network and matched clothes generation method based on attribute generative adversarial network

    CN110909754A

  • Multi-mode-based generative compatible garment matching scheme generation method and system

    CN111861672A

  • Sketch guidance-based paired clothing image generation method

    CN113298906A

  • Information processing device, data extraction method, and data extraction program

    JP2020098521A

  • A method for providing a fashion item recommendation service to users using swipe gestures

    JP2022515617A