Information processing device, information processing system, and learning device
The information processing device uses visual and audio data to accurately recommend clothing items based on customer characteristics and preferences, addressing the limitations of existing technologies by providing more personalized fashion recommendations.
Patent Information
- Application Number
- JP2023187443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2038-11-02
AI Technical Summary
Existing technologies struggle to accurately recommend clothing items to customers based on their characteristics and preferences, often relying on incomplete or irrelevant data such as physical features or personal measurements.
An information processing device that uses a camera, microphone, and image and voice feature extraction units to identify face and body areas, extract feature amounts, and input them into a trained estimation model to propose clothing items that match the customer's characteristics and preferences.
The system achieves higher accuracy in recommending clothing items by utilizing a combination of visual and audio data to understand customer preferences, leading to more personalized and effective fashion recommendations.
Smart Images

Figure 0007672008000001 
Figure 0007672008000002 
Figure 0007672008000003
Abstract
Description
[Technical field]
[0001] The present invention relates to a technique for proposing clothing items appropriate to a customer from among a plurality of clothing items. [Background technology]
[0002] In a store that sells clothing items, many clothing items are displayed, making it difficult for a prospective purchaser to find the clothing item he or she desires.
[0003] For example, JP 2017-215667 A (Patent Document 1) addresses the problem of not being able to easily recommend recommended products to a customer who visits a store based on the items the customer possesses or the sales products the customer is looking at, and discloses a configuration that uses a photograph taken of items the customer is wearing or sales products displayed in the store where the customer is located to suggest product information about recommended products of a type corresponding to owner information of the items, etc., depicted in the photograph.
[0004] International Publication No. 2003 / 069526 (Patent Document 2) discloses a fashion advising system that includes a first database device that outputs data on fashion content that suits input physical characteristics, and a second database device that outputs data on stores that offer the fashion content based on the fashion content data output from the first database device.
[0005] JP 2001-502090 A (Patent Document 3) relates to a method for fashion shopping by a customer, and specifically discloses a method for helping a customer select an appropriate fashion to purchase based on data related to the customer. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] JP 2017-215667 A [Patent Document 2] International Publication No. 2003 / 069526 [Patent Document 3] Special Publication No. 2001-502090 Summary of the Invention [Problem to be solved by the invention]
[0007] The configuration disclosed in Patent Document 1 recommends products that are different in type from the item shown in the photograph and have design elements such as color, shape, pattern, etc. that match the item, or when the owner information is second owner information, the recommended products are, for example, products that are the same type as the item shown in the photograph and have design elements such as color, shape, pattern, etc. that match the item, and is not one that recommends products according to the customer's preferences.
[0008] The configuration disclosed in Patent Document 2 is primarily focused on determining fashion content that suits a customer's physical characteristics when those characteristics are input, and is not intended to provide fashion that corresponds to the customer's preferences.
[0009] The configuration disclosed in Patent Document 3 acquires personal information including measurements of bust, hips, waist, arm length, height, and center front to assist in the selection of clothing items to be purchased. However, personal information is required to suggest clothing items, and the configuration is not suitable for general-purpose use. It is a composition.
[0010] The present invention aims to provide a technique for suggesting clothing items suitable for a customer from among a plurality of clothing items with higher accuracy based on feature quantities representing the customer's characteristics. [Means for solving the problem]
[0011] According to an aspect of the present invention, there is provided an information processing device that suggests clothing items suitable for a customer from among a plurality of clothing items based on a feature amount representing the customer's features. The information processing device includes a camera for capturing an image of the customer, a microphone for collecting voice, a region specifying unit for specifying a face region representing the customer's face and a body region representing the customer's body in an image obtained by capturing an image of the customer with the camera, an image feature extracting unit for extracting a first feature amount from the face region of the image and a second feature amount from the body region of the image, a voice feature extracting unit for extracting a third feature amount from a portion of the voice corresponding to the customer's speech among the voice collected by the microphone, a trained estimation model that receives the first feature amount, the second feature amount, and the third feature amount and outputs as an estimation result a possibility that each of the plurality of clothing items is a clothing item to be suggested, and a display unit for displaying the clothing items suitable for the customer based on the estimation result. The estimation model is generated by a learning process using a training dataset, and the training dataset includes multiple pieces of training data in which images of other customers and audio utterances of the other customers are labeled with clothing items purchased by the other customers.
[0012] Before collecting audio through the microphone, the display unit may display a list of categories indicating classifications of clothing items, and may display a message encouraging the customer to select by voice one of the categories displayed in the list.
[0013] The region identifying unit may identify a portion representing clothing worn by the customer as the body region.
[0014] Each of the plurality of clothing items may belong to one of a plurality of predetermined categories. The information processing device may further include a voice analysis unit for identifying a category selected by the customer from among the plurality of categories based on a voice uttered by the customer. The display unit may display, among the clothing items displayed based on the estimation result, clothing items belonging to a category identified by the voice analysis unit and clothing items not belonging to the identified category in different display modes.
[0015] According to another aspect of the present invention, an information processing system includes an information processing device that inputs a feature amount representing a customer's feature into a trained estimation model to suggest a clothing item suitable for the customer from among a plurality of clothing items, and a learning device for generating the estimation model. The information processing device includes a camera for capturing an image of the customer, a microphone for collecting voice, a region specifying unit for specifying a face region representing the customer's face and a body region representing the customer's body in an input image obtained by capturing an image of the customer with the camera, an image feature extracting unit for extracting a first feature amount from the face region of the input image and a second feature amount from the body region of the input image, and a voice feature extracting unit for extracting a third feature amount from a portion of the voice corresponding to the customer's speech among the voice collected by the microphone. The estimation model is trained to receive the input of the first feature amount, the second feature amount, and the third feature amount and output, as an estimation result, a possibility that each of the plurality of clothing items is a clothing item to be suggested. The information processing device further includes a display unit for displaying the clothing item suitable for the customer based on the estimation result. The learning device includes an acquisition unit for acquiring a learning dataset. The learning data set includes a plurality of learning data in which learning images obtained by capturing images of other customers and learning voices uttered by the other customers are labeled with clothing items purchased by the other customers. The data acquisition system includes an area identification unit for identifying within an image a face area representing the face of another customer and a body area representing the body of another customer; an image feature extraction unit for extracting a first learning feature from the face area of the training image and extracting a second learning feature from the body area of the training image; an audio feature extraction unit for extracting a third learning feature from a portion of the training audio corresponding to the speech of another customer; and a learning unit for optimizing the estimation model so that an estimation result output by inputting the first learning feature, the second learning feature, and the third learning feature extracted from the training data into the estimation model approaches the purchase history of the clothing item labeled in the training data.
[0016] According to yet another aspect of the present invention, there is provided a learning device for receiving an input of a feature amount representing a customer's feature and generating an estimation model used to propose a clothing item suitable for the customer from among a plurality of clothing items. The learning device includes an acquisition unit for acquiring a learning dataset. The learning dataset includes a plurality of learning data in which an image obtained by capturing an image of a customer and a voice uttered by the customer are labeled with clothing items purchased by the customer. The learning device includes a region identification unit for identifying a face region representing the customer's face and a body region representing the customer's body in the image, an image feature extraction unit for extracting a first feature amount from the face region of the image and a second feature amount from the body region of the image, a voice feature extraction unit for extracting a third feature amount from a portion of the voice corresponding to the customer's utterance, and a learning unit for optimizing the estimation model such that an estimation result output by inputting the first feature amount, the second feature amount, and the third feature amount extracted from the learning data into the estimation model approaches the purchase history of the clothing items labeled in the learning data.
[0017] According to yet another aspect of the present invention, there is provided a trained estimation model that is used to receive input of features representing characteristics of a customer and to suggest clothing items appropriate to the customer from among a plurality of clothing items. The estimation model is generated by a learning process using a training dataset. The training dataset includes a plurality of training data in which an image obtained by capturing an image of a customer and a voice uttered by the customer are labeled with clothing items purchased by the customer. The learning process includes the steps of: identifying, for each of the training data, a face region representing the customer's face and a body region representing the customer's body in the image; extracting a first feature from the face region of the image and a second feature from the body region of the image; extracting a third feature from a portion of the voice corresponding to the customer's speech; and optimizing the estimation model such that an estimation result output by inputting the first feature, the second feature, and the third feature into the estimation model approaches the purchase history of the clothing items labeled in the training data.
[0018] According to yet another aspect of the present invention, there is provided a method for collecting learning data used to train an estimation model used to propose clothing items suitable for the customer from among a plurality of clothing items upon receiving input of feature quantities representing the characteristics of the customer. The method for collecting learning data includes the steps of: acquiring an image obtained by imaging the customer and a voice including the speech of the customer; inputting a plurality of feature quantities extracted from the image and the voice into a trained estimation model to generate a proposal of clothing items suitable for the customer; generating identification information; issuing a medium for encouraging the purchase of the clothing items, including the generated proposal of clothing items and the generated identification information; associating the generated identification information with the image and the voice; acquiring the identification information contained in the medium and the clothing items purchased by the customer; associating the identification information acquired from the medium with the clothing items purchased by the customer; and associating the image and the voice with the clothing items purchased by the customer using the identification information as a key, and saving the image and the voice with the clothing items purchased by the customer as learning data used to train the estimation model. Effect of the Invention
[0019] According to the present invention, clothing items suited to a customer can be suggested with higher accuracy from among a plurality of clothing items based on feature quantities representing the customer's characteristics. [Brief description of the drawings]
[0020] [Figure 1] 1 is a schematic diagram showing an example of the exterior of a store in which a clothing suggestion system according to an embodiment of the present invention is installed; [Diagram 2] 5 is a diagram for illustrating processing in a display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Diagram 3] 5 is a diagram for illustrating processing in a display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Figure 4] 11 is a diagram for explaining a customer using a coupon output from a display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Diagram 5] 11 is a diagram for illustrating a process of generating a learning data set in the clothing suggestion system according to the present embodiment. FIG. [Figure 6] 1 is a schematic diagram showing an example of a system configuration of a clothing suggestion system according to an embodiment of the present invention. [Figure 7] FIG. 2 is a schematic diagram showing an example of a hardware configuration of a display terminal constituting the clothing suggestion system according to the present embodiment. [Figure 8] FIG. 2 is a schematic diagram showing an example of a hardware configuration of a POS terminal constituting the clothing suggestion system according to the present embodiment. [Figure 9] FIG. 2 is a schematic diagram showing an example of a hardware configuration of a management device constituting the clothing suggestion system according to the present embodiment. [Figure 10] FIG. 2 is a schematic diagram showing an example of a functional configuration of a display terminal constituting the clothing suggestion system according to the present embodiment. [Figure 11] 11 is a diagram for illustrating processing details in a suggested item estimation function of the display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Figure 12]12 is a diagram for explaining the area specifying process performed by the area specifying module shown in FIG. 11. FIG. [Figure 13] 12 is a diagram for explaining a process of section specification by the section specification module shown in FIG. 11. FIG. [Figure 14] FIG. 12 is a schematic diagram illustrating an example of a network configuration of the estimation model illustrated in FIG. [Figure 15] 11 is a diagram for illustrating the processing contents of a display control function 150 and a coupon issuance control function of the display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Figure 16] 11 is a diagram for illustrating processing contents in an image and audio saving function 170 of the display terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Figure 17] 10 is a flowchart showing a processing procedure for an item estimation process in the display terminal constituting the clothing suggestion system according to the present embodiment. [Figure 18] FIG. 2 is a schematic diagram showing an example of a functional configuration of a POS terminal constituting the clothing suggestion system according to the present embodiment. [Figure 19] 11 is a diagram for illustrating processing contents in a sales information storage function 250 of the POS terminal constituting the clothing suggestion system according to the present embodiment. FIG. [Figure 20] 10 is a flowchart showing a processing procedure for sales management processing in the POS terminal constituting the clothing suggestion system according to the present embodiment. [Figure 21] FIG. 11 is a diagram for illustrating an overview of a learning phase in the clothing suggestion system according to the present embodiment. [Figure 22] FIG. 2 is a schematic diagram showing an example of a functional configuration of a management device constituting the clothing suggestion system according to the present embodiment. [Figure 23] FIG. 11 is a diagram illustrating the processing content of a learning data set generating function 350 of the management device constituting the clothing suggestion system according to the present embodiment. [Figure 24] FIG. 11 is a diagram illustrating the process performed by a learning function 360 of the management device constituting the clothing suggestion system according to the present embodiment. [Diagram 25]It is a flowchart showing the processing procedure of the learning process in the management device constituting the clothing proposal system according to the present embodiment. [Figure 26] It is a schematic diagram showing an example of the system configuration of the clothing proposal system according to Modification Example 1 of the present embodiment. [Figure 27] It is a diagram for explaining the item proposal screen displayed on the display terminal of the clothing proposal system according to Modification Example 2 of the present embodiment. [Figure 28] It is a diagram for explaining the processing contents in the display control function and the coupon issuance control function of the display terminal constituting the clothing proposal system according to Modification Example 2 of the present embodiment. [Figure 29] It is a diagram for explaining the processing contents in the proposed item estimation function of the display terminal constituting the clothing proposal system according to Modification Example 3 of the present embodiment. [Diagram 30] It is a diagram for explaining the processing contents in the proposed item estimation function of the display terminal constituting the clothing proposal system according to Modification Example 4 of the present embodiment. [Diagram 31] It is a schematic diagram showing an example of the use of the clothing proposal system according to Modification Example 5 of the present embodiment. [Diagram 32] It is a schematic diagram showing an implementation example of the clothing proposal system according to Modification Example 5 of the present embodiment.
Embodiments for Carrying Out the Invention
[0021] Embodiments of the present invention will be described in detail with reference to the drawings. For the same or corresponding parts in the drawings, the same reference numerals are given and the description thereof will not be repeated.
[0022] <A. Outline of the Clothing Proposal System> First, as a typical example of the information processing system according to the present invention, the outline of the clothing proposal system 1 according to the present embodiment will be described.
[0023] In this specification, "clothing" generally refers to clothing (apparel) and accessories (decorations) worn by a person. "Clothing item" is a term that refers to any product included in clothing. For ease of explanation, "clothing item" may also be referred to simply as "item."
[0024] In this specification, the term "customer" refers to a user who has some intention to purchase a clothing item. In the following description, a customer who visits a store is also referred to as a "shop visitor." In addition, a customer who uses the system according to the present embodiment via a mobile terminal is also referred to as an "Internet user."
[0025] Fig. 1 is a schematic diagram showing an example of the exterior of a store in which a clothing suggestion system 1 according to the present embodiment is installed. Fig. 2 and Fig. 3 are diagrams for explaining the processing in a display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment.
[0026] As shown in Fig. 1, assume that a customer (hereinafter also referred to as "customer 40") enters store 30. A display terminal 100, which is an example of an information processing device, is placed near the entrance of store 30. Display terminal 100 includes a relatively large display 102, and a human presence sensor 128, a camera 130, and a microphone 132 placed near display 102. A printer 120 is placed below display 102.
[0027] When a customer 40 approaches the display terminal 100 (FIG. 2(a)), the human presence sensor 128 detects the approach, and a category selection reception screen 50 is displayed on the display 102 (FIG. 2(b)). In this state, the customer 40 is captured by the camera 130 of the display terminal 100. That is, the display terminal 100 acquires an image showing the customer 40 (hereinafter, also referred to as a "captured image 136").
[0028] The category selection reception screen 50 displays a list of one or more categories. In addition, a message saying "Please select a category by voice" is displayed to encourage the customer 40 to speak.
[0029] Thereafter, the microphone 132 of the display terminal 100 starts collecting voice, and when the customer 40 utters a voice indicating the desired category ("jacket" in the example shown in FIG. 2) (FIG. 2(c)), an item suggestion screen 52 is output to the display 102 (FIG. 3(a)). At this time, the display terminal 100 acquires the voice uttered by the customer 40 (hereinafter also referred to as "collected voice 138").
[0030] In this way, before collecting audio by microphone 132, display 102 displays a list of categories indicating the classification of clothing items, and displays a message encouraging customer 40 to select by voice one of the categories displayed in the list.
[0031] The item proposal screen 52 includes a list display 54 of clothing items estimated to be "recommended" based on the preferences of the store customer 40. The items displayed in the list on the item proposal screen 52 are determined based on an estimation result obtained by executing an item estimation process using a trained model as described below. In this manner, the display terminal 100, which is an example of an information processing device, suggests clothing items suitable for the customer from among a plurality of clothing items based on feature amounts (typically, the captured image 136 and the collected voice 138) that represent the characteristics of the customer.
[0032] The item suggestion screen 52 further includes a coupon issue button 56. In response to pressing the coupon issue button 56, the printer 120 outputs a coupon 10.
[0033] The coupon 10 output from the printer 120 includes, in addition to the discount amount display 12, a list display 14 corresponding to the list display 54 included in the item suggestion screen 52, and a map 16 showing the location within the store of each item included in the list display 14 (Figure 3(b)).
[0034] Furthermore, the coupon 10 includes an identification image 18 such as a QR code (registered trademark) indicating a coupon ID described later. By using the coupon ID indicated by the identification image 18, a training dataset used for training the estimation model is generated.
[0035] FIG. 4 is a diagram for explaining a customer 40 who uses the coupon 10 output from the display terminal 100 constituting the clothing proposal system 1 according to the present embodiment. The customer 40 can enjoy shopping while referring to the content printed on the coupon 10 (FIG. 4(a)). Since a discount is applied by presenting the coupon 10, usually, the customer 40 presents the coupon 10 output from the display terminal 100 at the time of checkout (FIG. 4(b)).
[0036] FIG. 5 is a diagram for explaining the generation process of the training dataset in the clothing proposal system 1 according to the present embodiment. Referring to FIG. 5, the captured image 136 and the collected voice 138 acquired at the display terminal 100 and the information on the purchased item (hereinafter also referred to as "sales information 218") are associated with each other via the coupon 10 (specifically, the coupon ID 166). In this way, the associated captured image 136 and collected voice 138 and the sales information 218 are used as a training dataset for training the estimation model.
[0037] As described above, in the clothing proposal system 1 according to the present embodiment, items are proposed based on the preferences of the customer 40 at the time of entry, and an estimation model for proposing items can be trained using the information on the items actually purchased by the customer 40.
[0038] <B. Hardware Configuration Example of Clothing Proposal System> Next, a system configuration example of the clothing proposal system 1 according to the present embodiment will be described. First, after explaining the overall configuration example of the clothing proposal system 1, a hardware configuration example of the main devices included in the clothing proposal system 1 will be described.
[0039] (b1: System Configuration Example) Fig. 6 is a schematic diagram showing an example of a system configuration of a clothing suggestion system 1 according to the present embodiment. Referring to Fig. 6, the clothing suggestion system 1 includes one or more display terminals 100, one or more POS terminals 200, and a management device 300, which are connected via a local network 2.
[0040] The display terminal 100 is typically placed near the entrance of the store 30, and suggests clothing that matches the tastes of the customer. More specifically, the display terminal 100 captures an image of the customer and collects the voice of the customer. The display terminal 100 inputs the image (hereinafter also referred to as "captured image") and voice (hereinafter also referred to as "collected voice") of the customer into a trained model, thereby calculating the degree of suitability (hereinafter also referred to as "score") for the customer's taste for each item being sold. The display terminal 100 suggests items with high scores to the customer. The display terminal 100 can also issue a coupon on which the suggested item is printed to the customer.
[0041] The display terminal 100 can also transmit captured images and collected sounds to the management device 300 upon request.
[0042] The POS terminal 200 executes transaction processing for the items that the customer wishes to purchase. The POS terminal 200 generates information on the purchased items (sales information) and can also transmit this information to the management device 300 upon request.
[0043] The management device 300 is responsible for managing and updating the trained model used by the display terminal 100. More specifically, the management device 300 acquires captured images and collected audio from the display terminal 100, and also acquires sales information from the POS terminal 200. The management device 300 then generates a training dataset from the acquired captured images and collected audio and the acquired sales information. The management device 300 uses the generated training dataset to perform training of the trained model (which may include both new training and additional training).
[0044] The trained model generated or updated by the management device 300 is transmitted to the display terminal 100.
[0045] (b2: display terminal 100) 7 is a schematic diagram showing an example of a hardware configuration of the display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment. The display terminal 100 may be realized by using a general-purpose computer.
[0046] Referring to FIG. 7, the display terminal 100 includes, as its main hardware elements, a display 102, a processor 104, a memory 106, a network controller 108, a storage 110, a printer 120, an optical drive 122, a touch detection unit 126, a human presence sensor 128, a camera 130, and a microphone 132.
[0047] The display 102 outputs a category selection reception screen 50, an item proposal screen 52, etc. The display 102 is configured, for example, with an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display.
[0048] The processor 104 is a computing entity that executes various programs, as described below, to execute processes necessary to realize the display terminal 100. The processor 104 is configured, for example, with one or more CPUs (Central Processing Units) or GPUs (Graphics Processing Units). A CPU or GPU having multiple cores may also be used.
[0049] The memory 106 provides a storage area for temporarily storing program code, a work memory, etc., when the processor 104 executes a program. As the memory 106, for example, a volatile memory device such as a dynamic random access memory (DRAM) or a static random access memory (SRAM) may be used.
[0050] The network controller 108 transmits and receives data to and from any information processing device including the management device 300 via the local network 2. The network controller 108 may be adapted to support any communication method such as Ethernet (registered trademark), wireless LAN (Local Area Network), Bluetooth (registered trademark), etc.
[0051] The storage 110 stores an OS (Operating System) 112 executed by the processor 104, an application program 114 for implementing a functional configuration as described below, a trained model 116, and an item image 118 for generating the item proposal screen 52. As the storage 110, for example, a non-volatile memory device such as a hard disk or an SSD (Solid State Drive) may be used. Furthermore, the storage 110 may include The information storage device may store captured images of customers and collected voices uttered by the customers.
[0052] A part of the libraries and functional modules required when executing application program 114 on processor 104 may use libraries or functional modules provided as standard by OS 112. In this case, application program 114 alone does not include all of the program modules required to realize the corresponding functions, but by installing it in the execution environment of OS 112, it is possible to realize a functional configuration as described below. Therefore, even a program that does not include such a part of the libraries or functional modules may be included in the technical scope of the present invention.
[0053] The printer 120 issues a coupon with the recommended items printed on it to the customer. The printer 120 can use any printing method, such as electrophotography, inkjet, or thermal paper.
[0054] The optical drive 122 reads and writes data such as programs stored on an optical disk 124, such as a CD-ROM (Compact Disc Read Only Memory) or a DVD (Digital Versatile Disc). The optical disk 124 is an example of a non-transitory recording medium, and is distributed in a state in which an arbitrary program is stored in a non-volatile manner. The display terminal 100 according to the present embodiment can be configured by the optical drive 122 reading the program from the optical disk 124 and installing it in the storage 110. Therefore, the subject of the present invention can be the program itself installed in the storage 110 or the like, or a recording medium such as the optical disk 124 that stores a program for realizing the functions and processes according to the present embodiment.
[0055] FIG. 7 shows an optical recording medium such as an optical disk 124 as an example of a non-transient recording medium, but the present invention is not limited to this. Alternatively, a semiconductor recording medium such as a flash memory, a magnetic recording medium such as a hard disk or a storage tape, or a magneto-optical recording medium such as an MO (Magneto-Optical disk) may be used.
[0056] Alternatively, the program for realizing the display terminal 100 may not only be stored in any recording medium as described above and distributed, but may also be distributed by downloading it from a server device or the like via the Internet or an intranet.
[0057] The touch detection unit 126 is arranged in association with the display 102, and detects an input operation to the display 102. The touch detection unit 126 can employ any detection method, such as a capacitive method, a resistive film method, or an ultrasonic surface acoustic wave method.
[0058] The human presence sensor 128 detects the approach of a customer to the display terminal 100 using infrared rays or the like.
[0059] Camera 130 is a device that captures images of customers, and is arranged near the display area of display 102, etc., and configured to include in its field of view customers who are directly facing display 102. Camera 130 may capture images of its field of view continuously at a predetermined cycle, or may capture images in response to commands issued from processor 104, etc.
[0060] Microphone 132 is a device for collecting voice, and is disposed near the display area of display 102 so that it can collect voices emitted by customers. It is preferable that microphone 132 collects only the voices of customers directly facing display 102, and therefore it is preferable that microphone 132 has a sharp directivity.
[0061] 7 shows a configuration example in which a general-purpose computer (processor 104) executes an application program 114 to realize the display terminal 100, but all or part of the functions required to realize the display terminal 100 may be realized using a hard-wired circuit such as an integrated circuit. For example, they may be realized using an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array).
[0062] (b3: POS terminal 200) Fig. 8 is a schematic diagram showing an example of a hardware configuration of a POS terminal 200 constituting the clothing suggestion system 1 according to the present embodiment. Referring to Fig. 8, the POS terminal 200 includes, as main hardware elements, a display 202, a processor 204, a memory 206, a network controller 208, a storage 210, a printer 220, an optical drive 222, a touch detection unit 226, an optical reader 228, an input unit 230, and a payment processing unit 232.
[0063] The display 202 displays information necessary for the transaction process of the items, etc. The display 202 is configured, for example, with an LCD or an organic EL display.
[0064] The processor 204 is a computing entity that executes various programs, as described below, to execute processes required to realize the POS terminal 200. The processor 204 is, for example, configured with one or more CPUs. A CPU having multiple cores may also be used.
[0065] The memory 206 provides a storage area for temporarily storing program code, a work memory, etc., when the processor 204 executes a program. As the memory 206, for example, a volatile memory device such as a DRAM or an SRAM may be used.
[0066] The network controller 208 transmits and receives data to and from any information processing device including the management device 300 via the local network 2. The network controller 208 may be configured to communicate with any communication device, such as Ethernet, wireless LAN, Bluetooth, etc. The method may be adapted to correspond to the above-mentioned method.
[0067] The storage 210 stores an OS 212 executed by the processor 204, an application program 214 for implementing a functional configuration as described below, item information 216 including the price and attribute information of each item required for accounting, and sales information 218 which is information on purchased items. For example, a non-volatile memory device such as a hard disk or SSD may be used as the storage 210.
[0068] A part of the libraries and functional modules required when the application program 214 is executed by the processor 204 may be a library or a functional module provided as standard by the OS 212. In this case, the application program 214 alone does not include all of the program modules required to realize the corresponding functions, but by installing it in the execution environment of the OS 212, it is possible to realize a functional configuration as described below. Therefore, even a program that does not include such a part of the libraries or functional modules may be included in the technical scope of the present invention.
[0069] The printer 220 issues receipts on which the results of the transaction are printed, etc. The printer 220 can use any printing method, such as electrophotography, inkjet, or thermal paper.
[0070] The optical drive 222 reads out information such as a program stored in an optical disk 224 such as a CD-ROM or a DVD. The optical disk 224 is an example of a non-transient recording medium, and is distributed in a state in which an arbitrary program is stored in a non-volatile manner. The POS terminal 200 according to this embodiment can be configured by the optical drive 222 reading out the program from the optical disk 224 and installing it in the storage 210. Therefore, the subject of the present invention can be the program itself installed in the storage 210 or the like, or a recording medium such as the optical disk 224 that stores a program for realizing the functions and processes according to this embodiment.
[0071] FIG. 8 shows an optical recording medium such as an optical disk 224 as an example of a non-transient recording medium, but is not limited to this. Semiconductor recording media such as flash memory, magnetic recording media such as hard disks or storage tapes, and magneto-optical recording media such as MO may also be used.
[0072] Alternatively, the program for realizing the POS terminal 200 may not only be stored in any recording medium as described above and distributed, but may also be distributed by downloading it from a server device or the like via the Internet or an intranet.
[0073] The touch detection unit 226 is arranged in association with the display 202, and detects an input operation to the display 202. The touch detection unit 226 can employ any detection method, such as a capacitive method, a resistive film method, or an ultrasonic surface acoustic wave method.
[0074] The optical reader 228 optically reads information on an item tag attached to an item, a QR code included in a coupon, etc. The optical reader 228 can employ any detection method such as a laser scanning method or an image sensing method.
[0075] The input unit 230 accepts input operations of the amount, the type, etc. As the input unit 230, for example, a register key, a keyboard, a mouse, a touch panel, a pen, etc. may be used.
[0076] The payment processor 232 includes mechanisms necessary for cash payments as well as mechanisms necessary for electronic payments such as credit cards. More specifically, with respect to cash payments, the payment processor 232: It includes a cash storage section for storing bills and coins, a sales management section for managing sales amounts, etc. The settlement processing section 232 includes a mechanism for reading information stored in a credit card and exchanging settlement information with a settlement center or the like with regard to electronic settlement.
[0077] 8 shows an example of a configuration in which a general-purpose computer (processor 204) executes an application program 214 to realize the POS terminal 200, but all or part of the functions required to realize the POS terminal 200 may be realized using a hardwired circuit such as an integrated circuit. For example, it may be realized using an ASIC, an FPGA, or the like.
[0078] (b4:Management device 300) Fig. 9 is a schematic diagram showing an example of a hardware configuration of a management device 300 constituting the clothing suggestion system 1 according to the present embodiment. Referring to Fig. 9, the management device 300 includes, as main hardware elements, a display 302, a processor 304, a memory 306, a network controller 308, a storage 310, and an input unit 330.
[0079] The display 302 displays information necessary for processing in the management device 300. The display 302 is configured with, for example, an LCD or an organic EL display.
[0080] The processor 304 is a computing entity that executes various programs as described below to execute processes necessary for implementing the management device 300. The processor 304 is, for example, configured with one or more CPUs or GPUs. A CPU or GPU having multiple cores may also be used. In the management device 300, it is preferable to employ a GPU or the like that is suitable for learning processing for generating a trained model.
[0081] The memory 306 provides a storage area for temporarily storing program code, a work memory, etc., when the processor 304 executes a program. As the memory 306, for example, a volatile memory device such as a DRAM or an SRAM may be used.
[0082] The network controller 308 transmits and receives data to and from any information processing device including the display terminal 100 and the POS terminal 200 via the local network 2. The network controller 308 may be compatible with any communication method, such as Ethernet, wireless LAN, or Bluetooth.
[0083] The storage 310 stores an OS 312 executed by the processor 304, an application program 314 for realizing a functional configuration as described below, a preprocessing program 316 for generating a training dataset 324 from image / audio information 320 and sales information 322, and a training program 318 for generating a trained model 326 using the training dataset 324.
[0084] The image / audio information 320 is made up of the captured image 136 and collected audio 138 acquired from the display terminal 100. The sales information 322 is made up of the sales information 218 acquired from the POS terminal 200. The process of acquiring the image / audio information 320 and the sales information 322 will be described in detail later.
[0085] The learning dataset 324 is a training dataset in which sales information 322 is added as a label (or tag) to image / audio information 320. The trained model 326 is an estimated model obtained by executing a learning process using the learning dataset 324.
[0086] The storage 310 may be, for example, a non-volatile memory device such as a hard disk or SSD.
[0087] A part of the libraries and functional modules required when the application program 314, the pre-processing program 316, and the learning program 318 are executed by the processor 304 may be a library or a functional module provided as standard by the OS 312. In this case, the application program 314, the pre-processing program 316, and the learning program 318 do not each include all of the program modules required to realize the corresponding function individually, but by installing them in the execution environment of the OS 312, it is possible to realize a functional configuration as described below. Therefore, even a program that does not include such a part of the libraries or functional modules may be included in the technical scope of the present invention.
[0088] The application program 314, the preprocessing program 316, and the learning program 318 may be stored and distributed in a non-transient recording medium, such as an optical recording medium such as an optical disk, a semiconductor recording medium such as a flash memory, a magnetic recording medium such as a hard disk or a storage tape, and a magneto-optical recording medium such as an MO, and may be installed in the storage 310. Therefore, the subject of the present invention may be the program itself installed in the storage 310 or the like, or a recording medium storing a program for realizing the functions and processing according to the present embodiment.
[0089] Alternatively, the program for realizing the management device 300 may not only be stored in any recording medium as described above and distributed, but may also be distributed by downloading it from a server device or the like via the Internet or an intranet.
[0090] The input unit 330 accepts various input operations. As the input unit 330, for example, a keyboard, a mouse, a touch panel, a pen, etc. may be used.
[0091] 9 shows an example of a configuration in which a general-purpose computer (processor 304) executes an application program 314, a pre-processing program 316, and a learning program 318 to realize the management device 300, but all or part of the functions required to realize the management device 300 may be realized using a hardwired circuit such as an integrated circuit. For example, it may be realized using an ASIC, an FPGA, or the like.
[0092] (b5: Integrated configuration / cloud configuration) As a typical example, Figures 6 to 9 show a configuration in which the display terminal 100, the POS terminal 200, and the management device 300 each have a processor to realize the functions for which they are responsible. However, the present invention is not limited to this, and an integrated configuration in which the functions required to realize the clothing suggestion system 1 are realized by fewer computing entities may also be adopted.
[0093] As an example of such an integrated configuration, the functions performed by the display terminal 100 and the POS terminal 200 may be realized in the management device 300, and the display terminal 100 and the POS terminal 200 may provide only a user interface like a so-called thin client.
[0094] Furthermore, the management device 300 may also be realized by a plurality of computers connected via a computer network cooperating explicitly or implicitly. When a plurality of computers cooperate, some of the computers may be unspecified computers on the network, called so-called cloud computers.
[0095] A person skilled in the art will be able to realize the clothing proposal system 1 according to this embodiment by appropriately using the technology corresponding to the era in which the present invention is implemented.
[0096] <C. Functions and Processes of the Display Terminal 100> Next, the functions and processes of the display terminal 100 constituting the clothing proposal system 1 according to this embodiment will be described. In the clothing proposal system 1, the display terminal 100 is responsible for the operation phase of proposing clothing using a learned model (estimation model), and also for a part of the learning phase for constructing the learned model.
[0097] (c1: Functional Configuration of the Display Terminal 100) FIG. 10 is a schematic diagram showing an example of the functional configuration of the display terminal 100 constituting the clothing proposal system 1 according to this embodiment. Each function shown in FIG. 10 may typically be realized by the processor 104 of the display terminal 100 executing the OS 112 and the application program 114 (both shown in FIG. 7).
[0098] Referring to FIG. 10, the display terminal 100 has, as a functional configuration, a proposed item estimation function 140, a display control function 150, a coupon issuance control function 160, and an image and audio storage function 170.
[0099] The suggested item estimation function 140 accepts as input a captured image 136 obtained by capturing an image of a customer using a camera 130 and collected audio 138 obtained by collecting audio emitted by the customer using a microphone 132, and inputs these into the trained model 116 to output an estimation result.
[0100] The display control function 150 receives the estimation result from the suggested item estimation function 140 and generates a screen that suggests clothing according to the tastes of the customer.
[0101] The coupon issuance control function 160 receives information on the item that the display control function 150 has proposed to the customer, generates a coupon ID 166, and issues a coupon 10 on which the proposed item and the coupon ID 166 are printed.
[0102] The image and audio saving function 170 assigns the coupon ID 166 generated by the coupon issuance control function 160 to the captured image 136 and collected audio 138 received as input by the suggested item estimation function 140 and saves them. The captured image 136 and collected audio 138 (to which the coupon ID 166 has been assigned) saved by the image and audio saving function 170 are transmitted to the management device 300 and used in a learning process for generating a trained model, as described below.
[0103] (c2: Suggested item estimation function 140) Next, the suggested item estimation function 140 of the display terminal 100 shown in FIG. 10 will be described in detail.
[0104] 11 is a diagram for explaining the processing contents of the suggested item estimation function 140 of the display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment. With reference to FIG. 11, the display terminal 100 includes, as the suggested item estimation function 140, an area identification module 141, size adjustment modules 142 and 143, a section identification module 144, and a resampling module 145.
[0105] The area specifying module 141 analyzes the subject (customer) included in the captured image 136 to specify a face area and a body area. That is, the area specifying module 141 specifies a face area representing the customer's face and a body area representing the customer's body in an image obtained by capturing an image of the customer with the camera 130. The area specifying module 141 extracts a face area partial image 147 and a body area partial image 148 corresponding to the specified face area and body area from the captured image 136 and outputs them. do.
[0106] Typically, the region specifying module 141 specifies the face region and the body region by extracting facial features such as the eyes and nose, and skeletal features such as the hands and feet. In this case, the region specifying module 141 may specify a part representing the clothing worn by the customer as the body region.
[0107] Fig. 12 is a diagram for explaining the process of region identification by region identification module 141 shown in Fig. 11. Referring to Fig. 12, region identification module 141 extracts a region including the face of the customer as face region partial image 147, and extracts a region below the face of the customer as body region partial image 148.
[0108] The face area partial image 147 is considered to include attribute information such as the gender and age of the customer, and the body area partial image 148 is considered to include information regarding the customer's current clothing (i.e., information indicating clothing preferences).
[0109] 11 again, the face region partial image 147 extracted by the region specifying module 141 from the captured image 136 is output to the size adjustment module 142. Similarly, the body region partial image 148 extracted by the region specifying module 141 from the captured image 136 is output to the size adjustment module 143.
[0110] In the size adjustment modules 142 and 143, the face region partial image 147 and the body region partial image 148 are converted into features (feature vectors) having predetermined dimensions and are provided to the estimation model 1400. Here, since the image sizes of the face region partial image 147 and the body region partial image 148 extracted by the region identification module 141 may vary, the size adjustment modules 142 and 143 standardize the image sizes.
[0111] More specifically, the size adjustment module 142 adjusts the face area partial image 147 from the area identification module 141 to an image with a predetermined number of pixels, and then inputs the pixel values of each pixel constituting the adjusted image to the estimation model 1400 as face area features 1410.
[0112] Similarly, the size adjustment module 143 adjusts the body region partial image 148 from the region identification module 141 to an image with a predetermined number of pixels, and inputs the pixel values of each pixel constituting the adjusted image to the estimation model 1400 as body region features 1420.
[0113] In this way, the size adjustment modules 142, 143 extract face region features 1410 (first features) from the face region partial image 147 (face region of the image) and extract body region features 1420 (second features) from the body region partial image 148 (body region of the image).
[0114] The section identification module 144 identifies a section of the voice uttered by the customer included in the collected voice 138, and extracts and outputs the specific section voice 149. Typically, the section identification module 144 extracts the specific section voice 149 by analyzing the temporal change in the voice indicated by the collected voice 138 and identifying a section where the amplitude or frequency has changed with respect to the noise components around the display terminal 100.
[0115] Fig. 13 is a diagram for explaining the process of section identification by the section identification module 144 shown in Fig. 11. Referring to Fig. 13, the section identification module 144 identifies a section showing a significant change with respect to the preceding and following temporal changes among the temporal changes of the voice shown in the collected voice 138 as a speech section by a customer, and extracts it as a specific section voice 149.
[0116] The specific section voice 149 is a voice in which the customer speaks the desired category, and therefore includes information for identifying the desired category. Furthermore, the specific section voice 149 is considered to include information indicating the customer's current feeling.
[0117] In this manner, the section identification module 144 and the resampling module 145 extract the audio feature 1430 (third feature) from the part of the audio collected by the microphone 132 that corresponds to the customer's speech.
[0118] 11 again, the specific section voice 149 extracted by the resampling module 145 from the collected voice 138 is output to the resampling module 145. In the resampling module 145, the specific section voice 149 is converted into a feature amount (feature amount vector) having a predetermined dimension and provided to the estimation model 1400. Here, since the duration of the voice of the specific section voice 149 identified by the section identification module 144 may vary, the resampling module 145 normalizes the number of voice samples.
[0119] More specifically, the resampling module 145 samples the time waveform of the audio represented by the specific section audio 149 from the section identification module 144 at a predetermined number of samples, and inputs the amplitude value at each sampling point into the estimation model 1400 as an audio feature 1430.
[0120] The estimation model 1400 is constructed based on the trained model 116 that specifies the network structure and corresponding parameters. Face region features 1410, body region features 1420, and voice features 1430 are input to the estimation model 1400, whereby a calculation process defined by the estimation model 1400 is executed, and a score for each item is calculated as the estimation result 1450. Here, the score for each item is a value indicating the respective possibility that each clothing item is a clothing item to be proposed.
[0121] The estimation model 1400 is generated by a learning process using a learning dataset as described below. As described below, the learning dataset includes multiple pieces of learning data in which images of other customers and voices uttered by the other customers are labeled with clothing items purchased by the other customers.
[0122] In this way, estimation model 1400, which is a trained estimation model, receives face area features 1410 (first features), body area features 1420 (second features), and voice features 1430 (third features), and outputs the likelihood (score) that each of the multiple clothing items is a clothing item to be suggested as estimation result 1450.
[0123] (c3: Estimated model 1400) FIG. 14 is a schematic diagram illustrating an example of a network configuration of the estimation model 1400 illustrated in FIG. 11. Referring to FIG. 14, the estimation model 1400 is classified as a DNN (Deep Neural Network). The estimation model 1400 includes preprocessing networks 1460, 1470, and 1480 classified as a CNN (Convolutional Neural Network), an intermediate layer 1490, an activation function 1492 corresponding to an output layer, and a Softmax function 1494.
[0124] The pre-processing networks 1460, 1470, and 1480 are expected to function as a kind of filter for extracting effective features for calculating the estimation result 1450 from the face region features 1410, body region features 1420, and voice features 1430, which have relatively large degrees. Each of the pre-processing networks 1460, 1470, and 1480 has a configuration in which convolutional layers (CONV) and pooling layers (Pooling) are alternately arranged. The number of convolutional layers and pooling layers does not have to be the same. Also, on the output side of the convolutional layers, An activation function such as ReLU (rectified linear unit) is placed.
[0125] More specifically, the pre-processing network 1460 processes the face region feature quantity 1410 (x 11 ,x 12 ,···,x 1r ) and outputs internal features that indicate attribute information such as the gender and age of the customer. 21 ,x 22 ,···,x 2s ) and outputs internal features that indicate information about the customer's current attire (i.e., information indicating the customer's clothing preferences). 31 ,x 32 ,···,x 3t ) and outputs internal features that indicate information for identifying the category and information indicating the customer's current feeling.
[0126] The intermediate layer 1490 is made up of a fully connected network having a predetermined number of layers, and sequentially combines the outputs from each of the pre-processing networks 1460, 1470, and 1480, node by node, using weights and biases determined for each node.
[0127] An activation function 1492 such as ReLU is placed on the output side of the intermediate layer 1490, and finally, the estimation result 1450 (y 1 ,y 2 ,···,y N ) is output.
[0128] In the learning phase described below, the parameters of each element that constitutes the network of the estimation model 1400 are optimized.
[0129] (c4: Display control function 150 and coupon issuance control function 160) Next, the display control function 150 and the coupon issue control function 160 of the display terminal 100 shown in FIG. 10 will be described in detail.
[0130] 15 is a diagram for explaining the processing contents of the display control function 150 and the coupon issuance control function 160 of the display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment. With reference to FIG. 15, the display terminal 100 includes a display control module 152 as the display control function 150.
[0131] The display control module 152 receives the estimation result 1450 calculated by the suggested item estimation function 140, and generates the item proposal screen 52 using the item images 118 corresponding to the items having the top scores in the estimation result 1450. The display control module 152 outputs the generated item proposal screen 52 to the display 102. That is, the display 102 displays clothing items appropriate for the customer based on the estimation result 1450.
[0132] The item image 118 includes an image of each item associated with the identification information of the item. The display control module 152 extracts necessary images from the images included in the item image 118 based on the estimation result 1450.
[0133] Furthermore, display terminal 100 includes, as coupon issuance control function 160, coupon issuance control module 162 and coupon ID generation module 164. Coupon issuance control module 162 accepts suggested items from display control module 152 and coupon ID 166 from coupon ID generation module 164, and issues coupon 10 on which the information is printed, from a printer.
[0134] The coupon ID generation module 164 generates a coupon ID 166, which is unique identification information, by any method. The coupon ID 166 may be displayed as a coupon in the form of a QR code or the like. 10. In this case, the coupon ID generation module 164 may randomly generate a predetermined number of character strings and generate a QR code corresponding to the generated character strings. As will be described later, the coupon ID 166 is used as a key for generating a learning data set 324.
[0135] (c5: Image and audio storage function 170) Next, the image / audio storage function 170 of the display terminal 100 shown in FIG. 10 will be described in detail.
[0136] Fig. 16 is a diagram for explaining the processing contents of the image and audio saving function 170 of the display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment. Referring to Fig. 16, the display terminal 100 includes, as the image and audio saving function 170, an association module 172 and an image and audio storage unit 174.
[0137] In response to the coupon issuance control module 162 (see FIG. 15) issuing the coupon 10, the association module 172 receives the coupon ID 166 attached to the issued coupon 10, and associates the received coupon ID 166 with the captured image 136 and collected voice 138 used to issue the coupon 10. The association module 172 stores the associated coupon ID 166, the captured image 136, and the collected voice 138 in the image and voice storage unit 174 as a single unit.
[0138] The image / audio storage unit 174 is realized by using at least a part of the storage area provided by the memory 106 or the storage 110 (see FIG. 7 for both). The image / audio storage unit 174 stores data in units of a data set, each data set being made up of the coupon ID 166, the captured image 136, and the collected audio 138.
[0139] (c6: Processing procedure) Next, an item estimation process executed in the display terminal 100 constituting the clothing suggestion system 1 will be described.
[0140] Fig. 17 is a flowchart showing the procedure of an item estimation process in the display terminal 100 constituting the clothing suggestion system 1 according to the present embodiment. Typically, each step shown in Fig. 17 may be realized by the processor 104 of the display terminal 100 executing the OS 112 and the application program 114 (both of which refer to Fig. 7).
[0141] 17, first, display terminal 100 determines whether a visitor has been detected (step S100). In step S100, typically, whether a visitor is present is determined based on a detection result from human presence sensor 128 (see FIG. 7). If a visitor is not detected (NO in step S100), the process of step S100 is repeated.
[0142] When a visitor is detected (YES in step S100), the display terminal 100 displays a category selection reception screen (see FIG. 2) on the display 102 (step S102).
[0143] Next, the display terminal 100 captures an image 136 by capturing an image of a customer facing the display terminal 100 with the camera 130 (step S104). At the same time, the display terminal 100 starts collecting voice (step S106). Then, the display terminal 100 judges whether or not it has detected the customer's speech based on the collected voice (step S108). In step S108, as shown in FIG. 13, it is judged whether or not a section showing a significant change in the voice over time (a speech section) has appeared. If the customer's speech is not detected, If the answer is NO in step S108, the processes of steps S106 and S108 are repeated. In this manner, the display terminal 100 executes the process of acquiring the captured image 136 obtained by capturing an image of the customer and the collected voice 138 including the speech of the customer.
[0144] When speech of a customer is detected (YES in step S108), the display terminal 100 extracts the face area partial image 147 and the body area partial image 148 from the captured image 136 (step S110). Then, the display terminal 100 adjusts the size of the extracted face area partial image 147 and the body area partial image 148 to extract the face area feature amount 1410 and the body area feature amount 1420 (step S112).
[0145] Additionally, the display terminal 100 extracts the speech section included in the collected voice 138 as a specific section voice 149 (step S114). Then, the display terminal 100 resamples the extracted specific section voice 149 to extract the voice feature quantity 1430 (step S116).
[0146] The display terminal 100 inputs the face region feature 1410 and the body region feature 1420 generated in step S112, and the voice feature 1430 generated in step S116 to the estimation model 1400 to generate an estimation result 1450 (step S118).
[0147] In this manner, the display terminal 100 inputs a plurality of feature amounts extracted from the captured image 136 and the collected voice 138 into the trained estimation model 1400 and executes a process of generating a proposal of clothing items suited to the customer.
[0148] The display terminal 100 generates and outputs the item proposal screen 52 based on the items with the highest scores in the estimation result 1450 generated in step S118 (step S120).
[0149] The display terminal 100 determines whether or not the coupon issue button has been pressed (step S122). If the coupon issue button has not been pressed (NO in step S122), the process of step S122 is repeated.
[0150] When the coupon issue button is pressed (YES in step S122), the display terminal 100 generates a coupon ID 166 and issues a coupon 10 on which the suggested item and the coupon ID 166 are printed (step S124). That is, the display terminal 100 executes a process to generate the coupon ID 166, which is identification information, and executes a process to issue the coupon 10, which is a medium. The coupon 10 includes the generated clothing item suggestion and the generated coupon ID 166 (identification information), and also displays the details of a discount to encourage the purchase of the clothing item.
[0151] Finally, the display terminal 100 stores the captured image 136 and the collected voice 138 in association with the coupon ID 166 (step S126). That is, the display terminal 100 executes a process of associating the generated coupon ID 166 (identification information) with the captured image 136 and the collected voice 138.
[0152] This completes the item estimation process for one store customer. (c7:Short summary) The display terminal 100 according to the present embodiment can propose clothing according to the preferences of a customer by providing an estimation model 1400 with a face region feature amount 1410 generated from a partial image 147 of the customer's face region, a body region feature amount 1420 generated from a partial image 148 of the customer's body region, and an audio feature amount 1430 generated from the customer's specific interval audio 149.
[0153] In addition, the display terminal 100 according to the present embodiment can provide a platform for generating a learning dataset used in a learning phase as described later by issuing a coupon 10 including a coupon ID 166.
[0154] <Functional Configuration of D.POS Terminal 200> Next, the functions and processes of the POS terminal 200 that constitutes the clothing proposal system 1 according to the present embodiment will be described. In the clothing proposal system 1, the POS terminal 200 mainly takes charge of a part of the learning phase for constructing a learned model.
[0155] (d1: Functional Configuration of POS Terminal 200) FIG. 18 is a schematic diagram showing an example of the functional configuration of the POS terminal 200 that constitutes the clothing proposal system 1 according to the present embodiment. Each function shown in FIG. 18 may typically be realized by the processor 204 of the POS terminal 200 executing the OS 212 and the application program 214 (both shown in FIG. 8).
[0156] Referring to FIG. 18, the POS terminal 200 has, as a functional configuration, a settlement function 240 and a sales information storage function 250.
[0157] The settlement function 240 is in charge of the settlement process when a customer purchases an item. More specifically, the settlement function 240 calculates the amount, discount amount, payment amount, etc. of the item to be purchased based on the information of the item tag 20 attached to the item to be purchased and the coupon ID 166 read from the coupon, and executes the settlement process. The settlement function 240 outputs sales information 218 indicating the item for which the settlement has been made.
[0158] The sales information storage function 250 assigns the coupon ID 166 read from the coupon 10 to the sales information 218 output from the accounting function 240 and stores the same. The sales information 218 (with the coupon ID 166 assigned) stored by the sales information storage function 250 is transmitted to the management device 300 and used in the learning process for generating a trained model, as described below.
[0159] (d2: Sales information storage function 250) Next, the sales information storage function 250 of the POS terminal 200 shown in FIG. 18 will be described in detail.
[0160] 19 is a diagram for explaining the processing contents of the sales information storage function 250 of the POS terminal 200 constituting the clothing suggestion system 1 according to the present embodiment. Referring to FIG. 19, the POS terminal 200 includes, as the sales information storage function 250, an association module 252 and a sales information storage unit 254.
[0161] In response to the sales information 218 and coupon ID 166 being input from the accounting function 240 (see FIG. 18), the association module 252 accepts the coupon ID 166 that was attached to the coupon 10 used when purchasing the item related to the sales information 218, and associates it with the sales information 218. The association module 252 stores the associated coupon ID 166 and sales information 218 together in the sales information storage unit 254.
[0162] The sales information storage unit 254 is realized by using at least a part of the storage area provided by the memory 106 or the storage 110 (see FIG. 7 for both). The sales information storage unit 254 stores data in units of a data set consisting of the coupon ID 166 and the sales information 218.
[0163] The sales information 218 typically stores the number of items sold for each item type (item 1, item 2, . . . , item N).
[0164] (d3: Processing procedure) Next, a sales management process executed in the POS terminal 200 constituting the clothing suggestion system 1 will be described.
[0165] Fig. 20 is a flowchart showing the procedure of sales management processing in the POS terminal 200 constituting the clothing suggestion system 1 according to the present embodiment. Typically, each step shown in Fig. 20 may be realized by the processor 204 of the POS terminal 200 executing the OS 212 and the application program 214 (both of which refer to Fig. 8).
[0166] 20, first, the POS terminal 200 determines whether the coupon 10 has been read by the optical reader 228 (see FIG. 8) (step S200). If the coupon 10 has been read (YES in step S200), the POS terminal 200 acquires the coupon ID 166 of the read coupon 10 (step S202). On the other hand, if the coupon 10 has not been read (NO in step S200), the process of step S202 is skipped.
[0167] Next, the POS terminal 200 determines whether the item tag 20 attached to the item to be purchased has been read by the optical reader 228 (see FIG. 8) (step S204). If the item tag 20 has been read (YES in step S204), the POS terminal 200 adds the item information of the read item tag 20 to the sales information 218 (step S206).
[0168] Then, the POS terminal 200 judges whether or not an instruction to end reading of the item tag has been given (step S208). If an instruction to end reading of the item tag has not been given (NO in step S208), the process from step S204 onwards is repeated.
[0169] When an instruction to finish reading the item tag is given (YES in step S208), the POS terminal 200 calculates the payment amount based on the presence or absence of the coupon 10 and the current sales information 218 (step S210). Then, the POS terminal 200 executes a settlement process for the payment amount calculated in step S210 (step S212).
[0170] Subsequently, the POS terminal 200 determines whether the coupon ID 166 has been acquired (step S214). That is, in step S200, it is determined whether the coupon 10 has been read.
[0171] If the coupon ID 166 has been acquired (YES in step S214), the POS terminal 200 saves the sales information 218 in association with the coupon ID 166 (step S216). On the other hand, if the coupon ID 166 has not been acquired (NO in step S214), the process of step S216 is skipped. Thus, the sales management process for one customer is completed.
[0172] (d4: small parenthesis) The POS terminal 200 according to the present embodiment executes a settlement process for the items purchased by the customer, reads the coupon ID 166 assigned to the coupon 10 presented at that time, and saves it in association with the purchased items. The saved information on the purchased items (sales information 218) is used to generate a learning dataset used in a learning phase as described later.
[0173] <E. Overview of the learning phase> Next, an overview of the learning phase in the clothing proposal system 1 according to the present embodiment will be described.
[0174] The clothing recommendation system 1 according to this embodiment generates a training data set 324 by matching the captured images 136 and collected voice 138 stored in the display terminal 100 with the sales information 218 stored in the POS terminal 200 for the same store customer, and uses the generated training data set 324 to train an estimation model.
[0175] Fig. 21 is a diagram for explaining an outline of the learning phase in the clothing suggestion system 1 according to the present embodiment. With reference to Fig. 21, the display terminal 100 transmits the captured image 136 and the collected voice 138 associated with the coupon ID 166, which are acquired during the execution of the item estimation process, to the management device 300 (sequence SQ1). Similarly, the POS terminal 200 transmits the sales information 218 associated with the coupon ID 166, which are acquired during the execution of the sales management process, to the management device 300 (sequence SQ2).
[0176] The management device 300 generates a learning data set 324 by associating the captured image 136 and collected voice 138 transmitted from the display terminal 100 with the sales information 218 transmitted from the POS terminal 200 using the coupon ID 166 as a key (sequence SQ3). In other words, sequence SQ3 corresponds to a pre-processing for generating the learning data set 324.
[0177] The management device 300 uses the generated training dataset 324 to train or additionally train the estimation model, thereby generating a trained model 326 (sequence SQ4). Then, the management device 300 transmits the generated trained model 326 to each of the display terminals 100 (sequence SQ5). The display terminals 100 store the trained model 326 transmitted from the management device 300 as the trained model 116. That is, the trained model 116 of the display terminal 100 is set or updated.
[0178] As shown in FIG. 21, in the clothing recommendation system 1 according to the present embodiment, since the coupon ID 166 assigned to the coupon 10 can be used to combine the information obtained by each of the display terminal 100 and the POS terminal 200, a learning dataset 324 for improving the estimation accuracy of the estimation model can be easily generated without imposing a burden on the customer.
[0179] <F. Functional Configuration of Management Device 300> Next, the functions and processes of the management device 300 that constitutes the clothing recommendation system 1 according to the present embodiment will be described. In the clothing recommendation system 1, the management device 300 will mainly be responsible for a part of the learning phase for constructing a learned model.
[0180] (f1: Functional Configuration of Management Device 300) FIG. 22 is a schematic diagram showing an example of the functional configuration of the management device 300 that constitutes the clothing recommendation system 1 according to the present embodiment. Each function shown in FIG. 22 may typically be realized by the processor 304 of the management device 300 executing the OS 312, the application program 314, the preprocessing program 316, and the learning program 318 (all shown in FIG. 9).
[0181] Referring to FIG. 22, the management device 300 has, as a functional configuration, an imaging image / collected voice / sales information acquisition function 340, a learning dataset generation function 350, and a learning function 360.
[0182] The imaging image / collected voice / sales information acquisition function 340 acquires the imaging image 136 and the collected voice 138 associated with the coupon ID 166 stored in the display terminal 100, and the sales information 218 associated with the coupon ID 166 stored in the POS terminal 200. These data will be used as a learning dataset. That is, the imaging image / collected voice / sales information acquisition function 340 of the management device 300 corresponds to a configuration for acquiring a learning dataset.
[0183] As a method of acquiring data from the display terminal 100 and the POS terminal 200, for example, some command may be given to the display terminal 100 and the POS terminal 200 so that the display terminal 100 and the POS terminal 200 transmit data, or the management device 300 may access the display terminal 100 and the POS terminal 200 to acquire data. Alternatively, the display terminal 100 and the POS terminal 200 may transmit data to the management device 300 at predetermined intervals.
[0184] The learning dataset generation function 350 generates a learning dataset 324 from the captured images 136 and collected audio 138 associated with the coupon ID 166 obtained from the display terminal 100, and the sales information 218 associated with the coupon ID 166 obtained from the POS terminal 200.
[0185] The learning function 360 generates a trained model 326 by learning an estimation model using the training dataset 324 generated by the training dataset generation function 350. The generated trained model 326 is transmitted to the display terminal 100.
[0186] (f2: Learning dataset generation function 350) Next, the learning data set generation function 350 of the management device 300 shown in FIG. 22 will be described in detail.
[0187] Fig. 23 is a diagram for explaining the processing contents of the learning data set generating function 350 of the management device 300 constituting the clothing suggestion system 1 according to the present embodiment. With reference to Fig. 23, in relation to the learning data set generating function 350, the management device 300 compares the captured image 136 and the collected voice 138 associated with the coupon ID 166 acquired from the display terminal 100 with the sales information 218 associated with the coupon ID 166 acquired from the POS terminal 200, and associates data having the same coupon ID 166.
[0188] FIG. 23 shows, as an example, a data set of a captured image 136 and a collected voice 138 to which "01", "02", and "03" are respectively assigned as the coupon ID 166, and sales information 218 to which "02", "03", and "08" are respectively assigned as the coupon ID 166. Of these, for the data to which "02" and "03" are assigned as the coupon ID 166, the captured image 136, the collected voice 138, and the sales information 218 are all complete. These three types of data can be determined as learning data (relationship between input information and the correct value of the estimation result). A learning data set 324 can be generated by generating learning data for each of the multiple coupon IDs 166.
[0189] At this time, the sales information 218 is used as a label (tag) for adaptation to a learning process to be described later. That is, in the learning dataset 324, the captured image 136 (learning image) obtained by capturing an image of a given customer and the collected voice 138 (learning voice) uttered by the given customer are labeled with the clothing item (sales information 218) purchased by the given customer.
[0190] (f3: learning function 360) Next, the learning function 360 of the management device 300 shown in FIG. 22 will be described in detail.
[0191] Fig. 24 is a diagram for explaining the processing contents of the learning function 360 of the management device 300 constituting the clothing suggestion system 1 according to the present embodiment. With reference to Fig. 24, the management device 300 includes, as the learning function 360, an area specifying module 141, size adjustment modules 142 and 143, a section specifying module 144, and a resampling module 145. These modules are substantially the same as the modules that the display terminal 100 has as the suggested item estimation function 140. Therefore, detailed description of these modules will not be repeated.
[0192] Furthermore, the management device 300 includes a parameter optimization module 362 as the learning function 360. The parameter optimization module 362 generates the trained model 326 by optimizing model parameters 364 for defining the estimation model 1400.
[0193] The parameter optimization module 362 optimizes the model parameters 364 using each set (learning data) of the captured images 136 , the collected voices 138 , and the sales information 218 included in the learning dataset 324 .
[0194] More specifically, the parameter optimization module 362 generates face area features 1410, body area features 1420, and voice features 1430 from each pair of captured images 136 and collected voices 138 included in the learning dataset 324, and inputs them to the estimation model 1400 to calculate the estimation result 1450. Then, the parameter optimization module 362 calculates an error by comparing the estimation result 1450 output from the estimation model 1400 with the corresponding sales information 218 (correct label), and optimizes (adjusts) the values of the model parameters 364 according to the calculated error.
[0195] That is, the parameter optimization module 362 corresponds to a learning unit, and optimizes the estimation model 1400 so that the estimation result 1450 outputted by inputting the face area feature 1410 (first feature), the body area feature 1420 (second feature), and the voice feature 1430 (third feature) extracted from the learning data (the captured image 136 and the collected voice 138 are labeled with the sales information 218) into the estimation model 1400 approaches the purchase history (sales information 218) of the clothing item labeled in the learning data. In other words, the parameter optimization module 362 adjusts the model parameters 364 so that the estimation result 1450 calculated when the features are extracted from the captured image 136 and the collected voice 138 included in the learning data and inputted into the estimation model 1400 matches the corresponding sales information 218.
[0196] Using a similar procedure, the model parameters 364 of the estimation model 1400 are repeatedly optimized based on each learning data (captured images 136, collected audio 138, and sales information 218) contained in the learning dataset 324, thereby generating a learned model 326.
[0197] The parameter optimization module 362 can use any optimization algorithm to optimize the values of the model parameters 364. More specifically, the optimization algorithm can be, for example, SGD (Stochastic Gradient Descent). ), Momentum SGD (SGD with inertia term), AdaGrad, RMSprop, AdaDelta, Adam (Adaptive moment estimation), and other gradient methods can be used.
[0198] Each element of the estimation result 1450 output from the estimation model 1400 is normalized as a probability. When outputting as a rate, it is preferable to also normalize the number of sales (see FIG. 19) for each item type (item 1, item 2, . . . , item N) included in the sales information 218.
[0199] The estimation model 1400 in which the model parameters 364 have been optimized by the parameter optimization module 362 corresponds to the trained model 326 and is transmitted to the display terminal 100.
[0200] (f4: Processing procedure) Next, a learning process executed in the management device 300 constituting the clothing suggestion system 1 will be described.
[0201] Fig. 25 is a flowchart showing the procedure of learning processing in the management device 300 constituting the clothing recommendation system 1 according to the present embodiment. Typically, each step shown in Fig. 25 may be realized by the processor 304 of the management device 300 executing the OS 312, the application program 314, the pre-processing program 316, and the learning program 318 (all of which refer to Fig. 9).
[0202] 25, the management device 300 acquires the captured image 136 and the collected voice 138 to which the coupon ID 166 is assigned from the display terminal 100 (step S300). In addition, the management device 300 acquires the sales information 218 to which the coupon ID 166 is assigned from the POS terminal 200 (step S302). That is, the management device 300 executes a process of acquiring the coupon ID 166 (identification information) included in the coupon 10, which is a medium, and the clothing item (sales information 218) purchased by the customer.
[0203] Then, the management device 300 generates a learning data set 324 by associating the captured image 136 and the collected voice 138 with the sales information 218 using the coupon ID 166 as a key (step S304). That is, the management device 300 executes a process of associating the coupon ID 166 (identification information) acquired from the coupon 10 as a medium with the clothing item purchased by the customer (sales information 218), and further executes a process of associating the captured image 136 and the collected voice 138 with the sales information 218 using the coupon ID 166 as a key and storing the association as learning data to be used for training the estimation model 1400.
[0204] The management device 300 selects one data set (learning data) from the generated learning data set 324 (step S306).
[0205] The management device 300 extracts the face region partial image 147 and the body region partial image 148 from the captured image 136 of the selected data (step S308). Then, the management device 300 adjusts the size of the extracted face region partial image 147 and the body region partial image 148 to extract the face region feature amount 1410 and the body region feature amount 1420 (step S310).
[0206] In this way, the management device 300 executes a process of identifying a face region representing the customer's face and a body region representing the customer's body in the captured image 136 of each learning data. Then, the management device 300 executes a process of extracting face region feature amount 1410 (first feature amount) from the face region of the captured image 136, and extracting body region feature amount 1420 (second feature amount) from the body region of the captured image 136.
[0207] Additionally, the management device 300 extracts the speech section included in the collected voice 138 of the selected data as the specific section voice 149 (step S312). Then, the management device 300 resamples the extracted specific section voice 149 to extract the voice feature quantity 1430 (step S314). In this way, the management device 300 extracts the speech section corresponding to the customer's speech from the collected voice 138. Then, the processing is performed to extract a speech feature 1430 (third feature) from the speech of the portion that corresponds to the speech.
[0208] The management device 300 inputs the face region feature 1410 and the body region feature 1420 generated in step S310, and the voice feature 1430 generated in step S314, to the estimation model 1400 to generate an estimation result 1450 (step S316).
[0209] The management device 300 optimizes the model parameters 364 of the estimation model based on the error between the sales information 218 of the selected data and the estimation result 1450 generated in step S316 (step S318).
[0210] In this way, the management device 300 optimizes the estimation model 1400 so that the estimation result 1450 output by inputting the face region feature amount 1410 (first feature amount), the body region feature amount 1420 (second feature amount), and the voice feature amount 1430 (third feature amount) into the estimation model 1400 approaches the purchase record (sales information 218) of the clothing item labeled in the learning data.
[0211] Then, the management device 300 determines whether all of the learning data set 324 generated in step S304 has been processed (step S320). If not all of the learning data set 324 has been processed (NO in step S320), the processes from step S306 onward are repeated.
[0212] If all of the learning data set 324 has been processed (YES in step S320), the management device 300 transmits the learned model 326 defined by the current model parameters 364 to each display terminal 100 (step S322). Thus, the learning process is completed.
[0213] (f5: small parenthesis) The management device 300 according to the present embodiment can easily generate the learning data set 324 by associating the captured image 136 and the collected voice 138 acquired from the display terminal 100 with the sales information 218 acquired from the POS terminal 200 using the coupon ID 166 as a key. By using such a learning data set 324, it becomes possible to construct an estimation model or perform additional learning of the learned model 326. Thereby, the proposal accuracy of clothing can be improved.
[0214] <G. Modification Example> In the above-described embodiment, as a typical example, the clothing proposal system 1 in which the display terminal 100, the POS terminal 200, and the management device 300 are arranged in a single store 30 is illustrated. However, the present invention is not limited to this, and various modifications are possible. Hereinafter, several modification examples will be described.
[0215] (g1: Multi-store cooperation: Modification Example 1) As a modified example, the management device 300 may be configured to manage a plurality of stores.
[0216] Fig. 26 is a schematic diagram showing an example of a system configuration of a clothing suggestion system 1A according to the first modification of the present embodiment. Referring to Fig. 26, each of the stores 30A and 30B has one or more display terminals 100 and one or more POS terminals 200. Each store 30 is connected to the same management device 300 via a wide area network 4.
[0217] The management device 300 obtains necessary information (captured image 136, collected voice 138, and sales information 218) from the display terminal 100 and the POS terminal 200 of the store 30A, and obtains necessary information from the display terminal 100 and the POS terminal 200 of the store 30B. Based on the collected information, the management device 300 generates a trained model common to both stores, or a trained model for each store.
[0218] By adopting the configuration shown in FIG. 26, the number of management devices 300 can be reduced and a larger number of learning data sets can be obtained, thereby improving the estimation accuracy of the trained model.
[0219] (g2: Item proposals by category: Variation 2) Since the estimation model 1400 (see FIG. 11) according to the above-described embodiment receives as input the speech feature 1430 corresponding to one of the categories displayed on the category selection receiving screen 50, basically, an item belonging to the spoken category will have a relatively high score in the output estimation result 1450. Each of the multiple clothing items will belong to one of multiple predetermined categories (product categories).
[0220] However, when there are many items belonging to other categories that were purchased at the same time as an item belonging to the selected category, items belonging to other categories having relatively high scores may be mixed in the estimation result 1450. In such a case, items belonging to categories other than the selected category are also suggested on the item suggestion screen 52.
[0221] Fig. 27 is a diagram for explaining an item proposal screen displayed on the display terminal 100 of the clothing proposal system 1 according to the modified example 2 of the present embodiment. As shown in Fig. 27(a), when an item belonging to another category has a relatively high score in the estimation result 1450, the list display 54 of the item proposal screen 52 includes the item (reference symbol 54M) belonging to the other category.
[0222] An item proposal screen 52 may be displayed that may include items belonging to such other categories, but as shown in FIG. 27(b), items belonging to the selected category and items belonging to other categories may be proposed in different display formats.
[0223] 27(b) includes a list display 54 of items belonging to the category selected by voice by the customer 40, and a list display 55 of items belonging to categories other than the category selected by voice by the customer 40. In list display 55, a message such as "How about this one too" is also displayed, indicating that the item is suitable for suggestion based on past sales performance, even though it is in a category different from the selected category.
[0224] Fig. 28 is a diagram for explaining the processing contents of the display control function 150A and the coupon issuance control function 160 of the display terminal 100 constituting the clothing suggestion system 1 according to the modified example 2 of the present embodiment. Referring to Fig. 28, the display terminal 100 has a display control module 152A, a voice analysis module 154, and category-item correspondence information 156 as the display control function 150A.
[0225] The voice analysis module 154 performs voice analysis on the collected voice 138 uttered by the customer 40 to identify the category selected by the customer 40 through voice. That is, the voice analysis module 154 identifies the category selected by the customer from among a plurality of categories based on the voice uttered by the customer. Note that the voice analysis method used by the voice analysis module 154 can use any known algorithm. The category identified by the voice analysis module 154 is provided to the display control module 152A.
[0226] The display control module 152A displays the estimated result calculated by the suggested item estimation function 140. The display control module 152A receives the estimation result 1450 and identifies items having high scores in the estimation result 1450. The display control module 152A refers to the category-item correspondence information 156 and determines whether or not each of the items having high scores in the estimation result 1450 belongs to a category identified by the voice analysis module 154. The display control module 152A then refers to the item images 118, and adds images of items that belong to a category identified by the voice analysis module 154 to the list display 54, and adds images of items that belong to a category other than the category identified by the voice analysis module 154 to the list display 55, thereby generating an item proposal screen 52A. The generated item proposal screen 52A is displayed on the display 102.
[0227] 27(b) is provided by executing the above-described processing in the display control module 152A. That is, the display 102 displays, in different display modes, clothing items that belong to the category identified by the voice analysis module 154 and clothing items that do not belong to the identified category among the clothing items displayed based on the estimation result 1450. By adopting such a display mode, it is possible to encourage the customer 40 to purchase items other than the category selected by the customer 40.
[0228] As other processes and functions are substantially the same as those described with reference to FIG. 15, detailed description will not be repeated.
[0229] (g3: Network: Variation 3) In the above embodiment, the estimation model 1400 to which the face region feature 1410, the body region feature 1420, and the voice feature 1430 are input is exemplified, but an estimation model to which additional information can be input may also be adopted.
[0230] Fig. 29 is a diagram for explaining the processing contents of the suggested item estimation function 140 of the display terminal 100 constituting the clothing suggestion system 1 according to the third modification of the present embodiment. Fig. 29 shows an estimation model 1400A that accepts meteorological information such as weather and temperature as input feature values 1440. In this way, by adding input information, the estimation accuracy can be improved.
[0231] When a feature quantity is added to be input to the estimation model 1400A, the information included in the learning data set used to train the estimation model 1400A is also increased in accordance with the input feature quantity.
[0232] 29 shows weather information as a typical example, the additional information to be input is not limited to this, and any information that is estimated to have some relevance to the determination of the items to be suggested can be used. For example, other weather information such as wind speed and sunshine hours, time information such as date and day of the week, and information on the congestion level of the store may be used.
[0233] (g4: Network: Variation 4) In the above embodiment, the estimation model 1400 to which the face region feature 1410, the body region feature 1420, and the voice feature 1430 are input is exemplified, but an estimation model in which some information is substituted may be adopted.
[0234] Fig. 30 is a diagram for explaining the processing contents of the suggested item estimation function 140 of the display terminal 100 constituting the clothing suggestion system 1 according to the fourth modified example of the present embodiment. Fig. 30 shows an example of a configuration in which an input feature 1442 indicating a category is input instead of the voice feature 1430. The input feature 1442 may be generated by performing voice analysis on the collected voice 138 uttered by the store customer 40 to identify the category selected by the store customer 40 by voice.
[0235] Alternatively, when a customer 40 selects a category by touching a portion corresponding to the category on the category selection reception screen 50 displayed on the display terminal 100, the category selected by the touch operation may be input as the input feature 1442.
[0236] Note that, while FIG. 30 shows an example in which the input feature 1442 indicating a category is input, by adopting the configuration shown in FIG. 28 described above, the input of the input feature 1442 may also be deleted.
[0237] (g5: Item suggestions using mobile devices: Variation 5) As a modified example, item suggestions as described above may be made on a personally owned mobile terminal instead of in a physical store.
[0238] Fig. 31 is a schematic diagram showing an example of use of clothing suggestion system 1B according to variant 5 of the present embodiment. Referring to Fig. 31, by installing an application for mobile terminal 500, functions similar to those of display terminal 100 can be realized on mobile terminal 500. An Internet user can receive suggestions for clothing as described above by executing the application on mobile terminal 500, taking an image of himself / herself using a camera mounted on mobile terminal 500, and vocalizing a desired category.
[0239] As an implementation for implementing the item estimation process according to this embodiment in mobile terminal 500, any form can be adopted.
[0240] FIG. 32 is a schematic diagram showing an implementation example of a clothing suggestion system according to the fifth modification of the present embodiment.
[0241] Fig. 32(a) shows an implementation example in which item estimation processing is realized by the mobile terminal 500 alone. As shown in Fig. 32(a), an application 510 is installed on the mobile terminal 500 from the server device 400. The application 510 has a suggested item estimation function 512, a display control function 514, and a coupon issuance control function 516. The suggested item estimation function 512, the display control function 514, and the coupon issuance control function 516 execute substantially the same processing as the suggested item estimation function 140, the display control function 150, and the coupon issuance control function 160 of the display terminal 100 (all of which refer to Fig. 10).
[0242] In the implementation example shown in Figure 32(a), a trained model 518 (substantially identical to the trained model 116 placed in the display terminal 100) is placed in the mobile terminal 500, so that the item estimation process can be executed even if communication with the server device 400 cannot be performed.
[0243] Fig. 32(b) shows an implementation example in which the server device 400 and the mobile terminal 500 cooperate to realize the item estimation process. As shown in Fig. 32(b), an application 520 is installed in the mobile terminal 500 from the server device 400. The application 520 has a feature generation function 522 and a display function 524. The feature generation function 522 extracts face area features 1410 and body area features 1420 from a captured image obtained by capturing an image of the Internet user, and generates voice features 1430 from collected voice 138 uttered by the Internet user and transmits them to the server device 400.
[0244] The display function 524 outputs the display contents from the server device 400 to the display of the mobile terminal 500 .
[0245] On the other hand, the server device 400 includes a suggested item estimation function 412, a display control function 414, and The display terminal 100 has a coupon issuance control function 416. The suggested item estimation function 412 corresponds to the suggested item estimation function 140 (see FIG. 11) of the display terminal 100 excluding the function of extracting features. The display control function 414 and the coupon issuance control function 416 execute substantially the same processes as the display control function 150 and the coupon issuance control function 160 (both see FIG. 10) of the display terminal 100.
[0246] 32(b), the trained model 518 (substantially the same as the trained model 116 arranged in the display terminal 100) is arranged in the server device 400, so that the trained model 518 can be appropriately updated in the server device 400. In addition, since the mobile terminal 500 only needs to extract features, it is possible to reduce resource consumption on the mobile terminal 500 side.
[0247] Fig. 32(c) shows an implementation example in which the server device 400 and the mobile terminal 500 cooperate to realize the item estimation process. As shown in Fig. 32(c), an application 530 is installed on the mobile terminal 500 from the server device 400. The application 530 has an image and audio transmission function 532 and a display function 524. The image and audio transmission function 532 transmits to the server device 400 a captured image obtained by capturing an image of the Internet user and a collected voice 138 uttered by the Internet user.
[0248] The display function 534 outputs the display contents from the server device 400 to the display of the mobile terminal 500 .
[0249] On the other hand, the server device 400 has a suggested item estimation function 412, a display control function 414, and a coupon issuance control function 416. The suggested item estimation function 412, the display control function 414, and the coupon issuance control function 416 execute substantially the same processes as the display control function 150 and the coupon issuance control function 160 of the display terminal 100 (see FIG. 10 for all of them).
[0250] 32(c), the trained model 518 (substantially the same as the trained model 116 arranged in the display terminal 100) is arranged in the server device 400, so that the trained model 518 can be appropriately updated in the server device 400. In addition, the mobile terminal 500 only needs to transmit the captured image 136 and the collected voice 138 to the server device 400 as they are, so that the consumption of resources on the mobile terminal 500 side can be reduced.
[0251] (g6:Other) It is obvious that the present invention is not limited to the above-mentioned modifications, and various modifications can be made in accordance with the spirit of the present invention. In addition, one or more of the above-mentioned modifications can be arbitrarily combined.
[0252] <H.まとめ> According to the clothing suggestion system 1 of this embodiment, by using face area feature 1410 generated from partial face area image 147 of the store visitor, body area feature 1420 generated from partial body area image 148 of the store visitor, and voice feature 1430 generated from specific section voice 149 of the store visitor as input information, it is possible to suggest clothing according to the tastes of the store visitor with higher accuracy.
[0253] Furthermore, according to the clothing suggestion system 1 of the present embodiment, the input information (captured image 136 and collected voice 138) acquired from the customer and the item actually purchased by the customer are associated with each other using the coupon ID 166 attached to the coupon 10, thereby generating a learning data set 324. By learning a constant model, the estimation accuracy can be continuously improved and the estimation model can be adapted even when new items are added.
[0254] Furthermore, since the clothing suggestion system 1 according to the present embodiment issues the coupon 10 for discounting the payment amount, customers have an incentive to actively use the coupon 10. As a result, it is possible to increase the possibility of collecting information for generating the learning dataset 324.
[0255] The embodiments disclosed herein should be considered to be illustrative and not restrictive in all respects. The scope of the present invention is defined by the claims, not by the description of the embodiments described above, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0256] 1, 1A, 1B Clothing recommendation system, 2 Local network, 4 Wide area network, 10 Coupon, 12 Discount amount display, 14, 54, 55 List display, 16 Map, 18 Identification image, 20 Item tag, 30, 30A, 30B Store, 40 Customer, 50 Category selection reception screen, 52, 52A Item suggestion screen, 56 Coupon issue button, 100 Display terminal, 102, 202, 302 Display, 104, 204, 304 Processor, 106,206,306 Memory, 108,208,308 Network controller, 110,210,310 Storage, 112,212,312 OS, 114,214,314 Application program, 116,326,518 Trained model, 118 Item image, 120,220 Printer, 122,222 Optical drive, 124,224 Optical disk, 126,226 Touch detector, 128 Human sensor, 130 Camera, 132 Microphone, 136 Captured image, 138 Collected voice, 140,412,512 Suggested item estimation function, 141 Region identification module, 142,143 Size adjustment module, 144 Section identification module, 145 Resampling module, 147 Face region partial image, 148 Body region partial image, 149 Specific section voice, 150, 150A, 414, 514 Display control function, 152, 152A Display control module, 154 Voice analysis module, 156 Category-item correspondence information, 160, 416, 516 Coupon issuance control function, 162 Coupon issuance control module, 164 Coupon ID generation module, 166 Coupon ID, 170 Image and voice storage function, 172, 252 Correspondence module, 174 Image and voice storage unit, 200 POS terminal, 216 Item information, 218, 322 Sales information, 228 Optical reader, 230, 330 Input unit, 232 Payment processing unit, 240 Accounting function, 250 Sales information storage function, 254 Sales information storage unit, 300 Management device, 316 Pre-processing program, 318 Learning program, 320 Voice information, 324 Learning data set, 340 Sales information acquisition function, 350 learning dataset generation function, 360 learning function, 362 parameter optimization module, 364 model parameters, 400 server device, 500 mobile terminal, 510, 520, 530 application, 522 feature generation function, 524, 534 display function, 532 image and audio transmission function, 1400, 1400A estimation model, 1410 face area feature, 1420 body area feature, 1430 audio feature, 1440, 1442 input feature, 1450 estimation result, 1460, 1470, 1480 preprocessing network, 1490 intermediate layer, 1492 activation function, 1494 Softmax function.
Claims
1. An information processing device that proposes clothing items suitable for a customer from among a plurality of clothing items based on a feature amount representing the characteristics of the customer, based on an estimation result of a trained estimation model, A camera for imaging the customer; an image feature extraction unit for extracting the feature amount from an image obtained by capturing an image of the customer with the camera; The machine learning for the estimation model is based on learning data in which a clothing item purchased using a coupon among the plurality of clothing items is associated as a correct answer of the output of the estimation model for an image of a customer to whom the coupon was issued, The information processing device, wherein the estimation model receives the features and outputs, as the estimation result, a respective possibility that each of the plurality of clothing items is a clothing item to be proposed.
2. the feature amount includes a first feature amount from a face region representing a face of the customer, and a second feature amount from a body region representing a body of the customer in the image; The information processing device includes: a region specifying unit for specifying the face region and the body region in the image; The information processing device according to claim 1 , further comprising a display unit for displaying clothing items according to the customer based on the estimation result.
3. The information processing device according to claim 2 , wherein the region specifying unit specifies a portion representing clothing worn by the customer as the body region.
4. An information processing system, an information processing device that inputs a feature quantity representing a customer's characteristic into a trained estimation model and proposes a clothing item suitable for the customer from among a plurality of clothing items based on an estimation result of the estimation model; a learning device for performing machine learning on the estimation model based on learning data in which a clothing item purchased using a coupon among the plurality of clothing items is associated with an output of the estimation model for an image of a customer to whom the coupon was issued, The information processing device includes: A camera for imaging the customer; an image feature extraction unit for extracting the feature amount from an input image obtained by capturing an image of the customer with the camera, The information processing system, wherein the estimation model is trained to receive the features and output, as the estimation result, the respective likelihoods that each of the plurality of clothing items is a clothing item to be proposed.
5. The learning device includes: an acquisition unit for acquiring a training dataset, the training dataset including a plurality of training data sets in which training images obtained by capturing images of other customers are associated with clothing items purchased by the other customers among the plurality of clothing items as correct answers of the output of the estimation model; and an image feature extraction unit for extracting learning features from the learning images; and a learning unit for optimizing the estimation model so that an estimation result outputted by inputting the learning features extracted from the learning data into the estimation model approaches a purchase history of a clothing item labeled in the learning data.
6. A learning device for receiving an input of a feature quantity representing a characteristic of a customer and generating an estimation model used for proposing a clothing item suitable for the customer from among a plurality of clothing items, comprising: an acquisition unit for acquiring a training dataset, the training dataset including a plurality of training data sets in which an image of a customer to whom a coupon has been issued is associated with a clothing item purchased by the customer using the coupon among the plurality of clothing items as a correct answer of an output of the estimation model; and an image feature extraction unit for extracting the feature amount from the image; and a learning unit for optimizing the estimation model so that an estimation result output by inputting the features extracted from the learning data into the estimation model approaches the purchase history of the clothing items labeled in the learning data.
Citation Information
Patent Citations
Systems and methods for fashion shopping
JP2001502090A
Image processing device, image processing method, and program
JP2013196224A
Image processing apparatus, image processing method, and image processing program
JP2014044475A
Information processing apparatus, information processing method, and program
JP2016181196A
Recommendation device, recommendation method, and program
JP2017215667A