Impression estimation system, impression estimation model learning method, and program
By clustering the human attributes of annotator, training data based on human sets is generated, which solves the problems of cumbersome preparation of training data and insufficient individual combination data in the existing technology, and improves the accuracy of the impression estimation model.
Patent Information
- Application Number
- JP2023044425
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-03-20
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-03-20
AI Technical Summary
When training an impression estimation model, the existing technology requires preparing a large amount of training data. If the training data of individual attribute combinations is insufficient, the accuracy of the model will be difficult to improve, and the process of preparing training data is very cumbersome.
By clustering the annotator's human attributes, people with similar sensitivity are combined into a human set, and training data based on this human set is generated, reducing the cumbersomeness of training data preparation and improving the accuracy of the model.
It effectively reduces the cumbersomeness of training data preparation, and improves the accuracy of the impression estimation model, which can better estimate the impression of different groups of people on the image.
Smart Images

Figure 0007673360000007 
Figure 0007673360000008 
Figure 0007673360000009
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an impression estimation system, an impression estimation model learning method, and a program. [Background technology]
[0002] Conventionally, methods for estimating the impression a person receives from an image have been considered. For example, Non-Patent Documents 1 and 2 describe a method using a convolutional neural network. Non-Patent Document 3 describes a method in which an impression estimation model learns training data that inputs a training image and a combination of personal attributes, which is a combination of personal attributes such as the gender and age of the annotator, and outputs annotation information. Non-Patent Document 3 describes six combinations of personal attributes. The method of Non-Patent Document 3 has higher estimation accuracy than the methods of Non-Patent Documents 1 and 2, as the impression estimation model estimates impressions according to the combination of personal attributes. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Hidemichi Suzuki, Atsuhiro Yamada, Kensuke Tobitani, Sho Hashimoto, and Noriko Nagata, “An automatic modeling method of kansei evaluation from product data using a CNN model expressing the relationship between impressions and physical features,” in Proceedings of the 21st International Conference on Human-Computer Interaction, pp. 86-94, July 2019. [Non-Patent Document 2] Natsuki Sunda, Kensuke Tobitani, Iori Tani, Yusuke Tani, Noriko Nagata, and Nobufumi Morita, “Impression estimation model for clothing patterns using neural style features,” in Proceedings of the 22nd International Conference on Human-Computer Interaction, pp. 689-697, July 2020. [Non-Patent Document 3] Mayu Nakamoto, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Mitsuru Nakazawa, Chae Yeongnam, Stenger Bjorn, “A study on product image impression estimation method considering customer attributes,” 2021 Institute of Electronics, Information and Communication Engineers General Conference, D-12-5, March 2021. Summary of the Invention [Problem to be solved by the invention]
[0004] However, when training the impression estimation model on more than the six combinations of person attributes described in Non-Patent Document 3, it is necessary to prepare a large amount of training data, which is very time-consuming. On the other hand, if there is little training data for each combination of person attributes, it is not possible to sufficiently improve the accuracy of the impression estimation model. For this reason, there is a demand for improving the accuracy of the impression estimation model while reducing the time required for preparing training data.
[0005] One of the objectives of the present disclosure is to improve the accuracy of the impression estimation model while reducing the effort required for preparing training data. [Means for solving the problem]
[0006] The impression estimation system according to the present disclosure includes a training image acquisition unit that acquires training images for training an impression estimation model that estimates the impression a person has from an image, an annotation information acquisition unit that acquires annotation information related to the impression an annotator has from the training images, an annotator person attribute acquisition unit that acquires person attributes of the annotator, a person attribute clustering execution unit that executes person attribute clustering on the person attributes based on the annotation information so that person attributes of similar sensitivities belong to the same person set, a training data generation unit that generates training data for the impression estimation model based on the training images, the person set to which the person attributes of the annotator belong, and the annotation information, and a learning execution unit that executes learning of the impression estimation model based on the training data. Effect of the Invention
[0007] According to the present disclosure, the accuracy of an impression estimation model is improved while reducing the effort required to prepare training data. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of an overall configuration of an impression estimation system. [Diagram 2] FIG. 13 is a diagram showing an example of an outline of a process for grouping person attribute combinations with similar sensitivities into a person set. [Diagram 3] FIG. 13 is a diagram illustrating an example of an annotation screen. [Figure 4] FIG. 13 is a diagram showing a detailed example of person attributes. [Diagram 5] FIG. 2 is a diagram illustrating an example of an outline of an impression estimation model. [Figure 6] FIG. 13 is a diagram showing an example of a search screen. [Figure 7] FIG. 2 is a diagram illustrating an example of functions realized by the impression estimation system. [Figure 8] FIG. 2 is a diagram illustrating an example of a training image database. [Figure 9] FIG. 2 is a diagram illustrating an example of an annotation database. [Figure 10]FIG. 2 is a diagram illustrating an example of a person attribute database. [Figure 11] FIG. 13 is a diagram showing an example of a person set database. [Figure 12] FIG. 2 is a diagram illustrating an example of a training database. [Figure 13] FIG. 13 is a diagram illustrating an example of an estimated image database. [Figure 14] FIG. 11 is a diagram illustrating an example of a process executed in the impression estimation system. [Figure 15] FIG. 11 is a diagram illustrating an example of a process executed in the impression estimation system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] [1. Overall configuration of the impression estimation system] An example of an embodiment of an impression estimation system according to the present disclosure will be described. Fig. 1 is a diagram showing an example of the overall configuration of the impression estimation system. For example, the impression estimation system 1 includes a crowdsourcing server 10, an annotator terminal 20, a learning terminal 30, a search server 40, and a searcher terminal 50. Each of the crowdsourcing server 10, the annotator terminal 20, the learning terminal 30, the search server 40, and the searcher terminal 50 is connected to a network N such as the Internet or a LAN.
[0010] The crowdsourcing server 10 is a server computer for a crowdsourcing service. For example, the crowdsourcing server 10 includes a control unit 11, a storage unit 12, and a communication unit 13. The control unit 11 includes at least one processor. The storage unit 12 includes a volatile memory such as a RAM, and a non-volatile memory such as a flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication.
[0011] The annotator terminal 20 is a computer of an annotator who is a crowdworker registered in the crowdsourcing service. The crowdsourcing service is a service that outsources work to an unspecified number of crowdworkers via a network N. For example, the crowdsourcing service may be provided by the same provider as other services such as electronic commerce services, travel reservation services, communication services, financial services, or payment services. In this case, users of the other services can perform work as crowdworkers after completing the registration procedure for the crowdsourcing service. Crowdworkers are users of the crowdsourcing service.
[0012] For example, the annotator terminal 20 is a personal computer, a tablet, or a smartphone. The annotator terminal 20 includes a control unit 21, a memory unit 22, a communication unit 23, an operation unit 24, and a display unit 25. The physical configurations of the control unit 21, the memory unit 22, and the communication unit 23 may be similar to those of the control unit 11, the memory unit 12, and the communication unit 13, respectively. The operation unit 24 is an input device such as a keyboard, a mouse, or a touch panel. The display unit 25 is a display such as a liquid crystal or organic EL.
[0013] The learning terminal 30 is a computer that executes learning of an impression estimation model described below. For example, the learning terminal 30 is a personal computer, a tablet, or a smartphone. The learning terminal 30 includes a control unit 31, a memory unit 32, a communication unit 33, an operation unit 34, and a display unit 35. The physical configurations of the control unit 31, the memory unit 32, the communication unit 33, the operation unit 34, and the display unit 35 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 24, and the display unit 25, respectively.
[0014] The search server 40 is a server computer for a search service. The search service is a service for searching for content such as web pages or images. For example, the search service may be one of the services in an e-commerce service or a travel reservation service. In this case, the search service searches for product pages in the e-commerce service or facility pages in the travel reservation service. For example, the search server 40 includes a control unit 41, a memory unit 42, and a communication unit 43. The physical configurations of the control unit 41, the memory unit 42, and the communication unit 43 may be similar to those of the control unit 11, the memory unit 12, and the communication unit 13, respectively.
[0015] The searcher terminal 50 is a computer of a searcher who uses a search service. For example, the searcher terminal 50 is a personal computer, a tablet, or a smartphone. The searcher terminal 50 includes a control unit 51, a memory unit 52, a communication unit 53, an operation unit 54, and a display unit 55. The physical configurations of the control unit 51, the memory unit 52, the communication unit 53, the operation unit 54, and the display unit 55 may be similar to those of the control unit 11, the memory unit 12, the communication unit 13, the operation unit 24, and the display unit 25, respectively.
[0016] The programs stored in the storage units 12, 22, 32, 42, and 52 may be supplied via the network N. Also, the programs stored in a computer-readable information storage medium may be supplied via a reading unit (e.g., an optical disk drive or a memory card slot) that reads the information storage medium, or an input / output unit (e.g., a USB port) that inputs and outputs data to and from an external device.
[0017] Furthermore, the impression estimation system 1 only needs to include at least one computer, and is not limited to the example of Fig. 1. For example, the impression estimation system 1 may include only the learning terminal 30, without including the crowdsourcing server 10, the annotator terminal 20, the search server 40, and the searcher terminal 50. In this case, the crowdsourcing server 10, the annotator terminal 20, the search server 40, and the searcher terminal 50 exist outside the impression estimation system 1. For example, the impression estimation system 1 may include the learning terminal 30 and a computer not shown in Fig. 1, or may include only a computer not shown in Fig. 1.
[0018] [2. Overview of the impression estimation system] In this embodiment, a case where learning of the impression estimation model is executed by the learning terminal 30 is taken as an example. The impression estimation model is a model that estimates the impression a person receives from an image. The image is not limited to a photograph, and may be a computer graphic. The image shows an object for which an impression is to be estimated. In this embodiment, a case where the object for which an impression is to be estimated is a carpet is taken as an example, but the object for which an impression is to be estimated may be another object. For example, the object for which an impression is to be estimated may be a product, an animal, a plant, a building, an interior, a meal, a landscape, or another object.
[0019] For example, the learning terminal 30 executes learning of the impression estimation model based on a machine learning technique. In this embodiment, the impression estimation model is a supervised learning model such as a neural network or LightGBM, but the machine learning technique may be any other technique. For example, the impression estimation model may be a semi-supervised learning or unsupervised learning model.
[0020] In this embodiment, an impression estimation model based on the impression estimation model of Non-Patent Document 3 is taken as an example. For example, when an image and a combination of personal attributes of a person who views the image are input, the impression estimation model estimates the impression that the person receives from the image. A combination of personal attributes is a group of multiple personal attributes. A personal attribute is information for classifying a person. A personal attribute may be called a classification or a category or other name. A personal attribute may be information that can classify a person from some point of view, such as gender, age, membership rank in a service such as an electronic commerce service, place of residence, annual income, marital status, whether or not the person has children, car ownership status, or occupation.
[0021] As mentioned above, when trying to make the impression estimation model learn a large number of combinations of person attributes, it takes a lot of time to prepare training data. Furthermore, if there is little training data for each combination of person attributes, the accuracy of the impression estimation model cannot be sufficiently improved. In this regard, it is considered that stereotypes exist in the tendency of impressions received from images. For example, it is considered that an image that a young woman finds "cute" is different from an image that an older man finds "cute". Conversely, even if there is some difference in age, people of the same sex and close in age may have similar sensitivities and find similar images to be "cute".
[0022] Therefore, in this embodiment, by grouping person attribute combinations with similar sensitivities into one set, the effort required for preparing training data is reduced while the accuracy of the impression estimation model is improved. Hereinafter, a set of person attribute combinations is referred to as a person set. At least one person attribute combination belongs to a person set. A person set can also be referred to as a cluster or group of person attributes.
[0023] 2 is a diagram showing an example of an outline of a process for grouping person attribute combinations with similar sensitivities into a person set. For example, the storage unit 32 of the learning terminal 30 stores a training image database DB1, an annotation database DB2, and a person attribute database DB3. At least one of the training image database DB1, the annotation database DB2, and the person attribute database DB3 may be stored in a computer other than the learning terminal 30 or an external storage medium.
[0024] The training image database DB1 is a database in which training images for the impression estimation model M to learn are stored. In this embodiment, since the object for which the impression is to be estimated is a carpet, training images showing various carpets are stored in the training image database DB1. For example, the learning terminal 30 acquires all or a part of a product image showing a carpet sold as a commodity in an electronic commerce service. The learning terminal 30 stores all or a part of the acquired product images in the training image database DB1 as training images. Details of the data stored in the training image database DB1 will be described later.
[0025] The annotation database DB2 is a database in which annotation information regarding impressions that an annotator has received from a training image is stored. The annotator annotates the training image by viewing the training image and responding with an impression. In this embodiment, an example is given in which the annotator annotates each of a plurality of training images, but the annotator may annotate only one training image. Details of the data stored in the annotation database DB2 will be described later. For example, when the annotator logs in to a crowdsourcing service, an annotation screen is displayed on the display unit 25.
[0026] FIG. 3 is a diagram showing an example of an annotation screen. For example, impression words and training images are displayed on the annotation screen SC1. Impression words are words that indicate impressions. For example, impression words are words such as "stylish," "cute," "luxury," or "modern." The annotator checks the impression words and training images displayed on the annotation screen SC1. The annotator answers whether or not the training image gives the impression indicated by the impression words displayed on the annotation screen SC1.
[0027] For example, if the annotator receives the impression indicated by the impression word from the training image, the annotator selects button B10. If the annotator does not receive the impression indicated by the impression word from the training image, the annotator selects button B11. If the annotator selects button B12, other impression words are displayed on the annotation screen SC1. If the annotator selects button B12, other training images may be displayed on the annotation screen SC1. The annotator makes annotations one after another from the annotation screen SC1.
[0028] Returning to FIG. 2, the person attribute database DB3 is a database in which the person attributes of the annotator are stored. For example, when the annotator registers with a crowdsourcing service, the annotator inputs his / her own person attributes. In this embodiment, the annotator inputs not only one person attribute but also a person attribute combination. The person attribute database DB3 stores the person attribute combination input by the annotator. A person attribute combination registered in a service other than the crowdsourcing service may also be used. Details of the data stored in the person attribute database DB3 will be described later.
[0029] FIG. 4 is a diagram showing an example of details of a person attribute. For example, the person attribute "gender" has two details, such as "male" and "female". The details can also be said to be attribute values of the person attribute. Similarly, for other person attributes, at least one detail exists for each person attribute. The number of details of each person attribute may be the same or different. The person attribute database DB3 stores the person attributes that the annotator has registered in the crowdsourcing service. The person attributes that the annotator has not registered in the crowdsourcing service may be treated as missing values.
[0030] Returning to FIG. 2, the learning terminal 30 classifies a set of images I into which similar training images are included based on the characteristics of each of the plurality of training images stored in the training image database DB1. i In the example of FIG. 2, i is a number from 0 to k. i k is an integer. i is any natural number. The image clustering itself can use various methods used in the image processing field, such as the k-means method. For example, the learning terminal 30 executes image clustering based on a supervised machine learning method, a semi-supervised machine learning method, or an unsupervised machine learning method. The learning terminal 30 may execute image clustering based on a method other than machine learning.
[0031] For example, the learning terminal 30 acquires an image set I from the annotation database DB2. i The annotation information of each of the multiple training images belonging to is acquired. In this embodiment, an example is taken of the case where the annotation information is a vector indicating whether or not the annotator received the impression indicated by each of the multiple impression words. For example, if there are n impression words (n is a natural number), the annotation information becomes an n-dimensional vector. The elements of the vector are either a first value (e.g., 1) indicating that the annotator received the impression indicated by the impression word, or a second value (e.g., 0) indicating that the annotator did not receive the impression indicated by the impression word.
[0032] For example, the learning terminal 30 calculates an image set I based on the following formula 1. i In the above, the impressions of multiple annotators s who have the same combination of person attributes a are averaged and normalized to obtain the impression word score vector V(I i ,a) is calculated.
[0033]
number
[0034] In Equation 1, v(s) is the annotation information of annotator s. i ,a) is the image set I i V(I i The closer the element of a) is to 1, the more likely annotator s with the person attribute combination a is to rank the image set I i This indicates that there is a high probability that the corresponding impression will be received from the training images belonging to all the image sets I i On the other hand, the impression word score vector V(I i The learning terminal 30 calculates the impression word score vector V(I i , a) are sequentially concatenated to obtain a feature vector V(a) corresponding to the combination of person attributes a.
[0035]
number
[0036] For example, the learning terminal 30 performs person attribute clustering, which is clustering of the person attribute combinations a, based on the feature vector V(a) so that the person attribute combinations a with similar sensitivities belong to the same person set. For the person attribute clustering itself, various clustering methods can be used. For example, the learning terminal 30 performs person attribute clustering based on the k-means method. In the example of FIG. 2, the person attribute combinations a are clustered in the row direction in the image set I i The image set I is arranged in columns. iThe impression word score vector V(I i ,a) is the image set I i They are arranged in order.
[0037] When executing person attribute clustering, the learning terminal 30 may delete dimensions of the feature vector V(a) by principal component analysis until the cumulative contribution rate reaches a predetermined percentage (e.g., 90 percent) in order to avoid the curse of dimensionality. The dimension deletion may be performed based on a method other than principal component analysis. The learning terminal 30 acquires a person set based on the execution result of the person attribute clustering. The learning terminal 30 generates training data for the impression estimation model M based on the person set.
[0038] FIG. 5 is a diagram showing an example of an outline of the impression estimation model M. In this embodiment, the impression estimation model M similar to that described in Non-Patent Document 3 above is taken as an example. Here, the training image and the estimated image input to the impression estimation model M at the time of estimation are not distinguished, and are simply referred to as image I. For example, when image I is input to the impression estimation model M, the impression estimation model M calculates the features of image I. In the example of FIG. 5, it is assumed that ResNet50 previously trained by ImageNet is used as the impression estimation model M.
[0039] For example, a person attribute combination a is expressed as a one-hot vector α(a) in which one element is 1 and the remaining elements are all 0, and all the person sets that include the person attribute combination a are expressed as the one-hot vector α(a). In the example of Figure 5, 2048-dimensional image features and the number of person sets k A The learning terminal 30 calculates the impression word score vector p(I,α(a)) of the image I for α(a) based on the following formulas 3 and 4.
[0040]
number
[0041]
number
[0042] In Equation 3, θ is a parameter set of the impression estimation model M. V(a) is a feature vector of the person set. σ is an activation function. e is a function representing the feature extractor. g is a function representing the regressor. h is a function representing the fusion layer. In the fusion layer, the vector α(a) representing the person set is embedded into a vector of the same dimension as the vector representing the features of image I. Following the fusion layer, impression word scores are estimated by two fully connected layers.
[0043] For example, in learning the impression estimation model M, three items are given: the image feature, the feature vector V(a) of the person set, and the average value y(I,α(a)) of the impression received by the annotator s of the person set to which the person attribute combination a of the annotator s belongs. Each element of y(I,α(a)) represents each impression word score. y(I,α(a)) and the loss function L are calculated based on the following formula 5 and formula 6.
[0044]
number
[0045]
number
[0046] In Equation 5, v(s) is the impression vector of annotator s. j ) is the set of people A among annotators s of image I. j is a set of annotators s having a combination a of personal attributes included in the above. The learning terminal 30 optimizes the parameter θ and the feature vector V(a) while keeping the parameters of the pre-trained ResNet50 fixed. The output of the impression estimation model M is the impression word scores for the 24 impression words. For example, the learning terminal 30 executes learning of the impression estimation model M so that the loss function L of the formula 6 becomes sufficiently small.
[0047] In this embodiment, when learning of the impression estimation model M is completed, the learning terminal 30 performs labeling of an image to be searched in the search service. Hereinafter, this image is referred to as an estimated image. For example, the learning terminal 30 inputs an estimated image and a vector of an arbitrary group of people into the trained impression estimation model M, and obtains an estimation result output from the trained impression estimation model M. The search server 40 uses the estimation result output from the trained impression estimation model M as an index when a searcher uses a search service. For example, when a searcher accesses the search server 40, a search screen is displayed on the display unit 55.
[0048] FIG. 6 is a diagram showing an example of a search screen. For example, a searcher inputs a search query in an input form F20. The search server 40 acquires a person set to which the searcher's personal attributes belong. The search server 40 executes a search using impression words associated with the acquired person set as an index. The search server 40 displays the search results on the searcher terminal 50. For example, an image that a person set to which a woman in her 20s belongs may feel "cute" may differ from an image that a person set to which a man in his 60s belongs may feel "cute." Therefore, the search server 40 executes a search according to the person set to which the searcher's personal attribute combination belongs. As shown in FIG. 6, even with the same search query, the search results may differ depending on the person set to which the searcher's personal attribute combination belongs.
[0049] As described above, the impression estimation system 1 performs person attribute clustering so that person attributes with similar sensitivities belong to the same person set. The impression estimation system 1 performs learning of the impression estimation model M based on training data. This reduces the effort required to prepare training data while improving the accuracy of the impression estimation model M. Details of the impression estimation system 1 will be described below. Note that hereinafter, symbols such as annotator a described in formulas 1 to 6 will be omitted.
[0050] [3. Functions realized by the impression estimation system] FIG. 7 is a diagram showing an example of functions realized by the impression estimation system 1. As shown in FIG.
[0051] [3-1. Functions realized by the crowdsourcing server] For example, the crowdsourcing server 10 includes a data storage unit 100, a display control unit 101, and a collection unit 102. The data storage unit 100 is realized by the storage unit 12. The display control unit 101 and the collection unit 102 are realized by the control unit 11.
[0052] [Data storage section] The data storage unit 100 stores data necessary for providing a crowdsourcing service. For example, the data storage unit 100 stores a training image database DB1, an annotation database DB2, and a person attribute database DB3. The data storage unit 100 may store other data necessary for the processing described in this embodiment. For example, the data storage unit 100 may store HTML data of an annotation screen, etc.
[0053] 8 is a diagram showing an example of the training image database DB1. For example, the training image database DB1 stores an image set ID of each of a plurality of image sets and a file name of each of a plurality of training images belonging to the image set. The training image database DB1 may store other information related to the training images. For example, the actual data of each of the plurality of training images and the number of annotators who annotated each of the plurality of training images may be stored in the training image database DB1.
[0054] The image set ID is an example of image set identification information that can identify an image set. Therefore, the description of the image set ID can be read as image set identification information. In the example of FIG. 2, 0 to k i The numerical value corresponds to the image set ID. The image set identification information is not limited to the image set ID as long as it is information capable of identifying the image set in some way. For example, the image set identification information may be the name or number of the image set. The image set identification information may be stored in a database other than the training image database DB1.
[0055] The file name of the training image is an example of training image identification information that can identify the training image. Therefore, any part that describes the file name of the training image can be read as training image identification information. The training image identification information may be any information that can identify the training image in some way, and is not limited to the file name of the training image. For example, the training image identification information may be a path name indicating the location where the training image is stored, or an ID assigned to the training image.
[0056] In this embodiment, annotation is performed after image clustering is performed. However, annotation may be performed before image clustering is performed. In this case, the training image database DB1 stored in the data storage unit 100 does not store an image set ID. The annotator screen SC1 displays training images regardless of the image set.
[0057] 9 is a diagram showing an example of the annotation database DB2. For example, the annotation database DB2 stores annotator IDs of each of a plurality of annotators, file names of training images annotated by the annotators, and annotation information related to impressions the annotators received from the training images. The annotation database DB2 may store other information related to annotations. For example, instead of annotator IDs, a combination of personal attributes of the annotators may be stored in the annotation database DB2.
[0058] The details of the file names and annotation information of the training images are as described above. The annotator ID is an example of annotator identification information that can identify an annotator. Therefore, the part explaining the annotator ID can be read as annotator identification information. The annotator identification information is not limited to the annotator ID as long as it is information that can identify the annotator in some form. For example, the annotator identification information may be the annotator's email address or phone number. When the provider of the crowdsourcing service is the same as the provider of the other service, the annotator identification information may be common to these services.
[0059] 10 is a diagram showing an example of the person attribute database DB3. For example, the person attribute database DB3 stores an annotator ID of each of a plurality of annotators and a person attribute combination of the annotator. The person attribute database DB3 may store other information related to the person attributes of the annotators. For example, the person set ID of the person attribute to which each of the person attribute combinations of the plurality of annotators belongs may be stored in the person attribute database DB3.
[0060] For example, if there are nine personal attributes, such as gender, age group, membership rank, place of residence, annual income, marital status, whether or not the person has children, car ownership status, and occupation, details of these nine personal attributes are stored as a personal attribute combination in the personal attribute database DB3. The personal attribute combination may be expressed as a nine-dimensional vector, or in a format other than a vector (for example, an array format). When an annotator changes a personal attribute registered in the crowdsourcing service, the crowdsourcing server 10 changes the personal attribute combination associated with the annotator ID of this annotator.
[0061] [Display control section] The display control unit 101 displays the annotation screen SC1 on each annotator terminal 20 of a plurality of annotators. For example, the display control unit 101 transmits display data of the annotation screen SC1 to the annotator terminal 20, thereby displaying the annotation screen SC1 on the annotator terminal 20. The display data may be data for displaying some screen on the annotator terminal 20, and may be in any data format. For example, when the annotation screen SC1 is displayed on a browser, the display data is HTML data. When the annotation screen SC1 is displayed on an application dedicated to the crowdsourcing service, the display data may be in an image data format such as JPEG.
[0062] For example, when the display control unit 101 receives a display request for displaying the annotation screen SC1 from each annotator terminal 20 of a plurality of annotators, the display control unit 101 acquires at least one training image based on the training image database DB1. The display control unit 101 may randomly select at least one training image, or may select at least one training image for which the number of annotation information pieces collected is relatively small. The display control unit 101 generates display data for the annotation screen SC1 based on the acquired at least one training image. The display control unit 101 transmits the generated display data to the annotator terminal 20 that transmitted the display request, thereby causing the annotation screen SC1 to be displayed on the annotator terminal 20.
[0063] [Collection Department] The collection unit 102 collects annotation information from the annotator terminal 20 of each of the multiple annotators. The collection unit 102 stores the annotator ID of the annotator, the file name of the training image annotated by the annotator, and the collected annotation information in the annotation database DB2. The collection unit 102 may obtain information required for generating annotation information (e.g., information indicating the selection result of buttons B10 and B11) instead of collecting annotation information from the annotator terminal 20. In this case, the collection unit 102 generates annotation information based on the information and stores it in the annotation database DB2.
[0064] [3-2. Functions realized by the annotator device] For example, the annotator terminal 20 includes a data storage unit 200, a display control unit 201, and an operation reception unit 202. The data storage unit 200 is realized by the storage unit 22. The display control unit 201 and the operation reception unit 202 are realized by the control unit 21.
[0065] [Data storage section] The data storage unit 200 stores data necessary for annotation. For example, the data storage unit 200 stores a browser or an application dedicated to a crowdsourcing service. Note that the annotator does not have to be a crowdworker of the crowdsourcing service. For example, the annotator may perform annotation free of charge, or may be gathered at a specific place and perform annotation regardless of the crowdsourcing service.
[0066] [Display control section] The display control unit 201 causes the display unit 25 to display the annotation screen SC1 based on the display data of the annotation screen SC1.
[0067] [Operation reception section] The operation receiving unit 202 receives an operation on the annotation screen SC1. The annotator terminal 20 generates annotation information based on the annotator's operation on the annotation screen SC1. The annotator terminal 20 transmits the generated annotation information to the crowdsourcing server 10. When the annotation information is generated by the crowdsourcing server 10 instead of the annotator terminal 20, the annotator terminal 20 only needs to transmit information required for generating the annotation information to the crowdsourcing server 10. This information is as described above.
[0068] [3-3. Functions realized on the learning device] For example, the learning terminal 30 includes a data storage unit 300, a training image acquisition unit 301, an image clustering execution unit 302, an annotation information acquisition unit 303, an annotator person attribute acquisition unit 304, an important person attribute identification unit 305, a person attribute clustering execution unit 306, a training data generation unit 307, a learning execution unit 308, an estimated image acquisition unit 309, an estimated person attribute acquisition unit 310, and an estimation execution unit 311. The data storage unit 300 is realized by the storage unit 32. The training image acquisition unit 301, the image clustering execution unit 302, the annotation information acquisition unit 303, the annotator person attribute acquisition unit 304, an important person attribute identification unit 305, a person attribute clustering execution unit 306, a training data generation unit 307, a learning execution unit 308, an estimated image acquisition unit 309, an estimated person attribute acquisition unit 310, and an estimation execution unit 311 are each realized by the control unit 31.
[0069] [Data storage section] The data storage unit 300 stores data necessary for learning the impression estimation model M. For example, the data storage unit 300 stores a training image database DB1, an annotation database DB2, a person attribute database DB3, a person collection database DB4, a training database DB5, and an estimated image database DB6. The training image database DB1, the annotation database DB2, and the person attribute database DB3 are similar to those stored in the data storage unit 100 of the crowdsourcing server 10.
[0070] 11 is a diagram showing an example of the person set database DB4. The person set database DB4 is a database in which information related to person sets is stored. For example, the person set database DB4 stores a person set vector for each of a plurality of person sets and at least one person attribute combination belonging to the person set. The person set database DB4 may also store other information related to the person set. For example, the number of person attribute combinations belonging to the person set may be stored in the person set database DB4.
[0071] The person set vector is an example of person set identification information capable of identifying a person set. Therefore, any description of the person set vector can be read as person set identification information. The person set identification information is not limited to the person set vector as long as it is information capable of identifying a person set in some way. For example, the person set identification information may be an ID, name, number, or sequence indicating a person set. When person attribute clustering is performed by the person attribute clustering execution unit 306 described below, a person set database DB4 is generated.
[0072] 12 is a diagram showing an example of the training database DB5. The training database DB5 is a database in which each of a plurality of training data to be learned by the impression estimation model M is stored. The training data includes an input portion input to the impression estimation model M and an output portion corresponding to the input portion. The input portion is a portion of the training data that is input to the impression estimation model M during learning. The output portion is a portion of the training data that should be output from the impression estimation model M when the input portion is input to the impression estimation model M. The output portion corresponds to a correct answer during learning.
[0073] For example, the input and output parts of the training data have the same format as the input and output parts of the impression estimation model M at the time of estimation. In this embodiment, a pair of a training image and a person set vector corresponds to the input part of the training data. The impression word score calculated by Equation 5 corresponds to the output part of the training data. When training data is generated by training data generation unit 307 described later, the generated training data is stored in training database DB5.
[0074] The input and output parts of the training data may be data corresponding to the respective formats of the data input to the impression estimation model M and the data output from the impression estimation model M, and are not limited to the example of this embodiment. For example, when an image is input to the impression estimation model M after undergoing some image processing, the image after image processing may be included in the input part of the training data. When the person group identification information is in a format other than the vector format, data in the other format may be included in the input part of the training data. Data in which the impression word score is binarized using a threshold value may correspond to the output part of the training data.
[0075] 13 is a diagram showing an example of the estimated image database DB6. The estimated image database DB6 is a database in which information related to estimated images is stored. For example, the estimated image database DB6 stores the file names of each of a plurality of estimated images, the person set vector, and estimation result information. The estimated image database DB6 may store other information related to estimated images. For example, instead of the person set vector, a person attribute combination may be stored.
[0076] The estimation result information is information on the estimation result by the impression estimation model M. The estimation result information is stored for each pair of an estimated image and a person set vector. In this embodiment, an impression word in which the impression word score output by the impression estimation model M is binarized with a threshold value and the value indicates "YES" is stored as the estimation result information. For example, if it is estimated that a person belonging to a person set indicated by a person set vector has an impression of "stylish" and "cute" for a certain estimated image, the estimation result information indicates two impression words, "stylish" and "cute". These two impression words are used as indexes during a search by a searcher who belongs to a person set indicated by a person attribute vector associated with the impression word.
[0077] The estimation result information may be information that indicates the estimation result by the impression estimation model M in some form, and is not limited to the example of this embodiment. For example, the estimation result information may be the impression word score output by the impression estimation model M, or information in which the impression words are binarized using a threshold value. In this case, impression words whose impression word scores are equal to or greater than the threshold value, or impression words whose information in which the impression words are binarized using the threshold value indicates "YES", are used as indexes during search.
[0078] In this embodiment, the data storage unit 300 stores an impression estimation model M in addition to the databases described above. The impression estimation model M includes a program for executing impression estimation and parameters adjusted by learning. The data storage unit 300 stores the program and parameters of the impression estimation model M. For example, the data storage unit 300 stores the impression estimation model M whose parameters are at initial values. When learning is executed by the learning execution unit 308 described below, the impression estimation model M whose parameters are at initial values is replaced with a learned impression estimation model M. The learned impression estimation model M is an impression estimation model M whose parameters have been adjusted.
[0079] [Training image acquisition section] The training image acquisition unit 301 acquires training images for training the impression estimation model M. For example, the training image acquisition unit 301 acquires each of a plurality of training images whose file names are stored in the training image database DB1. Image data of the training images is assumed to be stored in the data storage unit 300 or an external storage medium. The training image acquisition unit 301 may acquire all the training images whose file names are stored in the training image database DB1, or may acquire some of the training images.
[0080] [Image clustering execution part] The image clustering execution unit 302 executes image clustering on the training images so that similar training images belong to the same image set. For example, the image clustering execution unit 302 executes image clustering based on each of the multiple training images, and identifies which image set among the multiple image sets each of the multiple training images belongs to. Multiple training images that belong to the same image set are similar to each other.
[0081] As described above, the image clustering method may be a known method. The image clustering execution unit 302 calculates the feature amount of each of the multiple training images based on a known method, and executes image clustering so that training images with similar feature amounts belong to the same image set. The number of image sets may be specified in advance, or may not be specified in particular. The image clustering execution unit 302 updates the training image database DB1 based on the execution result of the image clustering. The image clustering execution unit 302 associates the file name of each of the multiple training images with the image set ID to which the training image belongs.
[0082] [Annotation information acquisition section] The annotation information acquisition unit 303 acquires annotation information related to the impression that the annotator received from the training image. That is, the annotation information acquisition unit 303 acquires annotation information related to the impression that each of the multiple annotators received from at least one training image. In this embodiment, since the annotation information is stored in the annotation database DB2, the annotation information acquisition unit 303 acquires each of the multiple annotation information stored in the annotation database DB2.
[0083] The annotation information acquisition unit 303 may acquire all of the annotation information stored in the annotation database DB2, or may acquire only a portion of the annotation information. If the annotation information is recorded in a computer or an external storage medium other than the learning terminal 30, the annotation information acquisition unit 303 may acquire the annotation information from the other computer or the external storage medium. For example, the annotation information acquisition unit 303 may acquire the annotation information from the crowdsourcing server 10.
[0084] [Annotator person attribute acquisition section] The annotator person attribute acquisition unit 304 acquires the person attributes of the annotator. In this embodiment, the annotator is a crowdworker registered in a crowdsourcing service, so the annotator person attribute acquisition unit 304 acquires the person attributes of the crowdworker registered in the service. For example, the annotator person attribute acquisition unit 304 acquires each of a plurality of person attributes stored in the person attribute database DB3.
[0085] In this embodiment, since not only one person attribute but a person attribute combination that is a combination of multiple person attributes is used, the annotator person attribute acquisition unit 304 acquires the person attribute combination. The annotator person attribute acquisition unit 304 acquires each of the multiple person attribute combinations stored in the person attribute database DB3. That is, the annotator person attribute acquisition unit 304 acquires the person attribute combination of each of the multiple annotators.
[0086] The annotator person attribute acquisition unit 304 may acquire all the person attribute combinations stored in the person attribute database DB3, or may acquire some of the person attribute combinations. When the person attribute combinations are recorded in a computer or an external storage medium other than the learning terminal 30, the annotation information acquisition unit 303 may acquire the person attribute combinations from the other computer or the external storage medium. For example, the annotation information acquisition unit 303 may acquire the person attribute combinations from the crowdsourcing server 10.
[0087] [Important Person Attribute Identification Department] The important person attribute specification unit 305 specifies at least one person attribute having a relatively high importance from among a plurality of person attributes. The importance is the degree of importance as a feature. In other words, the importance is the degree to which the attribute is related to an impression. The higher the importance, the stronger the influence on the impression. For example, if five of nine person attributes, namely gender, age, membership rank, place of residence, annual income, marital status, presence or absence of children, car ownership status, and occupation, strongly influence the impression, namely gender, age, place of residence, occupation, and annual income, the importance of these five attributes will be higher than the remaining four attributes.
[0088] For example, the important person attribute specification unit 305 calculates the importance for each person attribute based on impression estimation using a machine learning technique such as LightGBM. The model for calculating the importance may be a provisionally created impression estimation model M, or may be another model different from the impression estimation model M. The other model may be a known model used in general impression estimation. The calculation of the importance itself can use a known technique. For example, the important person attribute specification unit 305 may calculate the importance of each of the multiple person attributes based on a correlation coefficient, SHAP (SHapley Additive exPlanations), or a technique called Cohort Shapley.
[0089] For example, the important person attribute identification unit 305 calculates at least one of the average importance of each person attribute for the impression words (sometimes referred to as average feature importance) and the cumulative importance of each person attribute for the impression words (sometimes referred to as cumulative feature importance). In this embodiment, a case will be described in which the important person attribute identification unit 305 calculates both of these, but it is also possible to calculate only one of these. For example, the important person attribute identification unit 305 may identify a person attribute whose average importance and cumulative importance are equal to or greater than a threshold, or the important person attribute identification unit 305 may identify a person attribute whose average importance and cumulative importance are high.
[0090] [Person attribute clustering execution part] The person attribute clustering execution unit 306 executes person attribute clustering on the person attributes based on the annotation information so that person attributes with similar sensitivities belong to the same person set. In this embodiment, since not only one person attribute but a person attribute combination consisting of multiple person attributes is used, the person attribute clustering execution unit 306 executes person attribute clustering so that person attribute combinations with similar sensitivities belong to the same set.
[0091] In this embodiment, the person attribute clustering execution unit 306 executes person attribute clustering based on Formula 1 and Formula 2. The person attribute clustering execution unit 306 updates the person set database DB4 based on the execution result of the person attribute clustering. The person attribute clustering execution unit 306 associates each of a plurality of person attribute combinations with a person set vector to which the person attribute combination belongs.
[0092] For example, the person attribute clustering execution unit 306 acquires, for each personal attribute combination, an evaluation value related to the sensibility of the personal attribute combination based on annotation information of at least one annotator who has the personal attribute combination. In this embodiment, a case will be described in which the impression word score vector of Formula 1 corresponds to the evaluation value, but the evaluation value is not limited to the impression word score vector of Formula 1. For example, the person attribute clustering execution unit 306 may perform only averaging without normalizing Formula 1, or may perform only normalization without averaging Formula 1.
[0093] Alternatively, for example, the person attribute clustering execution unit 306 may acquire, as an evaluation value, annotation information of a randomly selected annotator for each person attribute combination, without performing averaging and normalization. The evaluation value may be information that indicates the sensibility of the person attribute combination in some form. The evaluation value may be in any format, and is not limited to the vector format as in this embodiment. For example, the evaluation value may be a single numerical value, information consisting of multiple numerical values that are not in vector format, an array, a matrix, or other formats.
[0094] For example, the person attribute clustering execution unit 306 performs person attribute clustering so that a plurality of person attribute combinations with similar evaluation values belong to the same person attribute. Similar evaluation values correspond to similar sensitivities. Similar evaluation values mean that the difference (difference) between evaluation values is small. A smaller difference between evaluation values corresponds to a similar sensitivities. For example, the person attribute clustering execution unit 306 performs person attribute clustering so that a plurality of person attribute combinations with a difference in evaluation value less than a threshold value belong to the same person attribute.
[0095] In this embodiment, the person attribute clustering execution unit 306 performs person attribute clustering so that the person attributes of annotators who received similar impressions from training images belonging to the same image set belong to the same person set. For example, since the impression word score vector of Formula 1 is calculated for each image set, the person attribute clustering execution unit 306 performs person attribute clustering based on the impression word score vector calculated for each image set. The person attribute clustering execution unit 306 performs person attribute clustering so that when evaluation values calculated based on annotation information of annotators who have viewed each of a plurality of training images belonging to a certain image set are similar to each other, the person attribute clustering execution unit 306 performs person attribute clustering so that combinations of person attributes belong to the same person attribute.
[0096] In this embodiment, the person attribute clustering execution unit 306 performs at least one of averaging and normalization of annotation information of each of a plurality of annotators who have the same person attribute, and performs person attribute clustering based on the execution result of at least one of the averaging and normalization. As described above, the case where the person attribute clustering execution unit 306 performs both averaging and normalization based on Equation 1 will be described, but only either averaging or normalization may be performed.
[0097] In this embodiment, the person attribute clustering execution unit 306 executes dimension deletion for at least one of the annotation information and the person attributes, and executes person attribute clustering based on the execution result of the dimension deletion. For example, the person attribute clustering execution unit 306 executes dimension deletion for at least one of the annotation information and the person attributes based on principal component analysis. The person attribute clustering execution unit 306 executes person attribute clustering after reducing the dimension of at least one of the annotation information and the person attributes by the dimension deletion. The deleted dimension is not used in the person attribute clustering.
[0098] In this embodiment, a case will be described in which the person attribute clustering execution unit 306 performs dimension deletion regarding person attributes, but the person attribute clustering execution unit 306 may also perform dimension deletion regarding annotation information, or the person attribute clustering execution unit 306 may perform dimension deletion of both annotation information and person attributes.
[0099] In this embodiment, the person attribute clustering execution unit 306 executes person attribute clustering for at least one person attribute identified by the important person attribute identification unit 305. The person attribute clustering execution unit 306 targets only the at least one person attribute identified by the important person attribute identification unit 305 among the multiple person attributes for the person attribute clustering. The person attribute clustering execution unit 306 does not target other person attributes among the multiple person attributes for the person attribute clustering.
[0100] As described above, the method of avoiding the curse of dimensionality is not limited to the principal component analysis. The person attribute clustering execution unit 306 may delete a dimension that relatively does not contribute to the person attribute clustering from at least one of the annotation information and the person attributes. For example, the person attribute clustering execution unit 306 may execute the dimension deletion of at least one of the annotation information and the person attributes based on t-SNE, UMAP (Uniform Manifold Approximation and Projection), LLE (Locally Linear Embedding), random projection, or kernel method.
[0101] [Training data generation part] The training data generation unit 307 generates training data for the impression estimation model M based on the training images, the person sets to which the person attributes of the annotators belong, and the annotation information. The training data generation unit 307 executes the following process for each of the multiple training images. For example, the training data generation unit 307 generates training data that includes a combination of a training image and a person set as an input part and includes annotation information as an output part. In this embodiment, the training data generation unit 307 averages the annotation information of each of the multiple annotators having the same person attributes, and generates training data based on the executed averaging.
[0102] For example, the training data generation unit 307 acquires, as an input portion of the training data, one of a plurality of training images and a person set to which the person attribute combination of the annotator who annotated the training image belongs. The training data generation unit 307 acquires, as an output portion of the training data, an impression word score vector obtained by averaging annotation information of each annotator for a plurality of person attribute combinations belonging to the person set. The training data generation unit 307 generates training data one after another by similar processing based on each of the plurality of training images. The training data generation unit 307 stores the generated training data in the training database DB5.
[0103] [Learning Execution Department] The learning execution unit 308 executes learning of the impression estimation model M based on the training data. For the learning itself, various techniques used in machine learning techniques can be used. The parameters of the impression estimation model M are adjusted by the learning. For example, a gradient descent method or an error backpropagation method may be used. The learning execution unit 308 executes learning of the impression estimation model M based on each of a plurality of training data stored in the training database DB5 such that when an input portion of the training data is input to the impression estimation model M, an output portion of the training data is output from the impression estimation model M.
[0104] [Estimated image acquisition section] The estimated image acquisition unit 309 acquires an estimated image to be estimated by the trained impression estimation model M. For example, the estimated image acquisition unit 309 acquires each of a plurality of estimated images whose file names are stored in the estimated image database DB6. The image data of the estimated images is assumed to be stored in the data storage unit 300. The estimated image acquisition unit 309 may acquire all the estimated images whose file names are stored in the estimated image database DB6, or may acquire some of the estimated images.
[0105] [Estimated person attribute acquisition unit] The estimated person attribute acquisition unit 310 acquires person attributes to be estimated by the trained impression estimation model M. For example, the estimated person attribute acquisition unit 310 acquires each of a plurality of person attributes stored in the person attribute database DB3. The estimated person attribute acquisition unit 310 may acquire all of the person attributes stored in the person attribute database DB3, or may acquire some of the person attributes.
[0106] In this embodiment, since a person attribute combination that is a combination of multiple person attributes is used instead of using only one person attribute, the estimated person attribute acquisition unit 310 acquires the person attribute combination. The estimated person attribute acquisition unit 310 acquires each of the multiple person attribute combinations stored in the person attribute database DB3. The estimated person attribute acquisition unit 310 may acquire all the person attribute combinations stored in the person attribute database DB3, or may acquire some of the person attribute combinations.
[0107] [Estimation execution department] The estimation execution unit 311 estimates an impression that a person having a person attribute to be estimated receives from the estimated image, based on the estimated image, a person set to which the person attribute to be estimated belongs, and the trained impression estimation model M. For example, the estimation execution unit 311 inputs, to the trained impression estimation model M, the estimated image acquired by the estimated image acquisition unit 309 and the person set to which the person attribute combination acquired by the estimated person attribute acquisition unit 310 belongs.
[0108] For example, when an estimated image and a group of people are input to the impression estimation model M, it calculates their feature amounts. The flow of the process for calculating the feature amounts is as described with reference to FIG. 5. The impression estimation model M outputs an impression word score based on the calculated feature amounts. The estimation execution unit 311 updates the estimated image database DB6 so that the estimated image and group of people input to the impression estimation model M are associated with estimation result information obtained by binarizing the impression word score output from the impression estimation model M using a threshold value.
[0109] [3-4. Functions realized by the search server] For example, the search server 40 includes a data storage unit 400 and a search unit 401. The data storage unit 400 is realized by the storage unit 42. The search unit 401 is realized by the control unit 41.
[0110] [Data storage section] The data storage unit 400 stores a person set database DB4 and an estimated image database DB6. The person set database DB4 and the estimated image database DB6 may be similar to those stored in the data storage unit 300 of the learning terminal 30. The data storage unit 400 may store a combination of person attributes of a searcher, or may store a person set to which the searcher belongs.
[0111] [Search section] The search unit 401 searches for an estimated image based on a search query input by a searcher. For example, the search unit 401 searches for an estimated image based on the search query input by the searcher and an impression word associated with a person set to which the person attribute combination of the searcher belongs.
[0112] [3-5. Functions realized on the searcher's terminal] For example, the searcher terminal 50 includes a data storage unit 500, a display control unit 501, and an operation reception unit 502. The data storage unit 500 is realized by the storage unit 52. The display control unit 501 and the operation reception unit 502 are realized by the control unit 51.
[0113] [Data storage section] The data storage unit 500 stores data necessary for using the search service. For example, the data storage unit 500 stores a browser or an application dedicated to the search service.
[0114] [Display control section] The display control unit 501 causes the display unit 55 to display the search screen SC2.
[0115] [Operation reception section] The operation receiving unit 502 receives an operation by a searcher. For example, the operation receiving unit 502 receives an input of a search query by the searcher.
[0116] [4. Processing performed by impression estimation system 1] Figures 14 and 15 are diagrams showing an example of processing executed by the impression estimation system 1. The processing in Figures 14 and 15 is executed by the control units 11, 21, 31, 41, and 51 operating in accordance with programs stored in the storage units 12, 22, 32, 42, and 52, respectively.
[0117] As shown in FIG. 14, the learning terminal 30 acquires each of the multiple training images stored in the training image database DB1 stored in the storage unit 32 (S1). The learning terminal 30 performs image clustering based on each of the multiple training images acquired in S1 (S2). In S2, the learning terminal 30 performs image clustering so that similar training images belong to the same image set. The learning terminal 30 associates each of the multiple training images that have been the subject of image clustering with the image set ID of the image set to which the training image belongs. The learning terminal 30 uploads the training image database DB1 to the crowdsourcing server 10 (S3).
[0118] A process for collecting annotation information is executed between the crowdsourcing server 10 and the annotator terminal 20 (S4). In S4, the annotator terminal 20 accesses the crowdsourcing server 10 and displays the annotation screen SC1 on the display unit 25. The annotator terminal 20 transmits the annotation information to the crowdsourcing server 10 based on the annotator's operation on the annotation screen SC1. The crowdsourcing server 10 stores the annotation information received from the annotator terminal 20 in the annotation database DB2.
[0119] A process for sharing annotation information is executed between the crowdsourcing server 10 and the learning terminal 30 (S5). In S5, the crowdsourcing server 10 transmits the annotation database DB2 stored in the storage unit 32 to the learning terminal 30. The learning terminal 30 records the annotation database DB2 received from the crowdsourcing server 10 in the storage unit 32. The learning terminal 30 acquires annotation information of each of the multiple annotators stored in the annotation database DB2 (S6). The learning terminal 30 acquires a person attribute combination of each of the multiple annotators stored in the person attribute database DB3 (S7).
[0120] The learning terminal 30 performs person attribute clustering based on the annotation information of each of the multiple annotators acquired in S6 and the person attribute combinations of each of the multiple annotators acquired in S7 (S8). In S8, the learning terminal 30 performs person attribute clustering so that the person attribute combinations of annotators who have received similar impressions from training images belonging to the same image attribute belong to the same person set. The learning terminal 30 associates each of the multiple person attribute combinations that are the subject of the person attribute clustering with the person set ID of the person set to which the person attribute belongs.
[0121] The learning terminal 30 generates each of the multiple training data based on the execution result of the person attribute clustering in S8 (S9). In S9, the learning terminal 30 generates an input portion of the training data for each pair of a training image and a person attribute based on the training image and the person attribute. The learning terminal 30 generates an output portion of the training data based on the impression that an annotator belonging to the person attribute gets from the training image. The learning terminal 30 stores each of the multiple training data in the training database DB5.
[0122] The learning terminal 30 executes learning of the impression estimation model M based on each of the multiple training data generated in S9 (S10). In S10, the learning terminal 30 adjusts the parameters of the impression estimation model M so that when an input portion of each of the multiple training data is input to the impression estimation model M, an output portion of the training is output from the impression estimation model M. The learning terminal 30 acquires each of the multiple estimated images stored in the estimated image database DB6 (S11).
[0123] 15, the learning terminal 30 acquires each of the multiple person attribute combinations stored in the person attribute database DB3 (S12). The learning terminal 30 estimates an impression of a person of the person attribute combination from the estimated image based on each of the multiple estimated images acquired in S11, each of the multiple person attribute combinations acquired in S12, and the trained impression estimation model M (S13). In S13, the learning terminal 30 inputs the estimated image and a person set to which the person attribute combination belongs to the impression estimation model M, and acquires the estimation result output from the impression estimation model M. The learning terminal 30 updates the estimated image database DB6.
[0124] A process is executed between the learning terminal 30 and the search server 40 to share the estimation results stored in the estimated image database DB6 (S14). In S14, the learning terminal 30 transmits the estimated image database DB6 to the search server 40. The search server 40 records the estimated image database DB6 received from the learning terminal 30 in the storage unit 42. A search is executed between the search server 40 and the searcher terminal 50 (S15).
[0125] [5. Summary of the embodiment] The impression estimation system 1 of the present embodiment performs person attribute clustering on person attributes based on annotation information so that person attributes with similar sensitivities belong to the same person set. The impression estimation system 1 generates training data for the impression estimation model M based on a training image, a person set, and annotation information. The impression estimation system 1 performs learning of the impression estimation model M based on the training data. As a result, by grouping person attributes with similar sensitivities into one person set, annotation information only needs to be collected for each person set, so that the effort of preparing training data can be reduced. For example, when the sensitivities of a man in his 30s and a man in his 40s are similar to each other, the annotation information of the man in his 30s and the annotation information of the man in his 40s are not collected separately, but are grouped into one person set and annotation information is collected for each person set, thereby reducing the effort of preparing training data. Furthermore, even if there is little annotation information for men in their 30s and men in their 40s, by grouping these into one person set and collecting annotation information on a person set basis, the amount of annotation information per person set can be increased, thereby improving the accuracy of the impression estimation model M.
[0126] Furthermore, in the impression estimation system 1, the annotator person attribute acquisition unit 304 executes person attribute clustering so that person attribute combinations with similar sensitivities belong to the same set. When attempting to execute impression estimation taking into consideration a person attribute combination consisting of multiple person attributes, rather than a single person attribute, the number of person attribute combinations tends to become enormous, but by acquiring a person set through person attribute clustering, the accuracy of the impression estimation model M is improved while reducing the effort required for preparing training data.
[0127] Furthermore, the impression estimation system 1 performs image clustering so that similar training images belong to the same image set. The impression estimation system 1 performs person attribute clustering so that the person attributes of annotators who received similar impressions from training images belonging to the same image set belong to the same person set. This allows the impression estimation system 1 to perform person attribute clustering taking into consideration whether similar impressions are received from similar training images, thereby improving the accuracy of the person attribute clustering.
[0128] Furthermore, the impression estimation system 1 performs at least one of averaging and normalization of annotation information of each of a plurality of annotators who have the same person attribute, and performs person attribute clustering based on the execution result of at least one of the averaging and normalization. This improves the accuracy of the person attribute clustering.
[0129] Furthermore, the impression estimation system 1 executes dimension deletion for at least one of the annotation information and the person attributes, and executes person attribute clustering based on the execution result of the dimension deletion. This makes it possible to execute person attribute clustering while avoiding the curse of dimensionality.
[0130] Furthermore, the impression estimation system 1 averages the annotation information of each of a plurality of annotators who have the same person attributes, and generates training data based on the executed averaging. This makes it possible to generate training data in which the impressions of a plurality of annotators who have similar sensitivities from a training image are averaged, thereby improving the accuracy of the impression estimation model M.
[0131] Furthermore, the impression estimation system 1 performs personal attribute clustering on at least one personal attribute that has a relatively high degree of importance, thereby making it possible to perform personal attribute clustering based on a personal attribute that is more important for impression estimation.
[0132] In addition, the impression estimation system 1 acquires personal attributes of crowdworkers registered in the crowdsourcing service, which allows for efficient collection of annotation information.
[0133] Moreover, the impression estimation system 1 estimates the impression that a person having the person attributes to be estimated receives from the estimated image, based on the estimated image, a person set to which the person attributes to be estimated belong, and the trained impression estimation model M. This improves the accuracy of impression estimation.
[0134] [6. Modifications] The present disclosure is not limited to the above-described embodiment, and may be modified as appropriate without departing from the spirit and scope of the present disclosure.
[0135] For example, in the embodiment, the annotation information of each of a plurality of annotators belonging to the same person attribute is averaged and combined into one, but the training data generation unit 307 may generate training data for each annotator based on a training image, a person set of the annotator, and the annotation information of the annotator. That is, annotation information by each annotator does not need to be combined into one. Although the embodiment differs in that training data is generated for each annotator, the learning itself based on the training data is the same as the embodiment.
[0136] The impression estimation system 1 of the modified example generates training data for each annotator based on a training image, a person set of the annotator, and annotation information of the annotator. This reduces the effort required to prepare training data and improves the accuracy of the impression estimation model M.
[0137] For example, the functions described as being realized by the learning terminal 30 may be realized by at least one computer in the impression estimation system 1, or the functions may be shared among multiple computers. In this case, the functions may be shared by each of the multiple computers transmitting its own processing results to the other computers. For example, the functions described as being realized by the learning terminal 30 may be realized by the crowdsourcing server 10 or the search server 40.
[0138] [7. Notes] For example, the impression estimation system according to the present disclosure may be configured as follows. (1) a training image acquisition unit that acquires training images for training an impression estimation model that estimates the impression a person has from an image; An annotation information acquisition unit that acquires annotation information regarding an impression received by an annotator from the training image; an annotator person attribute acquisition unit that acquires a person attribute of the annotator; a person attribute clustering execution unit that executes person attribute clustering on the person attributes based on the annotation information so that the person attributes having similar sensitivities belong to the same person set; a training data generation unit that generates training data for the impression estimation model based on the training image, the person set to which the person attribute of the annotator belongs, and the annotation information; a learning execution unit that executes learning of the impression estimation model based on the training data; An impression estimation system comprising: (2) the annotator person attribute acquisition unit acquires a combination of a plurality of the person attributes; the person attribute clustering execution unit executes the person attribute clustering such that the combinations of similar sensitivities belong to the same set. The impression estimation system according to (1). (3) The impression estimation system further includes an image clustering execution unit that executes image clustering on the training images such that similar training images belong to the same image set; the person attribute clustering execution unit executes the person attribute clustering such that the person attributes of the annotators who have received similar impressions from the training images belonging to the same image set belong to the same person set. An impression estimation system according to (1) or (2). (4) the person attribute clustering execution unit performs at least one of averaging and normalization of the annotation information of each of the multiple annotators who have the same person attribute, and performs the person attribute clustering based on a result of performing at least one of the averaging and the normalization; The impression estimation system according to any one of (1) to (3). (5) the person attribute clustering execution unit executes dimension deletion regarding at least one of the annotation information and the person attributes, and executes the person attribute clustering based on a result of the dimension deletion. The impression estimation system according to any one of (1) to (4). (6) the training data generation unit averages the annotation information of each of the multiple annotators who have the same person attribute, and generates the training data based on the average thus performed. The impression estimation system according to any one of (1) to (5). (7) the training data generation unit generates, for each of the annotators, the training data based on the training image, the person set of the annotator, and the annotation information of the annotator. The impression estimation system according to any one of (1) to (6). (8) The impression estimation system further includes an important person attribute identification unit that identifies at least one of the plurality of person attributes having a relatively high degree of importance, the person attribute clustering execution unit executes the person attribute clustering for the at least one person attribute; The impression estimation system according to any one of (1) to (7). (9) The annotator is a crowdworker registered with a crowdsourcing service, The annotator person attribute acquisition unit acquires the person attributes of the crowdworkers registered in the crowdsourcing service. The impression estimation system according to any one of (1) to (8). (10) The impression estimation system includes: an estimated image acquisition unit that acquires an estimated image that is to be estimated by the trained impression estimation model; an estimated person attribute acquisition unit that acquires the person attribute to be estimated by the trained impression estimation model; an estimation execution unit that estimates an impression that a person having the person attribute that is the estimation target receives from the estimated image based on the estimated image, the person set to which the person attribute that is the estimation target belongs, and the trained impression estimation model; The impression estimating system according to any one of (1) to (9), [Explanation of symbols]
[0139] 1 impression estimation system, N network, 10 crowdsourcing server, 11, 21, 31, 41, 51 control unit, 12, 22, 32, 42, 52 memory unit, 13, 23, 33, 43, 53 communication unit, 20 annotator terminal, 24, 34, 54 operation unit, 25, 35, 55 display unit, 30 learning terminal, 40 search server, 50 searcher terminal, 100 data storage unit, 101 display control unit, 102 collection unit, 200 data storage unit, 201 display control unit, 202 operation reception unit, 300 data storage unit, 301 training image acquisition unit, 302 image clustering execution unit, 303 annotation information acquisition unit, 304 annotator person attribute acquisition unit, 305 important person attribute identification unit, 306 person attribute clustering execution unit, 307 training data generation unit, 308 Learning execution unit, 309 estimated image acquisition unit, 310 estimated person attribute acquisition unit, 311 estimation execution unit, 400 data storage unit, 401 search unit, 500 data storage unit, 501 display control unit, 502 operation reception unit, DB1 training image database, DB2 annotation database, DB3 person attribute database, DB4 person set database, DB5 training database, DB6 estimated image database, SC1 annotation screen, SC2 search screen.
Claims
1. a training image acquisition unit that acquires training images for training an impression estimation model that estimates the impression a person has from an image; An annotation information acquisition unit that acquires annotation information regarding an impression received by an annotator from the training image; an annotator person attribute acquisition unit that acquires a person attribute of the annotator; a person attribute clustering execution unit that executes person attribute clustering on the person attributes based on the annotation information so that the person attributes having similar sensitivities belong to the same person set; a training data generation unit that generates training data for the impression estimation model based on the training image, the person set to which the person attribute of the annotator belongs, and the annotation information; a learning execution unit that executes learning of the impression estimation model based on the training data; An impression estimation system comprising:
2. the annotator person attribute acquisition unit acquires a combination of a plurality of the person attributes; the person attribute clustering execution unit executes the person attribute clustering such that the combinations of similar sensitivities belong to the same person set. The impression estimation system according to claim 1 .
3. The impression estimation system further includes an image clustering execution unit that executes image clustering on the training images such that similar training images belong to the same image set; the person attribute clustering execution unit executes the person attribute clustering such that the person attributes of the annotators who have received similar impressions from the training images belonging to the same image set belong to the same person set. The impression estimating system according to claim 1 .
4. the person attribute clustering execution unit performs at least one of averaging and normalization of the annotation information of each of the multiple annotators who have the same person attribute, and performs the person attribute clustering based on a result of performing at least one of the averaging and the normalization; The impression estimating system according to claim 1 .
5. the person attribute clustering execution unit executes dimension deletion regarding at least one of the annotation information and the person attributes, and executes the person attribute clustering based on a result of the dimension deletion. The impression estimating system according to claim 1 .
6. the training data generation unit averages the annotation information of each of the multiple annotators who have the same person attribute, and generates the training data based on the average thus performed. The impression estimating system according to claim 1 .
7. the training data generation unit generates, for each of the annotators, the training data based on the training image, the person set of the annotator, and the annotation information of the annotator. The impression estimating system according to claim 1 .
8. The impression estimation system further includes an important person attribute identification unit that identifies at least one of the plurality of person attributes having a relatively high degree of importance, the person attribute clustering execution unit executes the person attribute clustering for at least one of the person attributes; The impression estimating system according to claim 1 .
9. The annotator is a crowdworker registered with a crowdsourcing service, The annotator person attribute acquisition unit acquires the person attributes of the crowdworkers registered in the crowdsourcing service. The impression estimating system according to claim 1 .
10. The impression estimation system includes: an estimated image acquisition unit that acquires an estimated image that is to be estimated by the trained impression estimation model; an estimated person attribute acquisition unit that acquires the person attribute to be estimated by the trained impression estimation model; an estimation execution unit that estimates an impression that a person having the person attribute that is the estimation target receives from the estimated image based on the estimated image, the person set to which the person attribute that is the estimation target belongs, and the trained impression estimation model; The impression estimating system according to claim 1 or 2, comprising:
11. A training image for training an impression estimation model that estimates the impression a person has from an image; The person set is determined by performing person attribute clustering on the person attributes so that person attributes having similar sensitivities belong to the same person set based on annotation information on impressions that an annotator has received from the training images, the person set to which the person attributes of the annotator belong; The annotation information; An impression estimation system in which training data generated based on the an estimated image acquisition unit that acquires an estimated image that is to be estimated by the trained impression estimation model; an estimated person attribute acquisition unit that acquires the person attribute to be estimated by the trained impression estimation model; an estimation execution unit that estimates an impression that a person having the person attribute that is the estimation target receives from the estimated image based on the estimated image, the person set to which the person attribute that is the estimation target belongs, and the trained impression estimation model; An impression estimation system comprising:
12. A computer comprising: A training image acquisition step of acquiring training images to be trained into an impression estimation model that estimates the impression a person has from an image; An annotation information acquisition step of acquiring annotation information regarding an impression received by an annotator from the training image; an annotator person attribute acquisition step of acquiring a person attribute of the annotator; a person attribute clustering execution step of executing person attribute clustering on the person attributes based on the annotation information so that the person attributes having similar sensitivities belong to the same person set; a training data generation step of generating training data for the impression estimation model based on the training image, the person set to which the person attribute of the annotator belongs, and the annotation information; a learning execution step of learning the impression estimation model based on the training data; A method for training an impression estimation model that performs the above.
13. a training image acquisition unit that acquires training images to be trained into an impression estimation model that estimates the impression a person has from an image; an annotation information acquisition unit that acquires annotation information regarding an impression received by an annotator from the training image; an annotator person attribute acquisition unit that acquires a person attribute of the annotator; a person attribute clustering execution unit that executes person attribute clustering on the person attributes based on the annotation information so that the person attributes having similar sensitivities belong to the same person set; a training data generation unit that generates training data for the impression estimation model based on the training image, the person set to which the person attribute of the annotator belongs, and the annotation information; a learning execution unit that executes learning of the impression estimation model based on the training data; A program that makes a computer function as a
Citation Information
Patent Citations
Image analyzer, image analyzing method, image search system, and program
JP2013050857A
Learning apparatus, estimation apparatus, learning method, estimation method, program, and program of learned estimation model
JP2022013346A
Image evaluation predicting device and method
JP2022037575A
Support system, analyzer, analysis method, support device, support method, analysis program, and support program
JP2022046314A
Product search device, product search system, server system, and product search method
WO2015129334A1