Comment outputting device, comment outputting method, and comment outputting program
The comment output device addresses the issue of inaccurate comment selection by using latent attributes derived from multiple-question responses to determine user similarity, ensuring relevant and accurate comment display.
Patent Information
- Application Number
- JP2024060360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-16
AI Technical Summary
Existing comment output technologies fail to accurately classify users based on latent attributes, leading to inaccurate selection and display of comments, as they rely solely on explicit attributes without utilizing latent attributes derived from multiple-question responses.
A comment output device that determines a user's latent attribute identifier using two or more aggregate quantities calculated from answers to three or more predetermined questions, allowing for the selection and display of comments from users with matching latent attributes.
Enables users to view comments from others who are similar to them by accurately utilizing latent attributes, improving the relevance and accuracy of displayed comments.
Smart Images

Figure 2025157967000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a comment output device, a comment output method, and a comment output program for outputting comments such as word-of-mouth comments. [Background technology]
[0002] Known technologies for outputting comments such as word-of-mouth about objects, i.e., beauty products such as cosmetics and beauty services, health foods and medicines, health management services, etc., include those disclosed in Patent Documents 1 and 2.
[0003] Patent Document 1 discloses a technology for a cosmetics information analysis method that extracts only cosmetics users whose cosmetics usage information overlaps with that of a user, calculates the user similarity between the extracted cosmetics users and the user based on cosmetics user attribute information, displays images of the cosmetics users sorted based on the user similarity on the display means of the user terminal, and when a user selects an image of a cosmetics user, displays the cosmetics user attribute information (including word-of-mouth) of that cosmetics user on the display means of the user terminal.
[0004] Patent Document 2 discloses a technology for a comment output device having an output unit that outputs, for each other user's own information, only other comments that correspond to other information that satisfies a predetermined similarity condition for the own information. Here, the own information is information having one or more own comments and scores associated with each of the own comments, and the other information is information having one or more other comments of each other user and scores associated with each other comment. The technology described in Patent Document 2 aims to provide a comment output device that allows a user to appropriately view comments such as word-of-mouth from others that are similar to the user's own. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 4478479 [Patent Document 2] Patent No. 7001662 Summary of the Invention [Problem to be solved by the invention]
[0006] The technology described in Patent Document 1 extracts only cosmetics users whose cosmetics use overlap with the user, and calculates the user similarity between the cosmetics user and the user based on cosmetics user attribute information. Here, cosmetics user attribute information includes skin type, age, region, favorite brand, occupation, and user review ratings for the cosmetics used. For example, users set their "skin type" themselves on a basic information setting screen, choosing from normal skin, dry skin, oily skin, combination skin, sensitive skin, and atopic skin. However, classifying "skin type" is difficult using only "explicit attributes" obtained through direct questionnaires such as "Which of the following options best describes your skin type?" Accuracy is low unless "skin type" is classified using "latent attributes" expressed using two or more aggregate quantities obtained from responses to a questionnaire containing three or more questions, preferably several dozen questions. In the technology described in Patent Document 1, the cosmetics user attribute information that influences the decision to display or not display comments is the "explicit attributes" set or entered by the user. "Latent attributes," which would enable more accurate user classification, are not used as cosmetics user attribute information.
[0007] In the technology described in Patent Document 2, the difference between the score corresponding to the comment and the score corresponding to the other comment, and the difference between one or more evaluation words of the comment and the score of the other comment are calculated. Whether or not the information itself and other information satisfy the similarity condition is determined based on the number of matches with one or more evaluation words, etc. However, Patent Document 2 does not disclose the use of "latent attributes" expressed using two or more aggregate quantities obtained from answers to a questionnaire containing three or more questions, more preferably several tens of questions, as user attributes used to determine whether to display or hide a comment.
[0008] An object of one aspect of the present invention is to provide a comment output device that allows a user to appropriately view comments, such as reviews, of others similar to the user, by utilizing a latent attribute of the user that is determined using at least two aggregation quantities calculated using all of the user's answers to three or more predetermined questions.An object of another aspect of the present invention is to provide a comment output device that calculates a predicted value of the aggregation quantity of the user from an image of the user using machine learning, determines a latent attribute of the user from the thus obtained predicted values of at least two aggregation quantities, and allows the user to appropriately view comments, such as reviews, of others similar to the user, by utilizing the latent attribute.An object of yet another aspect of the present invention is to provide a comment output method that allows a user to appropriately view comments, such as reviews, of others similar to the user, by utilizing the latent attribute of the user.An object of yet another aspect of the present invention is to provide a comment output program that causes a computer to function as the comment output device. [Means for solving the problem]
[0009] The present invention has been made to solve such problems, and a first form of the present invention is a comment output device comprising: a receiving unit that receives a comment output instruction; a comment acquisition unit that acquires one or more comments together with the user identifier and latent attribute identifier of the user who made the comment; a comment selection unit that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction; an output unit that outputs only the comments selected by the comment selection unit; and an analysis unit that determines the latent attribute identifier of each user using two or more aggregate quantities that are calculated in advance using all of the user's answers to three or more predetermined questions.
[0010] According to this embodiment, a comment output device can be provided that allows a user to appropriately view comments such as word-of-mouth from others who are similar to the user, by utilizing the user's latent attribute identifier, which is determined using two or more aggregate quantities calculated using all of the user's answers to three or more predetermined questions.
[0011] A second aspect of the present invention is a method for determining the accuracy of the first aspect of the present invention, in which a large number of subjects are asked in advance a medical interview consisting of three or more predetermined questions, each answer of which can be expressed by one real number selected from a finite number of real numbers, and a first aggregate quantity S1 and a second aggregate quantity S2, both of which take real numbers, are determined as a function (S1, S2) = f(A) of the subject's answer A through statistical analysis of basic data, which is a collection of the subjects' answers to the plurality of questions, and a coordinate plane with the first aggregate quantity S1 on the horizontal axis and the second aggregate quantity S2 on the vertical axis is used as a discrimination surface, and a plurality of regions are arranged on the discrimination surface, and one latent attribute identifier is assigned to each region; a storage unit that stores the function f, the plurality of regions, and a correspondence relationship between each of the regions and a latent attribute identifier; The comment output device is configured to ask the user a questionnaire consisting of the three or more questions in advance, and the analysis unit determines the first aggregation amount T1 and the second aggregation amount T2 of the user from the user's answer B using the function f as (T1, T2) = f(B), and determines the latent attribute identifier corresponding to the area including the point (T1, T2) among the multiple areas arranged on the identification surface as the latent attribute identifier of the user.
[0012] According to this embodiment, the user is asked in advance a medical interview consisting of the three or more questions, and the analysis unit calculates the first aggregate amount T1 and the second aggregate amount T2 of the user from the answer B of the user. As described above, the latent attribute identifier corresponding to the area including point (T1, T2) among the multiple areas arranged on the identification surface is defined as the latent attribute identifier of the user. Therefore, by utilizing the latent attribute identifier thus defined, it is possible to provide a comment output device that allows a user to appropriately view comments such as word-of-mouth from others who are similar to the user.
[0013] Here, the "large number of subjects" in the basic data refers to, but is not limited to, approximately 1,000 or more subjects, preferably 3,000 or more subjects, and more preferably 5,000 or more subjects. The larger the scale of the basic data consisting of responses from subjects belonging to a sample population randomly sampled from the target customer population, the smaller the deviation between the relative position of each customer in the sample population as understood by the first aggregation amount and the second aggregation amount and the true relative position of that customer in the population, and the more accurately the latent attributes can be understood. It is preferable that at least one of the multiple regions intersects (overlaps) with any of the other regions. It is also more preferable that each of the multiple regions intersects (overlaps) with any of the other regions. By having regions intersect (overlap), it is possible to highly likely avoid the occurrence of an undesirable situation in which "two users whose first aggregation amount T1 and second aggregation amount T2 are both very similar in value do not have an opportunity to see each other's comments."
[0014] A third aspect of the present invention is the comment output device of the second aspect, wherein the two or more aggregate amounts are calculated using at least one statistical analysis method among factor analysis, principal component analysis, and collaborative filtering.
[0015] When the statistical analysis method is factor analysis, for example, k=1, 2, . . . , it can be determined that the k-th aggregation amount of each user is the factor score related to the k-th factor of the user. When the statistical analysis method is principal component analysis, for example, k=1, 2, . . . , it can be determined that the k-th aggregation amount of each user is the k-th principal component of the user. When the statistical analysis method is collaborative filtering, for example, where k=1, 2, ..., it can be determined that for each k, the k-th aggregate amount of each user is the k-th component of the user latent feature vector of that user. In this case, the inner product of the item latent feature vector related to each question (item) and the user latent feature vector related to each user approximately predicts the answer of that user to that question.
[0016] A fourth aspect of the present invention is the comment output device of the second aspect, wherein at least one of the plurality of areas arranged on the identification surface is an area whose boundary is partially configured by a curve.
[0017] In this embodiment, the multiple regions arranged on the identification surface are not only squares (rectangles) or triangles, but at least one of the multiple regions is a region whose boundary is partially formed by curves, so that the user's latent attributes of each user can be grasped with higher accuracy and the similarity of users can be determined more accurately, thereby providing a comment output device that allows users to appropriately view comments such as word-of-mouth from others who are similar to them.
[0018] A fifth aspect of the present invention includes a receiving unit that receives a comment output instruction, a comment acquiring unit that acquires one or more comments together with the user identifier and latent attribute identifier of the user who made the comment, a comment selecting unit that selects only comments of a user having a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction, an output unit that outputs only the comments selected by the comment selecting unit, and a supervised learning method for machine learning that predicts an aggregate amount calculated in advance using all of the user's answers to three or more predetermined questions from image data of the user, the aggregate amount being calculated in advance using all of the user's answers to three or more predetermined questions. and an analysis unit that determines an aggregate amount predicted value, which is a predicted value of the aggregate amount, using two or more of the predicted values.
[0019] According to this embodiment, the analysis unit uses a supervised learning technique in machine learning to obtain at least two aggregated amount prediction values from the user's image data, and determines a latent attribute identifier for the user using the aggregated amount prediction values, thereby providing a comment output device that allows a user to appropriately view comments such as word-of-mouth reviews from others who are similar to the user. As the supervised learning technique, various known techniques using, for example, a CNN (convolutional neural network) can be used.
[0020] A sixth form of the present invention is a comment output device according to any one of the first to fifth forms, wherein the comment is a comment on a beauty product, a health product, or a food product.
[0021] According to this embodiment, a comment output device can be provided that outputs comments on beauty products, health products, or food products, and allows users to appropriately view comments such as word-of-mouth reviews from others who are similar to them. The beauty products mentioned above include various products and services, such as cosmetics and beauty services, that are primarily concerned with skin care. These include various cosmetics that cover everything from daily skin care to specialized care, such as cleansers, moisturizers, serums, and masks. Beauty services also include facial treatments offered at beauty salons, and hair care, nail care, and massage services at hair salons. Health products include various products and services used to contribute to people's health, such as disease prevention, health management, and extending healthy lifespan, and can include health foods, medications, supplements, health management services, and health advice services. Health foods are foods that provide balanced nutrition and have specific health benefits, such as organic foods and functional foods. Medications include those intended for disease prevention and treatment, and include prescription and over-the-counter medications. Supplements are used to supplement nutrients that are difficult to obtain through daily meals, such as vitamins, minerals, and proteins. Health management services offer personal training, nutritional consultations, health checkups, chiropractic care, and other services, with programs tailored to individual health conditions. Health advice services support the maintenance and improvement of better health through expert counseling and information provision. These foods include grains, fruits, vegetables, meat, seafood, dairy products, fats and oils, sweets and processed foods, legumes, seeds and nuts, eggs, beverages, spices and herbs, and mushrooms.
[0022] A seventh form of the present invention is a comment output method realized by a receiving unit, a comment acquisition unit, a comment selection unit, an output unit, and an analysis unit, the comment output method including: a receiving step in which the receiving unit receives an instruction to output a comment; a comment acquisition step in which the comment acquisition unit acquires one or more comments together with the user identifier and latent attribute identifier of the user who made the comment; a comment selection step in which the comment selection unit selects only comments of users having a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction; an output step in which the output unit outputs only the comments selected by the comment selection unit; and an analysis step in which the analysis unit determines the latent attribute identifier of each user using two or more aggregation amounts calculated in advance using all of the user's answers to three or more predetermined questions, or determines two or more aggregation amount prediction values, which are predicted values when the aggregation amount is predicted from image data of the user using a supervised learning technique in machine learning.
[0023] According to this embodiment, it is possible to provide a comment output method that utilizes the above-mentioned latent attributes of a user, allowing the user to appropriately view comments such as word-of-mouth from others who are similar to the user.
[0024] An eighth aspect of the present invention is a comment output program that causes a computer to function as a receiving unit that receives a comment output instruction, a comment acquisition unit that acquires one or more comments together with the user identifier and latent attribute identifier of the user who made the comment, a comment selection unit that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction, an output unit that outputs only the comments selected by the comment selection unit, and an analysis unit that determines the latent attribute identifier of each user using two or more aggregation amounts that are calculated in advance using all of the user's answers to three or more predetermined questions, or an analysis unit that determines two or more aggregation amount prediction values, which are predicted values when the aggregation amount is predicted from the user's image data using a supervised learning method in machine learning.
[0025] According to this aspect, it is possible to provide a program for causing a computer to function as a comment output device that allows a user to appropriately view comments such as word-of-mouth comments from others who are similar to the user. [Effects of the Invention]
[0026] According to one aspect of the present invention, a comment output device can be provided that allows a user to appropriately view comments such as reviews by others similar to the user, using a latent attribute of the user that is determined using at least two aggregation quantities calculated using all of the user's answers to three or more predetermined questions. According to another aspect of the present invention, a comment output device can be provided that calculates a predicted value of the aggregation quantity of the user from an image of the user using machine learning, determines a latent attribute of the user from the thus obtained predicted values of at least two aggregation quantities, and uses the latent attribute to allow a user to appropriately view comments such as reviews by others similar to the user. According to yet another aspect of the present invention, a comment output method can be provided that allows a user to appropriately view comments such as reviews by others similar to the user. According to yet another aspect of the present invention, a comment output program can be provided that causes a computer to function as the comment output device. [Brief explanation of the drawings]
[0027] [Figure 1] FIG. 1 is an explanatory diagram showing the configuration of a comment output device and a terminal device connected to the comment output device via a network. [Figure 2] FIG. 2 is an explanatory diagram showing the configuration of a comment output device, and a terminal device and an analysis device connected to the comment output device via a network. [Figure 3] FIG. 3 is an explanatory diagram showing five areas A1, A2, . . . , A5 arranged on the skin type discrimination surface. [Figure 4] FIG. 4 is an explanatory diagram showing 25 rectangular regions arranged on a skin type discrimination surface and one region As having a curved boundary. [Figure 5]In FIG. 5, diagram (5A) is an explanatory diagram showing 16 rectangular regions arranged on a skin type identification surface and one region As having a curved boundary, and diagram (5B) is an explanatory diagram showing that two rectangular regions may intersect (overlap). [Figure 6] FIG. 6 is an explanatory diagram showing an example of n regions arranged on a discrimination surface. [Figure 7] FIG. 7 is a flow diagram showing an example of the flow of comment output. [Figure 8] FIG. 8 is a flow diagram showing an example of the flow of the medical interview. [Figure 9] FIG. 9 is a flow diagram showing an example of the flow of image diagnosis. [Figure 10] FIG. 10 is a flow diagram showing an example of the flow of a medical interview in a configuration in which an analysis device exists separately from a comment output device. [Figure 11] FIG. 11 is a flow diagram showing an example of the flow of image diagnosis in a configuration in which an analysis device exists separately from a comment output device. [Figure 12] FIG. 12 is an explanatory diagram showing four rectangular regions arranged on the skin type identifying surface. [Figure 13] FIG. 13 is a contour map showing the distribution of the proportion of people who are aware of having sensitive skin in terms of skin type discrimination. [Figure 14] FIG. 14 is a configuration diagram of a neural network having an input layer consisting of d neurons, an intermediate layer (hidden layer), and an output layer consisting of one neuron. [Figure 15] FIG. 15 is a contour map showing the distribution of the percentage of cleansing (rinse-off type) users in terms of makeup awareness discrimination. [Figure 16] In FIG. 16, FIG. 16A is a table showing the distribution of the percentage of people who are concerned about "scalp and hair" in terms of circulation discrimination, and FIG. 16B is a table showing the distribution of the percentage of people who are concerned about "diet" in terms of circulation discrimination. [Figure 17]In Figure 17, Figure (17A) is a table showing the distribution of the percentage of people who are concerned about "menopausal troubles" in terms of vague symptom identification, and Figure (17B) is a table showing the distribution of the percentage of people who are concerned about "menstrual pain" in terms of vague symptom identification. [Figure 18] FIG. 18 is a contour map showing the distribution of the percentage of people who are concerned about "circulatory problems" in terms of identifying irregular eating habits and unhealthy eating habits. [Figure 19] FIG. 19 is a table showing the distribution of the percentage of people who are aware of having sensitive skin in the oily / dry skin discrimination area. [Figure 20] FIG. 20 is a table showing the distribution of the percentage of people who are aware of sensitive skin in the CF oily / dry discrimination area. DETAILED DESCRIPTION OF THE INVENTION
[0028] (motivation) For example, consider a method for selecting and displaying only useful comments from other users to a user visiting a web review bulletin board about cosmetics and beauty services. In the past, in such cases, it was common to use user "segmentation." "Segmentation" refers to classifying users based on overt information (overt attributes) such as age, occupation, income, family structure, and cosmetics used. For example, a conventional method would be to select and display only comments from other users who belong to the same category as the visiting user based on overt information. However, segmentation based on explicit information has the drawback of not taking into account latent information (latent attributes) that a user is unaware of. For example, the present inventor knows that users with "oily and dry skin" are at high risk of having sensitive skin. Even if a user is aware that they have "oily and dry skin," they may actually have "sensitive skin" but be unaware of it. For such users, having "sensitive skin" is latent information that is not conscious of them. If the user could see comments from other users with "sensitive skin" or read descriptions of beauty products for "sensitive skin," this would likely encourage the user to realize that they have sensitive skin and lead to insights, which would be beneficial to the user.
[0029] Latent information generally refers to information such as a user's personality (sensitivity), preferences, and background. Both explicit and implicit information about a user are collectively called "persona information." Classifying users based on "persona information" is called "personalization." For example, if only comments from other users who belong to the same classification as the visiting user are selected and displayed based on user classification through personalization, the visiting user's experience using a review bulletin board will be improved. User explicit information can be collected relatively easily by means of questionnaires and the like. However, collecting user persona information is not easy because it requires the collector to be highly skilled. The feature of the present invention is that, instead of collecting persona information by a skilled collector, the explicit information of users, which is big data, is statistically analyzed to derive persona information for each user, and the persona information is then used to The information is used to classify users, and the classification is used in a comment output system, etc.
[0030] Next, embodiments of a comment output device, a comment output method, and a comment output program according to the present invention will be described in detail with reference to the drawings.
[0031] (Comment Output Device 1) 1, a comment output device 1 according to one embodiment of the present invention is configured to be connectable to a network 4 capable of two-way communication, such as the Internet or an intranet, and is used together with a terminal device 2 also configured to be connectable to the network 4. In response to a comment output instruction from a user of the terminal device 2, which is a communication device such as a smartphone, tablet terminal, or personal computer, the comment output device 1 outputs (transmits) a comment to the terminal device 2, and the output comment is displayed on a terminal output unit 22, such as a display of the terminal device 1. The comment output device 1 includes a receiving unit 12 that receives a comment output instruction, a comment acquisition unit 14 that acquires one or more comments along with the user identifier and latent attribute identifier of the user who posted the comment, a comment selection unit 15 that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction, an output unit 11 that outputs only the comments selected by the comment selection unit 15, a transmission unit 13 that transmits the output comments to the terminal device 2, an analysis unit 17 that determines each user's latent attribute identifier using two or more aggregate amounts that are calculated in advance using all of the user's answers to three or more predetermined questions, and a storage unit 16 that stores comments, each user's user identifier, latent attribute identifier, explicit attributes such as the user's answers to the questions, and other information. Note that in this specification, "identifier" may be referred to as "ID," "user identifier" as "user ID," and "latent attribute identifier" as "latent attribute ID," but the former are synonymous with the latter.
[0032] (Terminal Device) The terminal device 2 has a terminal input unit 21 such as a keyboard, mouse, touch panel, and smart glasses, goggles, watch, etc. for receiving user instructions to output comments, a terminal transmitting unit 24 for transmitting user instructions to output comments to the comment output device 1, a terminal receiving unit 23 for receiving comments transmitted from the comment output device 1, and a terminal output unit 22 such as a display for displaying comments received from the output device 1.
[0033] (Comment Output Device 1') 2, a comment output device 1′ according to another embodiment of the present invention is configured to be connectable to a bidirectional communication network 4 such as the Internet or an intranet, and is used together with a terminal device 2 and an analysis device 3 also configured to be connectable to the network 4. In response to a comment output instruction from a user using the terminal device 2, which is a communication device such as a smartphone, a tablet terminal, or a personal computer, the comment output device 1′ outputs (transmits) a comment to the terminal device 2, and the output comment is displayed on a terminal output unit 22 such as a display of the terminal device 1. Unlike the comment output device 1 described above, the comment output device 1′ according to this embodiment is characterized in that, unlike the comment output device 1, the comment output device 1′ has an analysis unit that determines each user's latent attribute identifier using two or more aggregate amounts calculated in advance using all of the user's answers to three or more predetermined questions, but instead transfers the task of determining each user's latent attribute identifier to the analysis device 3 via the network 4. The comment output device 1' includes a receiving unit 12 that receives a comment output instruction, a comment acquiring unit 14 that acquires one or more comments together with the user identifier and latent attribute identifier of the user who made the comment, and a latent attribute identifier that matches the latent attribute identifier of the user who sent the output instruction. The comment output device 1′ includes a comment selection unit 15 that selects only comments from users with an attribute identifier; an output unit 11 that outputs only the comments selected by the comment selection unit 15; a transmission unit 13 that transmits the output comments to the terminal device 2; a memory unit 16 that stores comments, each user's user identifier, latent attribute identifier, explicit attributes such as the user's answers to the questions, and other information; a manifest attribute transmission unit 18; and a latent attribute receiving unit 19. The manifest attribute transmission unit 18 transmits explicit attributes such as the user's answers to the questions stored in the memory unit 16 to the analysis device 3 along with the user identifier of the user. The analysis device 3 determines the user's latent attribute based on the received explicit attributes of the user, and transmits the determined latent attribute identifier to the comment output device 1′. The latent attribute receiving unit 19 of the comment output device 1′ receives the user's latent attribute identifier transmitted from the analysis device 3 along with the user identifier of the user, and stores them in the memory unit 16.
[0034] The terminal device 2 used together with the comment output device 1' of this embodiment has the same configuration as the terminal device 2 used together with the comment output device 1 described above.
[0035] (Analyzer 3) The analysis device 3 has a manifest attribute receiving unit 33 that receives manifest attributes such as a user's answers from the comment output device 1' along with the user's user identifier and stores them in a memory unit 31 described below, a memory unit 31 that stores the user's manifest attributes, etc., an analysis unit 32 that determines each user's latent attribute identifier using two or more aggregate amounts that are calculated in advance using all of the user's answers to three or more predetermined questions, and a latent attribute sending unit 34 that sends each user's latent attribute identifier to the comment output device 1' along with the user's user identifier.
[0036] (A variant of extracting latent attributes from images) The analysis unit 17 of the comment output device 1 and the analysis unit 32 of the analysis device 3 determine the latent attribute identifier of each user in advance using two or more aggregate amounts calculated using all of the user's answers to three or more predetermined questions. In each of these modified forms, the analysis unit 17 of the comment output device 1 or the analysis unit 32 of the analysis device 3 can determine the latent attribute identifier of each user in advance using two or more aggregate amount prediction values, which are prediction values when the aggregate amount calculated using all of the user's answers to three or more predetermined questions is predicted from the user's image data using a supervised learning technique in machine learning. Note that the user's image data is included in the user's manifest attributes.
[0037] (Correspondence between areas on the skin type identification surface and latent attribute ID) FIG. 3 is an explanatory diagram showing an embodiment in which a latent attribute identifier for a user is determined using two aggregated quantities calculated using all of the user's answers to three or more predetermined questions. This embodiment is an example of the application of the present invention to the fields of skin and beauty. A questionnaire consisting of three or more questions about skin and beauty is used to conduct a medical interview with a large number of subjects in advance, and the answers are used as basic data. The answers to each question are expressed as real numbers, such as a binary choice of "0 or 1" or a multiple-choice answer of "0, 1, 2, or 3." The basic data is subjected to factor analysis to extract an O factor, which represents the degree of oiliness, and a D factor, which represents the degree of dryness. The user's factor scores for the O factor and the D factor can be calculated from the user's answers to the same questionnaire. In Figure 3, the factor score of the user's O factor is converted by a non-decreasing function and normalized to a value between 0 and 1, and this is called the "oiliness level SO" of the user. The factor score of the user's D factor is converted by a non-decreasing function and normalized to a value between 0 and 1, and this is called the "dryness level SO" of the user. D In Figure 3, the coordinate plane with oiliness on the horizontal axis and dryness on the vertical axis is called the "skin type identification plane." The user's oiliness and dryness are represented by the coordinates (S0,S D) is expressed as a single point. In other words, the user's explicit attributes are expressed as two aggregate quantities related to the user, The oiliness level (first aggregate amount in this embodiment) and dryness level (second aggregate amount in this embodiment) are calculated, and the latent attribute of the user is represented by one point on the skin type discrimination surface. In Figure 3, the latent attribute of a certain user with an oiliness level of 0.80 and a dryness level of 0.55 is represented by point P on the skin type discrimination surface.
[0038] In the embodiment shown in FIG. 3, five regions A1, A2, A3, A4, and A5 are arranged on the skin type identification surface. These regions correspond to latent attribute IDs 1, 2, 3, 4, and 5, respectively. Of these, the four regions other than A5 have a rectangular (or square) shape, and the boundary between two adjacent regions belongs to the left or lower region. Part of the boundary of region A5 is formed by a curve, and it intersects (overlaps) with each of regions A2, A3, and A4. In the case of the user whose latent attribute is indicated by point P, point P belongs to both region A2 and region A5. Therefore, the latent attribute identifier of the user is "2,5."
[0039] [Table 1]
[0040] Table 1 shows the correspondence between latent attribute identifiers (latent attribute IDs), latent attribute descriptions, and regions on the skin type identification surface in the embodiment shown in Figure 3. For example, region A1 with a latent attribute ID of 1 represents "normal skin." Region A2 with a latent attribute ID of 2 represents "oily skin." Region A3 with a latent attribute ID of 3 represents "dry skin." Region A4 with a latent attribute ID of 4 represents "oily & dry skin." Region A5 with a latent attribute ID of 5 represents "high risk of sensitive skin" and overlaps with the other three regions A2, A3, and A4. For example, the user whose latent attributes are indicated by point P in FIG. 3 has "oily skin" and "high risk of sensitive skin."
[0041] (Embodiments including a 5x5 grid area) FIG. 4 is an explanatory diagram illustrating how to determine a user's latent attribute ID in an embodiment that is a modification of the embodiment shown in FIG. 3 . In this embodiment, the user's oiliness level (first aggregate amount in this modification) is determined by converting the factor scores of the oiliness factor into a real number between 0 and 5 using a non-decreasing function, and similarly, the dryness level (second aggregate amount in this modification) is determined by converting the factor scores of the dryness factor into a real number between 0 and 5 using a non-decreasing function. A grid-like area consisting of 5 × 5 = 25 squares is arranged on a skin type discrimination surface, which is a coordinate plane with oiliness level on the horizontal axis and dryness level on the vertical axis. The scale on the horizontal axis is represented by integers with the decimal point of the oiliness level rounded up. For example, a scale of 1 is assigned to a portion where the oiliness level is between 0 and 1 but not exceeding 1, a scale of 2 is assigned to a portion where the oiliness level is between 1 and 2 but not exceeding 1, and a scale of 5 is assigned to a portion where the oiliness level is between 4 and 5 but not exceeding 5. The same applies to the scale on the vertical axis related to dryness level. For example, the area corresponding to the square where the horizontal axis for oiliness is 4 and the vertical axis for dryness is 3 is denoted as area A04d3, and the area corresponding to the horizontal axis for oiliness is 4 and the vertical axis for dryness is 3 is denoted as area A04d3. Note that the line segment at the boundary between two adjacent areas belongs to the left or lower of the two areas. Furthermore, on the skin type identification surface, there is arranged area As which has overlapping portions with nine of the above 25 square areas, such as area A05d5, and whose boundary is partially composed of curved lines.
[0042] In the embodiment shown in Figure 4, the latent attribute IDs increase in number from the bottom left to the right and upwards. For example, the latent attribute ID corresponding to area A01d1 is 1, the latent attribute ID corresponding to area A02d1 is 2, the latent attribute ID corresponding to area A05d1 is 5, and generally, the latent attribute ID corresponding to area A0idj (where i and j are integers between 1 and 5) is i + 5 × (j - 1). Also, the latent attribute ID corresponding to area As, which means "high risk of sensitive skin," is 26.
[0043] In the embodiment shown in FIG. 4, two users who have very similar oiliness and dryness levels may be assigned different latent attribute IDs. For example, suppose two users P and Q have exactly the same oiliness level of 1.5, and their dryness levels are 2.9 and 3.1, respectively. On the skin type identification surface, user P belongs to area A02d3, so his latent attribute ID is 12. However, user Q belongs to area A02d4, so his latent attribute ID is 17. The latent attribute IDs of users P and Q are different. Because their latent attribute IDs are different, in the comment output device of this embodiment, these two users will never have the opportunity to see each other's comments. This is not desirable.
[0044] (Another embodiment including a 5x5 grid area) FIG. 5 is an explanatory diagram illustrating a method for determining a user's latent attribute ID in yet another embodiment. This embodiment is a further modification of the embodiment shown in FIG. 4 and solves the above-mentioned problem. This embodiment is similar to the embodiment shown in FIG. 4 in that 25 squares are considered on the skin type identification surface, but the method for determining the regions is different. As shown in FIG. 5A, in this embodiment, four adjacent squares are combined into one region. That is, for example, the region where the horizontal axis for oiliness is 1 or 2 and the vertical axis for dryness is 2 or 3 is defined as region A012d23 (corresponding latent attribute ID is 5), and the region where the horizontal axis for oiliness is 4 or 5 and the vertical axis for dryness is 1 or 2 is defined as region A045d12 (corresponding latent attribute ID is 4). In general, where a and b are both integers between 1 and 4, the region where the horizontal axis scale relating to oiliness is a or a+1 and the vertical axis scale relating to dryness is b or b+1 is defined as "Ao(a,a+1)d(b,b+1)" (the corresponding latent attribute ID is a+(b-1)×4). The latent attribute ID corresponding to region As, which means "high risk of sensitive skin," is 17. Of the 25 squares, 9 of the inner squares belong to four or more regions, the remaining 12 squares on the four sides belong to two or more regions, and one of the remaining four squares in the four corners belongs to two regions. Therefore, the majority of the squares belong to multiple regions. For example, in the example shown in Figure (5B), area Aо12d23 and area Aо23d12 share one square and intersect (overlap). In this embodiment, it is highly possible to avoid the occurrence of an undesirable situation in which "two users who have very similar oiliness and dryness levels do not have an opportunity to see each other's comments."
[0045] (more common form) FIG. 6 is an explanatory diagram illustrating how to determine a user's latent attribute ID in one embodiment of the present invention. The present invention is applicable not only to the fields of skin and beauty but also to general fields, and this embodiment is an example of such a case. A coordinate plane with the first aggregate amount obtained from answers to three or more questions on the horizontal axis and the second aggregate amount on the vertical axis is used as an identification surface, and multiple areas A1, A2,...,An (where n is a natural number greater than or equal to 2) are arranged on the identification surface, and one latent attribute ID is assigned to each area. That is, for j = 1, 2,...,n, area Aj is assigned j as the latent attribute ID. Here, it is preferable that the union of the multiple areas covers the range of possible values that the pair of the first aggregate amount and the second aggregate amount can take on on the identification surface. Furthermore, from the viewpoint of allowing a user to selectively view comments by other users with similar latent attributes to themselves and enabling the selection process of such comments to be performed quickly, the multiple areas are used. The regions may share portions with each other, and it is preferable that at least one of the regions share portions with one or more other regions, and it is more preferable that all of the regions share portions with one or more other regions. When a first aggregation amount of a certain user is T1 and a second aggregation amount is T2, the latent attribute identifier corresponding to the region including point (T1, T2) among the plurality of regions arranged on the discrimination surface is determined as the latent attribute ID of the user. For example, on the discrimination surface shown in Figure 6, the latent attribute ID of the user whose pair of values of the first aggregation amount and the second aggregation amount is represented by point Q is "1, 2".
[0046] In one embodiment of the present invention shown in Figure 6, each of the multiple regions can be arranged to include n points C1, C2,..., Cn arranged on the discrimination surface. For example, if r1, r2,..., rn are n positive real numbers, region A1 has a "distance" from point C1 of Region A is the set of points whose "distance" from point C is less than or equal to r1, region A2 is the set of points whose "distance" from point C is less than or equal to r2, and generally, for j = 1, 2,..., n, region Aj can be defined as the set of points whose "distance" from point Cj is less than or equal to rj. Defining regions using "distance" from a single point makes it easier to determine the latent attribute ID for each user. Here, the "distance" between two points on the discrimination surface may be defined as (1) the L2 distance (normal Euclidean distance), i.e., the square root of the sum of the square of the difference between the ordinates and the square of the difference between the abscissas; (2) the L1 distance, i.e., the sum of the absolute value of the difference between the ordinates and the absolute value of the difference between the abscissas; (3) the Lp distance, i.e., the pth root of the sum of the pth power of the absolute value of the difference between the ordinates and the pth power of the absolute value of the difference between the abscissas, where p is a real number greater than or equal to 1; (4) the L∞ distance, i.e., the greater of the absolute value of the difference between the ordinates and the absolute value of the difference between the abscissas; or (5) any distance, i.e., any distance defined so that the discrimination surface satisfies the axioms of metric space (non-degeneracy, symmetry, and triangle inequality).
[0047] (Comment output method) 7 is a flow diagram of an embodiment of a comment output method according to the present invention. When comment output starts, the terminal input unit 21 of the terminal device 2 is on standby to receive input from the user. In step S11, the terminal input unit 21 of the terminal device 2 receives an instruction from the user to output a comment. In step S12, the terminal transmission unit 24 of the terminal device 2 transmits the output instruction together with the user ID of the user to the comment output device 1. In step S13, the reception unit 12 of the comment output device 1 receives the output instruction and the user ID of the user. In step S14, the comment output device 1 checks whether the user's latent attribute has been registered in the storage unit 16. If it has been registered, the process proceeds to step S15; if it has not been registered, the process proceeds to step S2.
[0048] If the latent attribute has been registered, first, in step S15, the comment acquisition unit 14 of the comment output device 1 acquires one or more comments together with the user ID of the user who posted the comment. Next, in step S16, the comment selection unit 15 of the comment output device 1 selects only comments from users having latent attribute IDs that match the latent attribute ID of the user who sent the output instruction. Next, in step 17, the output unit 11 of the comment output device 1 outputs only the comments selected by the comment selection unit 15, and the transmission unit 13 of the comment output device 1 transmits the output comments to the terminal device 2. Next, the terminal reception unit 23 of the terminal device 2 receives the comments from the comment output device 1. Finally, the terminal output unit 22 of the terminal device 2 displays the received comments, and the comment output ends.
[0049] Below, the steps of the method of the present invention will be explained using the example of constructing a latent attribute ID for a user based on the user's answers to questions about skin and beauty, but the same applies when constructing a latent attribute ID for a user based on the user's answers to general questions, such as questions about health awareness or diet.
[0050] If the latent attribute has not been registered, first, in step S2, the user using the terminal device 2 is asked whether or not to take a skin diagnosis questionnaire. That is, a question about whether or not to take a diagnosis questionnaire is sent from the transmission unit 13 of the comment output device 1 to the terminal device 2, the terminal reception unit 23 of the terminal device 2 receives the question, and the terminal output unit 22 of the terminal device 2 displays the question. The terminal input unit 21 of the terminal device 2 accepts a reply from the user, and the terminal transmission unit 24 of the terminal device 2 transmits the reply from the user to the comment output device 1. The reception unit 12 of the comment output device 1 receives the reply from the user. If the reply is YES, the process proceeds to step S20 (skin diagnosis), and if NO, the process proceeds to step S3.
[0051] If the latent attribute has not been registered and the user does not wish to undergo a skin diagnosis interview, first, in step S3, the user using the terminal device 2 is asked whether or not to undergo an image diagnosis. That is, a question about whether or not to undergo an image diagnosis is sent from the transmission unit 13 of the comment output device 1 to the terminal device 2, the terminal reception unit 23 of the terminal device 2 receives the question, and the terminal output unit 22 of the terminal device 2 displays the question. The terminal input unit 21 of the terminal device 2 accepts a reply from the user, and the terminal transmission unit 24 of the terminal device 2 transmits the reply from the user to the comment output device 1. The reception unit 12 of the comment output device 1 receives the reply from the user. If the reply is YES, the process proceeds to step S30 (image diagnosis), and if NO, the process proceeds to step S18.
[0052] If the latent attribute has not been registered and the user does not undergo either the skin diagnosis interview or the image diagnosis, in step S18, a comment for the user who has not registered the latent attribute is output, and the comment output ends.
[0053] (Skin diagnosis flow (when the comment output device has an analysis unit)) If the user answers YES to the question of whether or not to accept the skin diagnosis questionnaire, a skin diagnosis questionnaire process (step S20) is performed. As shown in FIG. 8 , when the comment output device 1 has the analysis unit 17, in the skin diagnosis questionnaire process S20, first, in step S21, the transmission unit 13 of the comment output device 1 transmits questionnaire data to the terminal device 2 of the user. Next, in step S22, the terminal reception unit 23 of the terminal device 2 receives the questionnaire data. Next, in step S23, the terminal output unit 22 of the terminal device 2 displays the questionnaire. Next, in step S24, the terminal input unit 21 of the terminal device 2 acquires the user's answers to the questions in the questionnaire. Next, in step S25, the terminal transmission unit 24 of the terminal device 2 transmits the acquired answers of the user to the comment output device 1. Next, in step 26, the reception unit 12 of the comment output device 1 receives the user's answers and stores them in the storage unit 16. Next, in step S27, the analysis unit 17 of the comment output device 1 obtains at least two aggregation quantities by statistical analysis based on the received answers from the user. Next, in step S28, the analysis unit 17 determines the user's latent attribute ID based on the at least two aggregation quantities. Finally, in step S29, the analysis unit 17 stores the determined latent attribute ID of the user in the storage unit 16, and the skin diagnosis questioning process S20 is completed.
[0054] (Image diagnosis flow (when the comment output device has an analysis unit)) If the user answers YES to the question of whether or not to undergo diagnostic imaging, the diagnostic imaging process (step S30) is carried out. As shown in FIG. 9, when the comment output device 1 has the analysis unit 17, in the diagnostic imaging process S30, first, in step S31, the transmission unit 13 of the comment output device 1 transmits image upload screen data to the terminal device 2 of the user. Next, in step S32, the terminal reception unit 23 of the terminal device 2 receives the image upload screen data. Next, in step S33, the terminal output unit 22 of the terminal device 2 displays the image upload screen. Next, in step S34, the terminal input unit 21 of the terminal device 2 accepts the user's input for image upload. Next, in step S35, the terminal transmission unit 24 of the terminal device 2 transmits the user's image data to the comment output device 1. The comment output device 1 transmits the image data of the user to the comment output device 1. Then, in step S36, the receiving unit 12 of the comment output device 1 receives the image data of the user and stores it in the storage unit 16. Then, in step S37, the analysis unit 17 of the comment output device 1 calculates at least two aggregation amount prediction values based on the received image data of the user. Then, in step S38, the analysis unit 17 determines the latent attribute ID of the user based on the at least two aggregation amount prediction values. Finally, in step S39, the analysis unit 17 stores the determined latent attribute ID of the user in the storage unit 16, and the image diagnosis interview process S30 is completed. In this specification, "image data" of a user refers to, for example, image data of the user's face, hands, feet, arms, legs, abdomen, back, hair, tongue, gums, eyes, etc.
[0055] (Skin diagnosis flow (when there is an analysis device separate from the comment output device)) If the user answers YES to the question of whether or not to take the skin diagnosis questionnaire, the skin diagnosis questionnaire process (step S20) is performed. As shown in FIG. 10 , when an analysis device 3 is provided in addition to the comment output device 1, in the skin diagnosis questionnaire process S20, first, in step S210, the transmission unit 13 of the comment output device 1 transmits questionnaire data to the terminal device 2 of the user. Next, in step S220, the terminal reception unit 23 of the terminal device 2 receives the questionnaire data. Next, in step S230, the terminal output unit 22 of the terminal device 2 displays the questionnaire. Next, in step S240, the terminal input unit 21 of the terminal device 2 acquires the user's answers to the questions in the questionnaire. Next, in step S250, the terminal transmission unit 24 of the terminal device 2 transmits the acquired user's answers to the comment output device 1. Next, in step 260, the reception unit 12 of the comment output device 1 receives the user's answers and stores them in the storage unit 16. Next, in step S261, the explicit attribute sending unit 18 of the comment output device 1 sends the user's answer to the analysis device 3. Next, in step S262, the explicit attribute receiving unit 33 of the analysis device 3 receives the user's answer (explicit attribute). Next, in step S270, the analysis unit 32 of the analysis device 3 obtains at least two aggregate quantities by statistical analysis based on the received user's answer. Next, in step S280, the analysis unit 32 determines the user's latent attribute ID based on the at least two aggregate quantities. Next, in step S281, the latent attribute sending unit 34 of the analysis device 3 sends the determined latent attribute ID of the user to the comment output device 1. Finally, in step S290, the latent attribute receiving unit 19 of the comment output device 1 receives the user's latent attribute ID and stores it in the memory unit 16, thereby completing the skin diagnosis questioning process S20.
[0056] (Image diagnosis flow (when there is an analysis device separate from the comment output device)) If the user answers YES to the question of whether or not to undergo diagnostic imaging, the diagnostic imaging process (step S30) is performed. As shown in FIG. 11 , when an analysis device 3 is provided in addition to the comment output device 1, the diagnostic imaging process S30 begins with step S310, in which the transmitting unit 13 of the comment output device 1 transmits image upload screen data to the user's terminal device 2. Next, in step S320, the terminal receiving unit 23 of the terminal device 2 receives the image upload screen data. Next, in step S330, the terminal output unit 22 of the terminal device 2 displays the image upload screen. Next, in step S340, the terminal input unit 21 of the terminal device 2 accepts the user's input for image upload. Next, in step S350, the terminal transmitting unit 24 of the terminal device 2 transmits the user's image data to the comment output device 1. Next, in step S360, the receiving unit 12 of the comment output device 1 receives the user's image data and stores it in the storage unit 16. Next, in step S361, the explicit attribute sending unit 18 of the comment output device 1 sends the image data of the user to the analysis device 3. Next, in step S362, the explicit attribute receiving unit 33 of the analysis device 3 receives the image data (explicit attributes) of the user. Next, in step S370, the analysis unit 32 of the analysis device 3 calculates at least two predicted aggregation amounts based on the received image data of the user. Next, in step S380, the analysis unit 32 determines the latent attribute ID of the user based on the at least two predicted aggregation amounts. Next, in step S38 In step S391, the latent attribute sending unit 34 of the analysis device 3 sends the determined latent attribute ID of the user to the comment output device 1. Finally, in step S390, the latent attribute receiving unit 19 of the comment output device 1 receives the latent attribute ID of the user and stores it in the memory unit 16, completing the image diagnosis process S30.
[0057] [Table 2]
[0058] (Storage of information related to comments in storage unit) Information related to comments made by users is stored in a readable form in the storage unit 16 of the comment output device 1 together with the user identifier of the user. The storage method is not limited, but for example, as shown in Table 2, the information related to comments can be stored in a "comment table" in a relational database (RDB). In this case, the columns of the comment table include at least a comment ID, a comment, a target ID, and a user ID. Here, a "comment ID" is an identifier such as a number that uniquely identifies a comment. A "comment" is the body of the comment. A "target ID" is an identifier such as a number that uniquely identifies the target of the comment. For example, in a comment output device related to a cosmetics review site, the target is an individual cosmetic product. Furthermore, if the target ID of a certain comment is zero, it means that the comment is a general comment that does not specify a target. A "user ID" is the user identifier of the user who made the comment.
[0059] [Table 3]
[0060] (Storage of information related to the target in the memory unit) Information related to the subject of the comment is stored in advance in a readable form in the storage unit 16 of the comment output device 1, together with the subject's name (e.g., product name), subject's description, and preferably subject's image data. The storage method is not limited, but for example, as shown in Table 3, Comments can be stored in an "object table" in an RDB. In Table 3, the columns of the object table include an object ID, an object name (e.g., product name), an object description, and preferably image data of the object. Here, the "object ID" is an identifier such as a number that uniquely specifies the object. The "object name" is, for example, the name of a cosmetic product (product name). The "object description" is, for example, a description of the cosmetic product. The "object image data" is, for example, image data of a photo of the cosmetic product's exterior, or the URL of the image data.
[0061] [Table 4]
[0062] (Storage of user-related information in storage unit) Information related to the user is stored in advance in a readable form in the storage unit 16 of the comment output device 1 together with the user ID, user name, latent attribute ID, etc. The method of storage is not limited, but for example, as shown in Table 4, comments can be stored in a "user table" in an RDB. In Table 4, the columns of the user table include the user ID, user name, latent attribute ID, etc. Here, the "user ID" is an identifier such as a number that uniquely specifies a user. The "user name" is, for example, the user's name or handle name. The "latent attribute ID" is the user's latent attribute ID. In addition to the above, the columns of the user table can include explicit attributes such as gender, address, occupation, hobbies, age, and family composition. In another embodiment of the present invention, the memory unit may be configured to store a user's explicit attributes and latent attributes in multiple tables, and when reading, to link the multiple tables to read the user's explicit attributes and / or latent attributes.
[0063] (A method for storing latent attribute IDs that can be read quickly) In the present invention, overlapping latent attribute IDs are permitted for users. For example, in Table 4, the latent attribute IDs of a user with a user ID of 2 are 4 and 5, and the latent attribute IDs of a user with a user ID of 4 are 2 and 5. To store overlapping latent attribute IDs in a single column of an RDB, binary representation (hereinafter, represented as, for example,
[0101] = 5) is used. For example, if a user's latent attribute ID is 3, it is represented as
[0100] = 4; if another user's latent attribute IDs are 4 and 5, it is represented as
[11000] = 24; and if yet another user's latent attribute IDs are 2 and 5, it is represented as
[10010] = 18. Generally, if the latent attribute IDs are a and b, they are represented in binary representation as [the a-th and b-th digits from the right are 1, and the remaining digits are 0], and the value is represented as 2^(a-1) + 2^(b-1). By expressing latent attribute IDs in binary notation in this way, it is possible to quickly select comments from other users that have latent attribute IDs that match at least one of the latent attribute IDs of the user who issued the comment output instruction. This is because an RDB can quickly execute SQL statements that have the result of a bitwise AND operation as a conditional expression. In the technology described in Patent Document 1, comments by users who use cosmetics that overlap with the user who issued the comment output instruction are first selected, and then the similarity between the users is determined based on the content of the comments. Narrowing down the comments by cosmetics used before determining the similarity takes time. In the present invention, as described above, areas corresponding to each of multiple latent attribute IDs overlap with each other on the discrimination surface. As described above, each latent attribute ID is defined, and the process of selecting only comments by users that have the same latent attribute ID as the user's latent attribute ID can be performed quickly by bitwise AND operation. Therefore, without having to narrow down the comments in advance based on the cosmetics used, it is possible to perform a selection process (i.e., execute an SQL statement) on all comments at high speed, and select and output only comments by users that have the same latent attribute ID as the user's latent attribute ID.
[0064] [Table 5]
[0065] (Storage of evaluation information in storage unit) A "rating" refers to a user's rating of an object, and is expressed as a real number such as "0 or 1" or "an integer between 0 and 4." Rating information, along with a user ID and an object ID, is stored in a readable form in the storage unit 16 of the comment output device 1. While the storage method is not limited, for example, comments can be stored in a "rating table" in an RDB, as shown in Table 5. In Table 5, the columns of the rating table include rating ID, user ID, object ID, rating value, etc. Here, "rating ID" is an identifier such as a number that uniquely identifies a rating. "User ID" is an identifier such as a number that uniquely identifies a user. "Object ID" is an identifier such as a number that uniquely identifies an object. Storing information related to "rating" in the storage unit allows, when outputting a comment about an object, the comment to be output along with the user's rating value for the object, and allows the user reading the comment to change the order of comments to be output based on the rating value or select whether to output comments based on the rating value, depending on the comment output format settings of the user reading the comment. This allows a variety of options for the comment output format to be provided to the user reading the comment.
[0066] <Example 1> (Rectangular area) This Example 1 is an example in which latent attribute IDs are defined using rectangular regions on the skin type identification surface. It is confirmed that the risk of sensitive skin differs for each of the multiple rectangular regions. This serves as supporting evidence that the latent attributes of the user can be grasped using at least two aggregated values obtained from the user's answers (manifest attributes) to three or more questions.
[0067] Using a questionnaire containing 21 questions about skin and beauty, approximately 6,300 female subjects aged 20 to 69 were interviewed, and valid responses were obtained from 6,069. Each question was in the form of "Are you concerned about ~?", where "~" is the part that indicates "Poor makeup application," "Fine wrinkles," "Acne," "Dryness," "Rough skin," "Clean pores," "Acne scars," "Makeup smudges," "Dark circles," "Rough skin texture," "Prevents dryness," "Firms," "Moisturizes," "Blotches," "Sensitive / hypersensitive skin," "Firmness / elasticity / sagging," "Brightness / transparency," "Texture," "Pores," "Texture," and "Youthfulness" These valid responses were used as the basic data for statistical analysis using factor analysis. The responses were binary answers (YES or NO, i.e., 1 or 0) to the 21 questions above. The number of factors to be extracted was 5. This is because 1 out of the eigenvalues of the correlation matrix of the above responses This is because there are five factors that satisfy the above criteria. Maximum likelihood method was used to estimate the factor loads, and factors were rotated using the varimax method. Description of factor loads is omitted. Regression method was used to calculate factor scores. The score for each factor for each user or subject can be calculated by calculating the product of the answer value (1 or 0) of the user or subject, normalized to a mean of 0 and a standard deviation of 1, and the corresponding factor score coefficient for each question, and then adding up these values for the 21 questions.
[0068] [Table 6]
[0069] Table 6 shows the factor score coefficients obtained by the statistical analysis using the above-mentioned factor analysis method. The five factors are referred to as factors B1 to B5, in order. Factor B1 is a factor that is highly related to questions about aesthetics, such as "texture," "pores," "feel," and "brightness and transparency." Factor B2 is a factor that is highly related to questions related to aging, such as "firmness, elasticity, sagging," "youthfulness," and "fine wrinkles." Factor B3 is a factor that is highly related to questions related to oiliness, such as "acne" and "acne scars." Factor B4 is highly correlated with questions about makeup dissatisfaction (dissatisfaction with skin quality), such as "makeup comes off," "makeup doesn't stick well," and "rough skin texture," and is a factor related to base makeup. Factor B5 is a factor that is highly related to questions related to dryness, such as "prevents dryness," "moisturizes," and "dryness." In this example, the factor score of factor B3 is normalized by linear transformation so that the minimum value is 0 and the maximum value is 1, and the normalized value is called "oiliness." Also, the factor score of factor B5 is normalized by linear transformation so that the minimum value is 0 and the maximum value is 1, and the normalized value is called "dryness."
[0070] In addition to the 21 questions above, the subjects were also asked a two-choice question asking "Do you think you have sensitive skin?" (Answer value: 1 for YES, 0 for NO). Figure 12 is a coordinate plane with oiliness as the first aggregate quantity on the horizontal axis and dryness as the second aggregate quantity on the vertical axis. This is a graph showing the oiliness level x3 and dryness level x5 of 300 women randomly selected from the subjects on the skin type discrimination surface shown in Figure 12, represented by points with coordinates (x3, x5). Women who are aware of sensitive skin are plotted with a circle, and women who are not are plotted with an x. The skin type discrimination surface shown in Figure 12 has two lines drawn on it. One line is parallel to the vertical axis representing "oiliness level = 0.31." This line divides the skin type discrimination surface into two regions, left and right. The other line is parallel to the vertical axis representing "dryness level = 0.27." This line divides the skin type discrimination surface into two regions, top and bottom. When the two lines are joined, the skin type discrimination surface is divided into four rectangular regions: lower left, lower right, upper left, and upper right. However, of the four regions, points on the border between two adjacent regions belong to the left and / or bottom regions. The upper right region corresponds to "oily & dry skin," and a higher proportion of women are aware of having sensitive skin compared to the other three regions. Therefore, in this embodiment, if the oiliness and dryness of the subject fall within the upper right region on the skin type identification screen, the subject is predicted to have a "high risk of sensitive skin" (predicted value p=1), and if they do not fall within the upper right region, the subject is predicted to have a "not high risk of sensitive skin" (predicted value p=0). If the subject actually answers "I think I have sensitive skin," the actual value y=1, and if they do not answer in this way, the actual value y=0.
[0071] The proportion of subjects predicted as p=1 who actually have y=1 is called "precision." The proportion of subjects predicted as p=1 who actually have y=1 is called "recall." The harmonic mean of precision and recall is called "f1 score." The f1 score is a good indicator of prediction accuracy, taking into account the balance between increasing the probability of a correct positive prediction and reducing the probability of an incorrect negative prediction. The two lines above, "oilyness = 0.31" and "dryness = 0.27," were chosen to maximize the F1 score of the prediction, while varying the boundary positions of both oiliness and dryness from 0.01 to 0.99 in increments of 0.01. The maximum F1 score was 0.456. The percentage of subjects who self-identified as having sensitive skin was 0.220, and the F1 score of the prediction was approximately twice as high. As shown in Figure 3, placing several rectangular regions on the skin type identification screen, plotting points with coordinates representing the user's oiliness x3 and dryness x5, and considering which rectangular region the plotted points fall within, is useful for identifying latent attributes related to the user's skin type and cosmetic preferences.
[0072] Example 2 (Area with curved boundary) This Example 2 is an example in which the latent attribute ID is defined using a region with a curved boundary on the skin type identification surface. It is confirmed that by using a region with a curved boundary, the latent attribute of the user can be grasped more accurately than when a rectangular region is used. As in Example 1, we define the oiliness level x3 and dryness level x5 for each subject or user. On a skin type identification surface, which is a coordinate plane with oiliness level as the first aggregated quantity on the horizontal axis and dryness level as the second aggregated quantity on the vertical axis, we assume that there is a user whose pair of oiliness level and dryness level is represented by a point (x3, x5). The probability that the user believes that "they have sensitive skin" is considered to vary depending on the position on the skin type identification surface, i.e., the coordinates (x3, x5). This probability is represented as p(x3, x5). Figure 13 is a contour map of the probability p(x3, x5) estimated using the method described below. It can be seen that the probability p(x3, x5) generally increases as the oiliness level x3 increases and as the dryness level x5 increases. In FIG. 13, attention is focused on the contour line with a probability p(x3, x5) of 0.23. The region above this contour line, around and inside the outer square, is designated as As. The region As is a region with a curved boundary. Instead of the upper right rectangular region in Example 1, in Example 2, when the oiliness and dryness of the subject belong to the region As on the skin type discrimination surface, the subject is predicted to have a "high risk of sensitive skin" (prediction value p=1), and when they do not belong to the region As, the subject is predicted to have a "not high risk of sensitive skin" (prediction value p=0). The f1 score of this prediction is 0.518, which is 0.518 higher than the f1 score of the prediction in Example 1. This is larger than the core value of 0.456. That is, Example 2 has a higher prediction accuracy for the risk of sensitive skin than Example 1. It can be seen that using areas with curved boundaries as some of the areas to be arranged on the skin type identification surface is more useful in identifying latent attributes related to the skin type and cosmetic consciousness of the subject than using only rectangular areas. In this embodiment, the user's latent attributes can be expressed as shown in Table 1 using a total of five regions, for example, the four rectangular regions shown in FIG. 3 and the region As. These five regions have overlapping portions. In other words, the latent attribute "high risk of sensitive skin" is not independent of the four latent attributes "normal skin," "oily skin," "dry skin," and "oily & dry skin." For example, it is possible for a user to have both "oily & dry skin" and "high risk of sensitive skin."
[0073] Here, we provide additional information on how to estimate the probability p(x3, x5) and how to determine the probability threshold of 0.23. From the basic data consisting of 6,069 people's responses to 21 questions and whether or not they considered their skin sensitive (y = 1 or y = 0), 1,000 people's responses were randomly selected to serve as test data. Another 1,000 people were randomly selected from the remaining data to serve as CV data (cross-validation data), and the remaining 4,069 people's data served as training data. Consider the fully connected neural network shown in Figure 14, consisting of sigmoid neurons with a bias term. The input layer has two neurons, the middle layer has one layer consisting of 32 neurons, and the output layer has one neuron. The input layer receives oiliness x3 and dryness x5, and the output layer outputs the probability p = p(x3, x5; Θ). Here, Θ represents the weight of the connections between the neurons in the neural network. The predicted value of y is set to 1 or 0 depending on whether the probability p(x3, x5; Θ) is greater than the threshold h. The cost function is the sum of a binary cross-entropy term and a regularization term. The regularization term is calculated by multiplying the squared Frobenius norm of the weights Θ, excluding the weights related to the bias term, by a positive regularization parameter λ. While varying λ, the weights Θ that minimize the above cost function are determined using the training data, and λ is determined to minimize the cost function excluding the regularization term on the CV data. Next, the threshold h is determined to maximize the F1 score of the prediction of y on the CV data (h = 0.23 was obtained). Finally, using the λ determined above, the weights Θ that minimize the above cost function on the training data are determined, and the F1 score is calculated on the test data using the weights Θ and the h determined above.
[0074] <Example 3> (Cosmetic awareness and daily cosmetics use) Example 3 is an example in which latent attribute IDs are defined using areas on the cosmetic consciousness identification surface, using scores for other factors obtained by factor analysis of the same basic data as in Example 1. We confirm that the probability of continuing to use a specific cosmetic product differs depending on the position on the cosmetic consciousness identification surface. This provides supporting evidence that latent attributes related to a user's cosmetic consciousness can be identified using at least two aggregated amounts obtained from the user's answers (manifest attributes) to three or more questions.
[0075] In this Example 3, the factor score for factor B4 described in Example 1 was normalized by linear transformation so that the minimum value was 0 and the maximum value was 1, and this value was referred to as the "base makeup level." Furthermore, the factor score for factor B1 was normalized by linear transformation so that the minimum value was 0 and the maximum value was 1, and this value was referred to as the "aesthetics level." In addition to the 21 questions above, the subjects were also asked a multiple-choice question asking, "Do you regularly use cosmetics (rinse-off cleansers)?" (Answer values were 1 for YES and 0 for NO). Figure 15 is a graph showing the base makeup level x4 and aesthetics level x1 of 300 women randomly selected from the subjects, represented by points with coordinates (x4, x1) on a cosmetic awareness discrimination plane, which is a coordinate plane with base makeup level as the first aggregate quantity on the horizontal axis and aesthetics level as the second aggregate quantity on the vertical axis. Women who answered "I regularly use cosmetics (rinse-off cleansers)" are plotted with a circle, and women who did not are plotted with an x. As can be seen from Figure 15, the lower the level of base makeup and the lower the level of aesthetics, the more likely it is that the above cosmetics are used on a daily basis. It can be seen that the proportion of women who answer "yes" increases. In Figure 15, the probability p(x4, x1) that a woman plotted at the coordinates (x4, x1) on the cosmetic awareness discrimination surface answers that she uses the above-mentioned cosmetic product on a daily basis is estimated and shown in a contour map. The probability p(x4, x1) was estimated using the same method as in Example 2.
[0076] In Figure 15, we focus on the contour line where p(x4, x1) is 0.22. The area to the left (or below) of this contour line, both within and around the outer square, is designated Ae. This area Ae has a curved boundary. In this Example 3, if the subject's combination of base makeup level and aesthetic level falls within area Ae on the cosmetic consciousness identification surface, we predict that the subject "uses the above-mentioned cosmetics on a daily basis" (prediction value p = 1). If the combination does not fall within area Ae, we predict that the subject "does not use the above-mentioned cosmetics on a daily basis" (prediction value p = 0). The f1 score for this prediction is 0.474, which is higher than the average cosmetic use rate of the above-mentioned cosmetics, 0.261 (= 26.1%). In other words, Example 3 demonstrates that by arranging several areas, preferably areas with curved boundaries, on the cosmetic consciousness identification surface, it is possible to predict the user's daily cosmetics use to some extent, and thus to grasp latent attributes related to the user's cosmetic consciousness.
[0077] <Example 4> (Circulation discrimination and awareness of physical and mental condition, etc.) This fourth embodiment is an example in which latent attribute IDs are defined using multiple regions arranged on the circulation identification surface. It is confirmed that the risk of physical and mental disorders differs depending on the region on the circulation identification surface. This serves as supporting evidence that the latent attributes of the user can be grasped using at least two aggregated values obtained from the user's answers (manifest attributes) to three or more questions about their awareness of their physical and mental condition.
[0078] Using a questionnaire containing 30 questions about self-awareness of physical and mental condition, 6,500 female subjects aged 20 to 69 were interviewed, and valid responses were obtained from 5,161. Each question was in the form of "Are you concerned about ~?", where "~" is the part that indicates "Headache," "Stiff shoulders," "Lower back pain (cause unknown)," "Menstrual pain," "Irregular menstruation," "Gastritis / stomach pain," "Constipation," "Loss of appetite," "Overeating," "Continuous fatigue," "Swelling (in the feet)," "Coldness (in the feet)," "Anemia," "Low blood pressure," "Lack of exercise," "Irritability," "Frequent stress," "Tendency to feel depressed," "Irritability," "Insomnia," "Lack of sleep," "Hay fever," "Atopic constitution," "Easily irritated mouth (stomatitis)," "Bad breath," "Eye fatigue," "Decreased vision," "Dry eyes," "Facial hot flashes," and "Excessive facial sweating" These valid responses were used as the base data for statistical analysis using factor analysis. Responses were binary choices (YES or NO, i.e., 1 or 0) for the 30 questions above. The number of factors extracted was 15. This is because the number of eigenvalues in the correlation matrix of the responses above that were greater than or equal to 1 was 15. Maximum likelihood estimation was used to estimate factor loadings, and factor rotation was performed using the varimax method. The factor loadings are not shown here. Regression was used to calculate factor scores. The score for each factor for each user or subject was calculated by multiplying the user's or subject's response value (1 or 0), normalized to a mean of 0 and a standard deviation of 1, by the corresponding factor score coefficient for each question, and then adding these values together for all 30 questions.
[0079] [Table 7]
[0080] Table 7 shows the factor score coefficients obtained through the statistical analysis using the above factor analysis method. The 15 factors are referred to in order as factors T1 to T15. Below, we will interpret and name some of the factors with reference to concepts in traditional Chinese medicine. Factor T3 is a factor that is highly related to questions related to "water circulation," such as "swelling (feet)," "overeating," "constipation," and "lack of exercise." Factor T9 is a factor that is highly related to questions related to "flow of energy," such as "anemia," "lack of sleep," "irritability," and "swelling (feet)." Factor T15 is a factor that is highly related to questions related to "blood circulation," such as "low blood pressure," "anemia," and "sensitive to cold (feet)." Factor T6 is highly related to question items such as "bad breath," "decreased vision," "lower back pain (cause unknown)," and "hot flashes of the face," and also has a high (negative) correlation with question items such as "sensitive to cold (feet)" and "swelling (feet)," and is therefore a factor highly related to question items related to "recurrent chronicity" (i.e., whether the poor health is chronic and continues over time, or recurrent with periods of rest). Factor T7 is highly related to questions such as "hot flashes of the face," "tired eyes," and "swelling (feet)," and also has a high (negative) correlation with questions such as "poor vision," "lack of exercise," and "irregular menstruation," and is therefore a factor highly related to questions regarding "illness and disorder" (= whether the area and cause of the deterioration in physical condition are vague (disorder) or clear (illness)).
[0081] (Liquid factor score S L (Definition of For each user (or subject), the score of the liquid factor S L is defined by the following equation 1: (Formula 1) S L = (S B + S W ) / 2 where S B is the score of factor T15 (= blood factor) after transformation by a non-decreasing function, S W is the score of factor T3 (= water factor) after transformation by a non-decreasing function. That is, the score of the liquid factor S L is the score of the converted blood factor S B and the converted water factor score S W The liquid factor score S is defined as the arithmetic mean of L is considered to represent the "degree of circulation disorder of the body." Transformation using a non-decreasing function is not limited to, but for example, a linear transformation in which the maximum value of the factor score after transformation for subjects included in the basic data or training data is 1 and the minimum value is 0 is used.
[0082] The coordinate plane with the score of the liquid factor, the first aggregate quantity, on the horizontal axis and the score of the converted Qi factor, the second aggregate quantity, on the vertical axis, is called the circulation discrimination surface. Figure (16A) shows how the circulation discrimination surface is divided into 5 x 5 = 25 regions by dividing the closed section with the maximum and minimum values on the horizontal axis at both ends into five equally spaced intervals, and similarly dividing the vertical axis into five intervals. In addition to the 30 questions above, subjects were also asked a multiple-choice question asking "Are you concerned about your scalp and hair?" (Answer value: 1 for YES, 0 for NO). The percentage of subjects who answered YES to this question is entered in each of the 25 regions. Regions with five or fewer subjects are left blank. As shown in Figure 16A, the larger the first aggregated amount (score of the liquid factor), the larger the proportion of subjects who answered "YES." Also, the larger the second aggregated amount (score of the converted Qi factor), the larger the proportion of subjects who answered "YES." The proportion of subjects who answered "YES" was 23.1% of the total. In the 25 regions on the circulation discrimination surface, the proportion was lowest at 10.8% in the lower left corner and 65.0% in the upper right corner, with the ratio being approximately 6 times higher.
[0083] The subjects were further asked a multiple-choice question asking "Are you concerned about your diet?" (answer value: 1 for YES, 0 for NO). In Figure (16B), the proportion (percentage) of subjects who answered YES to the question is entered in each of the 25 areas on the discrimination surface. As shown in Figure (16B), the proportion of subjects who answered YES generally increased as the first aggregate amount (score of the liquid factor) increased, and the proportion of subjects who answered YES also increased as the second aggregate amount (score of the converted qi factor) increased. The proportion of subjects who answered YES was 0.01% of the total. In the 25 regions on the circular discrimination surface, the ratio is 12.6% in the lower left corner region and the maximum value of 66.7% in the upper right corner region, which is about 5 times higher.
[0084] This Example 4 shows that the subject's level of interest in the risk of physical and mental disorders such as scalp and hair disorders and their own dietary habits varies significantly depending on which area on the environmental discrimination surface the subject belongs to.It is thought that the user's latent attributes can be grasped by using at least two aggregated values obtained from the user's answers (manifest attributes) to three or more questions about their awareness of their physical and mental condition.
[0085] (Advantages of using factor analysis when determining health awareness related to circulation through interviews) In the present invention, when assessing health consciousness related to circulation through a medical interview, rather than calculating scores solely from the customer's responses (manifest attributes) to direct questions about circulation of qi, water, and blood, the number of direct questions is reduced, and factor analysis is used to calculate scores for each factor by aggregating all of the customer's responses to approximately 10 to 100, preferably approximately 15 to 40, questions about general health consciousness related to the body and mind. The scores for these factors are used to comprehensively and indirectly assess the customer's health consciousness (latent attributes) related to circulation, thereby enabling the customer's latent health-related attributes to be ascertained, even if the customer himself is not aware of them. In addition, the customer's health consciousness related to circulation is visually represented by the position on the circulation discrimination surface, allowing both the customer and the seller, etc., to quickly assess and share the customer's health consciousness related to circulation.
[0086] <Example 5> (Identification of vague symptoms and risk of menopausal troubles, etc.) This Example 5 is an example in which latent attribute IDs are defined using multiple areas arranged on the vague symptom identification surface. It is confirmed that the risk of menopausal troubles, etc. differs depending on the area on the vague symptom identification surface. This serves as supporting evidence that the latent attributes of the user can be grasped using at least two aggregated amounts obtained from the user's answers (manifest attributes) to three or more questions about their awareness of their physical and mental condition.
[0087] Continuing, we use the same factor analysis results as in Example 4. For each user (or subject), the score of factor T7 after transformation using a non-decreasing function is called the "disorder level," and the score of factor T6 after transformation using a non-decreasing function is called the "recurrence chronicity level." Both the disorder level and the recurrence chronicity level take integer values between 1 and 5, and are discretized into five levels.
[0088] (Percentage of people concerned about menopausal problems) The coordinate plane with the first aggregated quantity, the degree of disorder, on the horizontal axis and the second aggregated quantity, the degree of recurrence chronicity, on the vertical axis is called the "indefinite symptom identification surface." Figure 17A shows how the indefinite symptom identification surface was divided into 5 × 5 = 25 regions, and then the four regions at each corner were connected, and the two regions near each side were connected, resulting in a total division of the indefinite symptom identification surface into nine regions. In addition to the 30 questions above, subjects were also asked a multiple-choice question asking "Are you concerned about menopausal problems?" (Answer values were 1 for YES and 0 for NO). Figure 17A shows the percentage of subjects who answered YES to the question, along with the number of subjects in each region. As shown in Figure (17A), the larger the first aggregate amount (degree of disorder), the larger the proportion of subjects who answered "YES." Also, the larger the second aggregate amount (degree of recurrence and chronicity), the larger the proportion of subjects who answered "YES." The proportion of subjects who answered "YES" was 8.9% of the total. In the nine regions on the vague complaint identification screen, the proportion was 7% in the lower left corner region and 23% in the upper right corner region, a ratio of approximately three times.
[0089] (Percentage of people concerned about menstrual pain) The subjects were also asked a two-choice question asking, "Are you bothered by menstrual pain?" (answer value: YES) A scale of 1 was imposed for yes, and 0 for no. In Figure 17B, the proportion (percentage) of subjects who answered yes to the question is entered in each of the nine areas, along with the number of subjects belonging to each area. As shown in Figure 17B, the smaller the second aggregate amount (recurrence chronicity), the greater the proportion of subjects who answered yes. Also, when the second aggregate amount is small (the three areas in the bottom row), the smaller the first aggregate amount (disorder level), the greater the proportion of subjects who answered yes. The proportion of subjects who answered yes was 20.5% of the total. In the nine areas on the vague complaint identification screen, the proportion was 51% in the area in the lower left corner and 13% in the area in the center of the top row, a ratio of approximately four times.
[0090] This Example 5 shows that the risk of a subject suffering from physical and mental disorders such as menopausal troubles and the risk of a subject suffering from physical and mental disorders such as menstrual pain vary considerably depending on which region on the vague symptoms discrimination surface the subject belongs to. It is believed that the latent attributes of the user can be grasped by using at least two aggregated amounts obtained from the user's answers (manifest attributes) to three or more questions about their awareness of their physical and mental condition. The advantages of using the indefinite symptom discriminating surface in the present invention are similar to those of the circulation discriminating surface described above.
[0091] Example 6 (Irregular eating habits, unhealthy eating habits, and awareness of poor circulation) This Example 6 is an example in which latent attribute IDs are defined using multiple areas arranged on the irregular eating / unhealthy eating discrimination surface. It is confirmed that the risk of physical and mental disorders differs depending on the area on the irregular eating / unhealthy eating discrimination surface. This serves as supporting evidence that the latent attributes of the user can be grasped using at least two aggregated amounts obtained from the user's answers (manifest attributes) to three or more questions about their eating habits.
[0092] Using a questionnaire containing 25 questions about dietary habits, 6,500 female subjects aged 20 to 69 were interviewed, and valid responses were obtained from 5,597 of them. Each question was in the form of a two-choice question asking whether or not "... applies to you." Here, the "..." part means "Eats a lot of meat," "Eats a lot of seafood," "Eats a lot of vegetables," "Eats a lot of fruit," "Eats a lot of dairy products," "Eats a lot of beans and nuts," "Tries to drink a lot of water," "Eats a lot of fried foods," "Eats a lot of instant foods," "Eats a lot of fast food," "Tries to actively consume dietary fiber," "Tries to actively consume vinegar (table vinegar)," "Tries to actively consume fermented foods," "Tries to actively consume seasonal ingredients," "Tries to actively consume organic and pesticide-free foods," "Irregular meal times," "Eats a lot of snacks," "I have a picky eater (I have a lot of likes and dislikes)," "Avoids foods I'm allergic to," "Avoids caffeine," "I often skip breakfast," "I often skip lunch," "I often skip dinner," "I use supplements," "I use protein products." These valid responses were used as the base data for statistical analysis using factor analysis. Responses were binary choices (YES or NO, i.e., 1 or 0) for the 25 questions above. The number of factors extracted was eight. This is because the number of eigenvalues in the correlation matrix of the responses above that were greater than or equal to 1 was eight. Maximum likelihood was used to estimate the factor loads, and factors were rotated using the varimax method. The factor loads are not shown here. Regression was used to calculate factor scores. The score for each factor for each user or subject was calculated by multiplying the user's or subject's response value (1 or 0), normalized to a mean of 0 and a standard deviation of 1, by the corresponding factor score coefficient for each question, and then adding these values together for the 25 questions.
[0093] [Table 8]
[0094] Table 8 shows the factor score coefficients obtained by the above-mentioned factor analysis method. The eight factors are called factors E1 to E8 in order. Below, these factors will be interpreted and named. Factor E1 is a factor that is highly related to questions related to "health consciousness," such as "I try to actively consume dietary fiber," "I try to actively consume fermented foods," "I try to actively consume seasonal ingredients," and "I eat a lot of vegetables." Factor E2 is a factor (unhealthy eating factor) that is highly related to questions related to "not being healthy (tendency to overeat)," such as "often eating instant foods," "often eating fried foods," and "often eating fast food." Factor E3 is a factor (irregular eating factor) that is highly related to questions related to "irregular eating habits," such as "often skip breakfast," "irregular meal times," "often skip lunch," and "picky eating (has many likes and dislikes)." Factor E4 is highly related to question items such as "I limit my caffeine intake" and "I often eat fast food," and also has a high (negative) correlation with question items such as "I often eat meat" and "I often eat snacks," and is therefore a factor that is closely related to awareness of "preventing lifestyle-related diseases." Factor E5 is a factor that is highly related to questions related to "supplement orientation," such as "I use supplements" and "I use protein products." Factor E6 is highly related to question items such as "I often eat beans and nuts" and "I actively try to eat organic and pesticide-free foods," and also has a high (negative) correlation with question items such as "I actively try to eat dietary fiber," "I often eat meat," "I try to drink plenty of fluids," and "I often skip breakfast," and is therefore a factor that is closely related to the "convenience-oriented" mindset. Factor E7 is highly related to question items such as "I often eat meat," "I often eat fruit," and "I often eat seafood," and also has a high (negative) correlation with question items such as "I actively try to eat seasonal ingredients" and "I actively try to eat dietary fiber," and is therefore a factor that is closely related to awareness of "low processing (little cooking)." Factor E8 is composed of "I actively try to eat organic and pesticide-free foods," "I actively try to eat seasonal ingredients," "I often eat seafood," "I often eat meat," and "I don't like fried foods." It is a factor closely related to the awareness of "menu (three meals a day, few desserts and few snacks)" because it is highly correlated with questions such as "I eat a lot of dairy products" and "I eat a lot of snacks" and also (negatively) correlated with questions such as "I often eat dairy products" and "I often eat snacks."
[0095] For each user, the score of one of the eight factors or the amount obtained by transforming it with a non-decreasing function is taken as the first aggregated amount, and the score of another factor or the amount obtained by transforming it with a non-decreasing function is taken as the second aggregated amount.Consider a discrimination surface, which is a coordinate plane with the first aggregated amount on the horizontal axis and the second aggregated amount on the vertical axis, and arrange multiple areas on the discrimination surface while allowing overlapping.The latent attributes of the user can be expressed by determining whether the pair of the user's first aggregated amount and second aggregated amount belongs to each of the multiple areas on the discrimination surface.
[0096] (Irregular eating habits, unhealthy eating habits, and circulatory disorders) The scores of each subject's irregular eating factors are linearly transformed so that the maximum value in the basic data is 1 and the minimum value is 0, and then the values are discretized into 10 levels of 0.05, 0.15, 0.25, . . . 0.95, and these values are referred to as the "irregular eating level" of that subject. Furthermore, the scores of each subject's unhealthy eating factors are similarly discretized into 10 levels and these values are referred to as the "unhealthy eating level" of that subject. Figure 18 shows the irregular eating / unhealthy eating discrimination surface, which is a coordinate plane with the irregular eating level, the first aggregate quantity, on the horizontal axis and the unhealthy eating level, the second aggregate quantity, on the vertical axis. In addition to the 25 questions, subjects were also asked a multiple-choice question asking, "Are you concerned about poor circulation?" (Answer value: 1 for YES, 0 for NO). 300 subjects were randomly selected, and subjects who answered YES to the question were plotted with a circle (◯) on the irregular eating / unhealthy eating discrimination surface, while subjects who answered NO were plotted with a cross (×). Furthermore, the proportion of subjects who were concerned about poor circulation on the irregular eating / unhealthy eating discrimination surface was estimated using a prediction function expressed using a neural network, as in Example 2, and represented by contour lines in Figure 18. As shown in Figure 18, the larger the first aggregate (irregular eating) was, the higher the proportion of subjects who answered YES. Furthermore, when the first aggregate (irregular eating) was not small, the larger the second aggregate (unhealthy eating) was, the higher the proportion of subjects who answered YES. The proportion of subjects who answered YES was 26% of the total. In the irregular eating / unhealthy eating discrimination surface, the ratio has a minimum value of about 0.2 on the left side and a maximum value of about 0.75 in the upper right corner, and the ratio is about 3 times.
[0097] In Figure 18, we focus on the contour line with a probability (= the predicted value of the proportion of subjects) of 0.25. The area to the right of this contour line, both around and inside the outer square, is designated Af. Region Af has a curved boundary. If a subject's irregular eating and unhealthy eating scores fall within region Af on the irregular eating / unhealthy eating discrimination surface, the subject is predicted to be at "high risk of poor circulation" (prediction value p = 1). If they do not fall within region Af, the subject is predicted to be at "not high risk of poor circulation" (prediction value p = 0). The f1 score for this prediction is 0.42, which is 1.6 times the proportion of subjects who answered "YES" (0.26), and is therefore quite large. It has been found that arranging multiple regions on the irregular eating / unhealthy eating identification surface, and more preferably arranging not only rectangular regions but also regions with curved boundaries, and determining whether a pair of a user's first aggregate amount and second aggregate amount belongs to each region, is useful in understanding and expressing latent attributes related to the user's eating habits and physical and mental condition.
[0098] Example 7 (Latent attributes related to skin type and cosmetic consciousness by principal component analysis) In Example 1, the first and second aggregate amounts for each subject were calculated by factor analysis, but in Example 7, principal component analysis is used. For basic data consisting of the same 21 questions as in Example 1 and the answer values of 6,039 people to a question asking for their full age as an integer value, the answer values of the subjects to each question were normalized to a mean of 0 and a standard deviation of 1, and then principal component analysis was performed. Six principal components, from the first principal component to the sixth principal component, were extracted, and to make it easier to interpret each principal component, varimax rotation, which is a rotation in a six-dimensional number vector space, was performed. The jth question for a certain subject The value of the standardized answer to z j (where j = 0, 1, 2, . . . , 21. j = 0 corresponds to the question asking about the subject's age in years), and the k-th principal component (after varimax rotation) of the subject is q k Then, q k =Σz j ×P jk (where k=1, 2, . . . , 6, and the sum is taken over j). Here, P jk is the j-th component of the k-th principal component vector (after 6-dimensional Varimax rotation) in the 22-dimensional number vector space, and takes the values shown in Table 9 below.
[0099] [Table 9]
[0100] For simplicity, the "kth principal component after Varimax rotation" will be referred to simply as the "kth principal component." Each principal component will be interpreted and named based on Table 9. The first principal component is highly related to question items such as "sensitive / irritable skin," "dryness," and "rough skin," and is the principal component mainly related to "dryness." The second principal component is highly correlated with question items such as "fine wrinkles," "age," "firmness and elasticity (sagging)," and "age spots," and is a principal component mainly related to "aging." The third principal component is highly related to question items such as "providing moisture" and "preventing dryness," and is a component mainly related to "moisturizing." While the first principal component indicates that people are already concerned about dryness, the third principal component indicates a high level of awareness of preventing dryness (a desire for moisturizing). The fourth principal component is highly related to question items such as "acne," "acne scars," and "rough skin," and is a principal component mainly related to "oiliness." The fifth principal component is highly related to question items such as "texture," "pores," "feel," and "brightness and transparency," and is a principal component mainly related to "aesthetics." The sixth principal component is highly related to question items such as "makeup does not stick well," "makeup comes off," and "rough skin texture," and is a principal component mainly related to "base makeup."
[0101] For each user, one of the six principal components above or the amount obtained by transforming it with a non-decreasing function is the first aggregated amount, and another principal component or the amount obtained by transforming it with a non-decreasing function is the second aggregated amount. A discriminant plane is a coordinate plane with the first aggregated amount on the horizontal axis and the second aggregated amount on the vertical axis. By considering this, multiple regions are arranged on the discrimination surface with overlapping allowed, and determining whether the pair of the user's first aggregated amount and second aggregated amount belongs to each of the multiple regions on the discrimination surface, the latent attributes of the user can be expressed.
[0102] (Oily / dry skin and sensitive skin) In this Example 7, the fourth principal component of each subject is linearly transformed so that the maximum value is 1 and the minimum value is 0. The interval [0,1] is then divided into three equal-width intervals, and the value discretized into three levels (1, 2, 3) depending on which interval the linearly transformed fourth principal component belongs to is called the "oilyness level" of that subject. Similarly, the first principal component of each subject is discretized into three levels (1, 2, 3) and called the "dryness level" of that subject. Figure 19 shows the oily / dryness discrimination surface, which is a coordinate plane with the oilyness level, the first aggregated quantity, on the horizontal axis and the dryness level, the second aggregated quantity, on the vertical axis. The oily / dryness discrimination surface is divided into 3 × 3 = 9 regions based on the oiliness level and dryness level. In addition to the 21 questions above, the subjects were also asked a multiple-choice question asking, "Are you concerned about sensitive skin?" (Answer values were 1 for YES and 0 for NO). Figure 19 shows the number of subjects belonging to each of the nine regions, the number of subjects belonging to each region who answered YES to the question, and the percentage of subjects belonging to each region who answered YES. As shown in Figure 19, the larger the first aggregate amount (oiliness level), the higher the percentage of subjects who answered YES, and the larger the second aggregate amount (dryness level), the higher the percentage of subjects who answered YES. The percentage of subjects who answered YES was 22% of the total. In the oily / dry discrimination screen, the percentage ranged from a minimum of approximately 0.00 in the lower left region to a maximum of approximately 0.54 in the upper right region.
[0103] Arranging multiple overlapping regions on the oily / dry discrimination surface, and more preferably arranging not only rectangular regions but also regions with curved boundaries, and determining whether a pair of the subject's first aggregate amount and second aggregate amount belongs to each region is useful in understanding and expressing latent attributes related to the subject's skin type, makeup awareness, physical and mental condition, dietary habits, etc.
[0104] Example 8 (Latent attributes related to skin type and cosmetic preferences using collaborative filtering) In Example 1, the first and second aggregated quantities for each subject were calculated using factor analysis, but in Example 8, collaborative filtering is used. Using basic data consisting of the response values of 6,039 people to the same 21 questions as in Example 1, the response values (0 or 1) of each subject to each question were predicted using collaborative filtering. Hereinafter, m = 6,039, n = 21. There are a total of m × n = 126,819 response values, represented by a matrix X of size m × n, where each element takes the value 1 or 0. Randomly selecting 1,000 elements from X were used as cross-validation data (CV data), and randomly selecting 1,000 elements from the remaining 125,819 elements were used as test data, with the remaining 124,819 elements used as training data. The dimension of the latent feature was set to d = 5, and matrix X was predicted by the product U × V' of a user latent feature matrix U of size m × d and an item latent feature matrix V of size n × d (where V' represents the transposed matrix of V). The user latent feature matrix U and item latent feature matrix V were determined so as to minimize the cost function expressed as the sum of 1 / 2 the sum of the squares of the prediction errors and the regularization term on the training data. Here, the regularization term is calculated by multiplying λ / (2m) with λ as the regularization parameter, the sum of the square of the Frobenius norm of matrix U and the square of the Frobenius norm of matrix V multiplied by m / n. λ was varied as follows: ··, 30, 10, 3, 1, 0.3, 0.1, ··, and for each λ, U and V that minimize the cost function on the training data were found as described above. The calculated U and V were then used to calculate the value of the cost function (excluding the regularization term) on the CV data. λ was determined so as to minimize the value of the cost function (excluding the regularization term) on the CV data obtained in this way (λ * Once again, we use λ as the regularization parameter. * Using this, U and V were calculated to minimize the cost function on the training data, and the calculated U and V were used to predict the matrix X on the test data. * The root mean square error (RMSE) of the prediction error on the test data is 0.316.
[0105] To facilitate interpretation of the user latent feature matrix U and the item latent feature matrix V, we performed a varimax rotation, which is a rotation of a d-dimensional vector space, on the item latent feature matrix V, and then performed the same rotation on the user latent feature matrix U. Below, U and V after the varimax rotation will simply be referred to as the "user latent feature matrix U" and the "item latent feature matrix V." The item latent feature matrix V takes the values shown in Table 10 below.
[0106] [Table 10]
[0107] In Table 10, each column of matrix V is called an "item latent feature vector," and the kth column is represented by the symbol "IPFk" (where k = 1, 2, , d). Based on Table 10, we will interpret and name each item latent feature vector. IPF1 is highly related to question items such as "texture," "brightness / transparency," "feel," and "youthfulness," and is an item latent feature vector mainly related to "awareness of aging." IPF2 is highly related to question items such as "dirty pores" and "acne scars," and is an item latent feature vector mainly related to "oiliness." IPF3 is highly related to question items such as "fine wrinkles," "firmness, elasticity, and sagging," and "blemishes," and is an item latent feature vector mainly related to "makeup awareness." IPF4 is highly related to question items such as "sensitivity / irritability of skin" and "dryness," and is an item latent feature vector mainly related to "dryness." IPF5 is highly related to question items such as "prevent dryness," "dryness," "provide firmness," and "provide moisture," and is an item latent feature vector mainly related to "moisturizing."
[0108] [Table 11]
[0109] Table 11 shows a part of the user latent feature matrix U. In Table 11, each column of the matrix U is called a "user latent feature vector," and the k-th column is represented by the symbol "UPFk" (where k = 1, 2, , d). In collaborative filtering, each column of the matrices U and V corresponds to each other. That is, UPF1 is a user latent feature vector mainly relating to "awareness of aging," UPF2 is a user latent feature vector mainly relating to "oily skin," UPF3 is a user latent feature vector mainly relating to "consciousness of makeup," UPF4 is a user latent feature vector mainly relating to "dry skin," and UPF5 is a user latent feature vector mainly relating to "moisturizing." For example, it can be seen that a user with a user ID of 2 is highly conscious of dryness, and a user with a user ID of 3 is highly conscious of oiliness.
[0110] For each user, one of the d user latent features or the amount obtained by transforming it with a non-decreasing function is taken as a first aggregated amount, and another quantity or the amount obtained by transforming it with a non-decreasing function is taken as a second aggregated amount.Consider a discrimination surface, which is a coordinate plane with the first aggregated amount on the horizontal axis and the second aggregated amount on the vertical axis, and arrange multiple areas on the discrimination surface with overlapping allowed.By determining whether the pair of the user's first aggregated amount and second aggregated amount belongs to each of the multiple areas on the discrimination surface, latent attributes related to the user's skin type, makeup awareness, physical and mental condition, etc. can be expressed.
[0111] (CF: Oily / dry skin type and sensitive skin type) In this Example 8, the value obtained by linearly transforming the UPF2 (oily) of each subject so that the maximum value is 1 and the minimum value is 0 is referred to as the "oilyness level" of that subject. Also, the value obtained by linearly transforming the UPF4 (dry) of each subject so that the maximum value is 1 and the minimum value is 0 is referred to as the "dryness level" of that subject. Fig. 20 shows the CF oily / dryness discrimination surface, which is a coordinate plane with the oily level, which is the first aggregated amount, on the horizontal axis and the dryness, which is the second aggregated amount, on the vertical axis. In addition to the 21 questions above, the subjects were also asked a multiple-choice question asking, "Are you concerned about sensitive skin?" (Answer value: 1 for YES, 0 for NO). In Figure 20, for the test data, subjects who answered YES to the question are plotted with a circle (◯) and subjects who answered NO are plotted with a cross (×). The probability that a subject would answer YES to the question on the CF oily / dry discrimination surface was also estimated using a neural network in the same manner as in Example 2, and the probability contours were plotted. As shown in Figure 20, the larger the second aggregate amount (dryness), the higher the proportion of subjects who answered YES. Furthermore, when the second aggregate amount (dryness) was larger, the larger the first aggregate amount (oilyness), the higher the proportion of subjects who answered YES. The proportion of subjects who answered YES was 22% of the total. In the CF oily / dry discriminating surface, the ratio has a minimum value of about 0.0 in the area near the bottom edge and a maximum value of about 0.65 in the upper right corner.
[0112] A plurality of regions are arranged on the CF oily dry discrimination surface with overlapping, and more preferably, rectangular regions. It has been found that arranging regions with curved boundaries and determining whether a pair of the subject's first aggregated amount and second aggregated amount belongs to each region is useful in understanding and expressing the subject's latent attributes related to skin type, makeup awareness, physical and mental condition, dietary habits, etc.
[0113] In Examples 1 to 8, the response values used in determining the first and second aggregate amounts representing the user's latent attributes through statistical analysis are those for questions in a questionnaire that includes three or more questions, and the questionnaire includes only one of questions regarding skin type and cosmetic awareness, questions regarding physical and mental condition, and questions regarding dietary habits. However, in other embodiments of the present invention, the questionnaire may include two or all three of questions regarding skin type and cosmetic awareness, questions regarding physical and mental condition, and questions regarding dietary habits, or may include questions regarding the user's explicit attributes such as age. In Examples 1 to 8, factor analysis, principal component analysis, and collaborative filtering were used as statistical analysis methods to determine the first and second aggregate amounts from the values of the user's responses, but other methods may also be used, or these methods may be used in combination, or these methods may be used in combination with other methods. In Examples 1 to 8, the latent attributes of a user grasped and expressed by the method of the present invention are only one of latent attributes related to skin type, makeup preferences, physical and mental condition, and dietary habits. However, in other embodiments of the present invention, the latent attributes of a user grasped and expressed by the method of the present invention may be those that span two or three or more of the areas of skin type, makeup preferences, physical and mental condition, and dietary habits.
[0114] The present invention relates to a comment output device, etc., but inventions related to the present invention (hereinafter referred to as "the related inventions") include inventions related to a "skin type determination device," a "beauty consciousness determination device," a "health consciousness determination device," a "dietary lifestyle determination device," etc., which are substantially described in the specification of the present application. Below, the related inventions will be explained using inventions such as the "dietary lifestyle consciousness determination device" as examples.
[0115] The first aspect of the present related invention is: A dietary lifestyle awareness assessment device for assessing a customer's dietary lifestyle awareness, an analysis unit that statistically analyzes interview data that is responses from a large number of subjects to three or more questions about their dietary habits, and obtains a function for calculating a first aggregate amount and a second aggregate amount for each subject from the responses of the subject; a storage unit that stores the customer's answers to the three or more questions and the function calculated by the analysis unit; a dietary habit consciousness determination unit that calculates a first aggregate amount E1 and a second aggregate amount E2 of the customer using the answer of the customer and the function stored in the storage unit, and expresses the dietary habit consciousness of the customer by the position of a point (E1, E2) on a coordinate plane with the first aggregate amount on the horizontal axis and the second aggregate amount on the vertical axis as a dietary habit consciousness discrimination surface; The present invention is a dietary lifestyle awareness assessment device having the above features.
[0116] In this embodiment, when determining a customer's awareness of their dietary habits through a medical interview, the amounts to be plotted on the horizontal axis and the vertical axis of the discrimination surface are not calculated only from the customer's answers to direct questions about specific aspects of their dietary habits, but rather the number of questions directly asked is reduced, and all of the customer's answers to about 10 to 100 questions, preferably about 15 to 40 questions, about general aspects of their dietary habits are combined to calculate the first aggregate amount E1 and the second aggregate amount E2 of the customer using statistical analysis processing such as factor analysis, principal component analysis, and / or collaborative filtering, and the customer's dietary habits are determined comprehensively and indirectly based on the positions of points (E1, E2) on the dietary habits awareness discrimination surface. In addition, the position of the points (E1, E2) on the dietary lifestyle awareness identification plane allows the customer's dietary lifestyle awareness to be accurately determined and visually expressed, allowing both the customer and the seller, etc., to quickly grasp and share the determination result of the customer's dietary lifestyle awareness (dietary lifestyle awareness).
[0117] A second aspect of the present related invention is a dietary lifestyle awareness assessment device in accordance with the first aspect, wherein the statistical analysis processing of the medical interview data is performed using at least one of factor analysis, principal component analysis, and collaborative filtering.
[0118] A third aspect of the present related invention is a method for determining a customer's dietary lifestyle consciousness, which is realized by an analysis unit, a storage unit, and a dietary lifestyle consciousness determination unit, comprising: an analysis step in which the analysis unit statistically analyzes interview data, which are answers from a large number of subjects to three or more questions about their dietary habits, to obtain a function for calculating a first aggregate amount and a second aggregate amount for each subject from the answers from the subjects; a storage step in which the storage unit stores the customer's answers to the three or more questions and the function calculated by the analysis unit; a dietary habit consciousness determination step in which the dietary habit consciousness determination unit calculates a first aggregate amount E1 and a second aggregate amount E2 of the customer using the answer of the customer and the function stored in the storage unit, and expresses the dietary habit consciousness of the customer by the position of a point (E1, E2) on a coordinate plane with the first aggregate amount on the horizontal axis and the second aggregate amount on the vertical axis as a dietary habit consciousness identification plane; The method for assessing dietary habits awareness comprises the steps of:
[0119] The fourth aspect of the present related invention is a program for determining a customer's eating habits awareness, which comprises: an analysis unit that statistically analyzes interview data that is responses from a large number of subjects to three or more questions about their dietary habits, and obtains a function for calculating a first aggregate amount and a second aggregate amount for each subject from the responses of the subject; a storage unit that stores the customer's answers to the three or more questions and the function calculated by the analysis unit; a dietary habit consciousness determination unit that calculates a first aggregate amount E1 and a second aggregate amount E2 of the customer using the answer of the customer and the function stored in the storage unit, and expresses the dietary habit consciousness of the customer by the position of a point (E1, E2) on a coordinate plane with the first aggregate amount on the horizontal axis and the second aggregate amount on the vertical axis as a dietary habit consciousness discrimination plane. This is a dietary awareness assessment program that functions as a
[0120] The present invention is not limited to the above-described embodiments and examples, and it goes without saying that various modifications and design changes are included within the technical scope of the present invention as long as they do not deviate from the technical idea of the present invention. [Industrial Applicability]
[0121] According to the present invention, a comment output device can be provided that allows a user to appropriately view comments such as reviews from others who are similar to the user, by utilizing all of the user's answers to three or more questions and / or the user's latent attributes determined using at least two aggregate quantities calculated using the user's image. It is also possible to provide a comment output method using the device, and a comment output program for causing a computer to function as the device. According to the present invention, online or real stores such as osteopathic clinics, drug stores, beauty salons, clinics, aesthetic salons, cosmetics stores, and health food stores can provide users with review bulletin board sites, recommend products such as health foods, beauty products, medicines, and herbal medicines, and provide advice on physical and mental health more accurately than before. The present invention has wide industrial applicability. [Explanation of symbols]
[0122] 1. Comment output device 11 output section 12 receiving section 13 Transmission unit 14 Comment acquisition unit 15 Comment selection section 16 Memory section 17 Analysis section 18 Manifest attribute transmission section 19 Latent attribute receiver 2. Terminal Device 21 terminal input section 22 terminal output section 23 Terminal receiving section 24 Terminal transmitting section 3 Analyzer 31 Storage section 32 Analysis section 33 Manifest attribute receiving section 34 Latent attribute transmitting section 4 Network A1~A5, As, Ae, Af, An Areas on the discrimination surface Aо2d3, Aо45d12 Area on the discrimination surface Cn Points r placed on the discrimination surface n radius L1 input layer L2 hidden layer L3 output layer
Claims
1. a receiving unit that receives an instruction to output a comment; a comment acquisition unit that acquires one or more comments together with a user identifier and a latent attribute identifier of a user who made the comment; a comment selection unit that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who transmitted the output instruction; an output unit that outputs only the comments selected by the comment selection unit; an analysis unit that determines a latent attribute identifier for each user in advance using two or more aggregate amounts calculated using all of the user's answers to three or more predetermined questions; A comment output device comprising:
2. A large number of subjects are previously asked a medical interview consisting of three or more predetermined questions, each answer of which can be expressed by a single real number selected from a finite number of real numbers, and a statistical analysis of basic data, which is a collection of the subjects' answers to the multiple questions, determines a first aggregate quantity S1 and a second aggregate quantity S2, both of which take real numbers, as a function (S1, S2) = f(A) of the subject's answer A, and defines a coordinate plane with the first aggregate quantity S1 on the horizontal axis and the second aggregate quantity S2 on the vertical axis as a discrimination surface, arranges multiple regions on the discrimination surface, and assigns one latent attribute identifier to each region, a storage unit that stores the function f, the plurality of regions, and a correspondence relationship between each of the regions and a latent attribute identifier; The user is also asked a medical interview consisting of the three or more questions in advance, The comment output device of claim 1, wherein the analysis unit determines the first aggregation amount T1 and the second aggregation amount T2 of the user from the user's answer B using the function f by (T1, T2) = f(B), and determines the latent attribute identifier corresponding to the area including point (T1, T2) among the multiple areas arranged on the identification surface as the latent attribute identifier of the user.
3. The comment output device according to claim 2 , wherein the two or more aggregate amounts are calculated using at least one of factor analysis, principal component analysis, and collaborative filtering.
4. 3. The comment output device according to claim 2, wherein at least one of the plurality of regions arranged on the discrimination surface is a region whose boundary is partially formed by a curve.
5. a receiving unit that receives an instruction to output a comment; a comment acquisition unit that acquires one or more comments together with a user identifier and a latent attribute identifier of a user who made the comment; a comment selection unit that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who transmitted the output instruction; an output unit that outputs only the comments selected by the comment selection unit; an analysis unit that determines a latent attribute identifier of each user using two or more aggregate quantity prediction values, which are prediction values when an aggregate quantity calculated in advance using all of the user's answers to three or more predetermined questions is predicted from the image data of the user using a supervised learning method in machine learning; A comment output device comprising:
6. 6. The comment output device according to claim 1, wherein the comment is a comment about a beauty product, a health product, or a food product.
7. A comment output method realized by a receiving unit, a comment acquiring unit, a comment selecting unit, an output unit, and an analyzing unit, a receiving step in which the receiving unit receives an instruction to output a comment; a comment acquisition step in which the comment acquisition unit acquires one or more comments together with a user identifier and a latent attribute identifier of a user who has made the comment; a comment selection step in which the comment selection unit selects only comments of users having latent attribute identifiers that match the latent attribute identifier of the user who transmitted the output instruction; an output step in which the output unit outputs only the comments selected by the comment selection unit; an analysis step in which the analysis unit determines a latent attribute identifier for each user using two or more aggregate amounts calculated in advance using all of the user's answers to three or more predetermined questions, or determines a latent attribute identifier for each user using two or more aggregate amount prediction values, which are prediction values when the aggregate amount is predicted from the image data of the user using a supervised learning method in machine learning; A comment output method comprising:
8. Computer, a receiving unit that receives an instruction to output a comment; a comment acquisition unit that acquires one or more comments together with a user identifier and a latent attribute identifier of a user who made the comment; a comment selection unit that selects only comments from users having a latent attribute identifier that matches the latent attribute identifier of the user who transmitted the output instruction; an output unit that outputs only the comments selected by the comment selection unit; An analysis unit that determines a latent attribute identifier for each user using two or more aggregate amounts calculated in advance using all of the user's answers to three or more predetermined questions, or an analysis unit that determines a latent attribute identifier for each user using two or more aggregate amount prediction values, which are prediction values when the aggregate amount is predicted from the image data of the user using a supervised learning method in machine learning. A comment output program to function as.
Citation Information
Patent Citations
Cosmetic Evaluation Information Analysis System and Method
JP4478479B2
Dictionary construction device, information processing device, comment output device, evaluation word dictionary production method, information processing method, comment output method and program
JP7001662B2