Information processing device, information processing method, and program

The information processing device uses matrix decomposition on purchase data to extract purchasing patterns, addressing the inefficiencies of conventional user analysis by providing transparent and efficient purchasing preference type creation.

JP7767249B2Active Publication Date: 2025-11-11KK TOSHIBA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022144962
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2025-11-11
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

Conventional user analysis in corporate marketing relies heavily on analyst experience and knowledge, leading to a heavy workload and lack of transparency in formulating purchasing preference types, which vary between analysts.

Method used

An information processing device and method that utilizes matrix decomposition on purchase data to calculate user and product hidden state information, reducing the workload by extracting representative purchasing patterns and facilitating transparent, data-driven analysis.

Benefits of technology

The solution enables efficient and transparent creation of purchasing preference types by displaying characteristic purchase patterns, reducing analyst workload and enhancing the accuracy and consistency of user analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767249000001
    Figure 0007767249000001
  • Figure 0007767249000002
    Figure 0007767249000002
  • Figure 0007767249000003
    Figure 0007767249000003
Patent Text Reader

Abstract

To reduce the workload of user analysis using purchasing data.SOLUTION: An information processing device includes an acquisition unit, a state calculation unit, and an output control unit. The acquisition unit acquires a plurality of pieces of purchasing data including any of a plurality of pieces of user identification information for identifying a plurality of users, any of a plurality of pieces of product identification information for identifying a plurality of pieces of products, and performance information including at least one of a product price and a number of purchases. The state calculation unit uses the plurality of pieces of user identification information and the plurality of pieces of product identification information as row and column indexes, respectively, decomposes a purchase matrix whose element values are non-negative values calculated based on the performance information, and calculates user hidden state information indicating a relationship between the plurality of pieces of user identification information and hidden states related to purchase and product hidden state information indicating a relationship between the hidden states and the plurality of pieces of product identification information. The output control unit controls output of at least one of the user hidden state information and the product hidden state information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] In corporate marketing, user analysis is conducted to develop effective product development and promotions. In user analysis, hypotheses are formulated about the user's purchasing preference types, such as "sale-lovers" or "health-conscious," and detailed analysis is conducted through interviews or panel surveys. Purchasing preference types can provide insight into the psychological factors behind purchases, such as the user's purchasing motivations and intentions. This can have a significant impact in a variety of situations, including product recommendations, product development, and product lineup optimization.

[0003] On the other hand, in recent years, the accumulation and utilization of purchasing data has progressed, and technologies that utilize purchasing data to support the analysis of purchasing preference types have been proposed. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 5902325 [Non-patent literature]

[0005] [Non-Patent Document 1] LEE, Daniel D.; SEUNG, H. Sebastian. “Learning the parts of objects by non-negative matrix factorization.” Nature, 1999, 401.6755:788-791. Summary of the Invention [Problem to be solved by the invention]

[0006] However, in conventional technology, analysts design purchasing preference types through trial and error based on their experience and knowledge. This can lead to a heavy workload and time-consuming formulation and verification of hypotheses. Furthermore, hypotheses about purchasing preference types vary depending on the analyst, which can lead to problems such as a lack of transparency.

[0007] An object of the present invention is to provide an information processing device, an information processing method, and a program that can reduce the workload of user analysis using purchase data. [Means for solving the problem]

[0008] An information processing device according to an embodiment includes an acquisition unit, a state calculation unit, and an output control unit. The acquisition unit acquires multiple purchase data sets, each set including one of multiple user identification information sets identifying multiple users, one of multiple product identification information sets identifying multiple products, and performance information including at least one of the product price and purchase quantity. The state calculation unit performs matrix decomposition on a purchase matrix in which the multiple user identification information sets and the multiple product identification information sets are used as row and column indices, respectively, and non-negative values ​​calculated based on the performance information are used as element values, to calculate user hidden state information indicating a relationship between the multiple user identification information sets and hidden states related to purchases, and product hidden state information indicating a relationship between the hidden state and the multiple product identification information. The output control unit controls output of at least one of the user hidden state information and the product hidden state information. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram of an information processing apparatus according to a first embodiment. [Figure 2] 10 is a flowchart of an analysis support process according to the first embodiment. [Figure 3] FIG. 10 is a diagram showing an example of a data structure of a purchase history. [Figure 4] FIG. 10 is a diagram showing an example of product information. [Figure 5] FIG. 10 is a diagram showing an example of user information. [Figure 6] FIG. 10 is a diagram showing an example of a generated purchase matrix. [Figure 7] FIG. 10 is a diagram showing an example of product hiding state information. [Figure 8] FIG. 10 is a diagram showing an example of user hidden state information. [Figure 9] FIG. 10 is a diagram showing an example of display of product hiding state information. [Figure 10] FIG. 10 is a diagram showing an example of display of user hidden state information. [Figure 11] FIG. 10 is a block diagram of an information processing apparatus according to a second embodiment. [Figure 12] 10 is a flowchart of an analysis support process according to the second embodiment. [Figure 13] FIG. 10 is a diagram showing an example of a classification result into clusters. [Figure 14] FIG. 10 is a diagram showing an example of statistical information. [Figure 15] FIG. 10 is a diagram showing an example of displaying user hidden state information for each cluster. [Figure 16] FIG. 10 is a block diagram of an information processing apparatus according to a third embodiment. [Figure 17] 10 is a flowchart of an analysis support process according to the third embodiment. [Figure 18] FIG. 10 is a diagram showing an example of a notable user label. [Figure 19] FIG. 10 is a diagram showing an example of calculating an average value of user hidden state information. [Figure 20] FIG. 10 is a diagram showing an example of display of statistical information on the user hidden state information of a user of interest. [Figure 21] FIG. 10 is a block diagram of an information processing apparatus according to a fourth embodiment. [Figure 22] 10 is a flowchart of an analysis support process according to the third embodiment. [Figure 23] FIG. 10 is a diagram showing an example of known information. [Figure 24] FIG. 10 is a block diagram of an information processing apparatus according to a fifth embodiment. [Figure 25] 13 is a flowchart of an analysis support process according to the fifth embodiment. [Figure 26] FIG. 10 is a diagram showing an example of the cluster classification results and the notable user labels. [Figure 27] FIG. 10 is a diagram showing an example of a cluster ratio to all users. [Figure 28] FIG. 10 is a diagram showing an example of a cluster ratio for a user of interest. [Figure 29] FIG. 10 is a diagram showing an example of displaying cluster information with a large difference in cluster ratios. [Figure 30] FIG. 10 is a diagram showing an example of a screen displaying user information and product information. [Figure 31] FIG. 10 is a diagram showing an example of a screen on which the number of purchases is plotted for each cluster. [Figure 32] FIG. 1 is a hardware configuration diagram of an information processing apparatus according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of an information processing apparatus according to the present invention will be described in detail below with reference to the accompanying drawings.

[0011] As described above, a technology has been proposed that utilizes purchasing data to support the analysis of purchasing preference types. The purchasing data includes, for example, a purchasing history indicating the products purchased by each user in physical stores and online shopping. Because purchasing behavior is considered to strongly reflect a user's preferences, the purchasing data can be used to develop purchasing preference types. Developing purchasing preference types based on purchasing data reduces the analyst's workload and is expected to enable the development of purchasing preference types based on objective facts rather than solely on experience and knowledge.

[0012] As a technology to support the analysis of purchasing preference types using purchasing data, a technology has been proposed that quantitatively evaluates purchasing preference types using the degree of agreement between a user's purchasing preference type and their actual product purchase history, thereby supporting the design of purchasing preference types that have a high degree of agreement. By quantitatively evaluating the proposed purchasing preference types, this technology makes it possible to determine whether the purchasing preference types are appropriate and to update, integrate, or divide the purchasing preference types.

[0013] However, these technologies are related to the quantitative evaluation of existing purchasing preference types. Therefore, the creation of the basic purchasing preference types relies on the experience and knowledge of analysts, as in the past. This can lead to a heavy workload for analysts when creating the initial purchasing preference types, or to analysts overlooking purchasing preference types.

[0014] In the following embodiments, purchasing data is used to support the creation of purchasing preference types. For example, an analyst can present purchasing information for creating purchasing preference types.

[0015] (First embodiment) 1 is a block diagram showing an example of the configuration of an information processing device 100 according to the first embodiment. As shown in FIG. 1, the information processing device 100 includes an acquisition unit 101, a state calculation unit 102, an output control unit 111, and a storage unit 121.

[0016] The acquisition unit 101 acquires various information used in the information processing device 100. For example, the acquisition unit 101 stores a plurality of purchase data. The purchase data includes any one of a plurality of user identification information (hereinafter referred to as user ID) that identifies a plurality of users, any one of a plurality of product identification information (hereinafter referred to as product ID) that identifies a plurality of products, and performance information. The performance information is information that includes, for example, at least one of the price of the product and the number of units purchased (purchase quantity).

[0017] The acquisition unit 101 may acquire information in any manner, for example, a method of receiving information transmitted from an external device, or a method of reading information from a storage medium, etc., can be applied.

[0018] The state calculation unit 102 analyzes the purchase data to calculate information representing the hidden state of the purchase data. For example, the state calculation unit 102 calculates a purchase matrix from the purchase data, where multiple user IDs and multiple product IDs are used as row and column indices, respectively, and where element values ​​are non-negative values ​​calculated based on performance information. The non-negative element values ​​are, for example, non-negative values ​​indicating whether or not a purchase was made, prices, or purchase quantities.

[0019] The state calculation unit 102 performs matrix decomposition on the purchase matrix to calculate user hidden state information and product hidden state information. The user hidden state information indicates the relationship between multiple user IDs and hidden states related to purchases. The product hidden state information indicates the relationship between hidden states and multiple product IDs.

[0020] The output control unit 111 controls the output of various data used in the information processing device 100. For example, the output control unit 111 controls the output of at least one of user hidden state information and product hidden state information. Any output method by the output control unit 111 may be used, but for example, a method of displaying on a display device such as a liquid crystal display, a method of transmitting data to an external device (a server, another information processing device, etc.), a method of outputting to a recording medium using an image forming device such as a printer, etc. can be applied.

[0021] Each of the above units (acquisition unit 101, state calculation unit 102, and output control unit 111) is realized, for example, by one or more processors. For example, each of the above units may be realized by having a processor such as a CPU (Central Processing Unit) execute a program, that is, by software. Each of the above units may be realized by a processor such as a dedicated IC (Integrated Circuit), that is, by hardware. Each of the above units may be realized by a combination of software and hardware. When multiple processors are used, each processor may realize one of the units, or may realize two or more of the units.

[0022] The storage unit 121 stores various data used by the information processing device 100. For example, the storage unit 121 stores the purchase data acquired by the acquisition unit 101, the processing results of the other units, and the like.

[0023] The storage unit 121 can be configured from any commonly used storage medium such as a flash memory, a memory card, a RAM (Random Access Memory), an HDD (Hard Disk Drive), and an optical disk.

[0024] Next, a description will be given of an analysis support process performed by the information processing device 100 according to the first embodiment. Fig. 2 is a flowchart showing an example of the analysis support process according to the first embodiment.

[0025] The acquisition unit 101 acquires purchase data including a purchase history (step S101). The purchase history includes the user ID of the user who purchased the product and the product ID of the purchased product. FIG. 3 is a diagram showing an example of the data structure of the purchase history. As shown in FIG. 3, the purchase history includes the time of purchase, the user ID, the product ID, the number of products purchased (quantity), and the price.

[0026] The purchase data may further include information other than the purchase history. The information other than the purchase history is, for example, product information and user information. FIG. 4 is a diagram showing an example of product information. In the example of FIG. 4, the product information includes a product name and a product category for a product ID. FIG. 5 is a diagram showing an example of user information. In the example of FIG. 5, the user information includes a gender and an age group for a user ID.

[0027] Returning to FIG. 2, the state calculation unit 102 generates a purchase matrix representing the association between users and products from the purchase history (step S102). For example, the state calculation unit 102 generates a purchase matrix in which the user ID is used as a row index, the product ID is used as a column index, and non-negative values ​​are used as element values. The element values ​​are, for example, a non-negative value indicating whether or not a purchase was made, the price, and the number of purchases. The element values ​​may also be non-negative values ​​obtained by calculation using these values. For example, the element value may be a unit price indicating a value obtained by dividing the price by the number of purchases.

[0028] FIG. 6 is a diagram showing an example of a generated purchase matrix. The purchase matrix in FIG. 6 is an example in which the element value is the presence or absence of a purchase. The presence or absence of a purchase is set to a value of 1 if the user has purchased a product, and a value of 0 if the user has not purchased a product. FIG. 6 is also an example of a purchase matrix for 1,000 users and 1,000 types of products. Conversion from a purchase history (such as FIG. 3) to a purchase matrix (such as FIG. 6) can be achieved by simple data processing.

[0029] 2, the state calculation unit 102 performs matrix decomposition on the purchase matrix to calculate user hidden state information and product hidden state information (step S103). As a matrix decomposition method, Non-negative Matrix Factorization (NMF) technology can be used.

[0030] NMF is a technique for decomposing an N-by-M matrix Y (with non-negative values) into the product of an N-by-K matrix H (with non-negative values) and a K-by-M matrix U (with non-negative values). NMF performs matrix decomposition so that the values ​​of each element between matrix Y and the product HU of matrix H and matrix U are as close as possible. NMF performs matrix decomposition through iterative calculations and is known to be relatively computationally lightweight. Here, K indices in the decomposed matrix represent hidden states and are set to values ​​smaller than the values ​​of N and M. For example, if NMF is applied to N face images with M pixels, the face image can be decomposed into a matrix (K-by-M matrix) representing K facial features such as eyes and nose, and a matrix (N-by-K matrix) representing the feature weights for each image, enabling effective feature extraction.

[0031] For example, the state calculation unit 102 decomposes the purchase matrix Y according to NMF, and treats one of the two matrices obtained by the decomposition, matrix H, as user hidden state information and the other, matrix U, as product hidden state information.

[0032] FIG. 7 is a diagram showing an example of product hidden state information. FIG. 8 is a diagram showing an example of user hidden state information. In the examples of FIGS. 7 and 8, the number of hidden states is set to 10. The number of hidden states is set, for example, by an analyst. The number of hidden states may be determined from the purchase data to be analyzed, such as 1 / 100 of the number of users or the number of products.

[0033] If the number of hidden states is set to 10, for example, a 1000-by-1000 column purchase matrix like the one in Figure 6 is decomposed into a 1000-by-10 column matrix of user hidden state information (Figure 7) and a 10-by-1000 column matrix of product hidden state information (Figure 8).

[0034] The element values ​​of 1 or 0 in the purchase matrix Y (e.g., Figure 6) represent the presence or absence of a purchase. Therefore, the hidden states of the matrix U, which represents the product hidden state information obtained by matrix decomposition of the purchase matrix Y, can be interpreted as representing the purchasing patterns of products purchased by the same user. Furthermore, the matrix H, which represents the user hidden state information, can be interpreted as representing the weight of the hidden state (product purchasing pattern) for each user.

[0035] In the product hidden state information in Figure 7, column values ​​(element values) for each product are obtained for 10 hidden states. For example, for hidden state 1 (H1), the element values ​​for products with product IDs "I0001" and "I0003" are large. This can be interpreted as meaning that the product "I0001" and the product "I0003" tend to be purchased by the same user, and that this tendency has been extracted as a purchasing pattern linked to H1.

[0036] In the user hidden state information in Figure 8, column values ​​(element values) for each hidden state are obtained for 1,000 rows of users. For example, the element values ​​for hidden states H1 and H3 are large for a user with user ID "U0003." This can be interpreted as indicating that the weights of the purchasing patterns corresponding to H1 and H3 for the user with user ID "U0003" are large. For ease of explanation, hereinafter, a user with user ID "*" may be referred to as user * (e.g., user U0003).

[0037] By calculating product hidden state information and user hidden state information, it is possible to extract representative purchasing patterns from the purchasing data of individual products and to extract the extent to which each user fits the extracted purchasing pattern.

[0038] Returning to FIG. 2, the output control unit 111 outputs (displays) the product hidden state information and the user hidden state information to the analyst (step S104).

[0039] FIG. 9 is a diagram showing an example of displaying product hiding state information. In FIG. 9, products with large element values ​​are displayed for each hiding state of the matrix U of the product hiding state information. The output control unit 111 may display products with element values ​​equal to or greater than a certain value, or may display a certain number of products in descending order of element values. In the example of FIG. 9, the product names corresponding to the product IDs obtained from the product information included in the purchase data are displayed.

[0040] FIG. 10 is a diagram showing an example of displaying user hidden state information. FIG. 10 shows an example of a graph plotting the weight (vertical axis) of each hidden state, which is the element value of a matrix, against the hidden state (horizontal axis) for user U0003. In this example, the larger the weight of a hidden state, the more likely the user is to purchase products using the purchasing pattern corresponding to that hidden state. As shown in FIG. 10, user U0003 has large element values ​​(weights) for H1 and H3. When combined with the information in FIG. 9, it can be seen that the weights of the following two purchasing patterns are large. A purchasing pattern in which a product "Cup Noodles A" with a product ID of "I0001" and a product "Cup Noodles B" with a product ID of "I0003" are purchased. A purchasing pattern in which the product "Snack A" with product ID "I1000" is purchased

[0041] By displaying the product hidden state information and user hidden state information as shown in Figures 9 and 10, analysts can confirm what purchasing patterns exist and how each user has those purchasing patterns. This information is useful for planning purchasing preference types. For example, from the information in Figures 9 and 10, analysts can easily confirm that a purchasing demographic that purchases instant noodles and snacks, such as user U0003, exists in the target purchasing data.

[0042] In this way, the information processing device according to the first embodiment outputs a plurality of pieces of information obtained from purchase data using matrix decomposition. For example, by displaying product hidden state information and user hidden state information, analysts can easily confirm the characteristic purchase patterns of each user. In other words, the workload of user analysis (such as planning purchase preference types) using purchase data can be reduced.

[0043] (Second embodiment) The information processing device of the second embodiment classifies the user hidden state information of each user by similarity (or distance), classifies groups of users with characteristic purchasing patterns into multiple clusters, and displays statistical information for each of the multiple clusters.

[0044] 11 is a block diagram showing an example of the configuration of an information processing device 100-2 according to the second embodiment. As shown in FIG. 11, the information processing device 100-2 includes an acquisition unit 101, a state calculation unit 102, a classification unit 103-2, an output control unit 111-2, and a storage unit 121.

[0045] The second embodiment differs from the first embodiment in that a classification unit 103-2 is added and in the function of an output control unit 111-2. The other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing device 100 according to the first embodiment, so the same reference numerals are used and the description thereof will be omitted here.

[0046] The classification unit 103-2 classifies multiple user IDs included in the user hidden state information into multiple clusters using the similarity between the hidden state information. For example, the hidden state information for each user ID is expressed as a vector having element values ​​equal to the number of hidden states (for example, 10). The classification unit 103-2 performs clustering so that user IDs with a high similarity between vectors are classified into the same cluster. The similarity may be expressed, for example, as the distance between vectors. In this case, the smaller the distance, the greater the similarity.

[0047] Clustering can be achieved using common unsupervised clustering techniques, such as the K-means method, which allows users with similar hidden state information to be grouped into the same cluster.

[0048] The output control unit 111-2 differs from the output control unit 111 of the first embodiment in that it further includes a function of outputting statistical information of the user hidden state information for each cluster.

[0049] Next, the analysis support process by the information processing device 100-2 according to the second embodiment will be described with reference to Fig. 12. Fig. 12 is a flowchart showing an example of the analysis support process according to the second embodiment.

[0050] Steps S201 to S203 are the same as steps S101 to S103 in the information processing device 100 according to the first embodiment, and therefore a description thereof will be omitted.

[0051] The classification unit 103-2 classifies users (user IDs) into a plurality of clusters based on the similarity between pieces of user hidden state information (step S204).

[0052] Fig. 13 is a diagram showing an example of the classification result into clusters. Fig. 13 shows an example of the classification result in which the cluster IDs of the clusters classified for the user IDs are assigned. In this example, user ID "U0003" and user ID "U1000" have similar user hidden state information, so they are each assigned the cluster ID "C1" of the same cluster.

[0053] The number of clusters is set by an analyst, for example, and may be determined from the purchase data to be analyzed, such as 1 / 50 of the number of users or products.

[0054] 12, the output control unit 111-2 calculates and displays statistical information of the user hidden state information for each cluster (step S205). The statistical information includes the average value, variance value, and quantile value of the user hidden state information corresponding to the user IDs belonging to each cluster.

[0055] FIG. 14 is a diagram showing an example of statistical information. FIG. 14 shows an example in which the average value of user hidden state information for user IDs belonging to each cluster ID is used as statistical information. For ease of explanation, hereinafter, a cluster with a cluster ID of "*" may be referred to as cluster * (e.g., cluster C1). For example, if 100 users including user U0001 and user U0003 belong to cluster C1, the statistical information shown in FIG. 14 can be obtained by calculating the average value of the user hidden state information corresponding to these 100 users. Other statistical quantities such as variance and quantile values ​​can also be calculated in a similar manner.

[0056] Figure 15 shows an example of user hidden state information display for each cluster. Figure 15 shows an example of statistical information for cluster C1. In Figure 15, the solid line represents the average value, and the dashed line represents the quartile. It can be seen that the 100 users belonging to cluster C1 have large hidden state element values ​​for H1 and H3 on average, and the quartile range indicates that the element values ​​have little variation. By displaying the information in Figure 15 together with Figure 9, analysts can confirm the number of users and purchasing patterns corresponding to each cluster. For example, analysts can easily find a cluster, such as cluster C1, that contains 100 users who purchase instant noodles and snacks, with just a simple visual inspection.

[0057] In this way, the information processing device according to the second embodiment can output information for each cluster into which users are classified, thereby further reducing the workload of user analysis using purchase data.

[0058] (Third embodiment) The information processing apparatus according to the third embodiment highlights and outputs items for which there is a large difference in hidden state between the designated user of interest and all users.

[0059] 16 is a block diagram showing an example of the configuration of an information processing device 100-3 according to the third embodiment. As shown in FIG. 16, the information processing device 100-3 includes an acquisition unit 101-3, a state calculation unit 102, a difference calculation unit 104-3, an output control unit 111-3, and a storage unit 121.

[0060] The third embodiment differs from the first embodiment in that a difference calculation unit 104-3 is added and in the functions of an acquisition unit 101-3 and an output control unit 111-3. The other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing device 100 according to the first embodiment, so the same reference numerals are used and the description thereof will be omitted here.

[0061] The acquiring unit 101-3 differs from the acquiring unit 101 of the first embodiment in that it further acquires a designation of a user of interest, which represents a user of interest as a target for analysis among a plurality of users. For example, an analyst specifies conditions for the user of interest (conditions for user demographics). The acquiring unit 101-3 accepts the designation of the conditions and acquires a user who meets the conditions as the user of interest.

[0062] The difference calculation unit 104-3 calculates the difference between the hidden state for a plurality of users (for example, all users) and the hidden state corresponding to the user of interest.

[0063] The output control unit 111-3 differs from the output control unit 111 of the first embodiment in that it further has a function of outputting the user hidden state information of a user of interest whose calculated difference is larger than that of other users in a manner different from that of other users.

[0064] Next, the analysis support process by the information processing device 100-3 according to the third embodiment will be described with reference to Fig. 17. Fig. 17 is a flowchart showing an example of the analysis support process according to the third embodiment.

[0065] Steps S301 to S303 are the same as steps S101 to S103 in the information processing device 100 according to the first embodiment, and therefore a description thereof will be omitted.

[0066] The acquisition unit 101-3 assigns a label (a noted user label) to the noted user (step S304). For example, the acquisition unit 101-3 acquires the conditions for a noted user specified by an analyst or the like, and sets a user who meets the acquired conditions as a noted user. The conditions may be specified in any way, but are, for example, specified as follows: Users who purchased a certain product group Users with specific attributes, such as men in their 40s Users with characteristics in their purchasing data, such as monthly purchase amounts exceeding a certain value

[0067] The noted user may be designated in units of clusters. In this case, the information processing device 100-3 may include a classification unit 103-2 as in the second embodiment. The acquisition unit 101-3 may acquire, as the noted user, a user who belongs to the designated cluster from among the clusters classified by the classification unit 103-2.

[0068] Fig. 18 is a diagram showing an example of a noted user label. In the example of Fig. 18, a noted user is assigned a noted user label of "True," and other users are assigned a noted user label of "False."

[0069] 17, the difference calculation unit 104-3 calculates the difference between the hidden states for multiple users (for example, all users) and the hidden state corresponding to the user of interest. For example, the difference calculation unit 104-3 calculates statistical information of the user hidden state information of all users and statistical information of the user hidden state information of the user of interest, and calculates the difference between the two. The statistical information of the user hidden state information includes the mean value, variance value, quantile value, etc., as in the second embodiment.

[0070] 19 is a diagram showing an example of calculating the average value of user hidden state information for all users and the user of interest. In this example, 150 users of interest are acquired. The average value of user hidden state information for all 1000 users and the average value of user hidden state information for the 150 users of interest are calculated.

[0071] 17, the output control unit 111-3 displays the user hidden state information of the user of interest whose calculated difference is larger than that of the other users in a manner different from that of the other users (step S306). For example, the output control unit 111-3 displays the user hidden state information whose difference is larger in an emphasized manner.

[0072] FIG. 20 is a diagram showing an example of displaying statistical information on user hidden state information for a user of interest. Squares represent statistical information for all users, and circles represent statistical information for the user of interest. FIG. 20 shows an example of displaying user hidden state information in descending order of the difference in average value between all users and the user of interest, as a method of highlighting user hidden state information with large differences. By displaying H3 and H1, which have large differences, on the left side, analysts can immediately discover hidden states characteristic of the user of interest.

[0073] In this way, in the third embodiment, a user of interest can be specified, and output can be controlled according to the difference in hidden state between the specified user of interest and all users. This can further reduce the workload of user analysis using purchase data.

[0074] (Fourth embodiment) The information processing device of the fourth embodiment specifies known product hidden state information or user hidden state information, and reflects the known relationship between the product and the hidden state or the relationship between the user and the hidden state in the calculation of product hidden state information and user hidden state information for new purchasing data.

[0075] 21 is a block diagram showing an example of the configuration of an information processing device 100-4 according to the fourth embodiment. As shown in FIG. 21, the information processing device 100-4 includes an acquisition unit 101-4, a state calculation unit 102-4, an output control unit 111, and a storage unit 121.

[0076] In the fourth embodiment, the functions of an acquisition unit 101-4 and a state calculation unit 102-4 are different from those in the first embodiment. The other configurations and functions are the same as those in FIG. 1, which is a block diagram of the information processing device 100 according to the first embodiment, so the same reference numerals are used and the description thereof will be omitted here.

[0077] The acquisition unit 101-4 differs from the acquisition unit 101 of the first embodiment in that it further has the function of acquiring known information, which is at least one of previously obtained user hidden state information and previously obtained product hidden state information.

[0078] State calculation unit 102-4 differs from state calculation unit 102 of the first embodiment in that it performs matrix decomposition using known information as initial values.

[0079] Next, the analysis support process by the information processing device 100-4 according to the fourth embodiment will be described with reference to Fig. 22. Fig. 22 is a flowchart showing an example of the analysis support process according to the third embodiment.

[0080] The acquisition unit 101-4 acquires purchase data and known information (step S401). The known information may be the result of processing previously performed by the information processing device 100-4. For example, the acquisition unit 101-4 acquires product hidden state information as shown in FIG. 7 as the known information.

[0081] The state calculation unit 102-4 uses such known information as initial values ​​to perform matrix decomposition on the latest purchase data, thereby calculating how many hidden states corresponding to previously revealed purchase patterns the user of the latest purchase data has.

[0082] The known information may be set to reflect the analyst's knowledge. FIG. 23 is a diagram showing an example of known information set in this manner. For example, assume that knowledge has been obtained that a purchasing pattern shows a tendency for a product with a product ID of "I0001" and a product with a product ID of "I0003" to be purchased at the same time. Based on this knowledge, FIG. 23 shows that for the hidden state H1, the values ​​of "I0001" and "I0003" are set to 1 to link the two products. Furthermore, since other relationships are unknown, random initial values ​​are set for the other hidden states.

[0083] 21, the state calculation unit 102-4 performs matrix decomposition on the purchase matrix using the known information as an initial value, and calculates user hidden state information and product hidden state information (step S403). Step S404 is the same as step S104 in the first embodiment, and therefore a description thereof will be omitted.

[0084] In matrix decomposition, starting from an initial value, a process of updating the element values ​​of the matrix is ​​repeatedly executed. During matrix decomposition, the state calculation unit 102-4 may decompose the matrix while fixing (without updating) the known part of the hidden state information, or may decompose the matrix while also updating the known part. In the former case, the known purchasing patterns are not updated, so it is possible to calculate how past purchasing patterns are weighted for each user in the latest purchasing data. In the latter case, the known purchasing patterns are updated, so it is possible to update the purchasing patterns when there is a slight change in the purchasing pattern due to product replacement or the like.

[0085] By using known information as the initial value, it is possible to improve the accuracy of the matrix decomposition more than when, for example, random initial values ​​are used.

[0086] (Fifth embodiment) The information processing device according to the fifth embodiment has a function of classifying multiple users into clusters, as in the second embodiment, and a function of acquiring a designation of a noted user, as in the third embodiment. The information processing device according to this embodiment also calculates a cluster ratio, which is the ratio of the number of users in each cluster to the number of users in all clusters, for the noted user and all users. Furthermore, the information processing device according to this embodiment highlights and displays clusters with a large difference between the cluster ratio of all users and the cluster ratio of the noted user.

[0087] 24 is a block diagram showing an example of the configuration of an information processing device 100-5 according to the fifth embodiment. As shown in FIG. 24, the information processing device 100-5 includes an acquisition unit 101-3, a state calculation unit 102, a classification unit 103-2, a difference calculation unit 104-5, an output control unit 111-5, and a storage unit 121.

[0088] The acquisition unit 101-3 is the same as in the third embodiment, and the classification unit 103-2 is the same as in the second embodiment. In this embodiment, a difference calculation unit 104-5 is added, and the function of the output control unit 111-5 is changed. The other configurations and functions are the same as in FIG. 1, which is the block diagram of the information processing device 100 according to the first embodiment, so the same reference numerals are used and the description here will be omitted.

[0089] The difference calculation unit 104-5 calculates the difference between the cluster ratio for all users and the cluster ratio for the user of interest. For example, for each cluster CA (first cluster) included in a plurality of clusters, the difference calculation unit 104-5 calculates a cluster ratio RA that represents the ratio (first ratio) of the number of users belonging to cluster CA to the number of users belonging to all clusters for all users. Furthermore, the difference calculation unit 104-5 calculates a cluster ratio RB that represents the ratio (second ratio) of the number of users belonging to cluster CA to the number of users belonging to all clusters for the user of interest. Then, the difference calculation unit 104-5 calculates the difference between the cluster ratio RA and the cluster ratio RB.

[0090] Output control unit 111-5 further outputs information indicating a cluster for which the difference calculated by difference calculation unit 104-5 is larger than other clusters in a format different from that of the other clusters.

[0091] Next, the analysis support process by the information processing device 100-5 according to the fifth embodiment will be described with reference to Fig. 25. Fig. 25 is a flowchart showing an example of the analysis support process according to the fifth embodiment.

[0092] Steps S501 to S504 are the same as steps S201 to S205 (FIG. 12) in the information processing device 100-2 according to the second embodiment, and therefore description thereof will be omitted.

[0093] In step S505, similarly to step S304 (FIG. 17) in the third embodiment, the acquisition unit 101-3 assigns a noted user label to the noted user (step S505). As a result, a noted user label is assigned to each user in addition to the classification result into a cluster.

[0094] Fig. 26 is a diagram showing an example of the cluster classification results and notable user labels. In Fig. 26, each user is assigned the cluster ID of the cluster into which the user is classified and a notable user label. As in Fig. 18, the notable user label is assigned as "True" if the user is a notable user, and "False" if not.

[0095] Returning to FIG. 25, the difference calculation unit 104-5 calculates the cluster ratio RA of all users and the cluster ratio RB of the user of interest (step S506).

[0096] Fig. 27 is a diagram illustrating an example of a cluster ratio RA for all users. In Fig. 27, for example, the total number of users is 1000, and the number of users belonging to cluster C1 is 100, so the cluster ratio RA of cluster C1 is calculated to be 0.1.

[0097] Fig. 28 is a diagram showing an example of a cluster ratio RB for a user of interest. As shown in Fig. 28, the number of users belonging to each cluster is calculated for only users whose user of interest label is "True." In Fig. 28, for example, the total number of users of interest is 150, and the number of users of interest belonging to cluster C1 is 30, so the cluster ratio RB for cluster C1 is calculated to be 0.2.

[0098] 25, difference calculation section 104-5 calculates the difference between the cluster ratio R of all users and the cluster ratio R of the user of interest. For example, the difference in cluster ratio for cluster C1 is 0.2-0.1=0.1, which is the cluster ratio R of the user of interest minus the cluster ratio R of all users.

[0099] The output control unit 111-5 highlights and displays information (cluster information) of clusters with large differences in cluster ratios (step S508).

[0100] Figure 29 is a diagram showing an example of displaying cluster information for clusters with large differences in cluster ratios. In Figure 29, the cluster ratios for all users and the cluster ratio for the user of interest are displayed in a bar graph in descending order of the difference in cluster ratios. From the display in Figure 29, analysts can confirm that the ratios of clusters C20 and C1 for the user of interest are larger than for the whole population, and by combining this with the information in Figures 9 and 13, they can confirm what types of purchasing preferences the user segment of users of interest has.

[0101] (Display screen example) The output control units (output control unit 111, output control unit 111-2, output control unit 111-3, output control unit 111-5) of the first to fifth embodiments may aggregate and display user information, product information, and purchase time information, etc. This makes it easier for analysts to plan purchasing preference types.

[0102] Fig. 30 is a diagram showing an example of a screen displaying user information and product information in addition to the processing results in the second embodiment. In the example screen of Fig. 30, statistical values ​​of user information and product information are displayed in addition to the average value of the user hidden state information of cluster C1 shown in the second embodiment.

[0103] The user information shows the lift value, which is the ratio of the age and gender ratio of users belonging to cluster C1 to the age and gender ratio of all users. The product information shows the average price and product category of the user hidden state information with a large value.

[0104] By looking at the information in Figure 30, the analyst can easily see that the analysis target is 100 users who purchase instant noodles and snacks, and that a large proportion of them are men in their 40s and 50s. Furthermore, with background knowledge such as the standard prices of product categories, the analyst can infer purchasing preferences regarding price, such as luxury orientation and sale preference.

[0105] Fig. 31 is a diagram showing an example of a screen in which the number of purchases of a certain product of interest per hour is plotted for each cluster based on the processing results in the second embodiment. The example screen of Fig. 31 shows that users belonging to cluster C1 often purchase the product of interest that the analyst is focusing on around 6 p.m. on weekdays. Fig. 31 shows the time patterns of purchases of the product of interest in each cluster, and by combining this with the information in Fig. 30, the analyst can understand the temporal characteristics of each cluster and obtain information that is useful for planning purchasing preference types.

[0106] As described above, according to the first to fifth embodiments, the workload of user analysis using purchase data can be reduced.

[0107] Next, the hardware configuration of the information processing device according to the first to fifth embodiments will be described with reference to Fig. 32. Fig. 32 is an explanatory diagram showing an example of the hardware configuration of the information processing device according to the first to fifth embodiments.

[0108] The information processing device according to the first to fifth embodiments includes a control device such as a CPU 51, a storage device such as a ROM (Read Only Memory) 52 or a RAM 53, a communication I / F 54 that connects to a network and communicates, and a bus 61 that connects each part.

[0109] The programs executed by the information processing devices according to the first to fifth embodiments are provided in advance in the ROM 52 or the like.

[0110] The programs executed by the information processing devices according to the first to fifth embodiments may be configured to be provided as a computer program product by being recorded in an installable or executable file format on a computer-readable recording medium such as a CD-ROM (Compact Disk Read Only Memory), a flexible disk (FD), a CD-R (Compact Disk Recordable), or a DVD (Digital Versatile Disk).

[0111] Furthermore, the programs executed by the information processing devices according to the first to fifth embodiments may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Also, the programs executed by the information processing devices according to the first to fifth embodiments may be provided or distributed via a network such as the Internet.

[0112] The programs executed by the information processing devices according to the first to fifth embodiments can cause a computer to function as each unit of the information processing device described above. In this computer, the CPU 51 can read the programs from a computer-readable storage medium onto a main storage device and execute the programs.

[0113] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]

[0114] 100, 100-2, 100-3, 100-4, 100-5 Information processing device 101, 101-3, 101-4 Acquisition section 102, 102-4 State calculation unit 103-2 Classification section 104-3, 104-5 Difference calculation part 111, 111-2, 111-3, 111-5 Output control section 121 Storage section

Claims

1. an acquisition unit that acquires a plurality of purchase data including any one of a plurality of pieces of user identification information that identify a plurality of users, any one of a plurality of pieces of product identification information that identify a plurality of products, and performance information that includes at least one of the price and the number of purchases of the products; a state calculation unit that performs matrix decomposition on a purchase matrix in which the plurality of pieces of user identification information and the plurality of pieces of product identification information are used as row and column indices, and in which element values ​​are non-negative values ​​calculated based on the performance information, to calculate user hidden state information indicating a relationship between the plurality of pieces of user identification information and a hidden state related to purchases, and product hidden state information indicating a relationship between the hidden state and the plurality of pieces of product identification information; an output control unit that controls output of at least one of the user hidden state information and the product hidden state information; a classification unit that classifies the plurality of pieces of user identification information included in the user hidden state information into a plurality of clusters using similarities between the pieces of user hidden state information; a difference calculation unit that calculates, for each first cluster included in the plurality of clusters, a difference between a first ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for all users, and a second ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for a user of interest designated as a user of interest, the output control unit outputs information indicating a cluster in which the difference is larger than other clusters in a manner different from that of the other clusters. An information processing device comprising:

2. the output control unit outputs statistical information of the user hidden state information for each of the clusters. The information processing device according to claim 1 .

3. An acquisition unit that acquires multiple purchase data including any one of multiple user identification information that identifies multiple users, any one of multiple product identification information that identifies multiple products, and performance information that includes at least one of the price and purchase quantity of the products; a state calculation unit that performs matrix decomposition on a purchase matrix in which the plurality of pieces of user identification information and the plurality of pieces of product identification information are used as row and column indices, and in which element values ​​are non-negative values ​​calculated based on the performance information, to calculate user hidden state information indicating a relationship between the plurality of pieces of user identification information and a hidden state related to purchases, and product hidden state information indicating a relationship between the hidden state and the plurality of pieces of product identification information; an output control unit that controls output of at least one of the user hidden state information and the product hidden state information; a difference calculation unit that calculates a difference between the hidden state for the plurality of users and the hidden state corresponding to a target user designated as a target user among the plurality of users, the output control unit outputs the user hidden state information of the focused user whose difference is larger than that of other users in a manner different from that of the other users. Information processing device.

4. a classification unit that classifies the plurality of pieces of user identification information included in the user hidden state information into a plurality of clusters using similarities between the pieces of user hidden state information; The target user is a user identified by the user identification information included in a specified cluster among the plurality of clusters. The information processing device according to claim 3 .

5. An acquisition unit that acquires multiple purchase data including any one of multiple user identification information that identifies multiple users, any one of multiple product identification information that identifies multiple products, and performance information that includes at least one of the price and purchase quantity of the products; a state calculation unit that performs matrix decomposition on a purchase matrix in which the plurality of pieces of user identification information and the plurality of pieces of product identification information are used as row and column indices, and in which element values ​​are non-negative values ​​calculated based on the performance information, to calculate user hidden state information indicating a relationship between the plurality of pieces of user identification information and a hidden state related to purchases, and product hidden state information indicating a relationship between the hidden state and the plurality of pieces of product identification information; an output control unit that controls output of at least one of the user hidden state information and the product hidden state information; The acquisition unit acquires known information that is at least one of the user hidden state information previously acquired and the product hidden state information previously acquired, the state calculation unit performs the matrix decomposition using the known information as an initial value. Information processing device.

6. The element value is a non-negative value indicating whether or not a purchase has been made, the price, or the number of purchases.

6. The information processing device according to claim 1.

7. An information processing method executed by an information processing device, an acquisition step of acquiring a plurality of purchase data including any one of a plurality of pieces of user identification information that identify a plurality of users, any one of a plurality of pieces of product identification information that identify a plurality of products, and performance information that includes at least one of the price and the number of purchases of the products; a state calculation step of matrix decomposing a purchase matrix in which the plurality of pieces of user identification information and the plurality of pieces of product identification information are used as row and column indices, and element values ​​are non-negative values ​​calculated based on the performance information, to calculate user hidden state information indicating a relationship between the plurality of pieces of user identification information and hidden states related to purchases, and product hidden state information indicating a relationship between the hidden state and the plurality of pieces of product identification information; an output control step of controlling output of at least one of the user hidden state information and the product hidden state information; a classification step of classifying the plurality of pieces of user identification information included in the user hidden state information into a plurality of clusters using similarities between the pieces of user hidden state information; a difference calculation step of calculating, for each first cluster included in the plurality of clusters, a difference between a first ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for all users, and a second ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for a user of interest designated as a user of interest, the output control step outputs information indicating a cluster in which the difference is larger than other clusters in a manner different from that of the other clusters. Information processing methods.

8. On the computer, an acquisition step of acquiring a plurality of purchase data including any one of a plurality of pieces of user identification information that identify a plurality of users, any one of a plurality of pieces of product identification information that identify a plurality of products, and performance information that includes at least one of the price and the number of purchases of the products; a state calculation step of matrix decomposing a purchase matrix in which the plurality of pieces of user identification information and the plurality of pieces of product identification information are used as row and column indices, and element values ​​are non-negative values ​​calculated based on the performance information, to calculate user hidden state information indicating a relationship between the plurality of pieces of user identification information and hidden states related to purchases, and product hidden state information indicating a relationship between the hidden state and the plurality of pieces of product identification information; an output control step of controlling output of at least one of the user hidden state information and the product hidden state information; a classification step of classifying the plurality of pieces of user identification information included in the user hidden state information into a plurality of clusters using similarities between the pieces of user hidden state information; a difference calculation step of calculating, for each first cluster included in the plurality of clusters, a difference between a first ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for all users, and a second ratio of the number of users belonging to the first cluster to the number of users belonging to the plurality of clusters for a user of interest designated as a user of interest, the output control step outputs information indicating a cluster in which the difference is larger than other clusters in a manner different from that of the other clusters. program.

Citation Information

Patent Citations

  • Etching method for insulation film

    JP1984002325A

  • Cluster extraction device, cluster extraction method, and cluster extraction program

    JP2016031639A

  • Analysis device, analysis method and analysis program

    JP2016081371A

  • Data analyzer, method and program

    JP2016177485A