Information processing method, information processing device, information processing program, and computer-readable storage medium storing the information processing program.

By clustering matrix data with PLSA and applying t-SNE to visualize relative positional relationships, the method enhances the identification of factors causing a specific factor to have a value of 0 or 1, addressing the limitations of existing methods in capturing trend differences within latent classes.

JP7859173B2Active Publication Date: 2026-05-15MAZDA MOTOR CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAZDA MOTOR CORP
Filing Date
2022-04-27
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods struggle to clearly identify factors that cause a specific factor in matrix data to have a value of 0 or 1, particularly when analyzing large datasets using stochastic latent semantic analysis (PLSA) and t-distribution stochastic neighbor embedding (t-SNE), as low-dimensional vectors fail to capture differences in trends among factors belonging to the same latent class.

Method used

An information processing method that applies PLSA to cluster matrix data into a quasi-block-diagonal form, generates multidimensional vectors for each latent class, and uses t-SNE to reduce dimensions further, visualizing the relative positional relationships between low-dimensional vectors to identify factors contributing to the target factor's value.

Benefits of technology

This approach allows for clearer identification of factors influencing the value of a particular factor by highlighting differences through relative positional relationships, providing a more nuanced understanding of the factors contributing to the target factor's value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859173000002
    Figure 0007859173000002
  • Figure 0007859173000003
    Figure 0007859173000003
  • Figure 0007859173000004
    Figure 0007859173000004
Patent Text Reader

Abstract

To more clearly grasp a cause that the value of a specific object factor is 0 or 1 in matrix data analysis as compared with prior art.SOLUTION: An information processing method includes: a step S1 of applying a PLSA method to matrix data 49; a step S2 of generating a multidimensional vector using, as a basis, each of object persons ni who belong to a specific latent class z' in the matrix data 49 to which the PLSA method has been applied and using a value of a factor pj which is assigned to each of the object persons ni as a component of each basis; a step S3 of generating a low-dimensional vector by applying a t-SNE method to the multidimensional vector; and a step S4 of visualizing a relative positional relationship between a low-dimensional vector corresponding to an object factor pf belonging to the specific latent class z' and a low-dimensional vector corresponding to each comparison factor pj" belonging to other latent classes z"; and a cause analysis step S7 of analyzing a cause that a value of the object factor is 0 or 1, based on the visualized relative positional relationship.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technologies disclosed herein relate to an information processing method, an information processing device, an information processing program, and a computer-readable storage medium storing the information processing program. [Background technology]

[0002] As an analytical method for complex big data, the so-called stochastic latent semantic analysis (PLSA) is widely known. PLSA is a method that reduces the dimensionality of at least one of the row and column components by clustering the matrix being analyzed according to its latent variables.

[0003] For example, when PLSA is applied to a matrix where the row components are survey respondents, the column components are survey items, and the values ​​are the survey responses, each latent class classified by the latent variables will contain one or more survey respondents and one or more survey items.

[0004] In this case, assuming that the matrix includes survey items such as "Do you prefer manufacturer A or not?", "Do you prefer manufacturer B or not?", and "Do you prefer manufacturer C or not?", it becomes possible to provide interpretations for survey respondents belonging to the same latent class according to the survey items, such as "Respondents belonging to the first latent class prefer manufacturer A, respondents belonging to the second latent class prefer manufacturer B, respondents belonging to the third latent class prefer manufacturer C, and so on."

[0005] As a specific example of PLSA application, for instance, Patent Document 1 discloses the use of PLSA to cluster mined data in a method for acquiring, analyzing, and mining target data and / or information.

[0006] Furthermore, the so-called t-distribution stochastic neighbor embedding method (t-SNE) is known as an algorithm for reducing the dimensionality of complex high-dimensional data. By using t-SNE, it becomes possible to transform high-dimensional data into 2D or 3D data while maintaining the local relationships between data points.

[0007] As a specific application example of t-SNE, for example, Patent Document 2 below discloses that a processing unit uses t-SNE to generate a two-dimensional or three-dimensional space in which each element of reference data is mapped. [Prior art documents] [Patent Documents]

[0008] [Patent Document 1] Special Publication No. 2009-525514 [Patent Document 2] Japanese Patent Publication No. 2019-91454 [Overview of the project] [Problems that the invention aims to solve]

[0009] The inventors of this application attempted to perform an analysis using PLSA on a matrix formed by creating a matrix of multiple subjects to be analyzed and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects to be analyzed.

[0010] In this case, the latent classes obtained by PLSA are characterized by one or more analytes belonging to each latent class, and one or more factors belonging to the same latent class as those analytes. Here, the values ​​of factors belonging to a given latent class will be similar among analytes belonging to the same latent class, but they will not necessarily be exactly the same.

[0011] To explain using the example mentioned earlier, "the survey respondents classified as a latent class that prefers manufacturer A include not only respondents who actually own manufacturer A's products (respondents assigned a value of 1), but also respondents who do not own manufacturer A's products (respondents assigned a value of 0)."

[0012] The inventors of this application considered analyzing the factors that cause a specific factor belonging to a particular latent class (hereinafter referred to as the "target factor") to have a value of 0 or 1, based on the relationship between that target factor and other factors.

[0013] One possible method for performing such analysis is to generate a multidimensional vector with the same number of dimensions as the number of analytes belonging to a particular latent class by treating each factor constituting a row or column component of a matrix corresponding to that latent class as a distinct vector. In this case, by analyzing the Euclidean distance between the vector corresponding to the target factor and the vectors corresponding to other factors, it becomes possible to extract factors that have the same trend as the value of the target factor, or conversely, factors that have the opposite trend to the value of the target factor.

[0014] However, when the number of subjects being analyzed is large, the aforementioned multidimensional vectors are difficult to visualize and intuitively grasp. Therefore, it is conceivable to apply t-SNE to the multidimensional vectors corresponding to each factor belonging to a specific latent class to transform them into lower-dimensional vectors.

[0015] However, as mentioned above, the values ​​of factors belonging to the same latent class can be similar among the subjects analyzed who belong to the same latent class. Therefore, low-dimensional vectors corresponding to such factors alone are insufficient to capture the difference in trends compared to the low-dimensional vector corresponding to the target factor.

[0016] This disclosure is made in view of the above, and its purpose is to more clearly identify the factors that caused a particular target factor to have a value of 0 or 1 when analyzing matrix data obtained by creating a matrix of multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects. [Means for solving the problem]

[0017] A first aspect of this disclosure relates to an information processing method for analyzing matrix data obtained by creating a matrix of multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects, using a computer equipped with an arithmetic unit for executing a program, and for visualizing the results of the analysis.

[0018] Furthermore, according to a first aspect of the present disclosure, the information processing method includes: a PLSA step in which the arithmetic unit clusters the matrix data by applying the PLSA method to the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes; an analysis target extraction step in which the arithmetic unit generates a multidimensional vector in the clustered matrix data, with each analysis target belonging to a specific latent class as a base and the values ​​of the factors assigned to each analysis target as components of each base, including factors belonging to other latent classes other than the specific latent class; and an analysis target extraction step in which the arithmetic unit applies the t-SNE method to the multidimensional vector. The system includes: a dimensionality reduction step of compressing the multidimensional vector into a two-dimensional or three-dimensional low-dimensional vector; a vector visualization step in which, assuming one factor belonging to a specific latent class is the objective factor and other factors belonging to other latent classes are multiple comparison factors, the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors, based on each low-dimensional vector obtained in correspondence with the multidimensional vector; and a factor analysis step of analyzing the factors that caused the value of the objective factor to be 0 or 1 based on the relative positional relationship visualized by the vector visualization step.

[0019] In the PLSA step described above, a given matrix is ​​rearranged (clustered) so that it approaches a block-diagonal matrix. Hereafter, the matrix after clustering will also be called a "quasi-block-diagonal matrix". Unlike the non-block-diagonal portion of a block-diagonal matrix, the non-block-diagonal portion of a "quasi-block-diagonal matrix" can take non-zero values.

[0020] According to the first embodiment described above, the t-SNE method is applied to a multidimensional vector in which each analyte belonging to a specific latent class (hereinafter simply referred to as "specific class") is used as a base, and the values ​​of the factors assigned to each analyte are used as components of each base, thereby reducing the dimensionality of the multidimensional vector to a lower-dimensional vector.

[0021] Furthermore, when visualizing low-dimensional vectors, we compare the objective factor belonging to a specific class with other comparison factors. For the latter comparison factors, we select factors belonging to other latent classes, rather than factors belonging to the specific class.

[0022] In other words, as is evident from the fact that each block matrix obtained by clustering (or, in other words, rearranging matrices to approximate a block diagonal matrix) corresponds to a latent class, low-dimensional vectors corresponding to factors belonging to the same latent class tend to show similar trends (for example, the vector lengths tend to be similar, or the inner product tends to be close to 1). Here, "block matrix" refers to the matrix that constitutes the block diagonal portion of the quasi-block diagonal matrix.

[0023] In contrast, by comparing the objective factor belonging to a specific class with the comparative factors belonging to other latent classes, the differences between low-dimensional vectors can be clarified. With these differences clarified, by searching for comparative factors that are close to or farther from the objective factor, it becomes possible to more clearly identify the factors that may have triggered the objective factor's value to be 0 or 1 than before.

[0024] Furthermore, comparing with comparison factors belonging to other latent classes is equivalent to searching for comparison factors that, although classified into different latent classes at the PLSA stage, still strongly correlate with the target factor. This provides insights that cannot be obtained with conventional PLSA methods, and allows for a more multifaceted understanding of the factors that caused a particular target factor to have a value of 0 or 1.

[0025] Furthermore, according to a second aspect of this disclosure, in the vector visualization step, the calculation unit visualizes the Euclidean distance between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the plurality of comparison factors as the relative positional relationship, and in the factor analysis step, the calculation unit searches for the factors that contributed to the value of the objective factor from among the plurality of comparison factors based on the length of each Euclidean distance visualized by the calculation unit.

[0026] According to the second embodiment, the factor analysis step is configured to search for factors based on the length of each Euclidean distance. This configuration makes it possible to visualize and search in a more intuitively understandable format. This is advantageous in identifying factors where the value of a particular target factor is 0 or 1.

[0027] Furthermore, according to a third aspect of this disclosure, the objective factor may be a flag indicating whether or not a person possesses a particular article.

[0028] According to the third embodiment described above, the basis for constructing the multidimensional vector corresponding to the objective factor includes the subjects of analysis whose flag is set to 1, that is, the subjects of analysis who actually own the item. By performing the aforementioned visualization on the low-dimensional vector generated based on such a multidimensional vector, it becomes possible to grasp the factors that led to owning the item, or the factors that led not to owning it, from a wider range of perspectives than before.

[0029] Furthermore, according to a fourth aspect of this disclosure, in the vector visualization step, one of the plurality of comparison factors is set as the second objective factor, the calculation unit visualizes the Euclidean distance between the low-dimensional vector corresponding to the second objective factor and the low-dimensional vectors corresponding to each of the remaining comparison factors excluding the second objective factor, and in the factor analysis step, the calculation unit searches for factors that contributed to the value of the second objective factor from among the remaining comparison factors based on the length of each Euclidean distance visualized by the calculation unit, and analyzes the factors that caused the value of the objective factor to be 0 or 1 by comparing the factors that contributed to the value of the second objective factor with the factors that contributed to the value of the objective factor.

[0030] According to the fourth embodiment described above, by performing an analysis using a second objective factor, it becomes possible to grasp from a wider range of perspectives the factors that caused the value of the first objective factor to be 0 or 1. For example, if the second objective factor is a flag indicating whether or not the user owns a competing product of the first objective factor, then by exploring the factors that caused the second objective factor to be 0 or 1, it becomes possible to clarify the areas for improvement, selling points, etc., of the competing product. By analyzing these findings, it becomes possible to clarify (visualize) the areas that need improvement in the product corresponding to the first objective factor.

[0031] Furthermore, according to a fifth aspect of this disclosure, the information processing method, if the latent class to which the second objective factor belongs is called the second specific class, includes a second analysis target extraction step in which the calculation unit generates a second multidimensional vector in the matrix data after clustering, with each analysis target belonging to the second specific class as a base and the value of the factor assigned to each analysis target as a component of each base, including factors belonging to other latent classes other than the second specific class; a second dimensionality reduction step in which the calculation unit applies the t-SNE method to the second multidimensional vector to compress the second multidimensional vector into a two-dimensional or three-dimensional low-dimensional vector; and a comparison between the low-dimensional vector corresponding to the objective factor, the low-dimensional vector corresponding to the second objective factor, and the remaining comparison based on each low-dimensional vector obtained corresponding to the second multidimensional vector. The factor analysis step may include a second vector visualization step in which the calculation unit visualizes the Euclidean distance between each of the factors and the corresponding low-dimensional vector, and the factor analysis step may include a first search step in which the calculation unit searches for factors that contributed to the values ​​of the target factor and the second target factor from among the remaining comparison factors based on the length of each Euclidean distance visualized by the calculation unit in the vector visualization step, and a second search step in which the calculation unit searches for factors that contributed to the values ​​of the target factor and the second target factor from among the remaining comparison factors based on the length of each Euclidean distance visualized by the calculation unit in the second vector visualization step, and the factors that caused the value of the target factor to be 0 or 1 are analyzed by comparing the factors searched in the first search step and the factors searched in the second search step.

[0032] According to the fifth embodiment described above, an analysis is performed in which the latent class to which the objective factor (first objective factor) belongs is designated as a specific class, and then an analysis is performed in which the latent class to which the second objective factor belongs is changed to a specific class. The latter analysis is based on a multidimensional vector in which each analyte belonging to the same latent class as the second objective factor is used as a base, and the factor values ​​assigned to each analyte are used as components of each base. Therefore, the latter analysis and the former analysis differ in the breakdown of analytes used to generate the low-dimensional vector.

[0033] By comparing the results of two analyses with different compositions of the subjects analyzed, it becomes possible to understand the factors that resulted in the target factor being 0 or 1 from a more multifaceted perspective.

[0034] Furthermore, according to a sixth aspect of this disclosure, the objective factor is a flag indicating whether or not a predetermined article is owned, and the second objective factor is a different article from the predetermined article. possession A flag indicating whether or not the predetermined article and the other Goods This may be defined as articles of the same type but from different manufacturers.

[0035] According to the sixth embodiment, by exploring the factors that cause the second objective factor to become 0 or 1, it becomes possible to identify areas for improvement, selling points, etc., of the other article. This makes it possible to clarify areas for improvement, etc., of the predetermined article.

[0036] A seventh aspect of this disclosure relates to an information processing device that includes a computer equipped with an arithmetic unit for executing a program, analyzes matrix data obtained by matrixing a plurality of subjects for analysis and a plurality of factors consisting of values ​​of 0 or 1 assigned to each of the plurality of subjects for analysis, and visualizes the results of the analysis.

[0037] Furthermore, according to a seventh aspect of the present disclosure, the information processing apparatus includes a PLSA means for clustering the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes, by applying the PLSA method to the matrix data by the calculation unit, and a multidimensional vector in the clustered matrix data, where each analysis subject belonging to a specific latent class is a base, and the values ​​of the factors assigned to each analysis subject are components of each base, including factors belonging to other latent classes other than the specific latent class, generated by the calculation unit. The analysis target extraction means, the calculation unit applies the t-SNE method to the multidimensional vector to compress the multidimensional vector into a 2-dimensional or 3-dimensional low-dimensional vector, the calculation unit uses one factor belonging to a specific latent class as the target factor and other factors belonging to other latent classes as multiple comparison factors, and the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the target factor and the low-dimensional vector corresponding to each of the multiple comparison factors based on each low-dimensional vector obtained corresponding to the multidimensional vector, and the vector visualization means The system includes a factor analysis means for analyzing the factors that cause the value of the objective factor to be 0 or 1, based on the relative positional relationship visualized by the system.

[0038] According to the seventh embodiment described above, when analyzing matrix data, the factors that cause a particular objective factor to have a value of 0 or 1 can be identified more clearly than in the conventional method.

[0039] An eighth aspect of this disclosure relates to an information processing program that analyzes matrix data, which is obtained by creating a matrix of multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects, by having it run on a computer equipped with an arithmetic unit for executing the program, and visualizes the results of the analysis.

[0040] Furthermore, according to an eighth aspect of the present disclosure, the information processing program includes: a PLSA step in which the computer clusters the matrix data by applying the PLSA method to the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes; an analysis target extraction step in which the computer generates a multidimensional vector in the clustered matrix data, with each analysis target belonging to a specific latent class as a basis and the values ​​of the factors assigned to each analysis target as components of each basis, including factors belonging to other latent classes other than the specific latent class; and the computer then applies t-S to the multidimensional vector. By applying the NE method, the following steps are performed: a dimensionality reduction step in which the multidimensional vector is compressed into a 2-dimensional or 3-dimensional low-dimensional vector; a vector visualization step in which, assuming one factor belonging to a specific latent class is the objective factor and other factors belonging to other latent classes are multiple comparison factors, the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors, based on each low-dimensional vector obtained in correspondence with the multidimensional vector; and a factor analysis step in which, based on the relative positional relationship visualized by the vector visualization step, the factors that caused the value of the objective factor to be 0 or 1 are analyzed.

[0041] According to the eighth aspect described above, when analyzing matrix data, the factors that caused a particular objective factor to have a value of 0 or 1 can be identified more clearly than in the conventional method.

[0042] Furthermore, a ninth aspect of this disclosure relates to a computer-readable storage medium characterized by storing the information processing program.

[0043] According to the ninth embodiment described above, when analyzing matrix data, the factors that caused a particular objective factor to have a value of 0 or 1 can be identified more clearly than in the conventional method. [Effects of the Invention]

[0044] As explained above, this disclosure makes it possible to more clearly identify the factors that caused a particular objective factor to have a value of 0 or 1 when analyzing matrix data. [Brief explanation of the drawing]

[0045] [Figure 1] Figure 1 is a diagram illustrating the hardware configuration of an information processing device. [Figure 2] Figure 2 is a diagram illustrating the software configuration of an information processing device. [Figure 3] Figure 3 is a flowchart illustrating the steps of an information processing method. [Figure 4] Figure 4 is a flowchart illustrating the steps of the PLSA procedure. [Figure 5] Figure 5 is an example of matrix data. [Figure 6] Figure 6 is a diagram illustrating the basic concepts of the PLSA steps. [Figure 7] Figure 7 is an example of the clustering results. [Figure 8] Figure 8 is a flowchart illustrating the steps for the analysis target extraction step. [Figure 9] Figure 9 is a diagram illustrating the basic concept of the analysis target selection step. [Figure 10] Figure 10 is a flowchart illustrating the steps of the dimensionality reduction process. [Figure 11] Figure 11 is a flowchart illustrating the steps of the vector visualization process. [Figure 12] Figure 12 is a flowchart illustrating the steps of the factor analysis process. [Figure 13] Figure 13 illustrates the content that is visualized by the vector visualization step. [Figure 14] Figure 14 illustrates the content visualized by the second vector visualization step. [Figure 15] Figure 15 shows an example of a change in the display mode. [Modes for carrying out the invention]

[0046] The embodiments of this disclosure will be described below with reference to the drawings. Note that the following description is illustrative.

[0047] <Device configuration> Figure 1 is a diagram illustrating the hardware configuration of the information processing device (specifically, computer 1 that constitutes the information processing device) related to this disclosure, and Figure 2 is a diagram illustrating its software configuration.

[0048] As illustrated in Figure 1, computer 1 comprises a Central Processing Unit (CPU) 3 that controls the entire computer 1, a Read Only Memory (ROM) 5 that stores boot programs and the like, a Random Access Memory (RAM) 7 that functions as main memory, and a Hard Disk Drive (HDD) 9 as secondary storage. Note that a Solid State Drive (SSD) or the like can be used instead of the HDD 9 as secondary storage.

[0049] Of these elements, the CPU3 executes various programs. The CPU3 functions as the arithmetic unit in this embodiment. The RAM7 and HDD9 temporarily or continuously store the programs executed by the CPU3. The RAM7 and HDD9 each function as the storage units in this embodiment.

[0050] Computer 1 also includes a display 11, a graphics memory (Video RAM: VRAM) 13 for storing image data displayed on the display 11, and a keyboard 15 and mouse 17 as a human-machine interface. The display 11 can display the calculation results of the CPU 3 and functions as a display unit in this embodiment. Furthermore, computer 1 according to this embodiment can send and receive data with external devices via a communication interface 21.

[0051] As illustrated in Figure 2, the program memory of HDD9 stores the operating system (OS) 19, PLSA program 29A, analysis target extraction program 29B, dimensionality reduction program 29C, vector visualization program 29D, specific class change program 29E, factor analysis program 29F, application program 39, and the like.

[0052] Of these elements, the PLSA program 29A, the analysis target extraction program 29B, the dimensionality reduction program 29C, the vector visualization program 29D, the specific class change program 29E, and the factor analysis program 29F constitute the information processing program 29 in this embodiment.

[0053] Here, the information processing program 29 is a program for executing the information processing method described later, and is configured to cause the computer 1 to execute each step that constitutes the method. The information processing program 29 is pre-stored in a computer-readable storage medium 18.

[0054] In the program memory of HDD9, the PLSA program 29A, the analysis target extraction program 29B, the dimensionality reduction program 29C, the vector visualization program 29D, the specific class change program 29E, and the factor analysis program 29F are each activated in response to commands input from the keyboard 15, mouse 17, etc. At that time, the PLSA program 29A, etc. are loaded from HDD9 into RAM7 and executed by CPU3.

[0055] Meanwhile, the data memory of HDD9 stores the matrix data 49 to be analyzed and the cluster data 59 obtained by clustering the matrix data 49.

[0056] In addition, various data generated by executing the PLSA program 29A, the analysis target extraction program 29B, the dimensionality reduction program 29C, the vector visualization program 29D, the specific class change program 29E, and the factor analysis program 29F, as well as the execution results of the application program 39, are stored in the data memory of the HDD9 or in the RAM7 as main memory, as needed.

[0057] The following provides a detailed explanation of the specific methodologies for information processing.

[0058] <Methodology> Figure 3 is a flowchart illustrating the steps of an information processing method. The method illustrated in Figure 3 uses computer 1 to analyze matrix data 49 and visualize the analysis results.

[0059] As shown in Figure 3, the information processing method is carried out by sequentially executing the following steps: the PLSA step (step S1), the analysis target extraction step (step S2), the dimensionality reduction step (step S3), the vector visualization step (step S4), the specific class change step (step S6), and the factor analysis step (step S7). The results of both the vector visualization step and the factor analysis step can be displayed on the display unit 11.

[0060] Of these steps, the PLSA step is performed by the CPU3 executing the aforementioned PLSA program 29A. Similarly, the analysis target extraction step is performed by the CPU3 executing the analysis target extraction program 29B, the dimensionality reduction step is performed by the CPU3 executing the dimensionality reduction program 29C, the vector visualization step is performed by the CPU3 executing the vector visualization program 29D, the specific class change step is performed by the CPU3 executing the specific class change program 29E, and the factor analysis step is performed by the CPU3 executing the factor analysis program 29F.

[0061] When CPU3 executes the PLSA program 29A, etc., computer 1 functions as an information processing device comprising: PLSA means for executing PLSA steps; analysis target extraction means for executing analysis target extraction steps; dimensionality reduction means for executing dimensionality reduction steps; vector visualization means for executing vector visualization steps; specific class change means for executing specific class change steps; and factor analysis means for executing factor analysis steps.

[0062] The following describes each step that constitutes the information processing method in order.

[0063] (PLSA Step) Figure 4 is a flowchart illustrating the steps of the PLSA process. Figure 5 is an example of matrix data 49, Figure 6 is a diagram explaining the basic concept of the PLSA process, and Figure 7 is an example of the clustering results.

[0064] Here, the flowchart illustrated in Figure 4 shows the process performed in step S1 of Figure 3. That is, when the control process proceeds to step S1 in Figure 3, the CPU 3 will execute steps S11-S14 of Figure 4 in order.

[0065] The PLSA step is configured such that CPU3 applies the PLSA method to the matrix data 49 to cluster the matrix data so that it approaches a block-diagonal matrix consisting of block matrices, each labeled by a different latent class. In other words, the PLSA step is configured to cluster the matrix data so that it becomes the aforementioned quasi-block-diagonal matrix containing block matrices, each labeled by a different latent class.

[0066] Specifically, in step S11 of Figure 4, the CPU 3 reads the matrix data 49.

[0067] This matrix data 49 is for multiple analysis subjects n i (i=1,2,…,N) and the multiple subjects n i Multiple factors p, each consisting of a value of 0 or 1 assigned to it. j This is formed by creating a matrix of (j=1,2,…,M). Hereafter, "analysis subjects" will be simply referred to as "subjects". The number of subjects (=N) is preferably 100 or more, and more preferably 1000. The number of factors (=M) is preferably 100 or more, and more preferably 200 or more.

[0068] In particular, in this embodiment, the row components of the matrix data 49 correspond to each subject, and the column components correspond to each factor. That is, if the matrix of N rows and M columns corresponding to the matrix data 49 is represented as X, the (i, j) component of X is the value of the j-th factor p i in the i-th subject n j .

[0069] The plurality of factors p j may each be a flag indicating the property, state, etc. of each subject n i . The properties of each subject n i include the gender, personality, and preferences of each subject n i . The states of each subject n i include the age, health status, and a flag indicating whether each subject n i owns a predetermined item. The plurality of factors p j may each be the result of a questionnaire answered for each subject n i . As a provisional name, the properties and states of each subject n i may be collectively referred to as "values".

[0070] In particular, in this embodiment, as the preferences of each subject n i , the plurality of factors p j include a flag indicating whether the subject owns an item manufactured by manufacturer A (for example, an automobile manufactured by company A) (factor p1: "user A" in FIG. 5), a flag indicating whether the subject owns an item of the same type as this item and manufactured by manufacturer B (for example, an automobile manufactured by company B) (factor p2: "user B" in FIG. 5), and a flag indicating whether the subject owns an item of the same type as this item and manufactured by manufacturer C (for example, an automobile manufactured by company C) (factor p3: "user C" in FIG. 5).

[0071] Specifically, in the example shown in FIG. 5, the plurality of factors p j include a flag indicating the gender of the subject n i (factors p4 and p5), and a flag indicating the gender of the subject n iFlags indicating age (factors p6 and p7: "Younger generation" and "Older generation") and the number of subjects n i Flags indicating preferences (factor p8~factor p) 11 ) and are assumed to be included.

[0072] For more details, the example shown in Figure 5 involves multiple factors p j This includes a flag indicating whether or not to pursue design aesthetics (factor p8: "design aesthetics"), a flag indicating whether or not to pursue functionality (factor p9: "functionality"), and a flag indicating whether or not to seek a sense of luxury (factor p 10 : "Premium feel") and a flag (factor p) indicating whether or not to pursue a sense of value for money. 11 : "Design quality") and other factors are indicated. In addition, the quality of eyesight, cognitive ability, etc., are considered for the target group n i Factor p is a flag indicating the health status. j This includes whether or not the target group enjoys group activities, etc. i A flag indicating the characteristics of factor p j You can include it in that.

[0073] Note that there are multiple factors p j This includes target person n i It may include a flag that is selected alternatively, such as the age of the subject n. i It may include flags that allow for overlap, such as preferences. For example, as shown in factors p6 and p7 in Figure 5, for subjects n i For age, either the young adult or older adult flag will be 1, and the other flag will be 0. In contrast, as shown in factors p8 and p9 of the same figure, for the subject n i The preference can be expressed as follows: both flags for "pursuing design" and "pursuing functionality" may be set to 1; one flag may be set to 1; or both flags may be set to 0.

[0074] Subsequently, in step S12, following step S11, the CPU3, acting as the arithmetic unit, applies the PLSA method to the matrix data 49 read in step S11. By applying the PLSA method, the matrix data 49 is rearranged to approximate a block diagonal matrix, thereby creating multiple clusters Cl z The clusters will be arranged in (z=1,2,…,Z). Here, each cluster C is a block matrix (the block diagonal portion of a quasi-block diagonal matrix). z This will be labeled by the so-called latent class (latent variable) z.

[0075] PLSA corresponds to the expansion of the joint probability distribution P(i,j) of row element i and column element j in matrix data 49 using latent class z, and expressed as shown in equation (1) below.

[0076]

number

[0077] As shown in the above equation, both the row element i and the column element j can be represented using the latent class z. In other words, multiple subjects n i It is classified into one of the latent classes z, and at the same time, multiple factors p j Each will be classified into one of the latent classes z. These classifications are all performed alternately.

[0078] In the specific calculations, the so-called EM algorithm is used to maximize the log-likelihood of the joint probability distribution P(i,j) and perform clustering. As a result, as shown in the lower part of Figure 6, the number of subjects n i and factor p j These can be classified into the first cluster Cl1 (z=1), the second cluster Cl2 (z=2), the third cluster Cl3 (z=3), and so on. z The constituent components are arranged so that those with a high probability of having a flag of 1 are grouped together. In other words, as shown in the (1,2) component of the first cluster Cl1, each cluster Clz The constituent components are not necessarily always equal to 1.

[0079] Note that the lower part of Figure 6 is merely a general example. As mentioned above, if the number of subjects is set to 100 or more and the number of factors is set to 100 or more, each cluster Cl z n individuals belonging to this group i and factor p j The number of each can range from several tens to around one hundred.

[0080] Subsequently, in step S13 following step S12, CPU3 assigns the customer ID (target person n) to each latent class z. i The argument i) and factor ID (factor p) j The argument j) is distributed. This distribution allows the matrix data 49 to be classified into at least two or more clusters, as shown in Figure 6.

[0081] Here, CPU3, for example, based on the operation input from the operator of computer 1, generates one or more factors p that constitute each cluster Clz. j each cluster Cl z This can be set as the name (signboard). This name is set for each cluster Cl z Among the constituent factors, the factor p has the highest probability of the flag being 1. j You may select this option, or you may manually select it according to your needs.

[0082] For example, in Figure 6, the factor p has a high probability of the flag being 1. j For example, in the first cluster Cl1, p 16 and p 18 In the second cluster Cl2, p7 and p 11 The third cluster Cl3 is selected, and p2 and p 12 The following will be selected: signboard and factor p j Please also refer to Figure 7 for the selection.

[0083] On the other hand, if this disclosure is made for the purpose of "wanting to know the factors that determine whether or not someone owns a product from Company A," then the first cluster Cl1 to which factor p1 belongs may be named "A user" or "A-oriented," the third cluster Cl3 to which factor p2 belongs may be named "B user" or "B-oriented," and the second cluster Cl2 to which none of factors p1 to p3 belong may be named "Other users."

[0084] Furthermore, a flag (e.g., p1) indicating whether or not a specific item is represented is used as factor p j If included, the number of subjects n classified into each latent class z i For example, the first cluster Cl1 includes subjects n8, n 19 and n 21 As shown, this includes those who actually own the item (product of Company A) (those for whom the value of p1 is 1). Below are the names of such individuals n. i These may be referred to as "major A company users." In the example shown in Figure 6, the subjects are n8,n 19 ,n 21 These correspond to major users of Company A.

[0085] Furthermore, the subjects n are classified into each latent class z. i In addition to the aforementioned main A company users, this includes the target individuals n of the first cluster Cl1. 18 This also includes individuals who belong to the first cluster Cl1, labeled "User A," but do not own the item (those whose p1 value is 0). The following describes such individuals n. i This is sometimes referred to as a "potential A company user." In the example shown in Figure 6, the target person n 18 This corresponds to a potential user of Company A.

[0086] Furthermore, the subjects n are classified into each latent class z. i This includes not only major A Company users and potential A Company users, but also individuals who own A Company products but belong to clusters other than "A users," such as target person n1 in the second cluster Cl2, which is labeled as "other users." The following are examples of such target persons n iThis is sometimes referred to as an "active A company user." In the example shown in Figure 6, subject n1 corresponds to an active A company user.

[0087] In the following, for example, with respect to the third cluster Cl3 to which factor p2, which indicates whether or not a person owns a product from company B, belongs, the terms "major B company users," "potential B company users," and "actual B company users" may be used in the same way. The cluster Cl outside the diagram includes factor p3, which indicates whether or not a person owns a product from company C. z The same applies to this matter.

[0088] Next, in step S14, which follows step S13, the CPU 3 stores matrix data 49, in which customer IDs and factor IDs are rearranged for each latent class z, as cluster data 59 in RAM 7 or HDD 9. In other words, the cluster data 59 referred to here is none other than the matrix data 49 after clustering, as shown in the lower part of Figure 6. The stored cluster data 59 is then read as needed in subsequent processes such as the analysis target extraction step. Once step S14 is completed, the control process returns from the flow exemplified in Figure 4 and proceeds to step S2 in Figure 3.

[0089] (Step for selecting analysis targets) Figure 8 is a flowchart illustrating the steps of the analysis target selection process. Figure 9 is a diagram illustrating the basic concepts of the analysis target selection process.

[0090] Here, the flowchart illustrated in Figure 8 shows the process performed in step S2 of Figure 3. That is, when the control process proceeds to step S2 in Figure 3, the CPU 3 will execute steps S21-S24 of Figure 8 in order.

[0091] The step of selecting subjects for analysis involves identifying subjects n belonging to a specific latent class (hereinafter also simply referred to as "specific class") z' in the clustered matrix data 49, i.e., the cluster data 59. i Each of these is used as a base, and the target person n i Factor p assigned to each jA multidimensional vector X with the values ​​of each basis as its components. z’,j The factor p belonging to a latent class z'' other than the specific latent class z'. j It is configured to be generated by CPU3, including the following. Here, "z" represents all possible values ​​for "z" except "z'".

[0092] Specifically, in step S21 of Figure 8, the CPU 3 reads the cluster data 59. As mentioned above, the cluster data 59 represents matrix data 49 in which the customer ID and factor ID are sorted for each latent class z.

[0093] Next, in step S22, which follows step S21, the CPU 3 accepts the selection of a specific class z' from among multiple potential classes z. This selection may be accepted, for example, based on an operation input from the operator of computer 1, or it may be read from setting data containing the specific class z' in advance from HDD 9 or the like.

[0094] In this embodiment, as an example of a specific class z', the first cluster Cl1 shown in Figures 6 and 9, etc., i.e., "User A" with z=1, is selected.

[0095] Next, in step S23, which follows step S22, CPU3 uses customer IDs belonging to a specific class z' as a base and factors p assigned to each customer ID j A multidimensional vector X with the values ​​of each basis as its components. z’,j This generates a multidimensional vector X. z’,j The dimension d is the number of subjects n belonging to a specific class z'. i This will coincide with the number d.

[0096] Also, multidimensional vector X z’,j Factor p considered in the generation j This is one factor p belonging to at least one specific class z'. j And two or more factors p belonging to other latent classes z'' other than the specific class z'. j This includes the former factor p. jis referred to as "the first target factor p" f " or simply "target factor p" f ".

[0097] For example, in the example shown in FIG. 9, subjects n i that belong to the first cluster Cl1 as a specific class, such as n8, n 18 , n 19 and n 21 serve as the basis of the multi-dimensional vector X 1,j . In this case, the multi-dimensional vector X 1,j is four-dimensional (i.e., d = 4).

[0098] Also, when the factor p1 belonging to the first cluster Cl1 is taken as the target factor p f , the multi-dimensional vector X f corresponding to that target factor p 1,j [[ID=:30]]can be expressed, as shown in FIG. 9 X 1,1 =(1,0,1,1) . In this case, the target factor p f is a flag indicating whether or not the subject owns a product manufactured by Company A as a predetermined item.

[0099] On the other hand, for each factor p j corresponding to other potential classes z" other than the specific class z', that is, the second cluster Cl2, the third cluster Cl3, etc., the multi-dimensional vector X 1,j is, within the range shown in the figure, X 1,7 =(0,0,0,0), X 1,11 =(0,0,0,0), X 1,2 =(0,1,0,0), X 1,12 =(0,0,1,0), X 1,17 =(0,0,0,0) and can be expressed as such.

[0100] A point to note here is that "the multi-dimensional vector X z’,jThe factors p belonging to other latent classes z'' are the values ​​of each component of the vector. j The settings are configured to take the value of into account, but the subject n should be considered as the basis for the vector. i This always refers to the target person n belonging to a specific class z' rather than another latent class z''. i One example is "being selected from."

[0101] As will be explained in more detail later, this is how a multidimensional vector X z’,j By setting this up, the target individuals n belonging to a specific class z' can be identified. i This makes it possible to more clearly identify the factors that differentiate the state, properties, etc., than before.

[0102] Next, in step S24, which follows step S23, the CPU3 generates the multidimensional vector X in step S23. z’,j This is stored in RAM7 or HDD9. The multidimensional vector X that is then stored is then stored. z’,j This data is read as needed in subsequent processes such as the dimensionality reduction step. Once step S24 is completed, the control process returns from the flow illustrated in Figure 8 and proceeds to step S3 in Figure 3.

[0103] (Dimensional compression step) Figure 10 is a flowchart illustrating the steps of the dimensionality reduction process.

[0104] Here, the flowchart illustrated in Figure 10 shows the process performed in step S3 of Figure 3. That is, when the control process proceeds to step S3 in Figure 3, the CPU 3 will execute steps S31-S33 of Figure 10 in order.

[0105] The dimensionality reduction step involves CPU3 inputting a multidimensional vector X z’,j By applying the t-SNE method to the multidimensional vector X, z’,j a 2D or 3D low-dimensional vector x z’,j It is configured to perform dimensionality reduction.

[0106] Specifically, in step S31 of Figure 10, CPU3 generates a multidimensional vector X z’,j Load the information.

[0107] Next, in step S32, which follows step S31, the CPU3 reads the multidimensional vector X in step S31. z’,j The t-SNE method is applied to the d-dimensional multidimensional vector X. For example, in high-dimensional space, the distance between vector data is measured using a scale based on a multivariate normal distribution, and in low-dimensional space, the distance between vector data is measured using a scale based on a t-distribution with 1 degree of freedom. Then, CPU3 performs dimensionality reduction by minimizing the Kullback-Leibler divergence between the two distributions. By applying the t-SNE method, the d-dimensional multidimensional vector X z’,j This is a 2D or 3D low-dimensional vector x z’,j It can be converted to this.

[0108] In this embodiment, a two-dimensional low-dimensional vector x z’,j Assume that it has been transformed into a multidimensional vector X. z’,j Number of items (factors to consider p) j The number of x is a low-dimensional vector x z’,j It is the same number as [number of items].

[0109] Next, in step S33, which follows step S32, the CPU3 generates the low-dimensional vector x in step S32. z’,j This is stored in RAM7 or HDD9. The low-dimensional vector x stored in this way z’,j This data is read as needed in subsequent processes such as the vector visualization step. Once step S33 is completed, the control process returns from the flow illustrated in Figure 10 and proceeds to step S4 in Figure 3.

[0110] Note that the low-dimensional vector x obtained by the flow shown in Figure 10 is a low-dimensional vector. z’,j This includes the aforementioned first objective factor (objective factor) p f Corresponding low-dimensional vector x z’,j In addition, the second objective factor p, described below, may be used as needed. s Corresponding low-dimensional vector xz’,j This will include it.

[0111] (Vector visualization step) Figure 11 is a flowchart illustrating the steps of the vector visualization process. Figure 13 is an example of what is visualized by the vector visualization process.

[0112] In this embodiment, all information that can be visualized can be displayed on the display 11, which serves as the display unit. Furthermore, the display unit is not limited to the display 11 directly connected to the computer 1. A display indirectly connected to the computer 1 via a network or the like may also be used as the display unit.

[0113] Here, the flowchart illustrated in Figure 11 shows the process performed in step S4 of Figure 3. That is, when the control process proceeds to step S4 in Figure 3, the CPU 3 will execute steps S41-S46 of Figure 4 in order.

[0114] The vector visualization step involves a factor p belonging to a specific latent class z'. j The objective factor p f And other factors p belonging to other latent class z'' j Multiple comparison factors p j If so, then multidimensional vector X z’,j Each low-dimensional vector x obtained in correspondence z’,j Based on this, the objective factor p f Corresponding low-dimensional vector x z’,j and multiple comparison factors p j Low-dimensional vector x corresponding to each of " z’,j The CPU3 is configured to visualize the relative positions of the two elements.

[0115] In particular, in the vector visualization step according to this embodiment, a low-dimensional vector x corresponding to the objective factor p' is used. z’,j and multiple comparison factors p j Low-dimensional vector x corresponding to each of " z’,jThe CPU3 is configured to visualize the Euclidean distance between the two vectors. Alternatively, the dot product between the vectors may be calculated instead of, or in addition to, the Euclidean distance.

[0116] Specifically, in step S41 of Figure 11, CPU3 processes each low-dimensional vector x z’,j Load the information.

[0117] Next, in step S42, which follows step S41, CPU3 processes the first objective factor p f And, if necessary, a second objective factor p s The operation to select the first and second objective factors p f ,p s At least one of these may be determined when selecting the specific class z' as shown in step S22, or it may be determined during the processing shown in step S23.

[0118] Specifically, in step S42, CPU3 determines a factor p belonging to a specific class z'. j The first objective factor p f Select and two or more factors p belonging to the other class z'' j (multiple comparison factors p) j One of these is the second objective factor p s We select the second objective factor p. 2 The first objective factor is p 1 This will belong to a different latent class z. Preferably, the second objective factor p s The corresponding multidimensional vector X z’,j And the first objective factor p f The multidimensional vector X corresponding to the above z’,j These are mutually orthogonal.

[0119] Furthermore, the first objective factor p f As mentioned above, if a flag (factor p1) indicating whether or not one owns a specified item (product of Company A) is selected, the second objective factor p sPreferably, this serves as a flag indicating whether or not the user possesses an item different from the specified item (product of Company A). In this case, it is even more preferable that the item different from the specified item be an item of the same type but from a different manufacturer.

[0120] In the example shown in Figure 6, the second objective factor is p. s Therefore, a flag indicating whether or not the user owns the aforementioned B company product, i.e., factor p2 belonging to the third cluster Cl3, is selected.

[0121] Next, in step S43, which follows step S42, CPU3 processes the first objective factor p f Corresponding low-dimensional vector x z’,j And the second objective factor p s Corresponding low-dimensional vector x z’,j And the second objective factor p s Other comparison factors p j The corresponding low-dimensional vector x z’,j And are displayed (mapped) on the display unit, the display 11. When displaying, as shown in Figure 13, each low-dimensional vector x z’,j The starting point can be omitted, and only the ending point can be displayed.

[0122] As illustrated using Figures 5, 6, and 9, the first cluster Cl1 is selected as the specific class z' (i.e., z'=1), and the first objective factor p f Factor p1 of the first cluster Cl1 was selected, and the second objective factor p s The factor p2 of the third cluster Cl3 is selected, and a two-dimensional low-dimensional vector x 1,j Let's consider the case where the dimensions are compressed to the first objective factor p. f Please also refer to Figure 7 for the selection.

[0123] Note that the comparison factor p j " is a factor p belonging to a class other than the specific class z'. j It will be selected from among them. For example, in the above example, factor p belonging to the first cluster Cl1. j The comparison factor pj Please note that this will be excluded.

[0124] In this case, the low-dimensional vector x 1,j The display results are shown in Figure 13. “Pj” in the same figure refers to the factor p j Corresponding low-dimensional vector x 1,j This shows the position coordinates of the endpoint. In other words, “P1” and “P2” in Figure 13 represent the first objective factor p, respectively. f Corresponding low-dimensional vector x 1,1 The position coordinates of the endpoint and the second objective factor p s Corresponding low-dimensional vector x 1,2 This shows the position coordinates of the endpoint.

[0125] Furthermore, the comparison factor p in Figure 13 j "The first objective factor is p f p1 as the second objective factor p s p2 as such, and other factors p belonging to the first cluster Cl1 16 ,p 18 The remaining 16 factors p after removing and j This option is selected.

[0126] As shown in step S43, each low-dimensional vector x 1,j By visualizing the position coordinates of the endpoint, a low-dimensional vector x can be obtained as shown in the subsequent step S44. 1,j The Euclidean distance between the points is also visualized in a way that is intuitively understandable.

[0127] As illustrated in Figure 13, the second objective factor p s If this is added, CPU3 will use its second objective factor p s Corresponding low-dimensional vector x 1,j And the second objective factor p s The remaining comparison factor p after removing j Low-dimensional vector x corresponding to each of " 1,j The Euclidean distance between them may also be visualized.

[0128] In step S44, to help visualize the Euclidean distance, the first objective factor p is used, as shown by the dashed line in Figure 13. f P1 corresponding to and the second objective factor p s You may also display circles C11 and C12 centered on P2, which corresponds to the given point. In this case, the comparison factor p located inside each circle C11 and C12 is also shown. j " is the first objective factor p f and / or second objective factor p s It can be determined that they are in close proximity. On the other hand, the comparison factor p located outside each circle C11, C12 j " is the first objective factor p f and / or second objective factor p s Therefore, it can be determined that the distance is long. The two circles C11 and C12 may be of the same diameter.

[0129] Then, in step S45, which follows from step S44, CPU3 determines the first objective factor p f and the second objective factor p s relative short-range comparative factor p j Extract ".

[0130] Here, the relatively short-range comparison factor p j As such, the comparison factor p located inside the aforementioned circles C11 and C12 is defined as p j " can be extracted. In this process, the first objective factor p f or the second objective factor p s and the comparison factor p j Calculate the Euclidean distance to ".

[0131] Then, in step S46, which follows step S45, CPU3 processes the extraction results from step S45 and each extracted comparison factor p j The Euclidean distance related to " and is stored in RAM7 or HDD9. The low-dimensional vector x stored in this way z’,j This information is read as needed in subsequent processes such as the factor analysis step. Once step S46 is completed, the control process returns from the flow illustrated in Figure 11.

[0132] Here, the first objective factor p f and the second objective factor p s The first objective factor p f If only this is set, the control process proceeds from step S4 in Figure 3, skipping steps S5 and S6, to step S7.

[0133] On the other hand, the first objective factor p f and the second objective factor p s If both are set, the control process proceeds from step S4 to step S5 in Figure 3. In step S5, it is determined whether or not to perform the specific class change step (step S6) shown below. If this determination is YES, the process proceeds to step S7; otherwise, it proceeds to step S6 and executes the specific class change step.

[0134] Furthermore, the first objective factor p f and the second objective factor p s Even if both are set, the specific class change step may be omitted without making it mandatory.

[0135] (Specific class change step) Figure 14 illustrates the content visualized by the second vector visualization step.

[0136] In the specific class change step shown in step S6, a change is made to the specific class z'. Specifically, the second objective factor p s The latent class z to which it belongs is changed to a new specific class z'. This changes the first objective factor p s The latent class z to which it belongs is determined by the comparison factor p j It will be reset to "another latent class z from which it is extracted".

[0137] In the examples using Figures 5, 6, 9, and 14, the third cluster Cl3 is newly set to the specific class z', and the first cluster Cl1 is set to another latent class z''. The following describes the modified specific class (second objective factor p).s The latent class to which it belongs is sometimes referred to as the "second specific class" to distinguish it from the original specific class.

[0138] Once the modification of the specific class z' is complete, the control process returns to step S2 in Figure 3. The control process then repeats the analysis target extraction step (step S2), the dimensionality reduction step (step S3), and the vector visualization step (step S4), reflecting the modification of the specific class z'. At this point, the first and second objective factors p f ,p s This will be fixed to be the same as during the first execution.

[0139] Here, if we refer to each step that is executed again as the second analysis target extraction step, the second dimensionality reduction step, and the second vector visualization step, respectively, in the second analysis target extraction step, the CPU 3 generates a second multidimensional vector in the clustered matrix data 49, using each subject belonging to the second specific class as a basis, and the values ​​of the factors assigned to each subject as components of each basis, including factors belonging to other latent classes other than the second specific class.

[0140] In the example using Figure 14, the second multidimensional vector is "X 3,j This can be expressed as . The second multidimensional vector X in this case 3,j These are subjects n8,n belonging to the first cluster Cl1. 18 ,n 19 ,n 21 Instead of using the base, subjects n4,n belonging to the third cluster Cl3 12 ...will be the basis.

[0141] Next, in the second dimensionality reduction step, CPU3 applies the t-SNE method to the second multidimensional vector to reduce its dimensionality to a 2- or 3-dimensional low-dimensional vector.

[0142] In the example using Figure 14, etc., the low-dimensional vector is "x 3,jcan be expressed as "」. In this case, the low-dimensional vector x 3,j is obtained by applying t-SNE to a multi-dimensional vector with subjects n4, n 12 ,... belonging to the third cluster Cl3 as the basis.

[0143] Subsequently, in the second vector visualization step, for each low-dimensional vector x 3,j obtained corresponding to the second multi-dimensional vector, the CPU3 visualizes the Euclidean distance between the low-dimensional vector corresponding to the target factor (the first target factor), the low-dimensional vector corresponding to the second target factor, and the low-dimensional vectors corresponding to each of the remaining comparison factors.

[0144] The "low-dimensional vector corresponding to the first target factor" mentioned here refers to the one obtained corresponding to the second multi-dimensional vector. That is, in the case of the example using FIG. 14, etc., with subjects n4, n 12 ,... belonging to the third cluster Cl3 as the basis, and applying t-SNE to a multi-dimensional vector with the value of factor p1 as the components of each basis for the first target factor p f results in something corresponding to the "low-dimensional vector corresponding to the first target factor". The same applies to the "low-dimensional vector corresponding to the second target factor".

[0145] The display result of the low-dimensional vector x 3,j in the second vector visualization step is as shown in FIG. 14. The "Pj" shown in the figure refers to the position coordinates of the end point of the low-dimensional vector x j corresponding to the factor p 3,j . That is, "P1" and "P2" in FIG. 14 respectively represent the position coordinates of the end point of the low-dimensional vector x f corresponding to the first target factor p 3,1 and the position coordinates of the end point of the low-dimensional vector x s corresponding to the second target factor p 3,2 .

[0146] Also, as the comparison factor p j " in FIG. 14, it is the first target factor pf p1 as the second objective factor p s p2 as a factor, and other factors p belonging to the third cluster Cl3 12 ,p 17 The remaining 16 factors p after removing and j This option is selected.

[0147] Each low-dimensional vector x 3,j By visualizing the position coordinates of the endpoint, we can obtain a low-dimensional vector x after changing a specific class z'. 3,j The Euclidean distance (distance) between them is visualized in the same way as in the first vector visualization step.

[0148] In the second vector visualization step, similar to the first vector visualization step, to help visualize the Euclidean distance, the first objective factor p is used, as shown by the dashed line in Figure 14. f P1 corresponding to and the second objective factor p s Circles C31 and C32, each centered on P2 corresponding to the given point, may also be displayed. In this case, the comparison factor p located inside each circle C31 and C32 is shown. j " is the first objective factor p f or the second objective factor p s It can be considered to be in close proximity to one of the two.

[0149] Multidimensional vector X 3,j By changing the basis, the first objective factor p f or the second objective factor p s The comparison factor p is in close proximity to the given factor. j The breakdown of " may change from the example shown in Figure 13.

[0150] Finally, similar to the first vector visualization step, CPU3 calculates the first objective factor p f and the second objective factor p s relative short-range comparative factor p j Extract the ". Here, the relatively short-range comparison factor p j As such, the comparison factor p located inside or near the aforementioned circles C31 and C32 is defined as p j" can be extracted. In this process, the first objective factor p f or the second objective factor p s and the comparison factor p j You may also calculate the Euclidean distance to ".

[0151] (Factor analysis step) The factor analysis step shown in step S7 of Figure 3 is based on the relative positional relationships visualized by the vector visualization step, and the first objective factor p f It is configured to analyze the factors that caused the value to be 0 or 1.

[0152] Specifically, in the factor analysis step according to this embodiment, multiple comparison factors p are identified based on the length of each Euclidean distance visualized by the CPU3. j From among these, the objective factor (first objective factor) p f Factor p that contributed to the value j Explore.

[0153] The following details the factor analysis steps, explaining both the case where the specific class change step is not performed and the case where the specific class change step is performed.

[0154] 1. If the specific class change step is not performed In this case, the first objective factor p f Factor p that contributed to the value j The search for comparison factors p located inside or near circle C11 is performed. j The list L11 is used as the first objective factor p f This is done by arranging them in order of proximity to each other.

[0155] In the example shown in Figure 13, the first objective factor p f As such, factor p1 is used to indicate that the person belongs to the first cluster Cl1 and owns products from Company A. In this case, the first objective factor p fBeing close to factor p1 allows for the tentative interpretation that "it belongs to a different cluster than factor p1 (for example, B-oriented or other users), but is a factor that has a high correlation with factor p1 (a factor that could be a trigger for owning A company's products)."

[0156] Conversely, the first objective factor p f The fact that it is far from factor p1 allows for the tentative interpretation that "it belongs to a different cluster than factor p1, and yet is unlikely to be a factor that would lead to owning Company A's products."

[0157] That is, the low-dimensional vector x exemplified in Figure 13, etc. 1,j This is generated based on a basis that considers both the primary A Company users who actually own A Company products and the potential A Company users who belong to the first cluster Cl1 (i.e., who have similar values ​​to primary A Company users) but do not own A Company products. The low-dimensional vector x thus generated is 1,j By analyzing the distance between this factor and factor p1, which indicates that the person actually owns a product from company A, we can identify factors p belonging to clusters other than the first cluster Cl1. j From among these, the factor p that distinguishes major A company users from potential A company users. j In other words, factor p could be a factor that leads to owning a product from company A. j It becomes possible to explore.

[0158] In that case, factor p belonging to the first cluster Cl1 j p belonging to other clusters j By focusing on this, for example, a factor p that distinguishes major A company users from potential A company users from a different perspective than the values ​​held by major A company users. j This makes it possible to extract certain information. This allows us to encourage "awareness" from a different perspective than before.

[0159] Furthermore, the first objective factor p f In addition, the second objective factor p s If this setting is selected, CPU3 will use multiple comparison factors p based on the length of each Euclidean distance visualized.j From among these, the second objective factor p s We will search for factors that contributed to the value of [the variable].

[0160] This search involves comparison factors p located inside or near circle C12. j The list L12 is used as the second objective factor p s This can be done by arranging them in order of proximity to each other.

[0161] In the example shown in Figure 13, the second objective factor is p s As such, factor p2 is used to indicate that the individual belongs to the third cluster Cl3 and owns products from Company B. In this case, the second objective factor p s The fact that it is close to factor p1 allows for the tentative interpretation that "it is a factor that has a low correlation with factor p1 and a high correlation with factor p2 (a factor that could be a trigger for owning Company B's products)."

[0162] Furthermore, the second objective factor p s The fact that it is far from the company allows for the tentative interpretation that "it is unlikely to be a factor that would motivate someone to own a product from company B."

[0163] That is, the low-dimensional vector x illustrated in Figure 13 1,j As mentioned above, this is generated based on a base that considers both major and potential users of Company A. The low-dimensional vector x generated in this way 1,j By analyzing the distance between this factor and factor p2, which indicates that the person owns a product from company B rather than company A, we can identify factors p belonging to clusters other than the first cluster Cl1. j From among these, the factor p that distinguishes major A company users from potential A company users j In other words, factors p that led to owning products from other companies, even while possessing values ​​that favor Company A's products. j It becomes possible to explore.

[0164] First objective factor p f and the second objective factor p sBy setting both, the factor p that differentiates between major Company A users and potential Company A users can be obtained. j It becomes possible to comprehensively explore it.

[0165] 2. When passing through the specific class change step FIG. 12 is a flowchart illustrating the procedure of the factor analysis step (particularly, the procedure when passing through the specific class change step).

[0166] In this case, first, the CPU 3 searches (extracts) the comparison factor p that was searched before the change of the specific class z' in the same manner as the process when not passing through the specific class change step. j ” as the first and second target factors p f , p s and arranges them in ascending order of proximity to (step S71). This step S71 is based on the Euclidean distance obtained before the change of the specific class z', and searches for the factors that contributed to the values of the first target factor p f and the second target factor p s respectively, corresponding to the first search step.

[0167] After that, the CPU 3 searches (extracts) the comparison factor p that was searched after the change of the specific class z'. j ” as the first and second target factors p f , p s and arranges them in ascending order of proximity to (step S72). This step S72 is based on the Euclidean distance obtained after the change of the specific class z', and searches for the factors that contributed to the values of the first target factor p f and the second target factor p s respectively, corresponding to the second search step.

[0168] Specifically, in step S72, first, based on the content visualized after the change of the specific class z', the comparison factor p j ” located inside or near the circle C31 is listed up, and the list L31 is arranged in ascending order of the distance to the first target factor p f .

[0169] That is, the low-dimensional vector x illustrated in Figure 14 3,j This is generated based on a foundation that takes into account both major and potential B company users. Major and potential B company users may include existing A company users who have similar values ​​to major B company users but have come to own A company products instead of B company products.

[0170] Therefore, the low-dimensional vector x generated in this way 3,j By analyzing the distance to factor p1, which indicates that the person actually owns a product from company A, we can identify factors p belonging to clusters other than the third cluster Cl3. j From among them, the factor p that distinguishes whether or not someone is an active user of Company A. j , in other words, factors p that could be the trigger for owning Company A's products j It becomes possible to explore.

[0171] Furthermore, in step S72, based on the length of each Euclidean distance visualized by CPU3 after the change of a specific class z', multiple comparison factors p j From among these, the second objective factor p s We will search for factors that contributed to the value of [the variable].

[0172] Specifically, based on the content visualized after the change in a particular class z', the comparison factor p located inside or near circle C32 is determined. j The list L32 is used as the second objective factor p s Arrange them in order of proximity to each other.

[0173] That is, the low-dimensional vector x illustrated in Figure 14 3,j As mentioned above, this is generated based on a base that considers both major and potential B company users and may also include existing A company users. The low-dimensional vector x thus generated is 3,j By analyzing the distance between factor p2 and other factors belonging to clusters other than the third cluster Cl3, we can identify factors p j From among them, the factor p that distinguishes whether or not someone is an active user of Company A. jIn other words, factors p could be the reason why major B company users and potential B company users did not become actual A company users. j It becomes possible to extract it.

[0174] Subsequently, step S73, following step S72, is the step that explores the factor p explored in the first exploration step. j And the factor p discovered in the second search step j By comparing with the first objective factor p f It is configured to analyze the factors that caused the value to be 0 or 1.

[0175] Specifically, in step S73, CPU3 uses a comparison factor p common to both before and after the change of a particular class z'. j We search for " and the common comparison factor p that was thus searched for j The text " is displayed on the display 11, and its display mode is changed. The specific changes in the display mode include the common comparison factor p j The display color of " may be different, or a common comparison factor p j You may highlight or otherwise indicate ", or a common comparison factor p j You may underline the " or a common comparison factor p j You can change the font of the ". Figure 15 shows examples of changes in the display configuration. The examples shown in Figure 15 are based on the content visualized in Figures 13 and 14.

[0176] In other words, the upper part of Figure 15 shows the first objective factor p when a specific class is set to z'=1 and subjects belonging to the first cluster Cl1 (A-oriented) are used as the basis for each vector. f Four comparison factors close to p1 (see P19, P11, P12, and P3), and the second objective factor p s The four comparison factors closest to p2 (see P17, P15, P14, and P10) are shown, along with the comparison factors themselves. These comparison factors are listed in order from those closest to p1 or p2.

[0177] On the other hand, the lower part of Figure 15 shows the first objective factor p when the specific class is set to z'=3 and subjects belonging to the third cluster Cl3 (B-oriented) are used as the basis for each vector. f Four comparison factors close to p1 (see P15, P18, P19, and P9), and the second objective factor p s The four comparison factors closest to p2 (see P11, P16, P14, and P4) are shown, along with the comparison factors themselves. These comparison factors are listed in order from those closest to p1 or p2.

[0178] Here, the upper row and the first objective factor p f The display area and the lower section and the first objective factor p f In the display column, "P19" appears as a common comparison factor. This means that "whether the target is A-oriented or B-oriented, factor p 19 However, this can be tentatively interpreted as "it is highly likely to be a trigger for owning Company A's products." Since such factors are considered to be of higher importance compared to other factors, they are underlined as shown in Figure 15.

[0179] Similarly, the upper row and the second objective factor p s The display area and the lower section and the second objective factor p s In the display column, "P14" appears as a common comparison factor. This means that "whether the target is A-oriented or B-oriented, factor p 14 However, this can be tentatively interpreted as "this is likely to be the trigger for owning a product from Company B." Since such factors are considered to be of higher importance compared to other factors, they are underlined as shown in Figure 15.

[0180] Thus, a common comparison factor p is used before and after the change in a specific class z'. j By changing the display method of ", other comparison factors p j This allows users to visualize factors that are more important than others. This improves the usability of the information processing method.

[0181] Subsequently, the control process completes the flow shown in Figure 12 and returns from the flow shown in Figure 3.

[0182] <Effects, etc.> As described above, according to the embodiment, the target person n belonging to a specific class z' i Each of these is used as a base, and the target person n i Factor p assigned to each j A multidimensional vector X with the values ​​of each basis as its components. z’,j By applying the t-SNE method to the multidimensional vector X, z’,j to a low-dimensional vector x z’,j Dimensionality is reduced (see steps S1-S3 in Figure 3).

[0183] And, as explained using Figure 9, etc., the low-dimensional vector x z’,j In visualizing the first objective factor p belonging to a specific class z' f and other comparison factors p j When comparing with the latter, the comparison factor p j For this, we will use factors belonging to other latent classes z', rather than factors belonging to a specific class z'.

[0184] In other words, as is clear from the fact that each block matrix obtained by clustering corresponds to a latent class z, a low-dimensional vector x corresponds to a factor belonging to the same latent class z. z,j They show similar tendencies to one another.

[0185] In contrast, the objective factor p belonging to a specific class z' f And the comparison factor p belonging to another latent class z'' j By comparing with the low-dimensional vector x z’,j It becomes possible to clarify the differences between them. With the differences clarified in this way, the objective factor p f Comparison factor p that is close to or far from j By exploring the target factor p j This makes it possible to more clearly identify the factors that may trigger the value of " to become 0 or 1 than before.

[0186] Furthermore, the comparison factor p belongs to another latent class z'' in the first place. j "Comparing with the target factor p" means that although they were classified into different latent classes at the stage of performing the PLSA method, the target factor p f A comparison factor p that is strongly correlated with j This is equivalent to exploring ". This is knowledge that cannot be obtained with the conventional PLSA method, and a specific objective factor p f This allows us to understand the factors that caused the value to become 0 or 1 from a wider range of perspectives than before.

[0187] Furthermore, as explained using Figure 13, the factor analysis step is configured to explore factors based on the length of each Euclidean distance. This configuration makes it possible to visualize and explore factors in a more intuitively understandable format. This allows for the identification of a specific target factor p f This is advantageous in identifying the factors that caused the value to become 0 or 1.

[0188] Furthermore, as illustrated in Figure 9, the objective factor p f The basis for constructing the multidimensional vector corresponding to this is the number of people who actually own Company A's products (for example, n8 19 ,n 21 This will include the following. By performing the aforementioned visualization on the low-dimensional vectors generated based on these multidimensional vectors, it becomes possible to understand the factors that led to owning or not owning Company A's products from a wider range of perspectives than before.

[0189] Furthermore, as illustrated in Figure 9, the second objective factor p s By performing an analysis using this method, the first objective factor p s This allows us to understand the factors that caused the value to be 0 or 1 from a wider range of perspectives. For example, the second objective factor p s However, the first objective factor p s If it was a flag indicating whether or not the user owns a competing product, then the second objective factor p sBy exploring the factors that resulted in 0 or 1, it becomes possible to identify areas for improvement, selling points, etc., of the competing product. By analyzing these findings, the first objective factor p f This allows for the clarification (visualization) of areas for improvement in products that correspond to this.

[0190] Furthermore, as explained using Figure 12, etc., the objective factor (first objective factor) p f After performing an analysis where the latent class to which (z=1 in the diagram) belongs is designated as the specific class z', the second objective factor p s The analysis is performed by changing the latent class to which (z=3 in the example figure) belongs to a specific class z'. The latter analysis is performed using the second objective factor p s Subject n belonging to the same latent class z' i Each of these is used as a base, and the target person n i Factor p assigned to each j This analysis is based on multidimensional vectors with the values ​​of as components of each basis. Therefore, the latter analysis and the former analysis use the subject n used to generate the low-dimensional vector. i The breakdown will differ.

[0191] Therefore, subject n i By comparing the results of two analyses with different breakdowns, the target factor p f This allows us to understand the factors that caused the value to be 0 or 1 from a wider range of perspectives.

[0192] Other embodiments The above embodiment exemplifies an implementation by a single computer 1, but this disclosure is not limited to that example. The information processing method and information processing program 29 relating to this disclosure may be executed using multiple computers 1. Furthermore, the computer 1 in this disclosure also includes parallel computers such as supercomputers and PC clusters. [Explanation of Symbols]

[0193] 1. Computer (information processing device) 3 CPU (arithmetic unit) 7 RAM (memory section) 9 HDD (Storage Unit) 11. Display (Display Unit) 18 Storage medium 29 Information Processing Programs 49 Matrix Data S1PLSA Step S2 Step for extracting analysis targets S3 Dimensional Compression Step S4 Vector Visualization Step S7 Factor Analysis Step

Claims

1. An information processing method that uses a computer equipped with a program execution unit to analyze matrix data formed by matrixing multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects, and visualizes the results of the analysis, The calculation unit applies the PLSA method to the matrix data in a PLSA step, which clusters the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes. In the matrix data after clustering, the calculation unit generates a multidimensional vector in which each analysis subject belonging to a specific latent class is used as a basis, and the values ​​of the factors assigned to each analysis subject are used as components of each basis, including factors belonging to other latent classes other than the specific latent class, in the analysis subject extraction step. The calculation unit applies the t-SNE method to the multidimensional vector in a dimensionality reduction step, thereby compressing the multidimensional vector into a two-dimensional or three-dimensional low-dimensional vector. If one factor belonging to the aforementioned specific latent class is designated as the objective factor, and other factors belonging to the aforementioned other latent classes are designated as multiple comparison factors, the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors, based on each low-dimensional vector obtained in correspondence with the multidimensional vector. The vector visualization step includes a factor analysis step that analyzes the factors that caused the value of the objective factor to be 0 or 1, based on the relative positional relationships visualized by the vector visualization step. An information processing method characterized by the following:

2. In the information processing method described in claim 1, In the vector visualization step, the relative positional relationship is as follows: The calculation unit visualizes the Euclidean distance between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors. In the aforementioned factor analysis step, Based on the length of each Euclidean distance visualized by the calculation unit, the unit searches for the factor that contributed to the value of the objective factor from among the multiple comparison factors. An information processing method characterized by the following:

3. In the information processing method described in claim 2, The aforementioned objective factor is a flag indicating whether or not a person possesses a specified item. An information processing method characterized by the following:

4. In the information processing method described in claim 2, In the aforementioned vector visualization step, One of the aforementioned comparison factors is designated as the second objective factor. The calculation unit visualizes the Euclidean distance between the low-dimensional vector corresponding to the second objective factor and the low-dimensional vectors corresponding to each of the remaining comparison factors excluding the second objective factor. In the aforementioned factor analysis step, Based on the length of each Euclidean distance visualized by the calculation unit, the unit searches for the factors that contributed to the value of the second objective factor from among the remaining comparison factors. By comparing the factors that contributed to the value of the second objective factor with the factors that contributed to the value of the first objective factor, we can analyze the factors that caused the value of the objective factor to be 0 or 1. An information processing method characterized by the following:

5. In the information processing method described in claim 4, If the latent class to which the second objective factor belongs is referred to as the second specific class, then the second analysis target extraction step involves the calculation unit generating a second multidimensional vector in the matrix data after clustering, with each analysis target belonging to the second specific class as a base, and the factor values ​​assigned to each analysis target as components of each base, including factors belonging to other latent classes other than the second specific class. The calculation unit applies the t-SNE method to the second multidimensional vector in a second dimensionality reduction step, thereby reducing the second multidimensional vector to a two-dimensional or three-dimensional low-dimensional vector. The system includes a second vector visualization step in which the calculation unit visualizes the Euclidean distance between the low-dimensional vector corresponding to the objective factor, the low-dimensional vector corresponding to the second objective factor, and the low-dimensional vectors corresponding to each of the remaining comparison factors, based on each low-dimensional vector obtained in correspondence with the second multidimensional vector. In the aforementioned factor analysis step, A first search step in which, based on the length of each Euclidean distance visualized by the calculation unit in the vector visualization step, a factor is searched from the remaining comparison factors for the factor that contributed to the value of the objective factor and the second objective factor, respectively. A second search step is performed in which, based on the length of each Euclidean distance visualized by the calculation unit in the second vector visualization step, factors that contributed to the values ​​of the objective factor and the second objective factor are searched from among the remaining comparison factors. By comparing the factors explored in the first exploration step with the factors explored in the second exploration step, the factors that resulted in the target factor having a value of 0 or 1 are analyzed. An information processing method characterized by the following:

6. In the information processing method described in claim 4 or 5, The aforementioned objective factor is a flag indicating whether or not a person possesses a predetermined item. The second objective factor is a flag indicating whether or not the person possesses an item different from the predetermined item, The aforementioned specified article and the aforementioned other article are articles of the same type from different manufacturers. An information processing method characterized by the following:

7. An information processing device comprising a computer equipped with an arithmetic unit for executing a program, which analyzes matrix data obtained by matrixing multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects, and visualizes the results of the analysis, The calculation unit applies the PLSA method to the matrix data, and the PLSA means clusters the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes. The analysis target extraction means generates a multidimensional vector in the matrix data after clustering, where each analysis target belonging to a specific latent class is a basis, and the values ​​of the factors assigned to each analysis target are components of each basis, including factors belonging to other latent classes other than the specific latent class, which is generated by the calculation unit. The calculation unit applies the t-SNE method to the multidimensional vector, thereby compressing the multidimensional vector into a two-dimensional or three-dimensional low-dimensional vector. When one factor belonging to the aforementioned specific latent class is designated as the objective factor, and other factors belonging to the aforementioned other latent classes are designated as multiple comparison factors, the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors, based on each low-dimensional vector obtained in correspondence with the multidimensional vector. The system includes a factor analysis means that analyzes the factors that cause the value of the objective factor to be 0 or 1 based on the relative positional relationship visualized by the vector visualization means. An information processing device characterized by the following:

8. An information processing program that, when executed by a computer equipped with a program execution unit, analyzes matrix data formed by matrixing multiple subjects and multiple factors consisting of values ​​of 0 or 1 assigned to each of the multiple subjects, and visualizes the results of the analysis, To the aforementioned computer, The calculation unit applies the PLSA method to the matrix data in a PLSA step, which clusters the matrix data so that it approaches a block diagonal matrix consisting of block matrices labeled by different latent classes. In the matrix data after clustering, the calculation unit generates a multidimensional vector in which each analysis subject belonging to a specific latent class is used as a basis, and the values ​​of the factors assigned to each analysis subject are used as components of each basis, including factors belonging to other latent classes other than the specific latent class, in the analysis subject extraction step. The calculation unit applies the t-SNE method to the multidimensional vector in a dimensionality reduction step, thereby compressing the multidimensional vector into a two-dimensional or three-dimensional low-dimensional vector. If one factor belonging to the aforementioned specific latent class is designated as the objective factor, and other factors belonging to the aforementioned other latent classes are designated as multiple comparison factors, the calculation unit visualizes the relative positional relationship between the low-dimensional vector corresponding to the objective factor and the low-dimensional vector corresponding to each of the multiple comparison factors, based on each low-dimensional vector obtained in correspondence with the multidimensional vector. The process involves performing a factor analysis step, which analyzes the factors for which the value of the objective factor became 0 or 1, based on the relative positional relationships visualized by the vector visualization step. An information processing program characterized by the following features.

9. The system stores the information processing program described in claim 8. A computer-readable storage medium characterized by the following features.