Information processing device, information processing method, and program
The information processing device and method address the challenge of analyzing relationships between documents and items by arranging them in a k-dimensional space and calculating index values, facilitating accurate similarity selection and grouping.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- DAI NIPPON PRINTING CO LTD
- Filing Date
- 2022-07-25
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional technologies struggle to analyze relationships between documents and items, and accurately determine similarity based on item attributes and proximity in a virtual multidimensional space.
An information processing device and method that arranges items in a k-dimensional space, calculates an index value of relationships using a covariance matrix and Euclidean distance, and selects similar items based on these values.
Enables accurate selection of similar items based on their relationships, allowing for effective grouping and analysis of items within a category of interest.
Smart Images

Figure 0007859235000011 
Figure 0007859235000012 
Figure 0007859235000013
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing device, an information processing method, and a program. [Background technology]
[0002] In recent years, advancements in computer network technology have led to the circulation of vast amounts of data. This data is often related to various items, and it is expected that analyzing this data will reveal relationships and meanings between these items.
[0003] As an example of such technology, a technique has been developed that uses documents as items and presents important words within these documents (see Patent Document 1). [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2018-195108 [Overview of the project] [Problems that the invention aims to solve]
[0005] However, conventional technologies can present the words that make up a document, but they have the problem of being unable to analyze the relationships between documents using those words. Furthermore, it is difficult to analyze the relationships between multiple items, not just documents. Moreover, even if relationships can be analyzed, it is difficult to accurately determine similarity by placing items in a virtual multidimensional space based on their attributes and judging their similarity based on their proximity.
[0006] Therefore, the object of this disclosure is to provide an information processing device, an information processing method, and a program that can accurately select similar items based on the relationships between items belonging to a category of interest. In particular, the object of this disclosure is to provide an information processing device, an information processing method, and a program that, under the condition that an item group consisting of multiple items is set among the items belonging to a category of interest, can select items similar to the item group from among the items not included in this item group. [Means for solving the problem]
[0007] To address the above issues, this disclosure provides: An information processing device that uses a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items that are similar to a group of items, which is a set of multiple items, An item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, Similar item selection means calculates an index value of the relationship between an item not included in the aforementioned item group and the aforementioned item group, and selects similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship; It has, The aforementioned similar item selection means provides an information processing device that, as an index value of the relationship, projects all items onto the complementary space of the subspace spanned by the group of items in a multidimensional Euclidean space where all items are arranged, and uses the Euclidean distance from the group of items reduced to a single point on that complementary space.
[0008] Furthermore, this disclosure is, An information processing device that uses a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items that are similar to a group of items, which is a set of multiple items, An item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, Similar item selection means calculates an index value of the relationship between an item not included in the aforementioned item group and the aforementioned item group, and selects similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship; It has, The aforementioned similar item selection means provides an information processing device that generates a covariance matrix from the arrangement coordinate values of the item group, and uses the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix, as an index value of the relationship.
[0009] Furthermore, this disclosure is, An information processing device that arranges each item in a k-dimensional space using a set associated with each item belonging to a category of interest, selects similar items that are similar to a group of items, which is a set of multiple items, Among the items arranged in the aforementioned k-dimensional space, an item group is defined, which is a collection of multiple items. An index value is calculated for the relationship between items not included in the aforementioned item group and the aforementioned item group, and based on this index value, similar items that are similar to the aforementioned item group are selected from among the items not included in the aforementioned item group. When selecting the aforementioned similar items, the present invention provides an information processing method that uses the Euclidean distance from the item group reduced to a single point on that complementary space, after projecting all items onto the complementary space spanned by the subspace in the multidimensional Euclidean space where all items are arranged, as an index value of the relationship.
[0010] Furthermore, this disclosure is, An information processing device that arranges each item in a k-dimensional space using a set associated with each item belonging to a category of interest, selects similar items that are similar to a group of items, which is a set of multiple items, Among the items arranged in the aforementioned k-dimensional space, an item group is defined, which is a collection of multiple items. An index value is calculated for the relationship between items not included in the aforementioned item group and the aforementioned item group, and based on this index value, similar items that are similar to the aforementioned item group are selected from among the items not included in the aforementioned item group. The present invention provides an information processing method for selecting similar items, in which a covariance matrix is generated from the coordinate values of the item group, and the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix is used as an index value of the relationship.
[0011] Furthermore, this disclosure is, A program that, using a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items for a set of items, which is a group of items. Computers, Item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, This system functions as a similar item selection means, calculating an index value of the relationship between items not included in the aforementioned item group and the aforementioned item group, and selecting similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship. The aforementioned similar item selection means provides a program that, as an index value of the relationship, projects all items onto the complementary space of the subspace spanned by the group of items in a multidimensional Euclidean space where all items are arranged, and uses the Euclidean distance from the group of items reduced to a single point on that complementary space.
[0012] Furthermore, this disclosure is, A program that, using a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items for a set of items, which is a group of items. Computers, Item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, This system functions as a similar item selection means, calculating an index value of the relationship between items not included in the aforementioned item group and the aforementioned item group, and selecting similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship. The aforementioned similar item selection means provides a program that generates a covariance matrix from the arrangement coordinate values of the item group, and uses the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix, as an index value of the relationship. [Effects of the Invention]
[0013] According to this disclosure, it becomes possible to accurately select similar items based on the relationships between each item belonging to the category of interest. [Brief explanation of the drawing]
[0014] [Figure 1] This is a hardware configuration diagram of an information processing device according to one embodiment of the present disclosure. [Figure 2] A functional block diagram showing the configuration of an information processing device according to one embodiment of this disclosure. [Figure 3] This figure shows an example of information stored in a document database used in one embodiment of this disclosure. [Figure 4] This figure shows an example of information stored in a word database used in one embodiment of the present disclosure. [Figure 5] This is a flowchart showing the processing operation of an information processing device according to one embodiment of the present disclosure. [Figure 6] This figure shows the relationship between the histogram of the components of the triangular matrix in the relation matrix S and the adjustment function f(x). [Figure 7] This figure shows the relationship between the reduction coefficient s and the adjustment function f(x). [Figure 8] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is two-dimensional. [Figure 9] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is two-dimensional. [Figure 10] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is two-dimensional. [Figure 11] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is 3-dimensional. [Figure 12] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is 3-dimensional. [Figure 13] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is 3-dimensional. [Figure 14] This figure shows the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is 3-dimensional. [Figure 15] This figure shows the relationship between the set of items in k-dimensional space and the distance calculated using the first method. [Figure 16] This figure shows the relationship between the set of items in k-dimensional space and the distance calculated using the second method. [Figure 17] This figure shows the relationship between the set of items in k-dimensional space and the distance calculated using the third method. [Figure 18] This figure shows the relationship between the set of items in k-dimensional space and the distance calculated by the third method, by changing the value of p in (Equation 34). [Figure 19] This is a diagram used to explain scatter plots. [Modes for carrying out the invention]
[0015] Preferred embodiments of this disclosure will be described in detail below with reference to the drawings. <1.Device configuration> Figure 1 is a hardware configuration diagram of an information processing device 100 according to one embodiment of the present disclosure. The information processing device 100 according to this embodiment can be implemented in a general-purpose computer and, as shown in Figure 1, comprises a CPU (Central Processing Unit) 1, RAM (Random Access Memory) 2 which is the main memory of the computer, a large-capacity storage device 3 such as a hard disk, SSD (Solid State Drive), or flash memory for storing programs and data executed by the CPU 1, an instruction input I / F (interface) 4 such as a keyboard or mouse, a data input / output I / F (interface) 5 for data communication with external devices such as data storage media, a display unit 6 which is a display device such as a liquid crystal display, a GPU (Graphics Processing Unit) 7 which is a calculation processing unit specialized for graphics, and a frame memory 8 which holds images to be displayed on the display unit 6, and these are connected to each other via a bus. Since the calculation results by the GPU 7 are written to the frame memory 8, the GPU 7 and the frame memory 8 are often mounted on a video card equipped with an interface to the display unit 6 and installed in a general-purpose computer via a bus. The information processing device 100 may further include a communication unit that can communicate with other computers, etc., via a network such as the Internet.
[0016] In this embodiment, CPU1 may be a multi-core CPU. In this case, CPU1 has multiple CPU cores and is capable of parallel processing. In the example in Figure 1, only one RAM2 is shown, but each CPU core of CPU1 is configured to access one RAM2. Note that there may be multiple CPU1s. Furthermore, a multi-core CPU may be a CPU that has logically multiple CPU cores.
[0017] Figure 2 is a functional block diagram showing the configuration of the information processing device according to this embodiment. In Figure 2, 10 is a word database, 11 is a document database, 20 is an arithmetic processing unit, 21 is a common element matrix generation means, 22 is a relationship matrix generation means, 23 is an adjusted distance matrix generation means, 24 is a coordinate value calculation means, 27 is an item placement means, 25 is a morphological analysis means, 28 is an item group setting means, 29 is a similar item selection means, and 30 is an output means.
[0018] The common element matrix generation means 21 is a means for determining a common element matrix M for any two combinations of documents, where each document contains a word as an element. Each component of the common element matrix M is the number of common elements in the two sets. The relationship matrix generation means 22 is a means for determining a relationship matrix S from the common element matrix M. The adjusted matrix generation means 23 is a means for determining an adjusted matrix W by applying an adjustment function to the relationship matrix S. The coordinate value calculation means 24 is a means for calculating the coordinate values of each item based on the adjusted matrix W.
[0019] The item placement means 27 is a means for placing items in a k-dimensional space (where k is an integer of 2 or more) based on the calculated coordinate values of each item. The item group setting means 28 is a means for setting an item group, which is a collection of multiple items, from among the items arranged in the k-dimensional space. The similar item selection means 29 is a means for calculating an index value of the relationship between an item not included in the item group and the item group, and selecting similar items that are similar to the item group based on the index value of that relationship. The morphological analysis means 25 is a means for reading each document from the document database 11 and extracting words contained in each document, which is an item.
[0020] The common element matrix generation means 21, relationship matrix generation means 22, adjusted matrix generation means 23, coordinate value calculation means 24, item placement means 27, item group setting means 28, similar item selection means 29, and morphological analysis means 25 are included in the arithmetic processing unit 20 and are realized when the CPU 1 executes a program.
[0021] The output means 30 is a means for graphically outputting items placed at coordinate values in space by the item placement means 27, and is realized by a printer via the data input / output I / F 5 or a display unit 6 such as a display device.
[0022] The word database 10 is a database that stores words in association with word IDs that identify words, and is implemented by the storage device 3. The document database 11 is a database that stores documents consisting of text information in association with document IDs that identify documents, and is implemented by the storage device 3.
[0023] Figure 3 shows an example of information stored in the document database 11. In this embodiment, a set of words associated with each item (document) belonging to a category of interest (a predetermined group of documents managed in the word database 10, document database 11, etc.) is used as input data to process the placement of each item in a k-dimensional space. Here, the information stored in the document database 11, i.e., the set of documents, becomes the category of interest, and each document becomes each item. The set of words contained in each document becomes the set associated with each document.
[0024] The sets associated with each item refer to the sets that are linked to (associated with, linked to) each item, and include sets contained within each item and sets associated with each item. The sets contained within each item are, for example, the sets of words contained within a document if each item is a document. Documents are usually composed of sentences written in natural language. Therefore, if each item is a document, each item is associated with a sentence written in natural language. Furthermore, the sets associated with each item are, for example, the sets of elements (usually expressible as words) contained within a group if each item is a group of some kind.
[0025] For example, if the category of interest is movies, and each item belonging to that category is an individual movie, then the set associated with each item can be the set of actors who appeared in that movie. Similarly, if the category of interest is actors, and each item belonging to that category is an individual actor, then the set associated with each item can be the set of movies in which that actor appeared.
[0026] In these two examples, the relationship between a film and an actor can be transformed from a data structure where one is the object of interest and the other is an associated set to the reverse data structure. Furthermore, by applying the same transformation again, the original relationship can be restored. Such a relationship is generally called dual. The relationship between a document and a word can also be transformed into a dual relationship. That is, the word is considered the object of interest, and the set of documents containing each word is considered the associated set.
[0027] Furthermore, if the category of interest is a group of images, and each item belonging to that category is an individual image, then a set of words or phrases that describe the content represented in that image can be adopted as the set associated with each item.
[0028] As shown in Figure 3, the document database 11 stores the document name, author name, and document data associated with a document ID, which is document identification information that identifies a document. For the document data, it is sufficient that the storage address of the document data is recorded so that the document data can be retrieved by identifying the document ID. In the example in Figure 3, for example, it is shown that a document (work) titled "XXXXXX" by author "Mr. A" is registered with document ID "B001".
[0029] Figure 4 shows an example of information stored in the word database 10. As shown in Figure 4, the word database 10 stores words and the document IDs of the documents in which those words appear, associated with word IDs, which are word identification information that identifies words. In the example in Figure 4, the word registered as word ID "T0001" in the first line appears 5 times in the document identified by document ID "B001" and 3 times in the document identified by document ID "B002". Similarly, the word registered as word ID "T0002" in the second line appears 3 times in the document identified by document ID "B001" and 8 times in the document identified by document ID "B002".
[0030] Each word registered in the word database 10 is associated with a document ID, thus storing a set of words associated with each document belonging to the group of documents (registered in the document database 11) that is the category of interest. Because the word database 10 has the structure shown in Figure 4, it is also possible to identify the word IDs of all words that appear in a document by referencing it using the document ID. In the example in Figure 4, both the word registered as word ID "T0001" and the word registered as word ID "T0002" are included in both the set associated with document ID "B001" and the set associated with document ID "B002".
[0031] Each of the configurations shown in Figure 2 is actually implemented by installing a dedicated program on hardware such as a computer and its peripheral devices, as shown in Figure 1. In other words, the computer executes the contents of each configuration according to the dedicated program. In this embodiment, it is preferable that the CPU is a multi-core CPU. In this specification, "computer" means a device that has an arithmetic processing unit such as a CPU and a GPU and is capable of data processing, and includes not only general-purpose computers such as personal computers, but also portable terminals such as tablets equipped with a CPU and computers incorporated into various devices.
[0032] <2. Processing Operation> The information processing apparatus in this embodiment generates a scatter diagram in which each item is arranged by executing predetermined information processing. Next, together with the information processing method according to this embodiment, the processing operations of the information processing apparatus shown in FIGS. 1 and 2 will be described. FIG. 5 is a flowchart showing the processing operations of the information processing apparatus according to this embodiment. First, the morphological analysis means 25 reads each document from the document database 11 and extracts the words included in each document (step S10). Specifically, the morphological analysis means 25 executes morphological analysis on the extracted document and extracts words of a specific part of speech. As the specific part of speech, those specified in advance are used. For example, if the part of speech is specified as "noun", only noun words are extracted. The extracted words are registered in the word database 10 as shown in FIG. 4 together with the number of times they appear in the document.
[0033] Here, let the set of each document registered in the document database 11 be V, and each document be v j Let it be. Let the number of elements in the set V of documents be n. The set of documents is the category of interest, and the document v j is an item. And each document v j ∈V, that is, each document v j is an element of the set V. From each document v j , let the set of words of a specific part of speech extracted by morphological analysis be T j Let it be.
[0034] T j ={t j1 , t j2 , t j3 , ···}, j = 1, 2, 3, ···, n
[0035] For each word t ji (in the word database 10, managed by word ID), the value of the number of times that word appears in the document v j (in the word database 10, managed by document ID) can be obtained (see FIG. 4).
[0036] Next, for any two combinations of documents from the set of documents registered in the document database 11 as the category of interest, a common element matrix M is calculated using two sets of words contained in each document (step S20). Each component of the common element matrix M is the number of common elements in the two sets. This common element matrix M is an n x n square matrix. Furthermore, this common element matrix M is a symmetric matrix, that is, a matrix that is identical to its own transpose. The number of common elements is the number of elements (in this case, words) that appear in common in the sets related to the two documents. The number of common elements that appear in common in the sets related to the two documents vj and vi is m ji This is how it is expressed. In step S20, the specific processing is as follows: First, for all combinations of documents vj for j=1,2,3,...,n and documents vi for i=1,2,3,...,n from among the n documents, the number of common elements m ji Find the value of the number of common elements m. ji This is expressed by the following equation (Equation 11).
[0037] m ji =n(T j ∩T i )(j=1,2,3,···,n;i=1,2,3,···,n)…… (Equation 11) However, n() represents the number of elements in the set.
[0038] The process shown in (Equation 11) is obtained by identifying two document IDs and counting the number of word IDs recorded in association with both, when using the word database 10.
[0039] The common element matrix M is defined as having a common element number m, as shown in (Equation 12) below. ji Let be a matrix whose element is in the jth row and ith column.
[0040] M=[m ji ](j=1,2,3,···,n;i=1,2,3,···,n)…… (Equation 12)
[0041] Under these conditions, the number of common elements is m. ji and the number of common elements m ij Since they are equal, the common element matrix M is a symmetric matrix.
[0042] Next, the relationship matrix S is obtained from the common element matrix M (step S30). Here, the relationship is information that indicates the relationship between the two, and can be expressed using various indicators. For example, similarity, which indicates the degree to which the two are similar, and distance, which indicates the degree to which the two are far apart, can be used. Similarity increases as the two are closer, and distance increases as the two are farther apart. Therefore, similarity and distance can be converted to each other using a predetermined conversion formula.
[0043] Specifically, the process for finding the relation matrix S involves first, for each j,i, the document v j The set T is the set of words that appear in [the text]. j , document v i The set T is the set of words that appear in [the text]. i An index value s that represents the relationship with ji Calculate the index value s. ji This is the function f(m) shown in (Equation 13) below. jj, m ii, m ji ) using m jj, m ii, m ji It is calculated from the numerical values of the three common elements.
[0044] s ji =f(m jj, m ii, m ji )…… (Equation 13)
[0045] (Formula 13) shows s ji s is an index value of an asymmetric relationship. ji The function f(m jj, m ii, m jiThere are many possible specific forms for this. In this embodiment, as will be described later, we mainly use the asymmetric normalized self-information of the joint probability distribution of two variables, the asymmetric normalized mutual information of the joint probability distribution of two variables, and the regression coefficients of the joint probability distribution of two variables.
[0046] Index value s ji Of the numerical values for the number of common elements used to calculate the number of common elements, m jj is document v j The number of words of a specific part of speech, the number of common elements m ii is document v i This represents the number of specific words in the context. The number of common elements is m. ji , number of common elements m ij These are all documents v i and document v j Since it is a number that appears in common, m ji =m ij That is the case.
[0047] The relationship matrix S is defined by the relationship index value s, as shown in (Equation 14) below. ji Let be a matrix whose element is in the jth row and ith column.
[0048] S =[s ji ](j=1,2,3,···,n;i=1,2,3,···,n)…… (Equation 14)
[0049] The number of row and column elements of the relation matrix S is the same as that of the common element matrix M, and it is an asymmetric n x n square matrix corresponding to the number of documents n. In this embodiment, it is adjusted so that the distance between itself is 0. That is, s ii , s jj The diagonal elements of the relation matrix S are all zero. The details of this step S30 will be described later.
[0050] Once the relationship matrix S is obtained through the processing in step S30, a monotonically changing adjustment function is generated based on the relationship matrix S, and the adjustment matrix W is obtained by applying the adjustment function (step S40). Specifically, the adjustment matrix generation means 23 generates index values s of the relationships that constitute the relationship matrix S.ji Using the corresponding component w ji By calculating this, the adjustment matrix W is generated. Each component w of the adjustment matrix W ji This can be determined as a monotonically changing adjustment function using either a monotonically increasing adjustment function or a monotonically decreasing adjustment function. The monotonically changing adjustment function is generated based on the relation matrix obtained by the relation matrix generation means 22.
[0051] If the properties of the components of the relation matrix obtained in step S30 are to be used as is, a monotonically increasing adjustment function f(x) is used in step S40. On the other hand, if the properties of the components of the relation matrix obtained in step S30 are to be reversed, a monotonically decreasing adjustment function g(x) is used in step S40. As a monotonically increasing adjustment function f(x), for example, the ones shown in (Equation 15) and (Equation 16) below can be used. Also, as a monotonically decreasing adjustment function g(x), for example, the ones shown in (Equation 17) and (Equation 18) below can be used. The following (Equation 15) is an example of a monotonically increasing adjustment function f(x).
[0052] (Formula 15)
number
[0053] In (Equation 15), x is the index value of the relationship, s. ji Substitute this and let the component w be f(x). ji We obtain (Equation 15), where σ is the standard deviation of the components of the triangular matrix in relation matrix S, 2 σ is the variance of the components of the triangular matrix in the relation matrix S, and μ is the mean of the components of the triangular matrix in the relation matrix S. Here, a triangular matrix refers to an upper or lower triangular matrix in the relation matrix S, excluding the diagonal components. In this embodiment, the variance σ of the upper and lower triangular matrices in the relation matrix S is... 2Since the mean μ is the same in both cases, either an upper triangular matrix or a lower triangular matrix can be used as the triangular matrix. Thus, the monotonically increasing adjustment function f(x) shown in (Equation 15) is a function generated based on the relation matrix S.
[0054] exp() is an exponential function with base Napier's number, and the expression in parentheses following exp is the exponent. (Equation 15) is a normal distribution N(μ, σ) 2 The cumulative probability density function is the function obtained by integrating the probability density function of ) from negative infinity to x. In reality, it is not possible to set negative infinity, so the adjusted matrix generation means 23 sets a sufficiently small negative value to obtain the component w ji Calculate.
[0055] Figure 6 shows the relationship between the histogram of the components of the triangular matrix in the relation matrix S and the adjustment function f(x). In the example in Figure 6, the components of the triangular matrix are concentrated around 0. In such cases, each component s in the relation matrix S... ji By applying the adjustment function f(x) to the adjustment matrix W, each component w ji The value approaches a normal distribution. By applying the adjustment function f(x) in this way and obtaining the adjusted matrix W, it becomes possible to emphasize the cluster separation properties.
[0056] Furthermore, the adjustment matrix generation means 23 performs the following process according to (Equation 16) instead of the adjustment function f(x) shown in (Equation 15) above, to generate each component w of the adjustment matrix W. ji You may calculate this.
[0057] (Formula 16)
number
[0058] In (Equation 16), x is the index value of the relationship, s. ji Substitute this and let the component w be f(x). jiWe obtain the following. In equation 16, μ is the mean of the components of the triangular matrix in relation matrix S. The variance σ of the components of the triangular matrix in relation matrix S 2 It is not used. Thus, the monotonically increasing adjustment function f(x) shown in (Equation 16) is a function generated based on the relation matrix S. a is the coefficient obtained by dividing 2 by the standard deviation σ, so a = 2 / σ. (Equation 16) is the sigmoid function.
[0059] Comparing the case using the cumulative probability density function shown in (Equation 15) with the case using the sigmoid function shown in (Equation 16), the cumulative probability density function shown in (Equation 15) requires more computation due to the integration operation, while the sigmoid function shown in (Equation 16) requires less computation. Therefore, the adjusted matrix W can be obtained faster by calculating it using the sigmoid function shown in (Equation 16).
[0060] Furthermore, the following (Equation 17) is an example of a monotonically decreasing adjustment function g(x).
[0061] (Formula 17)
number
[0062] In (Equation 17), x is the index value of the relationship, s. ji Substitute this and let the component w be g(x). ji We obtain the following. As is clear from comparing (Equation 15) and (Equation 17), the adjustment function g(x) is obtained by subtracting the monotonically increasing adjustment function f(x) shown in (Equation 15) from 1. Similar to (Equation 15), in (Equation 17), σ is the standard deviation of the components of the triangular matrix in the relation matrix S, and σ 2 μ is the variance of the components of the triangular matrix in relation matrix S, and μ is the mean of the components of the triangular matrix in relation matrix S.
[0063] Here, a triangular matrix refers to an upper or lower triangular matrix obtained by removing the diagonal elements from the relation matrix S. In this embodiment, the variance σ of the upper and lower triangular matrices in the relation matrix S is... 2 Since the mean μ is the same in both cases, either an upper triangular matrix or a lower triangular matrix can be used as the triangular matrix. Thus, the monotonically decreasing adjustment function g(x) shown in (Equation 17) is a function generated based on the relation matrix S.
[0064] Although not shown in the diagram, when using a monotonically decreasing adjustment function g(x), the same applies to each component s in the relation matrix S as in the example shown in Figure 6. ji By applying the adjustment function g(x) to the adjusted matrix W, each component w ji The value approaches a normal distribution. By applying the adjustment function g(x) in this way and obtaining the adjusted matrix W, it becomes possible to emphasize the cluster separation properties.
[0065] Furthermore, the matrix to be adjusted means 23 performs a process according to the adjustment function g(x) shown in (Equation 18) below, instead of the adjustment function g(x) shown in (Equation 17) above, to adjust each component w of the matrix to be adjusted W. ji You may calculate this.
[0066] (Formula 18)
number
[0067] In equation (18), a = 2 / σ. In equation (18), x is the index value of the relationship s. ji Substitute this and let the component w be f(x). ji We obtain the following. As is clear from comparing (Equation 16) and (Equation 18), the adjustment function g(x) is obtained by subtracting the monotonically increasing adjustment function f(x) shown in (Equation 16) from 1. In (Equation 18) as well, μ is the mean of the components of the triangular matrix in the relation matrix S. Variance σ of the components of the triangular matrix in the relation matrix S 2It is not used. Thus, the monotonically decreasing adjustment function g(x) shown in (Equation 18) is a function generated based on the relation matrix S. a is the coefficient obtained by dividing 2 by the standard deviation σ, so a = 2 / σ. (Equation 18) is the sigmoid function.
[0068] By using the monotonically increasing adjustment function f(x) in (Equation 15) and (Equation 16), and the monotonically decreasing adjustment function g(x) in (Equation 17) and (Equation 18), the relational components s in the relational matrix S can be obtained. ji Adjust the relationship component after adjustment w ji This can be obtained. The adjustment functions shown in (Equation 15), (Equation 16), (Equation 17), and (Equation 18) all use the standard deviation σ of the components of the triangular matrix in relation matrix S and the mean μ of the components of the triangular matrix in relation matrix S. By adjusting the values of the relational components using an adjustment function in this way, the cluster segregation of each item can be emphasized.
[0069] In this embodiment, in order to further emphasize cluster separation, the adjusted matrix generation means 23 recalculates the variables by performing the following process according to (Equation 19).
[0070] (Formula 19)
number
[0071] In equation (19), r is a coefficient satisfying r≧1, and s is a reduction coefficient satisfying s0≦s≦1. Both r and s are real values.
[0072] Perform the process according to (Equation 19), and substitute the recalculated variables σ' and μ' into (Equation 15) and (Equation 17), respectively, to obtain the components w of the adjusted matrix W. ji The following is calculated. Furthermore, the recalculated variables a' and μ' are substituted for variables a and μ in (Equation 16) and (Equation 18) to obtain the values of each component w of the adjusted matrix W. jiCalculate it. By recalculating the variable in this way, the separability of the clusters can be further emphasized.
[0073] FIG. 7 is a diagram showing the relationship between the reduction coefficient s and the adjustment function f(x). In FIG. 7, when only the reduction coefficient s is changed, the corresponding adjustment function f(x) is shown. In the example of FIG. 7, the adjustment functions f(x) for three types of s = 1, 1 / 2, and 1 / 4 are shown. By setting the reduction coefficient s to a number less than 1, it becomes possible to further emphasize the cluster separability.
[0074] When using the distance as a component of the relational matrix S, by applying a monotonically increasing adjustment function f(x), each component w of the matrix W to be adjusted ji also becomes a distance. That is, in this case, the matrix W to be adjusted is a matrix of distances to be adjusted. When using the similarity as a component of the relational matrix S, by applying a monotonically increasing adjustment function f(x), each component w of the matrix W to be adjusted ji also becomes a similarity. That is, in this case, the matrix W to be adjusted is a matrix of similarities to be adjusted.
[0075] When using the distance as a component of the relational matrix S, by applying a monotonically decreasing adjustment function g(x), each component w of the matrix W to be adjusted ji becomes a similarity. That is, in this case, the matrix W to be adjusted is a matrix of similarities to be adjusted. When using the similarity as a component of the relational matrix S, by applying a monotonically decreasing adjustment function g(x), each component w of the matrix W to be adjusted ji becomes a distance. That is, in this case, the matrix W to be adjusted is a matrix of distances to be adjusted.
[0076] Next, based on the matrix W to be adjusted, coordinate values are calculated (step S50). Specifically, the coordinate value calculation means 24 calculates coordinate values for arranging each item in a k-dimensional space based on the matrix W to be adjusted. The specific processing method differs depending on whether each component w of the matrix W to be adjusted ji is a distance or a similarity, that is, whether the matrix W to be adjusted is a matrix of distances to be adjusted or a matrix of similarities to be adjusted.
[0077] First, each component w of the adjustment matrix W ji This section describes the case where the distance matrix W is the distance matrix. When the distance matrix W is the distance matrix, the coordinate value calculation means 24 applies a Young-Householder transformation to the distance matrix W to obtain a matrix Y, and then performs an operation by applying a multidimensional scaling method to decompose matrix Y into the form Y=XTX. This calculates the coordinate values for locating each item in k-dimensional space.
[0078] In this process, the coordinate value calculation means 24 first applies a Young-Householder transformation to the adjusted distance matrix W. The Young-Householder transformation is a transformation that uses a centering matrix Cn on an n x n square matrix to obtain matrix Y according to the following equation (formula 20).
[0079] (Formula 20) Y = -1 / 2·CnDCn
[0080] When n=3, Cn is given by the following (Equation 21).
[0081] (Formula 21)
number
[0082] Next, let the matrix Y be Y=X T We decompose it into the form X to obtain matrix X. This matrix X is a k x n coordinate matrix, and we can obtain the k-dimensional coordinates of n items.
[0083] Next, each component w of the adjusted matrix W ji We will now explain the case where the similarity is represented by the adjusted matrix W, i.e., the adjusted similarity matrix.
[0084] When the matrix W to be adjusted is the similarity matrix to be adjusted, in step S50, after the coordinate value calculation means 24 generates the graph Laplacian matrix P from the matrix W to be adjusted which is the similarity matrix to be adjusted, based on the graph Laplacian matrix P, it calculates the coordinate values for arranging each item in the k-dimensional space. A k×n graph Laplacian matrix P is generated from an n×n square matrix, and the k-dimensional coordinates of n items are obtained.
[0085] As described above, when the components of the relational matrix S obtained in step S30 are distances, if the monotonically increasing adjustment function f(x) is used in step S40, the components of the matrix W to be adjusted become distances. Conversely, when the monotonically decreasing adjustment function g(x) is used in step S40 when the components of the relational matrix S are distances, the components of the matrix W to be adjusted become similarities. On the other hand, when the components of the relational matrix S obtained in step S30 are similarities, if the monotonically increasing adjustment function f(x) is used in step S40, the components of the matrix W to be adjusted become similarities. Conversely, when the monotonically decreasing adjustment function g(x) is used in step S40 when the components of the relational matrix S are similarities, the components of the matrix W to be adjusted become distances.
[0086] And in step S50, for each component w of the matrix W to be adjusted ji when it is a distance, after performing the Young-Householder transformation on the matrix W to be adjusted which is the distance matrix to be adjusted to obtain the matrix Y, it performs an operation applying the multidimensional scaling method of decomposing the matrix Y into the form of Y = XTX to calculate the coordinate values for arranging each item in the k-dimensional space. On the other hand, for each component w of the matrix W to be adjusted ji when it is a similarity, after generating the graph Laplacian matrix P from the matrix W to be adjusted which is the similarity matrix to be adjusted, based on the graph Laplacian matrix P, it calculates the coordinate values for arranging each item in the k-dimensional space.
[0087] Whether the components of the relation matrix calculated in step S30 are distance or similarity, and whether the adjustment function used in step S40 is a simply increasing adjustment function f(x) or a simply decreasing adjustment function g(x), can be appropriately set depending on the type of data being handled and the purpose of visualization. Furthermore, whether to use the Young-Householder transform or the graph-Laplacian matrix in step S50 is set according to the type of components of the adjustment matrix W obtained in step S40.
[0088] <3. Details of Step S30> As described above, in this embodiment, in step S30, the number of common elements is m. ji and the number of common elements m ij The relationship indicator value s ji The function f(m jj, m ii, m ji The following are primarily used as the asymmetric normalized self-information of the joint probability distribution of two variables, the asymmetric normalized mutual information of the joint probability distribution of two variables, and the regression coefficients of the joint probability distribution of two variables. Of these, the asymmetric normalized self-information and asymmetric normalized mutual information are used to determine the number of common elements m ji and the number of common elements m ij The relationship indicator value s ji The approach to finding this will be explained below.
[0089] In this specification, the property indicating whether two objects are similar or dissimilar is defined as "relationship," and the degree to which two objects are similar to each other is defined as "similarity." Therefore, "distance" and "similarity," which are opposite concepts, are sub-concepts of "relationship," which represents the property indicating whether two objects are similar or dissimilar. First, as a general theory, we will list some distance indicators and similarity indicators between two given finite sets A and B. We will also explain how they can be derived. Indices can be classified into the following six types depending on whether they are distance (dissimilarity) or similarity, and whether they are symmetric or asymmetric.
[0090] (A1) Distance indicator, symmetry (A2) Similarity index, symmetry (A3) Distance indicator, asymmetric (from set A to set B) (A4) Distance indicator, asymmetric (from set B to set A) (A5) Similarity index, asymmetric (from set A to set B) (A6) Similarity index, asymmetric (from set B to set A)
[0091] However, (A4) can be mechanically obtained by swapping set A and set B in (A3), and (A6) can be similarly obtained from (A5), so it can be considered that there are essentially four types.
[0092] For each category, there are various possible indicators, and there isn't just one definitive set.
[0093] (B1) Indices derived from self-information (B2) Indices derived from mutual information (B3) Other various indicators
[0094] From here, we will describe the methods for deriving (B1) "indicators derived from self-information" and (B2) "indicators derived from mutual information." As preparation for this, we will define "self-information," which is the premise for (B1), and "entropy," which is the premise for (B2). For both, we will describe the cases with one random variable and the cases with two random variables.
[0095] Let U be a finite set (for example, all words recorded in word database 10), and let us assume that there are two subsets of U, A (words appearing in one document) and B (words appearing in the other document). Based on the number of elements in each set U, A, B, and A∩B (words appearing in both documents), we can define a joint probability distribution for random variables X and Y. Then, using the normalized index generated from the self-information of this joint probability distribution, we can use it as an index of the distance and similarity between A and B. (B2) For mutual information, by using "mutual information" instead of "self-information" in (B1), we can create six types of indices between the two sets through almost the same derivation process.
[0096] Here, we introduce random variables X and Y, corresponding to sets A and B, respectively. When α is an element of set A, the random variable X takes the value 1; otherwise, it takes the value 0. That is, α ∈ A → X = 1 If α ∈ A, then → X = 0 Similarly, we introduce a random variable Y for set B. While each variable can take only two values, 0 or 1, for generality, we assume that they can take m and q values, respectively. Based on self-information and mutual-information, six types of indicators can be created for each, resulting in a total of 12 types of indicators (indicator 1 to indicator 12, described below). These 12 types of indicators represent the relationships between sets.
[0097] First, let's explain self-information and entropy (single variable). For this purpose, we will set them up as follows.
[0098] Let there be a finite set X whose elements are x1, x2, x3, ..., xm. X = {x1, x2, x3, ..., xm} Let the random variable X take on elements of X as its value. Let x represent one element of X. x∈X The probability that the random variable X is x is denoted as P(X=x).
[0099] We define a certain amount of information h(X=x) that represents the information that "an event occurred in which the random variable X took the value x" as follows: h(X=x) = -log(P(X=x))
[0100] This h(X=x) is called "self-information" (or "selection information" or "self-entropy"). The base of the logarithm is typically 2. • Although h(X=x) is a dimensionless number, if we use 2 as the base of the logarithm, we may add the unit [bit]. Here, we omit the unit. For example, the self-information value of the information "an event with a probability of 1 / 8 occurred" is 3.
[0101] If H(X) is the mean (or expected value) of the self-information h(X=x) over all x∈X, then H(X) = Σ x∈X P(X=x)h(X=x). When taking the average, each event is weighted according to its probability of occurring. Expanding h(X=x) gives the following:
[0102] H(X) = -Σ x∈X P(X=x)log(P(X=x))
[0103] H(X) is called the "average information" (or "Shannon information" or "entropy of information") of X. Up to this point, we have defined self-information and entropy for the case of one variable. Next, we will discuss the case of two variables.
[0104] Given two sets X and Y, where X = {x1, x2, x3, ..., xm} and Y = {y1, y2, y3, ..., yq;}, the set Z shown below is called the "Cartesian product" of set X and set Y, and is expressed as follows.
[0105] Z = X × Y Z={(x,y); x∈X,y∈Y}
[0106] In the Cartesian product Z of sets X and Y, the probability that a random variable X takes the value x (∈X) and a random variable Y takes the value y (∈Y) is denoted as P(X=x,Y=y). This is called the joint probability or joint probability. If P(X=x,Y=y)=P(X=x)P(Y=y) holds for all x (∈X) and y (∈Y), then we say that the random variables X and Y are independent of each other. However, we will not assume independence below.
[0107] Here, the following equations (Equation 22-1) and (Equation 22-2) can be said. This is called "marginalization".
[0108] P(X=x)=Σ y∈Y P(X=x, Y=y)…… (Equation 22-1) P(Y=y) = Σ x∈X P(X=x,Y=y)…… (Equation 22-2)
[0109] In marginalization, each term being summed takes a non-negative value, so the following inequalities related to marginalization hold: (Equation 22-3) (Equation 22-4).
[0110] 0≦P(X=x,Y=y)≦P(X=x)…… (Formula 22-3) 0≦P(X=x,Y=y)≦P(Y=y)…… (Formula 22-4)
[0111] If we consider the Cartesian product set Z = X × Y of sets X and Y as a one-dimensional set with m × q elements, then the self-information and mean information can be defined naturally. When a random variable X takes the value x ∈ X and a random variable Y takes the value y ∈ Y, the self-information h(X=x, Y=y) can be defined as h(X=x, Y=y) = -log(P(X=x, Y=y)).
[0112] If H(X, Y) is the mean value of the self-information h(X=x, Y=y) over all x(∈X) and y(∈Y), then H(X, Y) = Σ x∈X Σ y∈Y P(X=x,Y=y)h(X=x,Y=y). If we write the summation as a single unit, H(X, Y)=Σ (x, y)∈X×Y P(X=x,Y=y) can also be written as h(X=x,Y=y). Expanding H(X=x,Y=y), we get H(X, Y)=-Σ (x, y)∈X×Y P(X=x,Y=y) log(P(X=x,Y=y)). This is called "joint entropy".
[0113] Up to this point, we have defined self-information and entropy, both for the one-variable and two-variable cases, respectively. From here, we will discuss the index derived from self-information, which corresponds to (B1) of the above (B1) to (B3). As preparation for this, we will first introduce the non-negative correlation assumption and then describe how to derive the three inequalities concerning self-information that hold under this assumption.
[0114] First, let's explain the properties of self-information. The following equations (Equation 23-1) and (Equation 23-2) hold true as the "first inequality concerning self-information".
[0115] h(X=x,Y=y)≧h(X=x)…… (Equation 23-1) h(X=x,Y=y)≧h(Y=y)…… (Formula 23-2)
[0116] We will prove the first inequality concerning self-information. As described above, the following inequality related to marginalization holds.
[0117] 0≦P(X=x,Y=y)≦P(X=x) …… (Formula 22-3) 0≦P(X=x,Y=y)≦P(Y=y)…… (Formula 22-4)
[0118] Due to the monotonically increasing nature of logarithmic functions, taking the logarithm of each side yields the same inequality, as shown below.
[0119] log(P(X=x,Y=y))≦log(P(X=x)) log(P(X=x,Y=y)) ≤ log(P(Y=y))
[0120] Multiplying both sides by (-1) reverses the inequality sign, resulting in the following:
[0121] -log(P(X=x,Y=y))≧-log(P(X=x)) -log(P(X=x,Y=y))≧-log(P(Y=y))
[0122] Proof complete. Therefore, the following equation holds true.
[0123] h(X=x,Y=y)≧h(X=x) h(X=x,Y=y)≧h(Y=y)
[0124] Here, we limit our consideration to the case where there is a non-negative correlation between event X = x and event Y = y (non-negative correlation assumption). P(X=x,Y=y)≧P(X=x)P(Y=y)…… (Formula 23-3)
[0125] Under the assumption of non-negative correlation (Equation 23-3), by taking the logarithm of both sides and multiplying by (-1), the inequality sign is reversed, and the following (Equation 23-4) holds as the "second inequality concerning self-information".
[0126] h(X=x,Y=y)≦h(X=x)+h(Y=y)…… (Formula 23-4)
[0127] Let's summarize the relationships regarding self-information. (Equations 23-1), (Equation 23-2), and (Equation 23-4) are shown again. h(X=x,Y=y)≧h(X=x)…… (Equation 23-1) h(X=x,Y=y)≧h(Y=y)…… (Formula 23-2) h(X=x,Y=y)≦h(X=x)+h(Y=y)…… (Formula 23-4) However, (Equation 23-4) is based on the non-negative correlation assumption (Equation 23-3). P(X=x,Y=y)≧P(X=x)P(Y=y)≧0…… (Formula 23-3) It is based on the following conditions.
[0128] Up to this point, we have derived three inequalities concerning self-information under the assumption of non-negative correlation. Using these, we can derive symmetric and asymmetric normalized self-information, but first we will describe how to derive the symmetric one. The above relationships concerning self-information (Equations 23-1), (Equations 23-2), and (Equations 23-4) show the lower and upper bounds of the possible values of the self-information h(X=x, Y=y). They can be combined and written as follows.
[0129] max(h(X=x,Y=y))≦h(X=x,Y=y)≦h(X=x)+h(Y=y)
[0130] When consolidating them into one, the following methods of summarization are also possible.
[0131] min(h(X=x,Y=y))≦h(X=x,Y=y)≦h(X=x)+h(Y=y)
[0132] However, the way min(h(X=x,Y=y)) is expressed is an inequality that is "too broad" with too much margin in the lower limit. Below, we will use max(h(X=x,Y=y)). However, using the latter allows us to derive the Google Distance described in the non-patent literature below. In other words, the index derived from (B1) self-information, which will be discussed below, is something different.
[0133] Non-patent literature: Rudi L. Cilibrasi and Paul MB Vitanyi. "The google similarity distance." IEEE Transactions on Knowledge and Data Engineering, Vol. 19, pp. 370-383, 2007.
[0134] The value of r can be defined by considering h(X=x,Y=y) as the point that divides the lower and upper bounds internally in the ratio r: 1-r. Alternatively, by looking at the internal division ratio in reverse, the value of r' can be defined by considering h(X=x,Y=y) as the point that divides the lower and upper bounds internally in the ratio 1-r':r'. Then, it is clear that r and r' satisfy the following equations.
[0135] 0≦r≦1,0≦r′≦1,r+r´=1
[0136] Rewriting the above equation, we get the following:
[0137] r=(h(X=x,Y=y)-max(h(X=x,Y=y))) / min(h(X=x,Y=y)) r´=(h(X=x)+h(Y=y)-h(X=x,Y=y)) / min(h(X=x,Y=y))
[0138] r represents distance, and r' represents similarity. Under the assumption of non-negative correlation, both r and r' take values between 0 and 1 (inclusive), but when there is a negative correlation, r' > 1 and r' < 0. As the joint probability approaches 0, they can become arbitrarily large or small. Here, we will call r and r' "normalized self-information," and respectively, d NSI (X=x, Y=y), sim NSI We will represent this as (X=x, Y=y).
[0139] To summarize the normalized self-information (symmetric), d NSI (X=x, Y=y), sim NSI (X=x, Y=y) can be expressed as follows:
[0140] d NSI (X=x,Y=y)=(h(X=x,Y=y)-max(h(X=x,Y=y))) / min(h(X=x),h(Y=y)) SIM NSI (X=x,Y=y)=(h(X=x)+h(Y=y)-h(X=x,Y=y)) / min(h(X=x,Y=y))
[0141] Here, d NSI distance, sim NSI This represents the degree of similarity. The above assumes non-negative correlation (Equation 23-3). Under the condition P(X=x,Y=y)≧P(X=x)P(Y=y)≧0, the following holds:
[0142] 0≦d NSI (X=x,Y=y)≦1 0≦sim NSI (X=x,Y=y)≦1
[0143] When there is a negative correlation, it will be as follows:
[0144] d NSI (X=x,Y=y)>1 SIM NSI (X=x,Y=y)<0
[0145] Up to this point, we have discussed the symmetric indices derived from (B1) self-information. Next, we will discuss the asymmetric ones. The relationships concerning self-information in (Equations 23-1), (Equation 23-2), and (Equation 23-4) above indicate the lower and upper bounds of the possible values of h(X=x,Y=y). In the above examples, symmetry was preserved by combining them using max(), but here we will treat them separately.
[0146] h(X=x)≦h(X=x,Y=y)≦h(X=x)+h(Y=y) h(Y=y)≦h(X=x,Y=y)≦h(X=x)+h(Y=y)
[0147] From here, as in the symmetric case, if we create a normalization index using the internal division ratio, we obtain the following asymmetric equation as the asymmetric normalized self-information.
[0148] d NSI (X=x→Y=y)=h(X=x,Y=y)-h(X=x) / h(Y=y) d NSI (Y=y→X=x)=h(X=x,Y=y)-h(Y=y) / h(X=x) SIM NSI (X=x→Y=y)=(h(X=x)+h(Y=y)-h(X=x,Y=y)) / h(Y=y) SIM NSI (Y=y→X=x)=(h(X=x)+h(Y=y)-h(X=x,Y=y)) / h(X=x)
[0149] Here, dNSI represents distance and simNSI represents similarity. Under the above non-negative correlation assumption (Equation 23-3) P(X=x,Y=y)≧P(X=x)P(Y=y)≧0, the following holds:
[0150] 0≦d NSI (X=x→Y=y)≦1 0≦d NSI (Y=y→X=x)≦ 0≦sim NSI (X=x→Y=y)≦1 0≦sim NSI (Y=y→X=x)≦1
[0151] Up to this point, we have described how to derive the index derived from self-information, which corresponds to (B1) in the aforementioned (B1) to (B3). Next, we will describe how to derive the index derived from mutual information, which corresponds to (B2). This can be done by using entropy instead of self-information in (B1), following a nearly identical path. As preparation, we will first describe how to derive the three inequalities related to entropy. For joint entropy, the "first inequality related to joint entropy" shown in (Equation 24-1) and (Equation 24-2) holds.
[0152] H(X, Y) ≥ H(X)…… (Equation 24-1) H(X, Y) ≥ H(Y)…… (Equation 24-2)
[0153] Prove the "First Inequality Regarding Joint Entropy". Prove that the value obtained by subtracting the right side from the left side in (Equation 24-1) is non-negative.
[0154] H(X, Y) - H(X) = -Σ (x,y)∈X×Y P(X = x, Y = y)log(P(X = x, Y = y)) + Σ x∈X P(X = x)log(X = x)
[0155] Here, from marginalization (Equation 22-1), P(X = x) = Σ y∈Y P(X = x, Y = y), so the following equation holds.
[0156] H(X, Y) - H(X) = -Σ (x,y)∈X×Y P(X = x, Y = y)log(P(X = x, Y = y)) + Σ x∈X {Σ y∈Y P(X = x, Y = y)}log(P(X = x)) = -Σ (x,y)∈X×Y P(X = x, Y = y)log(P(X = x, Y = y)) + Σ (x,y)∈X×Y P(X = x, Y = y)log(P(X = x)) = -Σ (x,y)∈X×Y P(X = x, Y = y)log(P(X = x, Y = y) / P(X = x)) + Σ x∈X {Σ y∈Y P(X = x, Y = y)}log(P(X = x))
[0157] Here, from marginalization (Equation 22-3), 0 ≤ P(X = x, Y = y) ≤ P(X = x). Therefore, all terms take non-negative values. Thus, the following equation holds.
[0158] H(X, Y) - H(X) ≥ 0
[0159] The same applies to (Equation 24-2) by swapping X and Y. This concludes the proof. Regarding bond entropy, the "second inequality concerning bond entropy" shown in (Equation 24-3) holds.
[0160] H(X, Y) ≤ H(X) + H(Y) ... (Equation 24-3)
[0161] Equality holds only when X and Y are independent random variables, that is, when P(X=x,Y=y)=P(X=x)P(Y=y). Next, we will prove this.
[0162] From marginalization (Equation 22-1), (Equation 22-2), P(X=x)=Σ y∈Y P(X=x, Y=y)…… (Equation 22-1) P(Y=y) = Σ x∈X P(X=x,Y=y)…… (Equation 22-2) Therefore, the following equation holds true.
[0163] H(X) = -Σ (x,y)∈X×Y P(X=x,Y=y)log(P(X=x)) H(Y) = -Σ (x,y)∈X×Y P(X=x,Y=y)log(P(Y=y))
[0164] Therefore, the following equation also holds true.
[0165] H(X)+H(Y)-H(X, Y)=-Σ (x,y)∈X×Y P(X=x,Y=y)log((P(X=x)P(Y=y) / P(X=x,Y=y))
[0166] Here, for any u > 0, log(u) ≤ u - 1, that is, -log(u) ≥ 1 - u, so the following equation holds.
[0167] H(X)+H(Y)-H(X, Y)≧ Σ (x,y)∈X×YP(X=x,Y=y)(1-P(X=x)P(Y=y) / P(X=x,Y=y))=0
[0168] This concludes the proof. Now, let's explain "mutual information". When random variables X and Y are not independent of each other, H(X, Y) and H(X)+H(Y) do not coincide. As shown in (Equation 24-3) above, H(X, Y)≦H(X)+H(Y)…… (Equation 24-3). The difference in the information between the two in (Equation 24-3) ((right side) - (left side)) is called "mutual information" and is expressed as follows. Mutual information is always a non-negative value.
[0169] I(X, Y)=H(X)+H(Y)-H(X, Y)... (Formula 24-4)
[0170] Let's explain "conditional entropy." The conditional self-information h(A|B) of event A, given that event B has occurred, is defined as follows.
[0171] h(A|B)=-log(P(A|B))
[0172] Given a random variable X, the expected value with respect to x of the conditional self-information h(X = x|B) = -log(P(X = x|B)) of event X = x under event B is called the "conditional entropy" and is expressed as follows:
[0173] H(X |B) = -Σ x∈X P(X = x|B)log(P(X = x|B))
[0174] Furthermore, given a random variable Y, the expected value of the conditional entropy H(X | Y=y) of event X = x, under the condition that event Y = y has occurred, can be expressed as follows. This is also called "conditional entropy".
[0175] H(X | Y) = Σ y∈Y P(Y=y) H(X | Y=y)
[0176] The following (Equation 24-5) and (Equation 24-6) hold as equations regarding conditional entropy.
[0177] H(X | Y)=H(X,Y)-H(Y)…… (Equation 24-5) H(Y | X)=H(X,Y)-H(X)……(Equation 24-6)
[0178] Summarizing the relationships of the average information quantity H and the mutual information quantity I up to this point, it becomes as follows.
[0179] H(X,Y) ≧H(Y) …… (Equation 24-1) H(X,Y) ≧H(Y) …… (Equation 24-2) H(X,Y) ≦ H(X)+H(Y) …… (Equation 24-3) I(X,Y) =H(X) +H(Y)-H(X,Y) …… (Equation 24-4) H(X|Y) =H(X,Y) -H(Y) …… (Equation 24-5) H(Y|X) =H(X,Y) -H(X) …… (Equation 24-6)
[0180] Up to this point, three inequalities regarding entropy have been derived. By using these, symmetric and asymmetric normalized mutual information quantities can be derived. First, the method of deriving the symmetric one will be described. (Equations 24-1) to (Equation 24-3) are reproduced below. H(X, Y)≧H(X) …… (Equation 24-1) H(X, Y)≧H(Y) …… (Equation 24-2) H(X, Y)≦H(X)+H(Y) …… (Equation 24-3)
[0181] (Equations 24-1) to (Equation 24-3) show the lower limit and the upper limit of the possible values of H(X, Y). Summing them up, they can be written as follows.
[0182] max(H(X), H(Y))≦H(X, Y)≦H(X)+H(Y)
[0183] We can define r by viewing H(X, Y) as the point that divides the lower and upper bounds internally in the ratio r:1-r. Alternatively, by inverting the internal division ratio, we can define r' by viewing H(X, Y) as the point that divides the lower and upper bounds internally in the ratio 1-r':r'. Then, it is clear that 0≦r≦1, 0≦r'≦1, and r+r'=1. Rewriting the above equation, we get the following.
[0184] r={H(X, Y)-max(H(X), H(Y))} / min(H(X),H(Y)) r´={H(X)+H(Y)-H(X, Y)} / min(H(X),H(Y))
[0185] The numerator of r' is the mutual information I(X, Y), but the presence of a denominator allows for normalization. Here, r represents distance and r' represents similarity. It is very similar to the normalization index for self-information, but in the case of mutual information, the assumption of non-negative correlation is not required. r and r' are generally called "normalized mutual information," and are expressed as d, respectively. NMI (X, Y), sim NMI This is represented as (X, Y).
[0186] To summarize the symmetrical normalized mutual information, it is called "normalized mutual information" d NMI (X, Y), sim NMI (X, Y) can be expressed as follows. In this case, the assumption of non-negative correlation is not necessary.
[0187] d NMI ={H(X, Y)-max(H(X), H(Y)) / min(H(X), H(Y)) SIM NMI =(H(X)+H(Y )-H(X, Y)) / min(H(X), H(Y))
[0188] The above d NMI (X, Y), sim NMI For (X, Y), 0≦d NMI (X, Y) ≤ 1, 0 ≤ sim NMI (X, Y)≦1 holds. Here, d NMI distance, sim NMI represents the degree of similarity. However, when X and Y are uncorrelated, d NMI It takes a maximum value of 1, sim NMI It takes a minimum value of 0, and when there is a negative correlation, it turns downward and upward again. That is, d NMI and sim NMI This represents how well the information from X was transmitted to Y, with no transmission occurring when there is no correlation and transmission occurring when there is a negative correlation.Up to this point, we have discussed the symmetric indicators derived from (B2) entropy (mutual information).Next, we will discuss the asymmetric ones.
[0189] (Equations 24-1) to (Equations 24-3) are reproduced below. H(X, Y)≧H(X) …… (Equation 24-1) H(X, Y)≧H(Y) …… (Equation 24-2) H(X, Y) ≤ H(X) + H(Y) …… (Equation 24-3)
[0190] This shows the lower and upper bounds of the possible values of H(X, Y). In the example above, symmetry was preserved by combining them using max(), but here they are treated separately.
[0191] H(X) ≤ H(X, Y) ≤ H(X) + H(Y) H(Y) ≤ H(X, Y) ≤ H(X) + H(Y)
[0192] From here, similar to the example above, if we create a normalization index using the internal division ratio, we obtain the following asymmetric equation as an asymmetric normalized mutual information quantity.
[0193] d NMI (X→Y)=(H(X, Y)-H(X)) / H(Y) d NMI (Y→X)=(H(X, Y)-H(Y)) / H(X) SIM NMI (X→Y)=(H(X)+H(Y)-H(X, Y)) / H(Y) SIM NMI (Y→X)=(H(X)+H(Y)-H(X, Y)) / H(X)
[0194] Since the assumption of non-negative correlation is not required here, the following holds: d NMI distance, sim NMI This represents the degree of similarity.
[0195] 0≦d NMI (X→Y)≦1 0≦d NMI (Y→X)≦1 0≦sim NMI (X→Y)≦1 0≦sim NMI (Y→X)≦1
[0196] Up to this point, we have described how to derive the index derived from entropy, which corresponds to (B2) in (B1) to (B3) mentioned above. Up to this point, (B1) and (B2) in (B1) to (B3) mentioned above have been defined using general random variables X and Y, but in order to finally reduce them to an index between two sets, we do the following.
[0197] The universal set U is a finite set, and it has two subsets, set A and set B (A⊂U, B⊂U). Now we define the probability distributions for sets. Sets X and Y are defined as follows.
[0198] X={0,1} Y={0,1}
[0199] Let X be the set of possible values x for a random variable X, and let Y be the set of possible values y for a random variable Y.
[0200] X = x (x ∈ X) Y=y(y∈Y)
[0201] Now, let's assume that the joint probability distribution of the probability variables X and Y is as follows.
[0202] P(X=0,Y=0)={n(U)-n(A)-n(B)+n(A∩B)} / n(U) P(X=1,Y=0)={n(A)-n(A∩B)} / n(U) P(X=0,Y=1)={n(B)-n(A∩B)} / n(U) P(X=1,Y=1)=n(A∩B) / n(U)
[0203] With this, the joint distribution P(X, Y) of random variables X and Y from sets A and B can be naturally defined. The distance and similarity between set A and set B can be defined as (Indicator 1) to (Indicator 12) below, using the distance and similarity between X and Y derived from the joint probability distribution P(X, Y) of random variables X and Y.
[0204] (Indicator 1) d NSI (A,B)=d NSI (X=1,Y=1) (Indicator 2) sim NSI (A,B) = sim NSI (X=1,Y=1) (Indicator 3) d NSI (A→B=d NSI (X=1→Y=1) (Indicator 4) d NSI (B→A=d NSI (Y=1→X=1) (Indicator 5) sim NSI (A→B) = sim NSI (X=1→Y=1) (Indicator 6) sim NSI (B→A=sim NSI (Y=1→X=1) (Indicator 7) d NMI (A,B)=d NMI (X,Y) (Indicator 8) sim NMI (A,B) = sim NMI (X,Y) (Indicator 9) d NMI (A→B=d NMI (X→Y) (Indicator 10) d NMI (B→A=d NMI (Y→X) (Indicator 11) sim NMI (A→B) = sim NMI (X→Y) (Indicator 12) sim NMI (B→A=sim NMI (Y→X)
[0205] As mentioned above, d NSI d NMI distance, sim NSI , sim NMI represents similarity. Therefore, of the above (Indicator 1) to (Indicator 12), (Indicator 1), (Indicator 3), (Indicator 4), (Indicator 7), (Indicator 9), and (Indicator 10) are indices of distance between two sets, and (Indicator 2), (Indicator 5), (Indicator 6), (Indicator 8), (Indicator 11), and (Indicator 12) are indices of similarity between two sets. The larger the value of the distance index, the greater the difference, and the smaller the value, the smaller the difference. The larger the value of the similarity index, the smaller the difference, and the smaller the value, the greater the difference. In other words, both the distance index and the similarity index are indices of the relationship as a larger concept. For example, if the value of the distance index is small, it can be judged that the similarity is high, so the distance index can also be used as an index of the relationship. Here, we will rewrite (Indicator 1) and (Indicator 2) and show the details. (Indicator 1) can be rewritten as follows.
[0206] (Indicator 1) d NSI (A,B)=d NSI (X=1,Y=1) =((h(X=1,Y=1)-max(h(X=1,Y=1))) / min(h(X=1),h(Y=1)) =(-log(P(X=1,Y=1))-max(-log(P(X=1)),-log(P(Y=1))) / min(-log(P(X=1)),-log(P( Y=1)))=(-log(n(A∩B))+log(min(n(A),n(B)))) / (-log(max(n(A),n(B)))+log(n(U)))
[0207] Furthermore, (Indicator 2) can be rewritten as follows:
[0208] (Indicator 2) SIM NSI (A,B) = sim NSI (X=1,Y=1) =((h(X=1+h(Y=1)-h(X=1,Y=1)) / min(h(X=1),h(Y=1)) =(-log(P(X=1)-log(P(Y=1))+log(P(X=1,Y=1))) / min(-log(P(X=1)),-log(P(Y=1))) =(-log(n(A))-log(n(B))+log(n(A∩B))+log(n(U))) / (-log(max(n(A),n(B)))+log(n(U)))
[0209] As a result, indicators 1 through 6 can be derived as follows.
[0210] (Indicator 1) d NSI (A,B) =(-log(n(A∩B))+log(min(n(A),n(B)))) / (-log(max(n(A),n(B)))+log(n(U))) (Indicator 2) sim NSI (A,B) =(-log(n(A))-log(n(B))+log(n(A∩B))+log(n(U))) / (-log(max(n(A),n(B)))+log(n(U))) (Indicator 3) d NSI (A→B) =(-log(n(A∩B))+log(n(A))) / (-log(n(B)))+log(n(U))) (Indicator 4) d NSI (B→A) =(-log(n(A∩B))+log(n(B))) / (-log(n(A)))+log(n(U))) (Indicator 5) sim NSI (A→B) =(-log(n(A))-log(n(B))+log(n(A∩B))+log(n(U))) / (-log(n(B))+log(n(U))) (Indicator 6) sim NSI (B→A) =(-log(n(A))-log(n(B))+log(n(A∩B))+log(n(U))) / (-log(n(A))+log(n(U)))
[0211] The above indicators (1) to (6) are normalized self-information values of the joint probability distributions of two variables. Of these, (1) and (2) are indicators of symmetric relationships, while (3) to (6) are indicators of asymmetric relationships. Therefore, any of (3) to (6), which are asymmetric normalized self-information values of the joint probability distributions of two variables, can be used as an indicator of an asymmetric relationship between two sets.
[0212] (Indicators 7) to (Indicators 12) are complex, so they will be expressed indirectly. First, probability is defined as follows. Note that A' is the complement of A, and B' is the complement of B.
[0213] P(A) = n(A) / n(U) P(B) = n(B) / n(U) P(A')=(n(U)-n(A)) / n(U) P(B')=(n(U)-n(B)) / n(U) P(A´∩B´)=(n(U)-n(A)-n(B)+n(A∩B)) / n(U) P(A∩B´)=(n(A)-n(A∩B)) / n(U) P(A´∩B)=(n(B)-n(A∩B)) / n(U) P(A∩B)=n(A∩B) / n(U)
[0214] Furthermore, entropy is defined as follows: H(A)=-P(A´)log(P(A´))-P(A)log(P(A)) H(B)=-P(B´)log(P(B´))-P(B)log(P(B)) H(A,B)=-P(A´∩B´)log(P(A´∩B´))-P(A∩B´)log(P(A∩B´))-P(A´∩B)log(P(A´∩B))-P(A∩B)log(P(A∩B))
[0215] Indicators (7) to (12) can be defined using the average information as follows.
[0216] (Indicator 7) d NSI (A,B) =(H(A,B)-max(H(A),(H(B))) / min(H(A),(H(B)) (Indicator 8) sim NMI (A,B) =(H(A)+H(B)-H(A,B)) / min(H(A),(H(B)) (Indicator 9) d NMI (A→B) =(H(A,B)-H(A)) / H(B) (Indicator 10) d NMI (B→A) =(H(A,B)-H(B)) / H(A) (Indicator 11) sim NMI (A→B) =(H(A)+H(B)-H(A,B)) / H(B) (Indicator 12) sim NMI (B→A) =(H(A)+H(B)-H(A,B)) / H(A)
[0217] The above indicators (7) to (12) are normalized mutual information of the joint probability distributions of two variables. Of these, indicators (7) and (8) are indicators of symmetric relationships, while indicators (9) to (12) are indicators of asymmetric relationships. Therefore, any of the asymmetric normalized mutual information of the joint probability distributions of two variables, (9) to (12), can be used as an indicator of an asymmetric relationship between two sets.
[0218] Up to this point, we have defined (B1) and (B2) of the aforementioned (B1) to (B3) using sets A and B. Next, we will discuss the Pearson coefficient and the regression coefficient as one of the various other indicators in (B3).
[0219] SIM P (A,B)=(n(A∩B)n(U)-n(A)n(B)) / (n(A)(n(U)-n(A))n(B)(n(U)-n(B))) 1 / 2
[0220] The Pearson coefficient uses the number of elements n(U) in the universal set U. It is also the correlation coefficient ρ of the probability distribution of the Cartesian product of random variables X and Y. The Pearson coefficient takes real values between -1 and 1. Furthermore, the Pearson coefficient takes a value of 0 when there is no correlation (independence). There is also an asymmetric version of the Pearson coefficient. This asymmetric version is called the regression coefficient and is defined as follows:
[0221] (Indicator 13) sim RC (A→B) =(n(A∩B)n(U)-n(A)n(B)) / (n(A)(n(U)-n(A)) (Indicator 14) sim RC (B→A) =(n(A∩B)n(U)-n(A)n(B)) / (n(B)(n(U)-n(B))
[0222] The above indicators (13) and (14), which are asymmetric versions of the Pearson coefficient, are the normalized mutual information of the joint probability distributions of two variables. Therefore, the asymmetric versions of the Pearson coefficient, which are regression coefficients of the joint probability distributions of two variables, can be used as indicators of the asymmetric relationship between two sets.
[0223] As described above, we have explained indices for the asymmetric relationship between two sets. In this embodiment, any of the following can be used as an indicator of the asymmetric relationship between two sets: (Indices 3) to (Indices 6), which are the asymmetric normalized self-information of the joint probability distributions of two variables; (Indices 9) to (Indices 12), which are the asymmetric normalized mutual information of the joint probability distributions of two variables; or (Indices 13) and (Indices 14), which are the normalized mutual information of the joint probability distributions of two variables. In particular, using pairs with the orientation of each set reversed, such as (A→B) and (B→A), makes the interdependence between the two items clear.
[0224] As an index of the asymmetric relationship between two sets, in addition to the index obtained as described above, other indexes can also be used. In this embodiment, when n(U) is the number of elements in the universal set that includes the set related to all items, n(A) is the number of elements in the set related to one item, n(B) is the number of elements in the set related to the other item, and n(A∩B) is the number of elements in the intersection of these two sets, three types of indexes are used to generate them.
[0225] The first type is the asymmetric normalized self-information of a joint probability distribution of two variables, generated by the number of elements n(U), n(A), n(B), and n(A∩B) in the set, as shown in (Indicator 3) to (Indicator 6). The second type is the asymmetric normalized mutual information of a joint probability distribution of two variables, generated by the number of elements n(U), n(A), n(B), and n(A∩B) in the set, as shown in (Indicator 9) to (Indicator 12). The third type is the regression coefficient of a joint probability distribution of two variables, generated by the number of elements n(U), n(A), n(B), and n(A∩B) in the set, as shown in (Indicator 13) and (Indicator 14).
[0226] For the asymmetric normalized self-information of the first type of joint probability distribution, (index 5) (index 6) (or (index 3) (index 4)) can be used. For example, when using (index 5) (index 6), the relation matrix generation means 22 performs processing according to (index 5) (index 6) in step S30. Specifically, based on the words extracted from the word database 10, the number of common elements m jj Let n(A) be the number of common elements, m ii Let n(B) be the number of common elements, m ji (=m ij Let n(A∩B) be the total number of words set in the word database 10, and perform the processing according to (Indicator 5) and (Indicator 6) above. Then, the sim calculated according to (Indicator 5) and (Indicator 6) NSI (A→B), sim NSI (B→A) Relationship index value s ji , s ij This is how we obtain the relation matrix S = [s ji The following equation is calculated: ](j=1,2,3,···,n; i=1,2,3,···,n).
[0227] For the second type of joint probability distribution, the asymmetric normalized mutual information can be (index 11) (index 12) (or (index 9) (index 10)). For example, when using (index 11) (index 12), the relation matrix generation means 22 performs processing according to the above (index 11) (index 12) in step S30. Specifically, based on the words extracted from the word database 10, the number of common elements m jj Let n(A) be the number of common elements, m ii Let n(B) be the number of common elements, m ji (=m ijLet n(A∩B) be the number of words set in the word database 10, and perform the processing according to the above (indicator 11) (indicator 12). However, since (indicator 11) (indicator 12) are actually formulas using the average information quantities H(A), H(B), and H(A,B), the formulas using the number of elements n(U), n(A), n(B), and n(A∩B) are used, using the probability definition formula and the entropy definition formula on the previous page. That is, the relation matrix generation means 22 performs the processing according to the formula for calculating the asymmetric normalized mutual information of the joint probability distribution using the number of elements n(U), n(A), n(B), and n(A∩B), and the calculated sim NMI (A→B), sim NMI (B→A) Relationship index value s ji , s ij This is how we obtain the similarity matrix S = [s ji The following equation is calculated: ](j=1,2,3,···,n; i=1,2,3,···,n).
[0228] For the regression coefficients of the third type of joint probability distribution, (indicator 13) and (indicator 14) can be used. In this case, the relation matrix generation means 22 performs processing according to (indicator 13) and (indicator 14) in step S30. Specifically, based on the words extracted from the word database 10, the number of common elements m jj Let n(A) be the number of common elements, m ii Let n(B) be the number of common elements, m ji (=m ij Let n(A∩B) be the number of words set in the word database 10, and then perform the processing according to (Indicator 13) and (Indicator 14) above. Then, the sim calculated according to (Indicator 13) and (Indicator 14) RC (A→B), sim RC (B→A) Relationship index value s ji , s ij This is how we obtain the relation matrix S = [s ji The following is generated: ](j=1,2,3,···,n;i=1,2,3,···,n).
[0229] In step S30, the relationship matrix generation means 22 generates the relationship matrix S. Then, in step S40, the adjustment matrix generation means 23 generates the adjustment matrix W. In step S50, the coordinate value calculation means 24 calculates coordinate values based on the adjustment matrix W. Then, in step S60, the item placement means 27 places the items corresponding to the calculated coordinate values. In practice, nodes including the names of items are assigned to predetermined coordinate values to generate data for output means such as a display device, and this data is output from output means 30 such as a display unit 6.
[0230] In this way, each item can be arranged in k-dimensional space. However, the embodiment is not limited to the above; other methods may be used as long as they can determine the coordinate values of each item in k-dimensional space.
[0231] <4. Steps S70, S80> Next, the item group setting means 28 sets the item group (step S70). The method for setting the item group can be any method that allows for the selection of items to be included in the item group. In other words, it is sufficient that an item group consisting of multiple items is identified by some method. For example, each item may be selected individually, or some conditions may be set and items that satisfy those conditions may be selected.
[0232] As a first example, if the category of interest is novels and the items are individual novels, then a collection of novels purchased by a person can be considered an item group. For example, if you have a document database like the one shown in Figure 3, you can set up an item group by storing the document ID of the accessed document (novel) in association with the user. In this case, it becomes possible to select novels with similar tastes from among those that are excluded from the item group, while taking into account all the novels belonging to this item group.
[0233] As a second example, if the items belonging to a category of interest are all the works held by a certain museum, then images of about 20 works from that collection can be displayed, and the user can select about 3 works they like from among them, and this set of preferred works can be made into an item group. For example, if there is an image database that manages works by image ID, an item group can be set up by remembering the image ID of the accessed work and associating it with the user. In this case, while taking into account all the works in the collection that belong to this item group, it becomes possible to select works with a similar taste from all the works in the collection, excluding those belonging to the item group.
[0234] In both the first and second examples, it becomes possible to set up an item group by storing the document ID or image ID that identifies the item in association with the user ID in the storage device 3 or the like within the information processing device 100.
[0235] Next, the similar item selection means 29 selects items similar to the group of items (step S80). In this embodiment, we use the complementary space of the entire k-dimensional space. First, let's explain complementary spaces. Given a vector space V over a field K and its subspace U, a subspace W of V is said to be the complementary space of U in V if both of the following two conditions (Condition 1) and (Condition 2) are satisfied. In (Condition 1), {0} is the zero vector space.
[0236] (Condition 1) U∩V={0} (Condition 2) U + W = V
[0237] Let's consider a two-dimensional space where k=2. Figures 8-10 show the relationship between the vector space V, the subspace U, and the complementary space W when the overall vector space V is two-dimensional. When the overall vector space V is two-dimensional, the following three examples can be given.
[0238] (Example 2-1) As shown in Figure 8, the X-axis is a subspace of the vector space V. Therefore, if we let the X-axis be a subspace U, then the complementary space W of subspace U is the Y-axis. In Figure 8, all three points lie on the X-axis, which is subspace U.
[0239] (Example 2-2) As shown in Figure 9, the Y-axis is a subspace of the vector space V. Therefore, if we let the Y-axis be a subspace U, then the complementary space W of subspace U is the X-axis. In Figure 9, all three points lie on the Y-axis, which is subspace U.
[0240] (Examples 2-3) As shown in Figure 10, a line passing through the origin is a subspace of space V. Therefore, if we call this line passing through the origin a subspace U, then the complementary space W of subspace U is a line passing through the origin and perpendicular to subspace U. In Figure 10, all three points lie on the line that is subspace U.
[0241] Let's consider a 3-dimensional space with k=3. Figures 11-14 show the relationships between the vector space V, the subspace U, and the complementary space W when the overall vector space V is 3-dimensional. When the overall vector space V is 3-dimensional, the following four examples can be given.
[0242] (Example 3-1) As shown in Figure 11, the XY plane is a subspace of the vector space V. Therefore, if we let the XY plane be a subspace U, then the complementary space W of subspace U is the Z axis. In Figure 11, all three points lie on the XY plane, which is subspace U.
[0243] (Example 3-2) As shown in Figure 12, the XZ plane is a subspace of the vector space V. Therefore, if we let the XZ plane be a subspace U, then the complementary space W of subspace U is the Y-axis. In Figure 12, all three points lie on the XZ plane, which is subspace U.
[0244] (Example 3-3) As shown in Figure 13, the YZ plane is a subspace of the vector space V. Therefore, if we let the YZ plane be a subspace U, then the complementary space W of subspace U is the X-axis. In Figure 13, all three points lie on the YZ plane, which is subspace U.
[0245] (Example 3-4) As shown in Figure 14, a plane passing through the origin is a subspace of the vector space V. Therefore, if we call this plane passing through the origin a subspace U, then the complementary space W of subspace U is a straight line passing through the origin and perpendicular to subspace U. In Figure 14, all three points lie on a plane passing through the origin, which is subspace U.
[0246] Two embodiments using complementary spaces are described below. <First Embodiment> The first embodiment is a method in which, in a multidimensional Euclidean space where all items are arranged, all items are projected onto the complementary space of the subspace spanned by the group of items, the Euclidean distance from the group of items reduced to a single point on that complementary space is calculated, and items are selected in order from those with the smallest Euclidean distance. This Euclidean distance is a type of index value of the relationship between the group of items. Relationship is information that indicates the relationship between two things and can be expressed by various indices. For example, similarity, which is the degree to which the two are similar, and distance, which is the degree to which the two are far apart, can be used. Similarity increases as the two are closer, and Euclidean distance increases as the two are farther apart. For this reason, similarity and Euclidean distance can be converted to each other by a predetermined conversion formula.
[0247] For example, let the number of all items in the k-dimensional space be n, and let the number of items belonging to the item group be m. In this case, naturally, m < n. Also, assume that the multi-dimensional vector space for arranging all items is a k-dimensional space. At this time, m < k < n must hold. In the k-dimensional space for arranging all items, the subspace spanned by the m-item group is (m - 1)-dimensional. For example, when the item group consists of three items, it suffices if the subspace containing these three points is a plane (= 2-dimensional).
[0248] The complementary space of the (m - 1)-dimensional subspace spanned by the item group is (n - m + 1)-dimensional. When all items are projected onto the complementary space, the points to which all items belonging to the item group move are reduced to the same point. Therefore, considering this on the complementary space, the distance between the item group and any one item ultimately becomes the distance between the single point to which the item group is reduced and the single point onto which the one item is projected. As that distance, the Euclidean distance can be used.
[0249] The specific processing by the similar item selection means 29 is as follows. First, the similar item selection means 29 performs principal component analysis on the item group in the k-dimensional space for arranging all items. Specifically, for the item group, a covariance matrix is obtained, and the covariance matrix is subjected to eigenvalue decomposition. The eigenvalues of the covariance matrix all take values of 0 or more, and the eigenvectors are orthogonal to each other.
[0250] Assuming that there is no dimensional degradation, the subspace spanned by the m items is (m - 1)-dimensional. Therefore, when the eigenvalues of the covariance matrix are arranged in descending order, there are (m - 1) positive values, and the rest are all 0.
[0251] The subspace spanned by the eigenvectors corresponding to the (m - 1) eigenvalues taking positive values is the subspace spanned by the items. The subspace spanned by the eigenvectors corresponding to the remaining (k - m + 1) zero eigenvalues is the complementary space of the above subspace.
[0252] We find (k-m+1) unit eigenvectors corresponding to eigenvalue 0, and take the inner product of each unit eigenvector with the configuration coordinate values of all items. This gives us (k-m+1) dimensional coordinate values. These are the coordinate values projected onto the complementary space. In the following equation (Equation 31), λ1, λ2, ..., λ_k are the eigenvalues of the covariance matrix arranged in descending order. v1, v2, ..., v_k are the eigenvectors corresponding to the eigenvalues λ1, λ2, ..., λ_k, respectively. Of these, v1, v2, ..., v_m-1 are the basis vectors of the subspace spanned by the items included in the item group, and v_m, v_m+1, ..., v_k are the basis vectors of the complementary space of that subspace.
[0253] (Equation 31)
number
[0254] When projected onto complementary space, all items belonging to a set of items obtain the same coordinate values; in other words, they are reduced to a single point. Therefore, the Euclidean distance between these two points can be used.
[0255] <Second Embodiment> In the second embodiment, as an index of similarity, a covariance matrix is generated from the arrangement coordinate values of the item group, and a vector is used that projects all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix. Defining distance in complementary spaces can be problematic. For example, let's say n=100 and k=10. When m=3, the dimension of the subspace spanned by the items is 2, and we can secure 8 dimensions as the complementary space. However, if m=7, the dimension of the subspace spanned by the item group becomes 6 dimensions, leaving only 4 dimensions for the complementary space. In this situation, the item group occupies too many dimensions, leaving little information in the complementary space. To avoid this situation, the upper limit of the dimensions occupied by the item group is restricted to h dimensions. In a k-dimensional vector space where all items are arranged, if the dimension of the subspace spanned by the m-1 item group is greater than h, then we find an h-dimensional subspace that forms a further subspace of the (m-1)-dimensional subspace spanned by the item group, and project all items onto its complementary space. The phrase "when the dimension of the subspace spanned by the m-1 item group is greater than h" can also be rephrased as "when the number of items m belonging to the item group is greater than (h+1)". In this case, since not all items belonging to the item group are reduced to a single point in the complementary space, it is necessary to define the distance between the point group and a single point.
[0256] The specific processing performed by the similar item selection means 29 is as follows. First, similar item selection means 29 performs principal component analysis on the item group in a k-dimensional space where all items are arranged, similar to the first embodiment. Specifically, it calculates a covariance matrix for the item group and performs eigenvalue decomposition on that covariance matrix. All eigenvalues of the covariance matrix are greater than or equal to 0, and the eigenvectors are orthogonal to each other.
[0257] For the covariance matrix obtained from the coordinate values of the set of items, the subspace spanned by the eigenvectors corresponding to the (m-1) positive eigenvalues is the subspace spanned by the items. Here, when m-1 > h, we want to find a further h-dimensional subspace of this subspace.
[0258] When the eigenvalues are arranged in descending order, the h largest eigenvalues are taken as belonging to the item group, and the subspace spanned by the eigenvectors corresponding to these h eigenvalues is a further h-dimensional subspace of the (m-1)-dimensional subspace spanned by the item group.
[0259] Consider the complementary space of this small subspace. For the remaining (kh) eigenvalues after taking the h largest eigenvalues, the subspace spanned by the corresponding eigenvectors is the complementary space of the small subspace mentioned above.
[0260] By finding (kh) unit eigenvectors corresponding to the (kh) items with smaller eigenvalues, and taking the inner product of the respective unit eigenvectors for the configuration coordinate values of all items, we can obtain (kh)-dimensional coordinate values. These are the coordinate values projected onto the complementary space of the smaller subspace. In this way, the coordinate values of the items are projected onto the complementary space. In the following equation (32), λ1, λ2, ..., λ_k are the eigenvalues of the covariance matrix arranged in descending order. v1, v2, ..., v_k are the eigenvectors corresponding to the eigenvalues λ1, λ2, ..., λ_k, respectively. Of these, v1, v2, ..., v_m-1 are the basis vectors of the h-dimensional subspace, which is a further subspace of the subspace spanned by the items included in the item group, and v_h+1, v_h+2, ..., v_k are the basis vectors of the complementary space of that subspace.
[0261] (Equation 32)
number
[0262] In the second embodiment, three methods can be used. Next, the first to third methods will be described in order. <First Method> The first method involves determining the distance from the centroid of the item group and selecting items in order from those with the smallest distance. In this case, the similar item selection means 29 calculates the centroid coordinates of the item group using the coordinates of each item in the item group. The centroid coordinates of the item group are calculated as the average value of the coordinates of each item that makes up the item group. Therefore, the similar item selection means 29 calculates the average coordinate value of the positional coordinate values of each item in the item group.
[0263] Figure 15 shows the relationship between an item group in k-dimensional space and the distance calculated by the first method. Figure 15 shows the case where k=2. In Figure 15, an item group consisting of three items A, B, and C is shown. In Figure 15, the center of the concentric circles is the centroid of the item group. In this case, the similar item selection means 29 selects items as similar items in order from those with coordinates closest to the center of the concentric circles. Specifically, the similar item selection means 29 calculates the Euclidean distance between each item other than the item group and the average coordinate value of the item group, and selects items with smaller Euclidean distances as similar items.
[0264] <Second Method> The second method involves finding the minimum distance from the items that make up the item group and selecting items in order from the one with the smallest minimum distance. In this case, the similar item selection means 29 selects similar items in order of proximity to each item in the item group. In the second method, an item is judged to be similar if it is close to any one of the items included in the item group.
[0265] Figure 16 shows the relationship between a group of items in k-dimensional space and the distances calculated by the second method. Figure 16 shows the case where k=2. In Figure 16, a group of items consisting of three items A, B, and C is shown. In Figure 16, each item is the center of a concentric circle. In this case, the similar item selection means 29 selects items as similar items in order from those with coordinates closest to the center of any of the concentric circles. Specifically, the similar item selection means 29 calculates the minimum Euclidean distance between each item outside the item group and each item in the item group, and selects items with smaller minimum Euclidean distances as similar items. Items with smaller minimum Euclidean distances to any of the items are selected as similar items. It is not necessary to specifically determine which item an item is closest to.
[0266] <Third Method> Using either the first method or the second method described above, it is possible to select items similar to the item group. However, in the first method, items that are not similar to individual items may be selected. Also, in the second method, items that are similar to a particular single item may be selected as being similar to the item group.
[0267] As a method to supplement the deficiencies of the first method and the second method, a third method is used. The third method uses the power mean. Here, the power mean will be explained. Suppose there is data x1, x2, x3, ···, xn consisting of n positive real values. Let p be a positive real constant. At this time, consider a certain type of average value μp represented by the following (Formula 33) using p.
[0268] (Formula 33)
Number
[0269] The average value defined by (Formula 33) is called the "power mean". Alternatively, it is also called the "powered mean" or the "exponential mean". Depending on the value of p, it becomes a familiar average value.
[0270] When p = 1, it becomes the normal average value (arithmetic mean). When p = 2, it becomes the root-mean-square (RMS). p = 2 corresponds to the first method. When p → +∞, it becomes the maximum value max(x1, x2, x3, ···, xn). When p → -∞, it becomes the minimum value min(x1, x2, x3, ···, xn). p → -∞ corresponds to the second method. As the third method, the power mean μp with 0 < p < 1 is used.
[0271] In this embodiment, when using the third method, regarding each item i in the item group and the target item, with the distance between them being di, the following (Formula 34) is used. (Formula 34) is obtained by replacing xi in the above (Formula 33) with di and μp with d, respectively.
[0272] (Formula 34)
number
[0273] Figure 17 shows the relationship between a group of items in k-dimensional space and the distance calculated by the third method. Figure 17 is for k=2, and shows the case where p=8 in (Equation 34). In Figure 17, a group of items consisting of three items A, B, and C is shown. In Figure 17, contour lines are formed with each item at its highest position, but the contour lines are not concentric circles. In this case, the similar item selection means 29 performs the process according to (Equation 34) above and selects similar items in order from the item with the smallest d. It is not necessary to specifically determine which item is closest to which item.
[0274] Figure 18 shows the relationship between a group of items in k-dimensional space and the distance calculated by the third method, by changing the value of p in (Equation 34). Figure 18 shows the case where k=2 and p=2 and 4 in (Equation 34). Figure 18(a) shows the case where p=2, and Figure 18(b) shows the case where p=4. In Figure 18, too, a group of items consisting of three items A, B, and C is shown. In Figure 18, as in Figure 17, contour lines are formed with each item at the highest position, but the contour lines are not concentric circles. However, the shape of the contour lines is different from that in Figure 17. In this case as well, the similar item selection means 29 performs the process according to (Equation 34) above and selects similar items in order from the item with the smallest d. It is not necessary to make any special determination as to which item is closest.
[0275] <4. Specific Examples> Next, we will explain with an example using specific data. We used 36 novels as documents. These 36 documents are registered in the document database 11, associated with document IDs that identify the documents. In step S10, the morphological analysis means 25 performs morphological analysis and extracts words. In this state, the extracted words are registered in the word database 10, associated with word IDs that identify the words and the document IDs of the documents in which they appear.
[0276] By identifying a document ID from the information stored in this word database 10, the words that appear in the document with that document ID can be identified. In step S20, the relationship matrix generation means 22 extracts two documents from the 36 documents and calculates the number of words that the two documents have in common as the number of common elements. This is calculated by referencing the word database 10 with the two document IDs and identifying the number of word IDs that are stored in common and associated with the two document IDs. By calculating for all combinations of two documents from the 36 documents, a 36 × 36 common element matrix is obtained, with the number of common elements as its components.
[0277] Then, in step S30, the relation matrix generation means 22 generates the relation matrix S, in step S40 the adjustment matrix generation means 23 generates the adjustment matrix W, and in step S50, the coordinate value calculation means 24 calculates coordinate values based on the adjustment matrix W.
[0278] In step S60, the item placement means 27 generates a spatial perspective view, which projects the three-dimensional space onto a two-dimensional plane, and three two-dimensional projection views, which project all combinations of the three dimensions into a scatter plot. In the spatial perspective view, each item is displayed as a small circle to save space, and in the two-dimensional projection views, the author's name and work title are displayed for each item. Figure 19 shows what such a scatter plot looks like.
[0279] By creating a scatter plot like the one shown in Figure 19, the spatial relationships between the artworks become immediately clear. It becomes possible to see which artworks are close to which others, and to grasp the overall structure showing the relationships between all the artworks.
[0280] As shown in Figure 19, by extracting the relationships between items based on the relevant data for each item belonging to the category of interest and displaying a scatter plot with each item as a node, it is possible to get an overview of the overall distribution of items belonging to the category of interest and to understand the structure of that category. Nodes on the graph that correspond to similar items are placed close to each other.
[0281] By focusing on a single item, we can identify conceptually similar items by extracting one or more other items located nearby. Furthermore, by viewing the overall picture, we can distinguish between central and peripheral groups of items within a category. Additionally, if the overall structure consists of multiple clusters, this method offers the advantage of visually understanding the cluster composition.
[0282] In step S70, the selection of items can be carried out using various methods. For example, items constituting the item group can be directly selected using a keyboard, mouse, touch panel, etc., via the instruction input I / F4. Alternatively, conditions can be entered via the instruction input I / F4, and items that satisfy those conditions can be set as an item group.
[0283] Furthermore, the information processing device 100 may be equipped with a communication function, allowing items to be selected based on instructions from an external device via a network such as the Internet. For example, the information of items recorded in the information processing device 100 may be made public on a server on the Internet. The access history of each item may then be stored for each user. In this case, the item group setting means 28 sets the accessed items as the above-mentioned item group. For example, if there is a document database as shown in Figure 3, the item group can be set by storing the document ID of the accessed document in association with the user. The similar item selection means 29 then selects items that the user has not yet accessed as similar items. In this way, the information processing device 100 according to this embodiment can be used to select items that are unknown to the user but that align with their past selection preferences. In other words, an item group is created when a specific user selects multiple items according to their preferences from all items belonging to a category of interest. Alternatively, item groups may be created by automatically acquiring data from product browsing and purchase history.
[0284] Preferred embodiments of the present disclosure have been described above, but the present disclosure is not limited to the above embodiments, and various modifications are possible. For example, in the above embodiments, documents written in natural language are used as each item belonging to the category of interest, and these documents are used as input data. Morphological analysis is performed on each document item to extract morphemes, and the set of extracted morphemes is used as the set associated with each item. However, sets associated with each item may be prepared in advance. For example, in the above embodiments, it is also possible to prepare only a word database 10 in advance without using the document database 11, and to calculate coordinate values that are useful for creating scatter plots, etc., using only the information recorded in the word database 10.
[0285] Furthermore, in the above embodiment, the category of interest was a group of documents, each item was a document written in natural language, and the input data was a set of words contained in the documents. However, it can be used for various other purposes. For example, a group of articles about multiple companies could be collected, these group of articles could be used as the category of interest, the monthly articles of each company could be used as items, and the set of words contained in each article could be used as input data to display the relationships between articles of different companies. Alternatively, for example, a description of multiple products could be collected, these group of descriptions could be used as the category of interest, each product could be used as an item, and the set of words contained in each product's description could be used as input data to display the relationships between different products.
[0286] Alternatively, images obtained by photographing paintings or other works of art may be used as each item belonging to the category of interest. These images may be used as input data, and by performing image analysis on each item, words that represent the content of the image may be identified. The set of identified words may then be considered as the set associated with each item. In this case, the information processing device may be configured to include an image database instead of the document database 11, and an image analysis means instead of the morphological analysis means 25.
[0287] The image analysis means, like the morphological analysis means 25, is included in the arithmetic processing unit 20 and is implemented by the CPU 1 executing a program. The image analysis means performs image analysis on each image and identifies words to tag as representing the content of the image. Various known methods can be used as the image analysis method executed by the image analysis means. For example, the image tagging software "Clarifai" from Clarifai, Inc. in the United States can be used as image analysis software. The identified words are registered in a word database in association with the image ID of the image. This word database stores the word and the image ID of the image to which the word is attached as a tag representing the content, in association with the word ID, which is word identification information that identifies the word. This word database is handled in the same way as the word database 10 in the above embodiment. The information processing device then calculates the coordinate values corresponding to each item (image) based on the information registered in the word database and creates a scatter plot. [Explanation of Symbols]
[0288] 1...CPU(Central Processing Unit) 2...RAM(Random Access Memory) 3...Storage device 4. Instruction Input Interface 5. Data Input / Output Interface 6...Display section 7. GPU 8...frame memory 10.. Word Database 11. Document Database 20... Processing Unit 21...Common element number matrix generation means 22. Relationship matrix generation means 23...Adjusted matrix generation means 24. Coordinate value calculation means 25...Morphological analysis means 27. Item placement methods 28. Item group setting means 29. Method for selecting similar items 30. Output means 100... Information Processing Device
Claims
1. An information processing device that uses a set of items related to each item belonging to a category of interest to arrange each item in a k-dimensional space, and selects similar items that are similar to a group of items, which is a set of multiple items, An item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, Similar item selection means calculates an index value of the relationship between an item not included in the aforementioned item group and the aforementioned item group, and selects similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship; It has, The aforementioned similar item selection means is an information processing device that, as an index value of the relationship, projects all items onto the complementary space of the subspace spanned by the group of items in a multidimensional Euclidean space on which all items are arranged, and uses the Euclidean distance from the group of items reduced to a single point on that complementary space.
2. An information processing device that uses a set of items related to each item belonging to a category of interest to arrange each item in a k-dimensional space, and selects similar items that are similar to a group of items, which is a set of multiple items, An item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, Similar item selection means calculates an index value of the relationship between an item not included in the aforementioned item group and the aforementioned item group, and selects similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship; It has, The aforementioned similar item selection means is an information processing device that generates a covariance matrix from the arrangement coordinate values of the item group, and uses the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of corresponding eigenvectors, starting from the smallest eigenvalues of the covariance matrix, as an index value of the relationship.
3. An information processing device that arranges each item in a k-dimensional space using a set associated with each item belonging to a category of interest, selects similar items that are similar to a group of items, which is a set of multiple items, Among the items arranged in the aforementioned k-dimensional space, an item group is defined, which is a collection of multiple items. An index value is calculated for the relationship between items not included in the aforementioned item group and the aforementioned item group, and based on this index value, similar items that are similar to the aforementioned item group are selected from among the items not included in the aforementioned item group. An information processing method for selecting similar items, wherein, as an index value of the relationship, all items are projected onto the complementary space of the subspace spanned by the group of items in the multidimensional Euclidean space where all items are arranged, and the Euclidean distance from the group of items reduced to a single point on that complementary space is used.
4. An information processing device that arranges each item in a k-dimensional space using a set associated with each item belonging to a category of interest, selects similar items that are similar to a group of items, which is a set of multiple items, Among the items arranged in the aforementioned k-dimensional space, an item group is defined, which is a collection of multiple items. An index value is calculated for the relationship between items not included in the aforementioned item group and the aforementioned item group, and based on this index value, similar items that are similar to the aforementioned item group are selected from among the items not included in the aforementioned item group. An information processing method for selecting similar items, wherein a covariance matrix is generated from the coordinate values of the arrangement of the items, and the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix is used as an index value of the relationship.
5. A program that, using a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items for a set of items, which is a group of items. Computers, Item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, This system functions as a similar item selection means, calculating an index value of the relationship between items not included in the aforementioned item group and the aforementioned item group, and selecting similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship. The aforementioned similar item selection means is a program that, as an index value of the relationship, projects all items onto the complementary space of the subspace spanned by the group of items in a multidimensional Euclidean space where all items are arranged, and uses the Euclidean distance from the group of items reduced to a single point on that complementary space.
6. A program that, using a set of items related to each item belonging to a category of interest, arranges each item in a k-dimensional space, and selects similar items for a set of items, which is a group of items. Computers, Item group setting means for setting an item group which is a collection of multiple items from among the items arranged in the k-dimensional space, This system functions as a similar item selection means, calculating an index value of the relationship between items not included in the aforementioned item group and the aforementioned item group, and selecting similar items from among the items not included in the aforementioned item group that are similar to the aforementioned item group based on the index value of the relationship. The aforementioned similar item selection means is a program that generates a covariance matrix from the arrangement coordinate values of the item group, and uses the Euclidean distance between the coordinate values obtained by projecting all items onto a subspace spanned by a certain number of eigenvectors corresponding to the smallest eigenvalues of the covariance matrix, as an index value of the relationship.