Method, system, and program for associating a plurality of items
The method addresses the challenge of associating and stratifying heterogeneous data by converting quantitative data into categorical values and using association rule mining, enhancing data analysis by identifying relevant relationships and patterns.
Patent Information
- Application Number
- JP2021013264
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-01-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing methods struggle to effectively associate and stratify heterogeneous data groups, particularly when dealing with quantitative and qualitative data, leading to inefficiencies in identifying meaningful relationships and patterns.
A method and system for associating and stratifying data by extracting items from heterogeneous data groups using a recursive iterative approach, converting quantitative data into categorical values through z-score or histogram-based transformations, and identifying relationships using association rule mining with scores like lift.
Enables the identification of meaningful relationships across heterogeneous datasets, enhancing data stratification and analysis by amplifying outlier information while reducing noise, thus improving the accuracy and relevance of data associations.
Smart Images

Figure 0007704373000031 
Figure 0007704373000032 
Figure 0007704373000033
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method, a system, and a program for associating a plurality of items. Further, the present disclosure also relates to a method, a system, and a program for stratifying a plurality of data by using a method, a system, or a program for associating a plurality of items.
Background Art
[0002] Clustering is known as a technique for grouping data (for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Means for Solving the Problems
[0004] In one aspect, the present disclosure provides a method or the like for associating a plurality of items. In one embodiment, the present disclosure provides a method or the like that enables associating at least one item extracted from each of a plurality of heterogeneous data groups. In another embodiment, the present disclosure also provides a method or the like for stratifying a plurality of data according to a specified relationship.
[0005] Examples of embodiments of the present disclosure include the following. (Item 1) A method for associating a plurality of items, receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items, and extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups. Identifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups A method including this. (Item 2) The relationship is The method according to item 1, including using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part. (Item 3) Identifying the relationship includes For each of the plurality of heterogeneous data groups Calculating a score when using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part, and Based on the score, determining at least one item to be the premise part and at least one item to be the conclusion part The method according to item 2, including this. (Item 4) The extracting includes Extracting at least one item among the plurality of items of each of the plurality of heterogeneous data groups that has an outlier The method according to any one of items 1 to 3, including this. (Item 5) The extracting includes extracting at least one item using a recursive iterative approach. The method according to any one of items 1 to 5, including this. (Item 6) The plurality of data groups include quantitative data. The method according to any one of items 1 to 5, including this. (Item 7) The method according to item 6, further including converting the quantitative data into data having values within a predetermined range. (Item 8) The method according to item 7, wherein the converting includes not using data within a threshold from the average value or the mode among the quantitative data. (Item 9) The method according to item 7 or item 8, wherein the converting includes setting the value of the average value or the mode as the lower limit value within the predetermined range, and approaching the upper limit value within the predetermined range as the distance from the average value or the mode increases. (Item 10) The method according to item 9, wherein the converting further includes setting a value that is more than a threshold away from the average value or the mode as the upper limit value within the predetermined range. (Item 11) The converting includes calculating a z-score from the quantitative data, dividing the z-score by a given value, obtaining a value by setting a value greater than 1 among the divided values to 1 and a value less than -1 to -1, taking the absolute value of a negative value among the obtained values, and is the method according to item 7. (Item 12) The converting includes converting the quantitative data into a histogram, dividing each of the plurality of bins of the histogram by the value of the bin with the highest frequency among the plurality of bins, subtracting the divided value from 1 and is the method according to item 7. (Item 13) A method for stratifying a plurality of data, comprising: stratifying the data within the plurality of data groups according to a relationship specified according to the method described in any one of items 1 to 12. The method includes. (Item 14) A system for associating a plurality of items, comprising: Receiving means for receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items, the receiving means, Extracting means for extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups, Specifying means for specifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups A system comprising (Item 14A) The system according to item 14, further including the features described in any one or more of items 1 to 13. (Item 15) A program for associating a plurality of items, the program being executed in a computer system including a processor, the program Receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items, Extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups, Specifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups A program that causes the processor to perform a process including (Item 15A) The program according to item 15, further including the features described in any one or more of items 1 to 13. (Item 16) A computer-readable storage medium storing a program for associating a plurality of items, the program being executed in a computer system including a processor, the program Receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items, Based on the plurality of data in each of the plurality of heterogeneous data groups, extracting at least one item from each of the plurality of heterogeneous data groups respectively; identifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups; A computer-readable storage medium that causes the processor to perform a process including the above. (Item 16A) The computer-readable storage medium according to Item 16, further including the features described in any one or more of Items 1 to 13.
[0006] In the present disclosure, it is intended that the above one or more features can be provided in combination in addition to the explicitly stated combinations. Further embodiments and advantages of the present disclosure will be recognized by those skilled in the art if understood by reading the following detailed description as needed.
Advantages of the Invention
[0007] The present disclosure can provide a method or the like for associating a plurality of items. Further, the present disclosure can also provide a method or the like for stratifying a plurality of data according to a specified relationship. In particular, the present disclosure can identify the relationship between a plurality of heterogeneous data groups and stratify them according to the identified relationship.
Brief Description of the Drawings
[0008]
Figure 1A
Figure 1B
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9A
Figure 9B
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Mode for Carrying Out the Invention
[0009] The following describes the present disclosure. Throughout this specification, it should be understood that singular expressions include the concept of their plural forms unless otherwise specified. Accordingly, singular articles (e.g., "a", "an", "the" in English) should be understood to include the concept of their plural forms unless otherwise specified. Also, the terms used in this specification should be understood to be used in the ordinary meaning commonly used in the art unless otherwise specified. Therefore, unless otherwise defined, all technical terms and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the art to which the present disclosure pertains. In case of contradiction, this specification (including definitions) shall prevail.
[0010] (Definitions) The terms and general techniques used in the present disclosure are described.
[0011] In this specification, "data group" refers to a collection of data. Even when it includes a single piece of data, it is called a data group. Each piece of data in the data group is labeled, which may be referred to as an item. That is, a data group has one or more items. The data may be a quantitative value or a qualitative value (e.g., binary values of 0 and 1). The data may be a continuous value, a discrete value, or a mixture of continuous and discrete values. For example, the data is expressed as a vector having a plurality of components, and each of the plurality of components of the vector corresponds to each of the plurality of items. For example, when a certain component has a value X, the data means that the value of the item corresponding to that component is X. For example, when a certain component has a value of 1, the data means that it has the item corresponding to that component, and when a certain component has a value of 0, it means that it does not have the item corresponding to that component.
[0012] In one example, the data group includes continuous-value data. For example, the data group may include data as shown in Table 1. [Table 1] In another example, the data group includes discrete value data. For example, the data group may include data as shown in Table 2.
Table 2
Table 3
[0013] As used herein, two data groups are “of the same type” means that all items match between the two data groups. That is, the first data group and the second data group are of the same type means that all of one or more first items of the first data group match all of one or more second items of the second data group.
[0014] As used herein, two data groups are “of different types” means that there are items that do not match between the two data groups. That is, the first data group and the second data group are of different types means that at least one of one or more first items of the first data group does not match at least one of one or more second items of the second data group.
[0015] As used herein, an “abnormal value” refers to a value with a low occurrence probability that is far from the average value or the most frequent value of the data. For example, an “abnormal value” can be a value with an occurrence probability of less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 3%, less than about 1%, less than about 0.5%, etc.
[0016] As used herein, “about” means ±10% of the numerical value that follows.
[0017] In this specification, the "relationship between items" refers to the relationship between the premise part and the conclusion part, that is, the relationship that if A (premise part) is true, then B (conclusion part) is true. The relationship between items includes, for example, but is not limited to, that event B also occurs when event A occurs, that item B is high / low when item A is high / low, etc.
[0018] (Preferred Embodiment) The preferred embodiments of the present disclosure will be described below. It should be understood that the embodiments provided below are for better understanding of the present disclosure, and the scope of the present disclosure should not be limited to the following description. Therefore, it is obvious that those skilled in the art can make appropriate modifications within the scope of the present disclosure with reference to the descriptions in this specification. It is also understood that the following embodiments of the present disclosure can be used alone or in combination.
[0019] In one embodiment, the method of the present disclosure utilizes an algorithm capable of finding the relationship between at least one item in each of a plurality of heterogeneous data groups. This algorithm is a kind of association rule mining that explores items with a co-occurrence tendency (frequent item set), but is different from the association rule mining that utilizes one data group in that it can associate a plurality of heterogeneous data groups. In the association rule mining of the present disclosure, a frequent item set of outliers can be explored from a plurality of heterogeneous data groups.
[0020] For example, in the association rule mining that utilizes one data group, the following Apriori algorithm is utilized.
[0021] First, n binary attributes I = {i1, i2, ···, i n} are taken as a set of "items". I is called an "item set". T = {t1, t2, ···, t mLet \(\{t_i\}\) be a set of \(m\) observations, where each \(t_i\) has \(I\). If one observation has \(k\) items (\(I\) has \(k\) 1s and \((n - k)\) 0s), this set of items is called a \(k -\)length item set. Association rules are defined as follows. X → Y where \(X\) (the antecedent) and \(Y\) (the consequent)
Chem.
[0022] Next, the Apriori algorithm counts the frequencies of 1 - length item sets (item sets containing only one item), and 1 - length item sets with low frequencies that do not satisfy a given minimum support are removed. Support is a score representing the frequency of \(X\), \(Y\), or the co - occurrence of \(X\) and \(Y\). In the case of \(X\rightarrow Y\), when the data is binary, support is expressed, for example, as follows.
Math.
[0023] Next, possible \((k + 1)-\)length item sets are generated from the \(k -\)length item sets, and those containing \(k -\)length item sets with support less than the given minimum support are removed. These processes are repeated until convergence is achieved. Using this procedure, frequent item sets (item sets with support higher than the given minimum support) are detected.
[0024] Next, association rules are generated by finding the antecedent and consequent within each frequent item set. Several scores (e.g., lift) are commonly used, and lift is expressed, for example, as follows.
Math.
[0025] For example, sales data analysis is used as an example to explain.
[0026] For example, assume that each of 10 consumers {t1, t2, t3, t4, t5, t6, t7, t8, t9, t 10} purchases at least one of four items (i1: diapers, i2: beer, i3: milk, i4: detergent). By analyzing the data at this time using association rule mining, the relationships that may exist among the four items (i1: diapers, i2: beer, i3: milk, i4: detergent) can be estimated. For example, when the first consumer t1 purchases diapers and beer, It1 = {1, 1, 0, 0} and when the second consumer t2 purchases diapers, beer, and milk, It2 = {1, 1, 1, 0} and when the third consumer t3 purchases milk and detergent, It3 = {0, 0, 1, 1} and when the fourth consumer t4 purchases diapers, beer, and detergent, It4 = {1, 1, 0, 1} and when the fifth consumer t5 purchases diapers, beer, milk, and detergent, It5 = {1, 1, 1, 1} and when the sixth consumer t6 purchases diapers, It6 = {1, 0, 0, 0} and when the seventh consumer t7 purchases beer and milk, It7 = {0, 1, 1, 0} and when the eighth consumer t8 purchases beer and detergent, It8 = {0, 1, 0, 1} and when the ninth consumer t9 purchases diapers, beer, milk, and detergent, It9 = {1, 1, 1, 1} and when the tenth consumer t 10 purchases diapers and milk, It 10 = {1, 0, 1, 0} results.
[0027] First, calculate the support when purchasing one item. The threshold value of the support to be set is 0.35. For example, Support(diaper) is obtained by dividing the number of people who purchased diapers (7) by the total number of people (10), resulting in 0.7. Similarly, Support(beer) = 0.7, Support(milk) = 0.6, and Support(detergent) = 0.5. It can be confirmed that all supports exceed the threshold value.
[0028] Next, calculate the support when purchasing two items. The possible combinations are six: {diaper, beer}, {diaper, milk}, {diaper, detergent}, {beer, milk}, {beer, detergent}, and {milk, detergent}. For example, Support(diaper, beer) is obtained by dividing the number of people who purchased diapers and beer (5) by the total number of people (10), resulting in 0.5. Similarly, Support(diaper, milk) = 0.4, Support(diaper, detergent) = 0.3, Support(beer, milk) = 0.4, Support(beer, detergent) = 0.4, and Support(milk, detergent) = 0.3. Since Support(diaper, detergent) and Support(milk, detergent) do not meet the support threshold, the combinations {diaper, detergent} and {milk, detergent} are excluded from the candidates for frequent item sets.
[0029] Next, calculate the support when purchasing three items. The possible combinations are {diapers, beer, milk}, {diapers, beer, detergent}, {diapers, milk, detergent}, and {beer, milk, detergent}, a total of 4 combinations. However, since the combinations of {diapers, detergent} and {milk, detergent} cannot be frequent item sets, only calculate the support for the one combination of {diapers, beer, milk} that does not include {diapers, detergent} or {milk, detergent}. For example, Support(diapers, beer, milk) is obtained by dividing the number of people who purchased diapers, beer, and milk (3) by the total number of people (10), resulting in 0.3. Since Support(diapers, beer, milk) does not meet the support threshold, {diapers, beer, milk} is removed from the candidates for frequent item sets. Through the above procedure, 4 frequent item sets ({diapers, beer}, {diapers, milk}, {beer, milk}, {beer, detergent}) are obtained.
[0030] Subsequently, create association rules from the frequent item sets. For example, the possible association rules from {diapers, beer} are · Diapers → Beer (People who buy diapers tend to buy beer together), · Beer → Diapers (People who buy beer tend to buy diapers together)
[0031] There are 2 such cases. Calculate a score (e.g., lift) to evaluate the validity of these association rules. For example, since support(diapers → beer) = 0.5, support(diapers) = 0.7, and support(beer) = 0.7, lift(diapers → beer) is approximately 1.02 according to [Equation 2]. The higher this value, the stronger the tendency to be bought together.
[0032] In this way, for example, if it is found that "diapers" and "beer" tend to be bought together, and a premise - conclusion relationship can be found between "diapers" and "beer", the relationship "People who buy diapers tend to buy beer together" can be obtained.
[0033] Thus, in the association rule mining that uses one data group, the association rules are generated within each frequent item set in one data group. The method described later in the present disclosure is different from the method of association rule mining that uses the one data group described above, and the association rules are generated across a plurality of different types of data groups. For example, the association rule is generated such that the premise part of the association rule is derived from one data set and the conclusion part of the association rule is derived from another data set. As a result, the generated association rules represent item sets derived from different data sets that are related to each other.
[0034] The method of the present disclosure uses, for example, the following algorithm to find the relationship between at least one item in each of a plurality of different types of data groups.
[0035] First, I1 = {i 1,1 , i 1,2 , …, i 1,p} of p attributes and I2 = {i 2,1 , i 2,2 , …, i 2,q} of q attributes are taken as a set of "items", and I1 and I2 are called "item sets". T1 = {t 1,1 , t 1,2 , …, t 1,m} and T2 = {t 2,1 , t 2,2 , …, t 2,m} are taken as a set of m observations, and each t1, t2 has I1, I2, respectively. Assume that T1 and T2 have the same number of observations and t 1,a and t 2,a (a ∈ {1, 2, …, m}) are associated with each other (for example, when associating a patient's medical record with a gene expression profile, t 1,a : medical record of patient ID a, t 2,a: It can be the gene expression profile of patient ID a). When T1 and / or T2 include quantitative attributes, preprocessing using the membership function described later can be performed to obtain values that can be handled by this algorithm.
[0036] Next, in this algorithm, frequent item sets are separately detected in T1 and T2 using a predetermined minimum support. When the data is a continuous value, support is expressed, for example, as follows.
Number
[0037] Next, the antecedent part is selected from the frequent item sets detected in T1, and the consequent part is selected from the frequent item sets detected in T2, and association rules are generated. Further, the antecedent part is selected from the frequent item sets detected in T2, and the consequent part is selected from the frequent item sets detected in T1, and association rules are generated. In order to limit the number of association rules to be output, several scores can be used. The score includes, for example, lift. Lift is expressed, for example, as follows.
Number
Number
[0038] For example, sales data analysis will be used as an example to explain.
[0039] For example, if each of 10 consumers {t1, t2, t3, t4, t5, t6, t7, t8, t9, t 10} satisfies at least one of six items (j1: 20 years old or younger, j2: 30s or 40s, j3: 50s or 60s, j4: 70 years old or older, j5: male, j6: female). Further, assume that at least one of four items (i1: diapers, i2: beer, i3: milk, i4: detergent) is purchased. By analyzing the data at this time using the method of the present disclosure, the relationship that may exist between the six items (j1: 20 years old or younger, j2: 30s or 40s, j3: 50s or 60s, j4: 70 years old or older, j5: male, j6: female) and the four items (i1: diapers, i2: beer, i3: milk, i4: detergent) can be estimated.
[0040] For example, if the first consumer t1 is a male in his 30s or 40s, Jt1 = {0, 1, 0, 0, 1, 0} and if the second consumer t2 is a female in her 20s or younger, Jt2 = {1, 0, 0, 0, 0, 1} and if the third consumer t3 is a female in her 50s or 60s, Jt3 = {0, 0, 1, 0, 0, 1} and if the fourth consumer t4 is a male in his 30s or 40s, Jt4 = {0, 1, 0, 0, 1, 0} and if the fifth consumer t5 is a female in her 30s or 40s, Jt5 = {0, 1, 0, 0, 0, 1} and if the sixth consumer t6 is a female in her 20s or younger, Jt6 = {1, 0, 0, 0, 0, 1} and if the seventh consumer t7 is a male 70 years old or older, Jt7 = {0, 0, 0, 1, 1, 0} and if the eighth consumer t8 is a male in his 20s or younger, Jt8 = {1, 0, 0, 0, 1, 0} and if the ninth consumer t9 is a male in his 30s or 40s, Jt9 = {0, 1, 0, 0, 1, 0} and for the 10th consumer t 10 when is a woman under 20 years old, Jt 10 = {1, 0, 0, 0, 0, 1} will be the case.
[0041] Also, when the 1st consumer t1 purchases diapers and beer, It1 = {1, 1, 0, 0} will be the case, and when the 2nd consumer t2 purchases diapers, beer, and milk, It2 = {1, 1, 1, 0} will be the case, and when the 3rd consumer t3 purchases milk and detergent, It3 = {0, 0, 1, 1} will be the case, and when the 4th consumer t4 purchases diapers, beer, and detergent, It4 = {1, 1, 0, 1} will be the case, and when the 5th consumer t5 purchases diapers, beer, milk, and detergent, It5 = {1, 1, 1, 1} will be the case, and when the 6th consumer t6 purchases diapers, It6 = {1, 0, 0, 0} will be the case, and when the 7th consumer t7 purchases beer and milk, It7 = {0, 1, 1, 0} will be the case, and when the 8th consumer t8 purchases beer and detergent, It8 = {0, 1, 0, 1} will be the case, and when the 9th consumer t9 purchases diapers, beer, milk, and detergent, It9 = {1, 1, 1, 1} will be the case, and for the 10th consumer t 10 when purchases diapers and milk, It 10 = {1, 0, 1, 0} will be the case.
[0042] First, extract the frequent item sets of data J. First, calculate the support for the characteristics of consumers. The support threshold to be set is 0.25. For example, Support(under 20 years old) is calculated by dividing the number of people under 20 years old (4) by the total number of people (10), resulting in 0.4. Similarly, Support(30 - 40 years old) = 0.4, Support(50 - 60 years old) = 0.1, Support(70 years old and above) = 0.1, Support(male) = 0.6, Support(female) = 0.4. Since Support(50 - 60 years old) and Support(70 years old and above) do not meet the support threshold, {50 - 60 years old} and {70 years old and above} are excluded from the candidates for the frequent item sets.
[0043] Next, calculate the support when two characteristics apply. The possible combinations are 6 cases: {under 20 years old, male}, {under 20 years old, female}, {under 20 years old, 30 - 40 years old}, {30 - 40 years old, male}, {30 - 40 years old, female}, {male, female}. For example, Support(under 20 years old, male) is calculated by dividing the number of males under 20 years old (1) by the total number of people (10), resulting in 0.1. Similarly, Support(under 20 years old, female) = 0.3, Support(under 20 years old, 30 - 40 years old) = 0, Support(30 - 40 years old, male) = 0.3, Support(30 - 40 years old, female) = 0.1, Support(male, female) = 0. Since Support(under 20 years old, male), Support(under 20 years old, 30 - 40 years old), Support(30 - 40 years old, female), Support(male, female) do not meet the support threshold, the combinations of {under 20 years old, male}, {under 20 years old, 30 - 40 years old}, {30 - 40 years old, female}, {male, female} are excluded from the candidates for the frequent item sets.
[0044] Next, calculate the support when three characteristics apply. The possible combinations are 4 cases: {under 20 years old, 30 - 40 years old, female}, {under 20 years old, 30 - 40 years old, male}, {under 20 years old, female, male}, {30 - 40 years old, male, female}. However, since there are no consumers who meet these combinations, they are excluded from the candidates for the frequent item sets.
[0045] By the above procedure, two frequent item sets ({Under 20s, female}, {30s - 40s, male}) are obtained.
[0046] Next, extract the frequent item sets of Data I. This is the same procedure and result as the above - mentioned example of sales data analysis, and four frequent item sets ({Diapers, Beer}, {Diapers, Milk}, {Beer, Milk}, {Beer, Detergent}) are obtained.
[0047] Subsequently, create association rules from the frequent item sets. The possible association rules when creating association rules with the frequent item set obtained from Data J (consumer characteristics) in the premise part and Data I (consumer - purchased goods) in the conclusion part are · {Under 20s, female} → {Diapers, Beer} (Females under 20s tend to buy diapers and beer together), · {Under 20s, female} → {Diapers, Milk} (Females under 20s tend to buy diapers and milk together), · {Under 20s, female} → {Beer, Milk} (Females under 20s tend to buy beer and milk together), · {Under 20s, female} → {Beer, Detergent} (Females under 20s tend to buy beer and detergent together), · {30s - 40s, male} → {Diapers, Beer} (Males in their 30s - 40s tend to buy diapers and beer together), · {30s - 40s, male} → {Diapers, Milk} (Males in their 30s - 40s tend to buy diapers and milk together), · {30s - 40s, male} → {Beer, Milk} (Males in their 30s - 40s tend to buy beer and milk together), · {30s - 40s, male} → {Beer, Detergent} (Males in their 30s - 40s tend to buy beer and detergent together) There are eight such cases.
[0048] Calculate a score (e.g., lift) to evaluate the validity of these association rules. For example, lift({30 - 40 years old, male} → {diapers, beer}) is 2 according to [Equation 2] because support({30 - 40 years old, male} → {diapers, beer}) = 0.3 (since t1, t4, t9 satisfy "30 - 40 years old", "male", "diapers", "beer"), support({30 - 40 years old, male}) = 0.3, and support({diapers, beer}) = 0.5. The higher this value, the stronger the relationship between the characteristics of consumers and the purchased products is considered to be.
[0049] In this way, for example, if it is found that "diapers" and "beer" tend to be bought together by "males" in the "30 - 40 years old" group, if a premise - conclusion relationship can be found between "30 - 40 years old" and "male" and "diapers" and "beer", the relationship "when the customer is a 30 - 40 - year - old male, there is a tendency to buy diapers and beer together" can be obtained.
[0050] The above algorithm enables the identification of related items within mutually related heterogeneous datasets.
[0051] When using the above - mentioned algorithm, the data to be analyzed may contain continuous values. For example, as shown in Figure 1A, the data to be analyzed may contain a mixture of continuous values and discrete values. Even in such a case, by using the method described later to convert the continuous values into values within a predetermined range, it is possible to apply the above - mentioned algorithm to the data to be analyzed.
[0052] In this example, in order to convert the continuous values into values within a predetermined range, the continuous values are expressed in fuzzy logic. Thereby, the continuous values can be converted into values within the predetermined range [0, 1].
[0053] For example, as shown in FIG. 1B, a continuous value such as temperature can be converted into a value within a predetermined range [0, 1] by being expressed in fuzzy logic. Here, a temperature of 13° C. is expressed as having "0.6 items of 'cool' and 0.4 items of 'cold'".
[0054] In this example, when converting a continuous value into a value within a predetermined range using fuzzy logic, a conversion formula can be used such that the value is converted into a higher value within the predetermined range as it deviates from the average value or the mode. This is important for the association rule mining of the present disclosure that can search for a set of frequently occurring items of outliers from a plurality of different data groups. Since an outlier is a value with a low occurrence probability that deviates from the average value or the mode, this conversion can amplify the information of the outlier without losing the information of the outlier included in the continuous value. A value close to the average value or the mode will be treated as information such as not having a value far from the average value or the mode, for example.
[0055] For example, in addition to or instead of converting to a higher value within the predetermined range as it deviates from the average value or the mode, data whose difference from the average value or the mode is within a threshold can be excluded. Thereby, unnecessary data for analysis that is not an outlier can be omitted, and the information of the outlier can be relatively amplified. The threshold can be any value. The threshold can be set according to the desired accuracy.
[0056] There are several methods for converting a continuous value into a value within a predetermined range using fuzzy logic.
[0057] For example, the Min-Max scaling method, the sigmoid function, and rank-based conversion can be mentioned.
[0058] For example, the Min-Max scaling method for a continuous value v is as follows.
Equation
[0059] For example, the formula for the sigmoid function for the continuous value v is as follows.
Equation
[0060] For example, the formula for the rank-based transformation for the continuous value v is as follows.
Equation
[0061] Techniques such as the Min-Max scaling method, the sigmoid function, and the rank-based transformation are not transformation formulas that convert to higher values within a predetermined range as they deviate from the average value or the mode. Rather, they tend to reduce the difference between the highest value and the lowest value and lack information on outliers. The inventors of the present disclosure have found that techniques such as the Min-Max scaling method, the sigmoid function, and the rank-based transformation are not preferable for the algorithm of the present disclosure. Then, the inventors of the present disclosure unexpectedly discovered some techniques that are preferable for the algorithm of the present disclosure. Since those techniques can be converted to higher values within a predetermined range as they deviate from the average value or the mode, they were suitable for the algorithm of the present disclosure.
[0062] In one example, the technique for converting a continuous value to a value within a predetermined range is the z-score-based transformation.
[0063] Figures 2A to 2B are diagrams showing the concept of the z-score-based transformation.
[0064] First, as shown on the left side of FIG. 2A, the quantitative value for a certain item is converted into a histogram having a plurality of bins. Here, the number of bins in the histogram can be any number of 2 or more. For example, the number of bins may be set according to the data, or may be fixed for a plurality of data. The number of bins may be specified by the user.
[0065] In the example shown in FIG. 2A, the value of miR-xxx is converted into a histogram having 10 bins.
[0066] Next, as shown on the right side of FIG. 2A, the histogram is converted into a standard normal distribution. The values in the standard normal distribution are such that 95% of them are in the range of [-2, 2].
[0067] Next, a z-score is calculated from the standard normal distribution. Then, a value is calculated by dividing the z-score by a given value (for example, 3 or 2). This value becomes an index in the z-score based conversion. For example, it is determined whether the absolute value of this value is greater than 1.
[0068] For example, when the absolute value of this value is greater than 1, the absolute value of this value is regarded as 1, and when the absolute value of this value is 1 or less, this value is used as it is. The value calculated in this way becomes a value within a predetermined range [-1, 1]. The distribution of such values is, for example, the distribution shown on the right side of FIG. 2B. The membership value is the value of [0, 1] and the absolute value of the value of [-1, 0], and becomes a value within the range of [0, 1].
[0069] The above-given value can be any value determined according to, for example, what percentage of the whole is to be converted to 1 or 0. The given value is preferably 3, and more preferably 2. When the given value is 3, about 0.3% of the whole will become 1 or 0 after conversion, and when the given value is 2, about 5% of the whole will become 1 or 0 after conversion. That is, when the given value is 2, a larger range will be converted to 1 or 0 than when the given value is 3, and the "value indicating ambiguity as to whether the item applies or not" distributed in the range of [0,1] will be reduced. This makes it easier to detect association rules by the association rule mining algorithm of the present disclosure, so it is useful.
[0070] Thus, the value after z-score-based conversion is suitable for use in the algorithm of the present disclosure. This is because the value after conversion sets all values that are far from the average value (0 in the above distribution) by more than the threshold to 1 or -1, not only the maximum and minimum values, and rather amplifies the information of outliers without missing it.
[0071] For a certain item, if a value higher than the average value (0 in the above distribution) is classified into the category "high" and a value lower than the average value is classified into the category "low", the value (membership value) in the category "high" will be a value within the range of [0,1], and the value (membership value) in the category "low" will be a value within the range of [0,1] by taking the absolute value. At this time, 0 is the value corresponding to the average value. The membership value of the continuous value v in each category can be expressed, for example, as follows.
Number
[0072] In the above transformation, values around the average value or the mode do not belong to either category "low" or category "high". As a result, values around the average value or the mode are not used in the association rule mining of the present disclosure. In this way, by omitting values around the average value or the mode as data unnecessary for analysis that are not outliers, the information of outliers can be relatively amplified.
[0073] In another example, a method for converting continuous values into values within a predetermined range is a histogram-based transformation.
[0074] Figures 3A to 3B are diagrams showing the concept of histogram-based transformation.
[0075] First, as shown on the left side of Figure 3A, quantitative values for a certain item are converted into a histogram having a plurality of bins. Here, the number of bins in the histogram can be any number of 2 or more. For example, the number of bins may be set according to the data, or may be fixed for a plurality of data. The number of bins may be specified by the user.
[0076] In the example shown in Figure 3A, the value of miR-xxx is converted into a histogram having 10 bins.
[0077] Next, as shown on the right side of Figure 3A, the value of each bin in the histogram is divided by the value of the bin with the highest frequency. In the example shown in Figure 3A, since the value of the bin with the highest frequency is 55, the value of each bin is divided by 55.
[0078] Next, the value after division by 1 is subtracted. The value calculated in this way becomes a value within a predetermined range [-1, 1]. The membership value is the absolute value of the value in [0, 1] and the value in [-1, 0], and becomes a value within the range of [0, 1]. For example, as shown in FIG. 3B, if a value higher than the bin with the highest frequency (i.e., the mode) is classified into the category "high" and a value lower than the bin with the highest frequency is classified into the category "low", the value (membership value) in the category "high" becomes a value within the range of [0, 1], and the value (membership value) in the category "low" becomes a value within the range of [0, 1]. At this time, 0 is the value corresponding to the mode.
[0079] In this way, the value after the histogram-based conversion is suitable for use in the algorithm of the present disclosure. This is because the converted value can prevent the use of data whose difference from the average value is within the threshold by setting the value within the threshold from the mode (i.e., the value within the bin with the highest frequency) to zero. As a result, unnecessary data that is not an outlier can be omitted, and the information of the outlier can be relatively amplified.
[0080] For a certain item, values higher than the upper limit b of the mode H are classified into the category "high", and values lower than the lower limit b of the mode L are classified into the category "low". The membership values of the continuous value v in each category can be expressed, for example, as follows.
Equation
[0081] In this way, even if the data to be analyzed is a continuous value, the algorithm of the present disclosure can be applied, and a plurality of items of a plurality of heterogeneous data groups can be related.
[0082] Note that all the embodiments described below show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions and connection forms of the components, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. In addition, among the components in the following embodiments, the components not described in the independent claims indicating the top-level concept are described as optional components.
[0083] FIG. 4 shows an example of the configuration of a computer system 100 for associating a plurality of items.
[0084] The computer system 100 may be, for example, a computer system (i.e., a server device) installed in a service provider, or a computer system (i.e., a user device) used by a user. Hereinafter, a computer system installed in a service provider will be described as an example.
[0085] The computer system 100 includes an interface unit 110, a processor unit 120, and a memory 130. The computer system 100 is connected to a database unit 200.
[0086] The interface unit 110 exchanges information with the outside of the computer system 100. The processor unit 120 of the computer system 100 can receive information from the outside of the computer system 100 via the interface unit 110 and can transmit information to the outside of the computer system 100. The interface unit 110 can exchange information in any form. The information terminal used by the first person and the information terminal used by the second person can communicate with the computer system 100 via the interface unit 110.
[0087] The interface unit 110 includes, for example, an input unit that enables information to be input into the computer system 100. It does not matter in what manner the input unit enables information to be input into the computer system 100. For example, when the input unit is a touch panel, the user may input information by touching the touch panel. Alternatively, when the input unit is a mouse, the user may input information by operating the mouse. Alternatively, when the input unit is a keyboard, the user may input information by pressing the keys of the keyboard. Alternatively, when the input unit is a microphone, the user may input information by inputting voice into the microphone. Alternatively, when the input unit is a camera, the information captured by the camera may be input. Alternatively, when the input unit is a data reading device, information may be input by reading information from a storage medium connected to the computer system 100. Alternatively, when the input unit is a receiver, the receiver may input information by receiving information from outside the computer system 100 via a network. In this case, the type of network does not matter. For example, the receiver may receive information via the Internet or via a LAN.
[0088] The interface unit 110 includes, for example, an output unit that enables information to be output from the computer system 100. It does not matter in what manner the output unit enables information to be output from the computer system 100. For example, when the output unit is a display screen, information may be output to the display screen. Alternatively, when the output unit is a speaker, information may be output by the sound from the speaker. Alternatively, when the output unit is a data writing device, information may be output by writing information to a storage medium connected to the computer system 100. Alternatively, when the output unit is a transmitter, the transmitter may output information by transmitting the information outside the computer system 100 via a network. In this case, the type of the network does not matter. For example, the transmitter may transmit information via the Internet or via a LAN.
[0089] The processor unit 120 executes the processing of the computer system 100 and controls the operation of the entire computer system 100. The processor unit 120 reads out a program stored in the memory unit 130 and executes the program. Thereby, it is possible to make the computer system 100 function as a system that executes desired steps. The processor unit 120 may be implemented by a single processor or may be implemented by a plurality of processors.
[0090] The memory unit 130 stores programs necessary for executing the processing of the computer system 100, data necessary for executing such programs, and the like. The memory unit 130 may store a program (for example, a program that realizes the processing shown in FIG. 6 described later) for causing the processor unit 120 to perform processing for associating a plurality of items and / or a program for causing the processor unit 120 to perform processing for stratifying a plurality of data. Here, it does not matter how the program is stored in the memory unit 130. For example, the program may be pre-installed in the memory unit 130. Alternatively, the program may be installed in the memory unit 130 by being downloaded via a network. In this case, the type of the network does not matter. The memory unit 130 may be implemented by any storage means. The memory unit 130 may include, for example, a non-transitory computer-readable storage medium.
[0091] The database unit 200 stores, for example, a plurality of data groups to be analyzed. Further, the database unit 200 may store data indicating the relationships between a plurality of items specified by the computer system 100. For example, the database unit 200 may store data stratified according to the specified relationships.
[0092] In the example shown in FIG. 4, the database unit 200 is provided outside the computer system 100, but the present invention is not limited to this. It is also possible to provide the database unit 200 inside the computer system 100. At this time, the database unit 200 may be implemented by the same storage means as the storage means implementing the memory unit 130, or may be implemented by storage means different from the storage means implementing the memory unit 130. In any case, the database unit 200 is configured as a storage unit for the computer system 100. The configuration of the database unit 200 is not limited to a specific hardware configuration. For example, the database unit 200 may be composed of a single hardware component, or may be composed of a plurality of hardware components. For example, the database unit 200 may be configured as an external hard disk device of the computer system 100, or may be configured as a storage on the cloud connected via a network.
[0093] FIG. 5A shows an example of the configuration of the processor unit 120.
[0094] The processor unit 120 includes a receiving means 121, an extracting means 122, and a specifying means 123.
[0095] The receiving means 121 is configured to receive information from the interface unit 110.
[0096] The receiving means 121 is configured to receive a plurality of different data groups. The receiving means 121 may receive, for example, a plurality of different data groups input to the computer system 100 via the interface unit 110 from the interface unit 110. For example, the plurality of data groups may be input to the computer system 100 via a network from a user device operated by a user, or a plurality of data groups stored in the database unit 200 may be input to the computer system 100.
[0097] The extraction means 122 is configured to extract at least one item from each of a plurality of heterogeneous data groups based on a plurality of data in each of the plurality of heterogeneous data groups. The extraction means 122 can extract, for example, at least one item in which the data has an abnormal value.
[0098] The extraction means 122 can extract at least one item using an arbitrary method. For example, at least one item can be extracted by a suitable method according to the processing speed, accuracy, machine performance, etc.
[0099] In one example, the extraction means 122 utilizes the following Apriori algorithm.
[0100] First, n binary attributes I = {i1, i2, ···, i n} are used as a set of "items". I is called an "item set". T = {t1, t2, ···, t m} is a set of m observations, and each t has I. When one observation has k items (I has k 1s and (n - k) 0s), this item set is called a k-length item set. The association rule is defined as follows. X → Y Here, X (antecedent part) and Y (consequent part) are
Chem.
[0101] Next, the Apriori algorithm counts the frequencies of 1-length item sets (item sets containing only one item), and 1-length item sets with low frequencies that do not satisfy a predetermined minimum support are removed. Support is a score representing the frequency of X, Y, or the co-occurrence frequency of X and Y. In the case of X → Y, support is expressed, for example, as follows.
Math.
[0102] Next, possible (k + 1)-length item sets are generated from the k-length item sets, and those including k-length item sets having a support smaller than a predetermined minimum support are removed. These processes are repeated until convergence is achieved. Using this procedure, frequent item sets (item sets having a support higher than a predetermined minimum support) are detected.
[0103] In another example, instead of the Apriori algorithm or its variants described above, the extraction means 122 may utilize algorithms such as eclat, VIPER, MAFIA, TM, FP-Growth, TFP, SSR, EXTRACT, etc. For example, the extraction means 122 can extract at least one item using a recursive iterative approach. The recursive iterative approach utilizes information on sets containing k items to search for sets containing (k + 1) items.
[0104] The specifying means 123 is configured to specify the relationship between each of at least one item extracted from each of a plurality of heterogeneous data groups.
[0105] The relationship can be, for example, a premise-conclusion relationship. For example, when specifying the relationship for two heterogeneous data groups, the specifying means 123 can specify a relationship in which either at least one item extracted from the first data group or at least one item extracted from the second data group is the premise part and the other is the conclusion part. For example, when specifying the relationship for n heterogeneous data groups, the specifying means 123 can specify a relationship in which at least one of at least one item extracted from the first data group, at least one item extracted from the second data group, ··· at least one item extracted from the nth data group is the premise part and the rest is the conclusion part.
[0106] The specifying means 123 can utilize, for example, the following algorithms to specify the relationship between the items extracted from each of the two data groups.
[0107] An association rule is generated with at least one item extracted from the first data group as the premise part and at least one item extracted from the second data group as the conclusion part. Further, an association rule is generated with at least one item extracted from the second data group as the premise part and at least one item extracted from the first data group as the conclusion part. In order to limit the number of association rules to be actually output from the plurality of generated association rules, several scores can be used. The scores include, for example, lift. Lift is expressed, for example, as follows. [Number] Here, X: At least one item extracted from the first data group, Y: At least one item extracted from the second data group, [Number]
[0108] An association rule with lift greater than or equal to a predetermined threshold can be output as a rule indicating the relationship between the two data groups.
[0109] According to the computer system 100, a relationship that could not be found by conventional grouping methods such as clustering can be found. According to the found relationship, for example, the computer system 100 can stratify the data in a plurality of data groups. As a result, the data can be stratified from a perspective that could not be found conventionally. This can lead to expanding the scope of various analyses.
[0110] FIG. 5B shows an example of the configuration of a processor unit 120' which is an alternative embodiment of the processor unit 120. The processor unit 120' can be used when at least one of the plurality of data groups has quantitative data.
[0111] The processor unit 120' includes a receiving means 121, a conversion means 124, an extraction means 122, and an identification means 123. The receiving means 121, the extraction means 122, and the identification means 123 are the same as those described with reference to FIG. 5A, and the description thereof is omitted here.
[0112] The conversion means 124 is configured to convert quantitative data into data having values within a predetermined range. The values within the predetermined range can be, for example, values within the range of [0, 1]. The conversion means 124 can be configured not to use data among the quantitative data whose difference from the average value or the mode value is within a threshold. Since data close to the average value or the mode value is not useful in identifying the relationship, the accuracy of the identified relationship can be improved by excluding such data. The threshold can be set to any value. The threshold can be a variable value or a fixed value.
[0113] The conversion means 124 can convert quantitative data such that, for example, the value of the average value or the mode value of the quantitative data is set as the lower limit value within the predetermined range (for example, "0" within the range of [0, 1]), and the value approaches the upper limit value within the predetermined range (for example, "1" of [0, 1]) as it is farther from the average value or the mode value. At this time, the conversion means 124 can set a value that is more than the threshold away from the average value or the mode value as the upper limit value within the predetermined range (for example, "1" of [0, 1]).
[0114] In one example, the conversion means 124 utilizes a z-score based conversion method. The z-score based conversion method is the method described above with reference to FIGS. 2A to 2B.
[0115] The value after the z-score based conversion is suitable for the processing by the extraction means 122 and the identification means 123. This is because the converted value sets all values that are not only the maximum value and the minimum value but also more than the threshold away from the average value to 1 or -1, amplifying rather than missing the information of the outlier values.
[0116] In another example, the conversion means 124 utilizes a histogram-based conversion method. The histogram-based conversion method is the method described above with reference to FIGS. 3A to 3B.
[0117] The values after histogram-based conversion are suitable for the processing by the extraction means 122 and the identification means 123. This is because, by setting the values within the threshold from the mode value to zero, it is possible not to use the data whose difference from the average value is within the threshold, thereby omitting unnecessary data that is not an outlier and relatively amplifying the information of the outlier.
[0118] Even when at least one of the plurality of data groups has quantitative data due to the conversion by the conversion means 124, the processing by the extraction means 122 and the identification means 123 can be performed, and a plurality of items of a plurality of heterogeneous data groups can be associated.
[0119] Note that each component of the computer system 100 described above may be constituted by a single hardware component or may be constituted by a plurality of hardware components. When constituted by a plurality of hardware components, the mode of connection of each hardware component is not limited. Each hardware component may be connected wirelessly or may be connected by wire. The computer system 100 of the present invention is not limited to a specific hardware configuration. It is also within the scope of the present invention to configure the processor unit 120 with an analog circuit instead of a digital circuit. The configuration of the computer system 100 of the present invention is not limited to that described above as long as its function can be realized.
[0120] FIG. 6 shows an example of a process 600 by the computer system 100 of the present disclosure. The process 600 can be executed in the processor unit 120 or the processor unit 120' of the computer system 100. Hereinafter, it will be described as being executed by the processor unit 120.
[0121] In step S601, the receiving means 121 of the processor unit 120 receives a plurality of heterogeneous data groups. The receiving means 121 may receive, for example, a plurality of heterogeneous data groups input to the computer system 100 via the interface unit 110 from the interface unit 110. For example, the plurality of data groups may be input to the computer system 100 via a network from a user device operated by a user, or a plurality of data groups stored in the database unit 200 may be input to the computer system 100.
[0122] In step S602, the extraction means 122 of the processor unit 120 extracts at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups. The extraction means 122 can extract, for example, at least one item for which the data has an abnormal value. The extraction means 122 can extract at least one item using any method.
[0123] In step S603, the specifying means 123 of the processor unit 120 specifies the relationship between at least one item extracted from each of the plurality of heterogeneous data groups. The relationship can be, for example, a premise - conclusion relationship. For example, when specifying the relationship for two heterogeneous data groups, the specifying means 123 can specify a relationship in which at least one item extracted from the first data group or at least one item extracted from the second data group is the premise part and the other is the conclusion part. For example, when specifying the relationship for n heterogeneous data groups, the specifying means 123 can specify a relationship in which at least one of at least one item extracted from the first data group, at least one item extracted from the second data group, ··· at least one item extracted from the nth data group is the premise part and the rest is the conclusion part.
[0124] In this way, a plurality of items of a plurality of heterogeneous data groups can be associated. Such an association can be used, for example, for data stratification.
[0125] For example, when at least one of a plurality of data groups has quantitative data, process 600 is executed by processor unit 120'. In this case, before step S602, step S6021 of converting the quantitative data into data with values within a predetermined range is performed.
[0126] In step S6021, conversion means 124 of processor unit 120' converts the quantitative data into data with values within a predetermined range. The values within the predetermined range can be, for example, values within the range of [0, 1]. The conversion means 124 can be configured not to use data among the quantitative data whose difference from the average value or the mode value is within a threshold. Since data close to the average value or the mode value is not useful in identifying the relationship, the accuracy of the identified relationship can be improved by excluding such data. The threshold can be set to any value. The threshold can be a variable value or a fixed value.
[0127] The conversion means 124 can convert the quantitative data such that, for example, the value of the average value or the mode value of the quantitative data is set as the lower limit value within the predetermined range (e.g., "0" within the range of [0, 1]), and the value approaches the upper limit value within the predetermined range (e.g., "1" of [0, 1]) as it moves away from the average value or the mode value. At this time, the conversion means 124 can set a value that is more than the threshold away from the average value or the mode value as the upper limit value within the predetermined range (e.g., "1" of [0, 1]).
[0128] The conversion means 124 can convert the quantitative data by using, for example, a z-score-based conversion method or a histogram-based conversion method.
[0129] By step S6021, even when at least one of a plurality of data groups has quantitative data, process 600 enables relating a plurality of items of a plurality of heterogeneous data groups.
[0130] Note that although the above-described processing has been described as being performed in a specific order, this order is exemplary, and the processing can be performed in any logically possible order. It is also understood that at least one step of the above-described processing can be omitted, and at least one step can be added to the above-described processing.
[0131] As described above, the present disclosure has been described with reference to preferred embodiments for ease of understanding. Hereinafter, the present disclosure will be described based on examples. However, the above description and the following examples are provided for illustrative purposes only and are not provided for the purpose of limiting the present disclosure. Therefore, the scope of the present disclosure is not limited to the embodiments or examples specifically described in this specification, but is limited only by the claims.
Example
[0132] (Example using an artificial dataset and a biological dataset) In this example, a gene expression profile dataset was used as the artificial dataset, and a hepatotoxicity dataset was used as the biological dataset to demonstrate the performance of various algorithms. Generally, patterns to be detected in the dataset were artificially added, and it was determined whether these patterns could be detected by the algorithm of the present disclosure.
[0133] (Method) An example of an algorithm The purpose of ARM is to find frequent item sets (items that tend to co-occur) and find association rules (patterns of co-occurrence between items in the frequent item set). The purpose of the algorithm of the present disclosure is to enable i) finding frequent item sets in a given dataset containing continuous values, and ii) finding association rules between frequent item sets detected in different given datasets, thereby enabling association of two heterogeneous datasets and focusing on a small number of items. The workflow of this algorithm is shown in FIG. 7.
[0134] · Apriori algorithm n binary attributes I = {i1, i2, …, i n} are regarded as a set of "items", and I is called an "item set". T = {t1, t2, …, t m} is a set of m observations, and each t has I. If one observation has k items (I has k 1s and (n - k) 0s), this item set is called a k-length item set. Association rules are defined as follows, where X (called the antecedent) and Y (called the consequent)
Chem.
[0135] First, the Apriori algorithm counts the frequencies of 1-length item sets (item sets containing only one item), and 1-length item sets with low frequencies that do not meet the user-specified minimum support are removed. Support is a score representing the frequency of X, Y, or the co-occurrence of X and Y. In the case of X → Y, support is expressed as follows.
Math.
[0136] Second, possible (k + 1)-length item sets are generated from k-length item sets, and those containing k-length item sets with support less than the user-specified minimum support are removed. These processes are repeated until convergence is achieved. Using this procedure, frequent item sets (item sets with support higher than the user-specified minimum support) are detected.
[0137] Third, by searching for the antecedent and consequent within each frequent item set, association rules are generated. Some scores (e.g., lift) are commonly used and are expressed as follows.
[0138] · Fuzzy Association Rule Mining Conventional ARM approaches assume that the input data contains categorical attributes. However, the data that the inventors actually handle can be quantitative or a mixture of quantitative and qualitative data. Therefore, for quantitative attributes, by providing a threshold for quantification to be converted into categorical attributes, when a quantitative attribute is converted into a categorical value (for example, when threshold 1 and threshold 2 are given and threshold 1 < threshold 2), the quantitative attribute is assigned to one of the following categories. i) ≤ threshold 1 ii) between threshold 1 and threshold 2 iii) threshold 2 ≤ This is called crisp data, and this procedure results in information loss. To solve this problem, fuzzy logic is introduced in the Apriori algorithm.
[0139] Fuzzy logic is defined as "a class of objects having a continuum of membership grades", and quantitative attributes are converted into several categories having "membership values" in the range of 0 to 1 (the quantitative attribute at threshold 1 is converted into 0.5 category i) and 0.5 category ii)). The concepts of union, intersection, and complement, which are used to calculate some important scores in association rule mining such as support, can also be extended to fuzzy sets.
[0140] · Function for calculating membership values To convert quantitative attributes into a fuzzy category set, the main problem is how to define the membership function, where the membership function is used to calculate the membership value. Since the membership value ranges from 0 to 1, generally, min-max scaling, sigmoid transformation, and rank-based transformation are used. However, these methods reduce the difference in membership values between the most appropriate category and the least appropriate category, obtaining data that is too fuzzy for the a priori algorithm. From this preliminary observation, the inventors designed a new membership function as described below.
[0141] · Histogram-based transformation (Fig. 7b) The frequency of quantitative data is transformed into a histogram with a user-specified number of bins. The quantitative attribute is transformed into three categories, "low", "average", and "high". These membership values for a quantitative attribute v are expressed as follows, where the frequency of the bin containing the quantitative attribute v is F v and the frequency of the highest bin is F H and the lower limit of the highest bin is b L and the upper limit of the highest bin is b H is.
Number
[0142] The sum of the membership values for the categories "low", "average", or "high" is supposed to be 1. However, the information within the category "average" is not used for association rule mining. This is because it is treated as a frequently occurring item and interesting association rules that include the category "low" or the category "high" cannot be detected.
[0143] · z-score-based transformation (Fig. 7c) The frequency of quantitative data was converted to the standard normal distribution to obtain z - scores. It is expected that 95% of the data lies in the range of - 2 to 2. Quantitative attributes are converted into three categories, "low", "average", or "high". The membership values for a quantitative attribute v are expressed as follows.
Number
[0144] Similar to the histogram - based conversion, the sum of the membership values for the categories "low", "average", and "high" is expected to be 1. However, the information within the category "average" is not used for association rule mining. This is because it is treated as a frequently occurring item and cannot detect interesting association rules that include the category "low" or the category "high".
[0145] · Other membership functions used for comparison The formula for min - max scaling for a quantitative attribute v is as follows.
Number
[0146] The formula for the sigmoid function for a quantitative attribute v is as follows.
Number
[0147] The formula for rank - based conversion for a quantitative attribute v is as follows.
Number
[0148] · Association of heterogeneous datasets In the conventional ARM approach, association rules are generated "within" each frequent item set. The aim of the inventors' method is to generate association rules such that the antecedent part of an association rule is derived from one data set and the consequent part of the association rule is derived from another data set, whereby the detected association rules represent item sets derived from different data sets that are related to each other. For this purpose, the inventors have developed a novel algorithm, as described below.
[0149] Let I1 = {i 1,1 , i 1,2 , …, i 1,p} of p attributes and I2 = {i 2,1 , i 2,2 , …, i 2,q} of q attributes be sets of "items", and I1 and I2 are called "item sets". Let T1 = {t 1,1 , t 1,2 , …, t 1,m} and T2 = {t 2,1 , t 2,2 , …, t 2,m} be sets of m observations, where each t1, t2 has I1, I2 respectively. Assume that T1 and T2 have the same number of observations and t 1,a and t 2,a (a ∈ {1, 2, …, m}) are associated with each other (e.g., t 1,a : medical record of patient IDa, t 2,a : gene expression profile of patient IDa). If T1 and / or T2 contain quantitative attributes, calculation of membership values for the categories "low" and "high" for those attributes is required as preprocessing.
[0150] First, the fuzzy Apriori algorithm separately detects frequent item sets in T1 and T2 using a user-specified minimum support. Support is expressed as follows.
Equation
[0151] Second, association rules are generated such that the premise part is selected from the frequently occurring item sets detected in T1, the conclusion part is selected from the frequently occurring item sets detected in T2, and vice versa. To limit the number of rules to be output, several scores can be used. For example, lift is expressed as follows.
Number
Number
[0152] This new algorithm enables the identification of related items within heterogeneous datasets that are related to each other.
[0153] Preprocessing and postprocessing of the hepatotoxicity dataset for experiments In the original biological data (hepatotoxicity data) used by the inventors for experiments, histological observations were described as "minimal", "mild", "moderate", or "marked". For the experiments of this example, these were converted to "0", "1", "2", or "3", respectively. Bushel, the author of the reference paper (Bushel, P.R., Wolfinger, R.D., & Gibson, G. (2007). Simultaneous clustering of gene expression data with clinical chemistry and pathological evaluations reveals phenotypic prototypes. BMC Systems Biology, 1(1), 15.), reported that 50 mg / kg body weight and 150 mg / kg of acetaminophen were subtoxic, and 1500 mg / kg body weight and 2000 mg / kg of acetaminophen were highly toxic (Bushel, et al., 2007). Therefore, the column of dose levels in Data 2 was converted to a binary attribute representing their toxicity levels (50 mg / kg body weight and 150 mg / kg: 0, 1500 mg / kg body weight and 2000 mg / kg: 1). Furthermore, the column of time points in Data 2 was converted to a binary attribute representing their toxicity levels (6, 18, 48 h: 0, 24 h: 1). This is because it was reported that 24 h was the peak of toxicity and the rats were in the recovery phase 48 h after acetaminophen treatment (Bushel, et al., 2007). In Data 1 (gene expression profile), genes were indicated by Agilent probe IDs. DAVID [(Dennis, et al., 2003)] was used to convert Agilent probe IDs to Entrez gene IDs and gene names.
[0154] Experiment The AI Bridging Cloud Infrastructure (ABCI) operating at the National Institute of Advanced Industrial Science and Technology (AIST; Japan) was used for the experiment.
[0155] Implementation The method used in this example is implemented in Python 3.0 and depends on the pandas, joblib, and os modules. In the detection of frequent item sets, the code of the apriori function in the Python module mlxtend was edited and modified to handle fuzzy logic.
[0156] Results The inventors conducted experiments using two artificial datasets and one real-world biological dataset to confirm the performance of the algorithm used in this example for detecting pairs of frequent item sets for subset combination. The algorithm used in this example can be applied to any pair of datasets. For simplicity, in this study, it is assumed that the dataset consists of one gene expression profile data (Data 1) and one clinical measurement data (Data 2).
[0157] · Artificial data (small) Artificial data was generated as shown in Figure 8. Generally, the gene expression profile data has 100 rows (e.g., 100 patients) and 200 columns (e.g., 200 genes), and random values were generated according to the standard normal distribution. The clinical measurement data was generated in the same procedure. The inventors added some irregular patterns randomly generated according to normal distributions with different means and standard deviations (S.D.) within these two matrices as the frequent item sets to be detected. The inventors evaluated the performance of the algorithm by checking whether the patterns in Table 4 were successfully detected.
Table 4
[0158] Using this dataset, the inventors compared five membership functions, min-max scaling, transformation using the sigmoid function, rank-based transformation, histogram-based transformation, and z-score-based transformation. The explanations of these methods can be found in the above (method) section. Table 5 summarizes this result. Using the specific settings the inventors tested, only the histogram-based function and the z-score-based function generated a pair of frequent item sets, and all three patterns the inventors added were included in the generated rules. The other three methods (min-max scaling, sigmoid, rank-based) did not detect any of these in the patterns the inventors tested.
Table 5
[0159] · Artificial data (large) Next, the inventors experimented with a large dataset. A pair of matrices with 1000 rows and 2000 columns was generated in the same procedure as the artificial data (small), and the inventors added the three patterns to be detected as shown in Figure 7b) and Table 4. Again, the histogram-based function and the z-score-based function successfully detected all three patterns the inventors generated in the output (Table 5).
[0160] · Real data (hepatotoxicity) Finally, the inventors conducted experiments using real-world biological datasets [(Bushel, et al., 2007)]. Acetaminophen, which is known to cause hepatotoxicity at high doses (5, 150, 1500, 20000 mg / kg body weight), was administered to 64 rats, which were sacrificed after 6, 18, 24, or 48 h. Liver gene expression profiles were obtained by Agilent microarray analysis, and 3116 genes were selected as those with significantly different expression levels due to acetaminophen treatment (Figure 9A, Data 1). Additionally, 48 histopathological observations and 10 clinical measurements were obtained from these rats (Figure 9A, Data 2). As described in the Methods section, the experimental conditions (dose and time point) were added to Data 2.
[0161] With the parameter settings tested by the inventors (minimum support: 0.02 for Data 1 and Data 2, minimum items: 10 for Data 1 and Data 2, lift: 4.8), 3986 pairs of association rules were generated. Among these, the association rules for the pairs with the highest explanatory power are shown in Figure 9B. The results demonstrate that high values of alkaline phosphatase (ALP), alanine aminotransferase (ALT), aspartate aminotransferase (AST), and total bile acid (TBA) accompanied by low values of cholesterol co-occur with histopathological observations such as LLL_Centrilob_Necrosis, LLL_Hepato_Hypertrophy, LML_Centrilob_Necrosis, and LML_Sinusoid_Cogestion, and these were related to the toxic dose and time point of acetaminophen treatment. These attributes in the results were paired in the antecedent with 10 probes ('A_43_P10003_High' (gene name: Hsph1), 'A_42_P717602_High' (gene name: Mat2a), 'A_42_484423_High' (gene name: Pgs1), 'A_43_P17455_Low' (gene name: Dnah9), 'A_43_P14864_High' (gene name: Dynll1), 'A_43_P16523_High' (gene name: Nomo1), 'A_43_P19279_Low' (gene name: Lyzl4), 'A_43_P12811_High' (gene name: Srm), 'A_42_P655825_High' (gene name: Smg9), 'A_42_P804499_High' (n.d.)). This result is consistent with the fact that high values of ALP, ALT, AST, TBA and low values of cholesterol are regarded as markers of hepatotoxicity and that high doses of acetaminophen cause hepatotoxicity.
[0162] It has been reported that an overdose of acetaminophen causes congestion of sinusoidal capillaries and centriolar necrosis [(Boyd and Bereczky, 1966)], which is caused by intermediate metabolites and has been reported to be fatal [(Prescott, 1980)]. In addition, acetaminophen also causes hepatocyte hypertrophy in rats [(Kishi, et al., 2020)]. Furthermore, some of the genes detected in the premise part indicated their involvement in liver injury. Methionine adenosyltransferase (Mat) is responsible for the biosynthesis of S-adenosylmethionine (AdoMet), and there are two isoforms in mammals, Mat1a and Mat2a. Mat2a is induced in response to liver injury and accelerates cell division and hepatocyte proliferation [(Martinez-Chantar, et al., 2002)]. AdoMet functions as a precursor of antioxidative glutathione (GSH) and a polyamine and is involved in cell proliferation and apoptosis. Its effects on GHS depletion and hepatocyte necrosis caused by acetaminophen treatment have been well studied [(Martinez-Chantar, et al., 2002)]. In addition, Srm plays an important role in polyamine synthesis. Under conditions of liver injury, downregulation of Mat1a and upregulation of Mat2a are observed, which results in a decrease in AdoMet, which has a protective effect against liver injury [(Lu and Mato, 2012)]. Database searches demonstrated that Hsph1 interacts with Mat2a [(Rouillard, et al., 2016)]. Pgs1 is responsible for the biosynthesis of phosphatidylglycerol and cardiolipin, which are located in the inner mitochondrial membrane, and its reactive oxygen species (ROS)-induced oxidation is associated with mitochondrial dysfunction [(Paradies, et al., 2014)]. AdoMet has been reported to prevent mitochondrial dysfunction induced by chronic alcohol treatment [(Bailey, et al., 2006)].
[0163] Collectively, these results indicate that dysregulation of methionine metabolism and decreased AdoMet are associated with hepatocyte necrosis, mitochondrial dysregulation, and cell proliferation induced by acetaminophen treatment. An overview of these reports is shown in Figure 10.
[0164] Discussion In this example, the inventors present a novel approach for finding attributes that are mutually related in paired data, which they also refer to as the subset combination approach. Instead of combining data from multiple perspectives to maximize mutual agreement, this approach finds the attributes of interest according to their co-occurrence. The advantage of this approach is that the co-occurrence statistics are easily computable, which makes the output interpretable. In addition, this approach can associate heterogeneous data in a data-driven manner without relying on prior knowledge and can be used for various purposes such as biomarker discovery, understanding the molecular basis of events, or patient stratification, depending on the input data.
[0165] (Example 2) Further analysis is performed using the following artificial example.
[0166] As shown in FIG. 11, two matrices (transactions_m, transactions_o) corresponding to a plurality of different data groups were created as artificial data for operation confirmation. Both of these matrices are 360×500 and are composed of random numbers generated from a normal distribution with a mean of 0 and a standard deviation of 1. In these matrices, the rows match, but the columns do not. Further, for rows 1 to 30 and columns 1 to 5 of transactions_m, they were replaced with random numbers generated from a normal distribution with a mean of -2 and a standard deviation of 0.5. Also, for rows 31 to 60 and columns 2, 4, 8, 10, 12 of transactions_m, they were replaced with random numbers generated from a normal distribution with a mean of 2 and a standard deviation of 0.5. Similarly, for rows 1 to 30 and columns 60, 70, 80, 90, 100 of transaction_o, they were replaced with random numbers generated from a normal distribution with a mean of 2 and a standard deviation of 0.5. Also, for rows 31 to 60 and columns 100, 200, 300, 400 of transactions_o, they were replaced with random numbers generated from a normal distribution with a mean of -2 and a standard deviation of 0.5. From such a data group, the purpose is to extract the relationship that "when columns 1 to 5 of transactions_m are outliers (lower than the mean or the most frequent value), columns 60, 70, 80, 90, 100 of transactions_o tend to be outliers (higher than the mean or the most frequent value)" and the relationship that "when columns 2, 4, 8, 10, 12 of transactions_m are outliers (higher than the mean or the most frequent value), columns 100, 200, 300, 400 of transactions_o tend to be outliers (lower than the mean or the most frequent value)".
[0167] Next, the distribution of the numerical values of the artificial data transactions_m created by the above method was visualized as a heatmap. The result is shown in FIG. 12. The lower the value, the bluer it is, around 0 it is red, and the higher the value, the whiter it is. The intentionally generated pattern (columns 1 to 5 of transactions_m are outliers (lower than the mean or the most frequent value), and columns 2, 4, 8, 10, 12 of transactions_m are outliers (higher than the mean or the most frequent value)) can be confirmed.
[0168] Next, as an expression for converting data into values within a predetermined range, the distribution of the converted numerical values obtained when using min-max scaling was visualized as a heatmap. The results are shown in FIG. 13. Each value is calculated for the converted values of "items lower than the average value or the mode" and "items higher than the average value or the mode", so the number of columns is doubled. The values of the items "the 1st to 5th columns of transactions_m are lower than the average value or the mode" and "the 2nd, 4th, 8th, 10th, and 12th columns of transactions_m are higher than the average value or the mode" are high (displayed in white in the heatmap). Overall, it is red (around a value of 0.5), which means that the fitting degree for each item is ambiguous.
[0169] Next, as an expression for converting data into values within a predetermined range, the distribution of the converted numerical values obtained when using the sigmoid function was visualized as a heatmap. The results are shown in FIG. 14. The way of viewing and interpreting the heatmap is the same as that of FIG. 13. Each value is calculated for the converted values of "items lower than the average value or the mode" and "items higher than the average value or the mode", so the number of columns is doubled. The values of the items "the 1st to 5th columns of transactions_m are lower than the average value or the mode" and "the 2nd, 4th, 8th, 10th, and 12th columns of transactions_m are higher than the average value or the mode" are high (displayed in white in the heatmap). Overall, it is red (around a value of 0.5), which means that the fitting degree for each item is ambiguous.
[0170] Next, as an expression for converting data into values within a predetermined range, the distribution of the converted numerical values obtained when using an expression based on the order of magnitude of numerical values (rank-based conversion) was visualized using a heatmap. The result is shown in FIG. 15. The method of viewing and interpreting the heatmap is the same as that in FIG. 13. For each value, the number of columns doubles to calculate the converted values for items "lower than the average value or the mode" and "higher than the average value or the mode". The values of the items "the 1st to 5th columns of transactions_m are lower than the average value or the mode" and "the 2nd, 4th, 8th, 10th, and 12th columns of transactions_m are higher than the average value or the mode" become higher (displayed in white in the heatmap). Overall, it is red (near the value of 0.5), meaning that the fitting degree for each item is ambiguous.
[0171] Next, as an expression for converting data into values within a predetermined range, the distribution of the converted numerical values obtained when using z-score-based conversion was visualized using a heatmap. The result is shown in FIG. 16. The method of viewing the heatmap is the same as that in FIG. 13. Overall, it is black (near the value of 0), meaning that it does not tolerate the ambiguity of the fitting degree of each item compared to the methods in FIGS. 12 to 15.
[0172] Next, as an expression for converting data into values within a predetermined range, the distribution of the converted numerical values obtained when using histogram-based conversion was visualized using a heatmap. The result is shown in FIG. 17. The method of viewing and interpreting the heatmap is the same as that in FIG. 16. Overall, it is black (near the value of 0), meaning that it does not tolerate the ambiguity of the fitting degree of each item compared to the methods in FIGS. 12 to 15.
[0173] (Example 3) (Example using food purchase data and medical check-up data)
[0174] For example, food purchase data can be input as the first data group, and medical check-up data can be input as the second data group to obtain an output from the algorithm of the present disclosure.
[0175] For example, food purchase data can be input as the first data group, and physical examination data of the same person can be input as the second data group to obtain an output from the algorithm of the present disclosure.
[0176] The food purchase data includes information on the purchase frequency of what foods (e.g., green and yellow vegetables, root vegetables, beef, pork, chicken, fish, alcoholic beverages, processed foods) were purchased and how many times in a month. The physical examination data includes body measurements, blood tests, urine tests, and fecal occult blood tests.
[0177] When the algorithm of the present disclosure is applied to these data, in one example, from the food purchase data, "high alcohol beverages, high processed foods" are extracted as a frequent item set. From the physical examination data, in one example, "high triglycerides, high LDL-cholesterol, high blood pressure" are extracted as a frequent item set.
[0178] When association rules are specified from the frequent item sets extracted in this way, in one example, a rule such as "when the purchase frequency of alcohol beverages and processed foods is high, there is a tendency for triglycerides, LDL-cholesterol, and blood pressure to be high" can be specified. This is considered to indicate that diet affects health (suggesting the possibility of improving health by improving diet).
[0179] As described above, the present disclosure has been illustrated using preferred embodiments of the present disclosure, but it is understood that the scope of the present disclosure should be interpreted only by the claims. It is understood that patents, patent applications, and other documents cited herein should be incorporated by reference into this specification as if the contents themselves were specifically set forth herein.
Industrial Applicability
[0180] The present disclosure is useful as providing a method and the like for associating a plurality of items. The present disclosure is also useful as providing a method and the like for stratifying a plurality of data according to a specified relationship.
Description of Signs
[0181] 100 computer system 110 interface unit 120 processor unit 130 memory unit
Claims
1. A method for associating a plurality of items, the method being executed in a computer system comprising a processor, the method comprising: the processor receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items; the processor extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; the processor identifying a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; wherein the relationship includes using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part.
2. Identifying the relationship for each of the plurality of heterogeneous data groups, calculating a score when using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part; determining at least one item to be used as the premise part and at least one item to be used as the conclusion part based on the score; The method according to claim 1, comprising the above.
3. A method for associating a plurality of items, the method being executed in a computer system comprising a processor, the method comprising: the processor receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items; the processor extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; the processor identifying a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; wherein the extracting includes extracting at least one item having an abnormal value among the plurality of items in each of the plurality of heterogeneous data groups. The method, comprising the above. **Claim 4** A method for associating a plurality of items, the method being executed in a computer system comprising a processor, the method comprising: the processor receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items; the processor extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; the processor identifying a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; comprising; wherein the extracting comprises extracting at least one item using a recursive iterative approach. **Claim 5** The method according to any one of claims 1 to 4, wherein the plurality of data groups include quantitative data. **Claim 6** The method according to claim 5, further comprising converting the quantitative data into data having values within a predetermined range. **Claim 7** The method according to claim 6, wherein the converting comprises not using data within the threshold difference from the average value or the mode value among the quantitative data. **Claim 8** The method according to claim 6 or claim 7, wherein the converting comprises setting the value of the average value or the mode value as the lower limit value within the predetermined range, and approaching the upper limit value within the predetermined range as the distance from the average value or the mode value increases. **Claim 9** The method according to claim 8, wherein the converting further comprises setting a value that is more than the threshold away from the average value or the mode value as the upper limit value within the predetermined range. **Claim 10** The converting comprises: calculating a z-score from the quantitative data; dividing the z-score by a given value; obtaining a value by setting values greater than 1 to 1 and values less than -1 to -1 among the divided values; taking the absolute value of negative values among the obtained values; The method according to claim 6, comprising. **Claim 11** The converting comprises: converting the quantitative data into a histogram; dividing each of the plurality of bins of the histogram by the value of the bin with the highest frequency among the plurality of bins; subtracting the divided value from 1; The method according to claim 6, comprising. **Claim 12** A method for stratifying a plurality of data, which is executed in a computer system including a processor, the method comprising: receiving, by the processor, a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items; extracting, by the processor, at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; identifying, by the processor, a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; stratifying, by the processor, the data in the plurality of data groups according to the identified relationship; A method including the above.
13. A system for associating a plurality of items, comprising: receiving means for receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items; extracting means for extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; identifying means for identifying a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; The relationship is: A system including using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part.
14. A program for associating a plurality of items, the program being executed in a computer system including a processor, the program comprising: receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data for a plurality of items; extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; identifying a relationship between at least one item extracted from each of the plurality of heterogeneous data groups; causing the processor to perform a process including the above, and the relationship is: A program comprising using at least one item extracted from one of the plurality of heterogeneous data groups as a premise part and at least one item extracted from another one of the plurality of heterogeneous data groups as a conclusion part.
15. A system for associating a plurality of items, comprising: Receiving means for receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items; Extracting means for extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; Specifying means for specifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups; The system is provided with the extracting means, The extracting means extracts at least one item having an abnormal value among the plurality of items in each of the plurality of heterogeneous data groups. A system that performs the above.
16. A program for associating a plurality of items, the program being executed in a computer system including a processor, and the program includes: Receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items; Extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; Specifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups; Causing the processor to perform a process including the above, and the extracting includes: Extracting at least one item having an abnormal value among the plurality of items in each of the plurality of heterogeneous data groups. A program that includes the above.
17. A system for associating a plurality of items, comprising: Receiving means for receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items; Extracting means for extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups; Specifying means for specifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups; A system comprising extraction means for extracting at least one item using a recursive iterative approach. **Claim 18**: A program for associating a plurality of items, the program being executed in a computer system comprising a processor, the program receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items, extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups, identifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups, causing the processor to perform a process including the above, and the extracting including extracting at least one item using a recursive iterative approach. **Claim 19**: A system for stratifying a plurality of data, comprising receiving means for receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items, extracting means for extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups, identifying means for identifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups, and means for stratifying the data in the plurality of data groups according to the identified relationship. A system comprising the above. **Claim 20**: A program for stratifying a plurality of data, the program being executed in a computer system comprising a processor, the program receiving a plurality of heterogeneous data groups, each of the plurality of heterogeneous data groups including a plurality of data about a plurality of items, extracting at least one item from each of the plurality of heterogeneous data groups based on the plurality of data in each of the plurality of heterogeneous data groups, identifying the relationship between at least one item extracted from each of the plurality of heterogeneous data groups, and stratifying the data in the plurality of data groups according to the identified relationship. A program causing the processor to perform a process including the above.
Citation Information
Patent Citations
Apparatus for supporting integration of heterogeneous database
JP2004086782A
Device and program for generating integrated log and recording medium
JP2010182194A
Table definition device and method
JP2016110646A
Clustering anatomical or physiological condition data
JP2019535046A