System, program and method for determining dependence between data
The data dependency determination system efficiently classifies and evaluates data dependencies using the Gini Coefficient, addressing computational inefficiencies in Bayesian estimation, enabling rapid and reliable dependency analysis across diverse data types.
Patent Information
- Application Number
- JP2023217006
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2043-12-22
AI Technical Summary
Existing data analysis methods, particularly Bayesian estimation, face challenges with high computational costs and inefficiencies in handling complex dependency relationships, especially in high-dimensional datasets, leading to difficulties in real-time data analysis and resource-constrained environments.
A data dependency determination system that classifies data into multiple classes, uses a bias evaluation method like the Gini Coefficient to quantify the bias in class distributions, and evaluates dependency possibilities, allowing for efficient determination of data dependencies with reduced computational requirements.
The system enables rapid and efficient determination of data dependencies, suppressing exponential calculation increases and providing reliable dependency results, even in high-dimensional datasets, while supporting various data types including categorical and time-series data.
Smart Images

Figure 2025099969000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data analysis technology.
Background Art
[0002] The following Patent Document 1 discloses a method for creating a Bayesian network model showing the dependency relationship between variables.
Prior Art Document
Patent Document
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As an analysis method for the dependency relationship between data, Bayesian estimation (Bayesian network) is widely used. Bayesian estimation may require multi-dimensional integration for the calculation of posterior probability, and in a complex dependency relationship model including a large number of parameters, the integration calculation may take a long time. In addition, for the calculation of Bayesian estimation, the Markov chain Monte Carlo (MCMC) method or other sampling methods are usually used, but since these require a large number of iterative calculations, the amount of calculation tends to become extremely large. Furthermore, in a high-dimensional dataset with many features and observation points, there is also a problem that a phenomenon called the "curse of dimensionality" occurs and the calculation efficiency deteriorates.
[0005] Analysis of the dependency relationship using Bayesian estimation is difficult to apply to real-time data analysis, online learning, etc. due to such high calculation costs. In addition, since Bayesian estimation consumes a large amount of computing resources (CPU, memory, and in some cases, GPU), it may take an extremely long time for its processing in an environment with limited resources.
[0006] In view of such problems, the problem to be solved by the present invention is to enable more efficient and faster determination of the dependency between data.
Means for Solving the Problem
[0007] To solve the above problems, the dependency determination system for data in the present invention includes an input means for first data and second data, which are two types of numerical data sets associated with each other, a first data value that is each data value of the first data, and a second data value that is each data value of the second data, a classification means for dividing them into a plurality of classes respectively, a distribution acquisition means for acquiring the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in a certain class of the first data, and a bias evaluation means for quantifying the bias of the number of occurrences between the classes of the second data. The gist is to be provided with these.
[0008] By classifying data sets associated with each other into classes respectively and determining the dependency between data from the bias of the number of occurrences of each class of the other data in a certain class of one data set, it is possible to suppress the exponential increase in the amount of calculation with respect to the increase in the number of variables and data, and obtain the determination result promptly.
[0009] More specifically, for example, if there is a dependency between the first data and the second data, that is, if the second data increases, decreases, or changes in conjunction with the increase, decrease, or change of the first data, the second data values associated with the first data values included in a certain class of the first data should be concentrated around relatively close values. That is, among all the classes of the second data, the second data values belonging to that specific class should be observed more frequently than the second data values belonging to other classes. Conversely, if the first data and the second data are unrelated, such a tendency will not appear, and it is highly likely to result in a more random result. Based on this concept, by determining the dependency from the classified data sets, it is possible to efficiently determine the dependency between data with less calculation amount.
[0010] At this time, it is preferable that the bias evaluation means quantifies the bias by the Gini Coefficient calculated from the histogram data obtained by cumulatively arranging the occurrence numbers of the respective classes of the second data in ascending order. In the present invention, since the occurrence numbers of the respective classes of the second data in any class of the first data are aggregated, data in the form of a histogram is consequently generated. Therefore, by applying the Gini Coefficient, which is used in economics to measure the inequality (bias) of income distribution from the histogram of the number of households and the cumulative income amount, to the quantification of the bias in this system, it becomes possible to quantify the bias of the application numbers of the respective classes of the second data in a simple and proven method.
[0011] Further, the dependency determination system of the present invention may further include a dependency possibility evaluation means for evaluating that the greater the bias or its average value, the higher the possibility that the increase or decrease of the second data depends on the increase or decrease of the first data, and the smaller the bias, the lower the possibility.
[0012] Further, the dependency determination system of the present invention has a plurality of candidates in the data set serving as the first data, and among the plurality of candidates, a first filter means for extracting a candidate in which the bias of the occurrence numbers or its average value between the classes of the second data quantified by the bias evaluation means is equal to or greater than a predetermined magnitude. For example, when the second data is a so-called target variable (result to be predicted, output value of the model) and the first data is an explanatory variable (input value for predicting the target variable), it can be considered that an explanatory variable with a low dependency possibility (the above bias) of the target variable has a low explanatory ability for the target variable. From this, by removing candidates with a low dependency possibility of the second data (target variable) among a plurality of candidates of the first data (explanatory variable), for example, it becomes possible to more efficiently perform machine learning or the like using these first data and second data thereafter.
[0013] Further, for a plurality of classes of the first data, the distribution acquisition means acquires, for each class, the number of occurrences of each class of the second data values associated with the first data values included in that class. The bias evaluation means preferably further includes averaging means for quantifying the bias in the number of occurrences among the classes of the second data for each class of the first data, and obtaining a weighted average of the bias using, as a weight, the total number of occurrences of each class of the second data for each class of the first data. By determining the dependency possibility between data for a plurality of classes of the first data, a more reliable determination result can be obtained.
[0014] At this time, the bias evaluation means may quantify the bias with the Gini coefficient calculated from the histogram data obtained by cumulatively arranging the number of occurrences of each class of the second data in ascending order for each class of the first data, and the averaging means may obtain the weighted average by the following formula.
Equation
[0015] Further, the distribution acquisition means further acquires, for any class in the entirety of the second data, the number of occurrences of each class to which these first data values belong from the first data values associated with the second data values included in that class, and the bias evaluation means further preferably quantifies the bias in the number of occurrences among the classes of the first data.
[0016] At this time, it is preferable that the dependency determination system of the present invention further includes an asymmetry determination means for evaluating the direction and degree of dependency between the first data and the second data from the bias in the number of occurrences between the classes of the second data quantified by the bias evaluation means and the bias in the number of occurrences between the classes of the first data.
[0017] Also at this time, in the dependency determination system of the present invention, there are a plurality of candidates for the data set serving as the first data. Among the plurality of candidates, the second filter means for extracting a candidate in which it is specified by the asymmetry determination means that the second data depends or the degree of dependency of the second data is equal to or greater than a predetermined degree may be further provided. The purpose and effect are the same as those of the first filter means described above.
[0018] By mutually swapping the first data and the second data (mutually swapping the target variable and the explanatory variable) and determining the dependency between the data by the same method, the direction and degree of dependency of these data can be specified. For example, even if a clear bias is found in the number of classes of the second data associated with any class of the first data, it is not clear whether the first data depends on the second data or the second data depends on the first data. By mutually swapping the positions of the first data and the second data and determining the dependency possibility, it becomes possible to specify the direction and degree of this dependency from the difference in these biases (dependency possibility).
[0019] In addition, in the dependency determination system of the present invention, the distribution acquisition means acquires, for each of a plurality of classes in the entirety of the first data, the number of occurrences of each class of the second data values associated with the first data values included in that class. The bias evaluation means quantifies the bias in the number of occurrences among the classes of the second data using the Gini coefficient calculated from the histogram data obtained by cumulatively arranging the number of occurrences of each class of the second data in ascending order for each of the plurality of classes of the first data. The averaging means further includes, for each of the plurality of classes of the first data, a weighted average of the bias of the second data, with the total number of the number of applications of each class of the second data as a weight. The distribution acquisition means further acquires, for each of a plurality of classes in the entirety of the second data, the number of occurrences of each class of the first data values associated with the second data values included in that class. The bias evaluation means further quantifies the bias in the number of occurrences among the classes of the first data using the Gini coefficient calculated from the histogram data obtained by cumulatively arranging the number of occurrences of each class of the first data in ascending order for each of the plurality of classes of the second data. The averaging means further obtains a weighted average of the bias of the first data, with the total number of the number of occurrences of each class of the first data as a weight for each of the plurality of classes of the second data. The asymmetry determination means may obtain the direction and degree of dependence between the first data and the second data by the following formula. [Number] totalGini A→B : Weighted average of the bias of the second data totalGini B→A : Weighted average of the bias of the first data Dependency: Direction and degree of dependence between the first data and the second data
[0020] In addition, the dependency determination system of the present invention may further include a visualization means for visualizing and displaying the direction of dependence evaluated by the asymmetry determination means.
[0021] Also, in order to solve the above problems, the dependency determination system for data in the present invention includes input means for first data and second data, which are two types of data sets associated with each other, and for each first data value that is a data value of the first data and each second data value that is a data value of the second data, when the data value is numerical data that can be classified, it is classified into classes and each class is used as a group, and when the data value is data that cannot be classified, each data value is used as a group respectively; grouping means; distribution acquisition means for acquiring the number of appearances of each group to which the second data values belong from the second data values associated with the first data values included in any one of the groups of the first data; and bias evaluation means for quantifying the bias in the number of appearances among the groups of the second data. This is the gist of the invention.
[0022] Also, in order to solve the above problems, the dependency determination system for data in the present invention includes input means for first data and second data, which are two types of data sets associated with each other, distribution acquisition means for acquiring the number of appearances of each data value of the second data from the data values of the second data associated with the same data value of the first data, and bias evaluation means for quantifying the bias in the number of appearances among the data values of the second data. This is the gist of the invention.
[0023] The dependency determination system for data in the present invention can determine the dependency between data by the same method not only for numerical data sets, but also when one of them is data that cannot be classified, such as categorical data, and even when both data are data that cannot be classified.
[0024] In addition, the dependency determination system for data may further include a time lag setting means capable of associating the first data and the second data, which are time-series measurement data, not with those having the same acquisition time, but with those having their acquisition times shifted. By applying the dependency determination system of the present invention to time-series data, it is possible to determine the presence or absence of a dependency relationship between the first data and the second data whose time axis is shifted from the first data. As a result, for example, data analysis for specifying signs of an event occurring becomes possible.
[0025] In addition, to solve the above problems, the data dependency determination program in the present invention causes a computer to function as an input means for first data and second data, which are two types of numerical data sets associated with each other, a first data value that is each data value of the first data, and a second data value that is each data value of the second data, a classification means for dividing each into a plurality of classes, and in any class of the first data, a distribution acquisition means for acquiring the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in that class, and a bias evaluation means for quantifying the bias of the number of occurrences between the classes of the second data. The intention, effect, and principle of each configuration of the present invention are the same as those of the above-described dependency determination system.
[0026] In addition, to solve the above problems, the data dependency determination method in the present invention includes a step of inputting first data and second data, which are two types of numerical data sets associated with each other, a step of dividing each of the first data value, which is each data value of the first data, and the second data value, which is each data value of the second data, into a plurality of classes, a step of acquiring the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in any class of the first data, and a step of quantifying the bias of the number of occurrences between the classes of the second data. The intention, effect, and principle of each step of the present invention are the same as those of the above-described dependency determination system.
[0027] At this time, in the distribution acquisition step, if the number of classes with an occurrence number of zero among all classes of the second data exceeds a predetermined ratio of the total number, it is preferable to return to the classification step and divide the first data value and the second data into a smaller number of classes. This is because the determination accuracy of the dependency can be improved by setting an appropriate number of classes.
Effect of the Invention
[0028] Thus, according to the present invention, it is possible to more efficiently and quickly determine the dependency between data.
Brief Description of the Drawings
[0029]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0030] <First Embodiment> Hereinafter, embodiments of the present invention will be described with reference to the drawings. A dependency determination system 11 (hereinafter simply referred to as "dependency determination system 11") for determining the presence or absence of dependency, the direction of dependency, and the degree of dependency between numerically associated data sets is a system for determining the presence or absence of dependency, the direction of dependency, and the degree of dependency between numerically associated data sets. Here, "dependency", "dependency", and "dependency relationship" between data sets mean a relationship in which the behavior of one data is affected by the other data, that is, one data increases or decreases or changes in conjunction with the increase or decrease or change of the other data. This does not necessarily have bidirectionality like a correlation relationship, and includes a relationship that affects only in one direction. Also, this is different from a causal relationship and does not directly show a cause-and-effect relationship. And the "degree of dependency" means the degree of that dependency. Also, the "possibility of dependency" is an index indicating the degree to which one data increases, decreases, or changes in response to an increase, decrease, or change in the other data. The direction of dependency cannot be determined by only the possibility of dependency in one direction.
[0031] (System Configuration) FIG. 1 is a block diagram showing the functional configuration of the dependency determination system 11. The dependency determination system 11 of this embodiment is configured by installing a dedicated application program on a general personal computer.
[0032] As shown in FIG. 1, the dependency determination system 11 of this embodiment mainly includes an input unit 21, a classification unit 31, a distribution acquisition unit 41, a bias evaluation unit 42, an averaging unit 43, a dependency probability evaluation unit 44, an asymmetry determination unit 51, and a visualization unit 91. These units are realized by the above application using the functions of a personal computer.
[0033] The dependency determination system 11 of this embodiment, that is, the above application program, does not require special hardware nor is it limited to specific hardware, and may be installed, for example, on a general smartphone or a single-board computer. Alternatively, it can be provided as a cloud service (SaaS) with a general web browser as a user interface. Or, each function of the above application program can be distributed among a plurality of server computers and these can be connected by a network to constitute a large-scale dependency determination system 11.
[0034] (Outline of Dependency Determination Method) FIG. 2 is a flowchart showing the flow of the dependency determination method by the dependency determination system 11. Although the details of each step will be described later, first, an overview of these general contents and the relationship with each unit provided in the dependency determination system 11 will be given.
[0035] When determining the dependency between data, the dependency determination system 11 of this embodiment first reads, by the input unit 21, a first data A and a second data B which are numerical data sets to be determined (step 10).
[0036] Thereafter, the dependency determination system 11 divides, by the classification unit 31, each data value of the first data A, i.e., the first data value Av, and each data value of the second data B, i.e., the second data value Bv, into classes An and Bn, respectively (step 20).
[0037] After that, the dependency determination system 11 starts analyzing the dependency of the second data B on the first data A. Here, in the following description, "A→B" means a relationship where a change in the first data A affects the second data B, that is, the second data B depends on the first data A, and "B→A" means a relationship where a change in the second data B affects the first data A, that is, the first data A depends on the second data B.
[0038] When determining the possibility of A→B dependency, first, the distribution acquisition means 41 obtains, from the second data values Bv associated with the first data value Av included in a certain class An of the first data A, the number of occurrences of each class Bn to which these second data values Bv belong (hereinafter, this number of occurrences is also referred to as "class frequency Bf") (step 31). Then, the bias evaluation means 42 quantifies the bias of the class frequency Bf among the classes Bn of the second data B (step 32). Here, "biased" class frequency Bf means that a specific class Bn of the second data B appears significantly more often than other classes Bn, and "not biased" means that each class Bn of the second data B appears randomly and it is difficult to recognize a significant difference. The distribution acquisition means 41 and the bias evaluation means 42 of this embodiment repeat the same processing (steps 31 and 32) for all classes An of the first data A.
[0039] After that, the dependency determination system 11 takes a weighted average of the bias values of the class frequency Bf in each class An of the first data A by the averaging means 43 (step 33). In this embodiment, this weighted average becomes a numerical value indicating the dependency possibility of the second data B on the first data A. If the second data B depends on the first data A, the bias becomes larger, and if there is no dependency relationship, the bias is likely to become smaller (resulting in a more random result). That is, the larger the bias, the higher the possibility that the increase or decrease of the second data B depends on the increase or decrease of the first data A, and the smaller the bias, the lower the possibility.
[0040] Since the dependency possibility of the second data B on the first data A can be evaluated by the above weighted average, at this point, the dependency possibility evaluation means 44 may evaluate the level of the dependency possibility (step 34). The dependency possibility evaluation means 44 may simply display the value of the above weighted average, or may process and display it in a form that is intuitively easy for the user to understand. In addition, in this embodiment, in order to further deeply analyze the dependency between these data sets, step 34 may be omitted.
[0041] After calculating the above weighted average, the dependency determination system 11 next determines the dependency possibility of B→A. Specifically, the positions of the first data A and the second data B are swapped, and the same processing as steps 31 to 34 is repeated (steps 41 to 44). That is, not only the dependency possibility of the second data B on the first data A but also the dependency possibility of the first data A on the second data B is obtained.
[0042] Thereafter, based on the weighted average (dependency possibility) of the biases in A→B and B→A, the asymmetry determination means 51 finally evaluates the direction and degree of the dependency between these data sets. Specifically, the weighted averages of the biases in A→B and B→A are compared, and it is evaluated that the dependency direction is from the smaller bias to the larger bias, that is, among A→B and B→A, the one with the larger bias is the dependency direction, and the difference in the bias represents the level of the dependency degree.
[0043] Thereafter, the dependency determination system 11 diagrams and displays the direction and degree of the dependency between these data sets evaluated by the asymmetry determination means 51 by the visualization means 91. Specifically, the dependency direction between the data sets is indicated by an arrow between the nodes, and the dependency degree is displayed numerically.
[0044] (Details of the Dependency Determination Method) Next, the dependency determination method by the dependency determination system 11 will be described in more detail using a specific example. FIG. 3 shows a part of a fictional dataset, the real estate price dataset, prepared for this explanation. The real estate price dataset is data in CSV format that summarizes indicators that can affect the value of real estate for each town.
[0045] In the above step 10 (input of dataset), the user of the dependency determination system 11 loads this real estate price dataset using the input means 21. In the following explanation, it is assumed that the user selects the item "crime rate per person" indicating the crime risk per person as the first data A, and the item "industrial land ratio" indicating the ratio of non-retail business land per town as the second data B. The items selected by the user here are arbitrary.
[0046] Incidentally, the input means 21 of this embodiment supports a CSV format data file as its input target, but other format data files such as JSON or XML may be made inputtable, or a dataset may be extracted from a database system (not shown) or downloaded from a data source on the Internet.
[0047] In the above step 20 (classification), the dependency determination system 11 divides the first data A and the second data B into classes An and Bn respectively by the classification means 31. FIG. 4 shows the states of the first data value Av and the second data value Bv classified by the classification means 31. The first data value Av and the second data value Bv in FIG. 4 are each divided into 10 classes An and Bn, and each data value Av and Bv is assigned a class number (class #) from 0 to 9. Incidentally, "class" may be read as "BIN".
[0048] FIG. 4 is a diagram for convenience of explaining the process of the dependency determination method by the dependency determination system 11, and the table data shown in FIG. 4 is not necessarily displayed to the user. The same applies to FIGS. 5 and 6 referred to later. Further, the classification means 31 may be executed by an explicit instruction of the user, that is, it may be executed manually, or may be automatically executed after the input of the data set by the input means 21 or after the selection of an item by the user. This (the fact that its execution does not depend on manual or automatic) also applies to the distribution acquisition means 41, the bias evaluation means 42, the averaging means 43, the dependency possibility evaluation means 44, the asymmetry determination means 51, and the visualization means 91.
[0049] FIG. 5 is a diagram showing the process in which the class frequency Bf is aggregated by the distribution acquisition means 41. After the classification of the first data A and the second data B, the dependency determination system 11 starts analyzing the dependency possibility of A→B. In the above step 31 (acquisition of class frequency), the distribution acquisition means 41 first obtains the class frequency Bf of each class Bn to which these second data values Bv belong from the second data values Bv associated with the first data value Av included in the class # "0" of the first data A.
[0050] FIG. 6 is an explanatory diagram showing how the bias evaluation means 42 quantifies the bias of the class frequency Bf in the class # "0" of the first data A. FIG. 6(a) shows the process of processing the class frequency Bf in quantifying the bias, and FIG. 6(b) is a reference diagram for explaining the concept when calculating the Gini coefficient from the processed class frequency Bf.
[0051] In step 32 (quantification of bias), the bias evaluation means 42 first sorts the class frequencies Bf aggregated by the distribution acquisition means 41 in ascending order. Then, the cumulative class frequency Bfc is obtained by accumulating the sorted class frequencies Bfa. Then, histogram data (FIG. 6(b)) with the cumulative class frequency Bfc on the vertical axis and the number of classes Bc on the horizontal axis is created from the cumulative class frequency Bfc. Here, the number of classes Bc is a numerical value obtained by simply counting the number of accumulated classes Bn, which is different from the class Bn.
[0052] Then, the bias evaluation means 42 calculates the area of the portion below the Lorenz curve in the plot area of the histogram data, assuming that there is a Lorenz curve connecting the tops (centers of the tops of each BIN) of the cumulative class frequencies Bfc at each class number Bc with a straight line. Specifically, the sum of the areas of trapezoids (total area) is obtained, where the cumulative class frequencies Bfc of adjacent class numbers Bc are used as the upper and lower bases, and 1, which is the difference between the adjacent class numbers Bc, is used as the height. For example, for the hatched portion in FIG. 6(b), the area of the trapezoid can be obtained by the following equation. (129 + 185) / 2 = 157
[0053] Then, the bias evaluation means 42 calculates the area of a right triangle (reference area) with the diagonal line from the origin in the same plot area (also called the perfect equality line) as the hypotenuse, and obtains the difference between the total area of the trapezoids and the reference area (area difference), that is, the area of the region partitioned by the equal distribution line and the Lorenz curve. Then, the Gini coefficient is calculated from this reference area and the area difference. Specifically, the number obtained by dividing the area difference by the reference area is defined as the Gini coefficient. The Gini coefficient quantifies the bias, with less bias as it approaches 0 and more bias as it approaches 1.
[0054] Here, if the second data B depends on the first data A, the second data values Bv associated with the first data values Av included in each class An of the first data A should be concentrated around relatively close values. That is, among all the classes Bn of the second data B, the second data values Bv belonging to a specific class Bn should be observed significantly more often than the second data values Bv belonging to other classes Bn. Conversely, if the second data B does not depend on the first data A, such a tendency will not appear, and a more random result is likely.
[0055] If this Gini coefficient is "1", the possibility that the second data B depends on the first data A is extremely high, and if it is "0", the possibility of dependence is almost non-existent. That is, as it approaches from "0" to "1", the dependence possibility increases, and as it approaches from "1" to "0", the dependence possibility decreases. Here, the criteria for determining whether the dependence possibility is "high" or "low" vary depending on the nature of the dataset and the use of its dependence possibility, etc., so it cannot be judged with a uniform fixed value. Therefore, when applying the obtained dependence possibility to some use, it is desirable to individually set a reference value according to the purpose and the nature of the dataset.
[0056] Also, in the dependence determination system 11 of this embodiment, since the class frequencies Bf of the second data B in each class An of the first data A are aggregated, as a result, data in the form of a histogram is generated. Therefore, by applying the Gini coefficient, which is used in economics to measure the inequality (bias) of income distribution from the histogram of the number of households and the cumulative income amount, to the quantification of the bias in this system, the dependence determination system 11 has realized quantifying the bias of the class frequencies Bf by a simple and proven method. Note that the evaluation of the bias is not limited to the method using the Gini coefficient, and other known methods such as entropy can also be used. In addition, as long as the quantification of the bias is possible, an original method may be adopted.
[0057] The distribution acquisition means 41 and the bias evaluation means 42 of this embodiment repeat steps 31 and 32 to obtain the Gini coefficient of the class frequencies Bf for all classes An of the first data A.
[0058] And in step 33 (averaging of the bias values), the dependence determination system 11 takes the weighted average of the Gini coefficients using the total number of class frequencies Bf in each class An of the first data A as the weight by the averaging means 43. Specifically, the averaging means 43 obtains the weighted average (totalGini A→B ) of the Gini coefficients according to the following formula.
Equation
[0059] Here, in this embodiment, the Gini coefficient of the class frequency Bf in all classes An of the first data A is obtained, and the reliability of the Gini coefficient is enhanced by taking its weighted average. However, the Gini coefficient does not necessarily have to be obtained in all classes An of the first data A, nor does the averaging method always have to be a weighted average. For example, even if only the Gini coefficient for class # "0" of the first data A is obtained, it is possible to determine the dependence possibility of A→B to some extent.
[0060] As described above, since the dependence possibility of the second data B on the first data A can be concluded by this weighted average, it is also possible to evaluate the level of this dependence possibility by the dependence possibility evaluation means 44 in step 34 (evaluation of dependence possibility). However, the direction of dependence cannot be determined solely by this dependence possibility. This dependence possibility only indicates the level of increase or decrease of the second data B in response to the increase or decrease and change of the first data A, that is, it only shows the data linkage rate of A→B. On the other hand, in applications where interest is directed to the linkage rate of this data (such as when using the first filter means 81 in the fourth embodiment described later), it is possible to obtain a certain determination result and effect only by this dependence possibility.
[0061] By the steps up to this point, the determination of the dependency possibility from A to B is completed. However, as described above, even if a high dependency possibility is obtained here, it is not clear at this point whether the first data A depends on the second data B or the second data B depends on the first data A. Therefore, in steps 41 to 44, the first data A and the second data B are swapped, and the dependency possibility from B to A is determined by the same method. From the difference in the weighted average (dependency possibility) of the Gini coefficients in both directions, the direction and degree of these dependencies can be specified. In this embodiment, for the sake of convenience of explanation, the determination of the dependency possibility from A to B and the determination of the dependency possibility from B to A are executed serially, but these may be executed in parallel.
[0062] After that, in step 50 (dependency asymmetry determination), the dependency determination system 11 calculates the direction and degree of the dependency between the first data A and the second data B by the asymmetry determination means 51. Specifically, the asymmetry determination means 51 calculates the direction and degree (Dependency) of the dependency by the following formula.
Equation
[0063] Here, if the calculation result by the asymmetry determination means 51 is greater than 1, it means A→B, that is, the second data B depends on the first data A. If it is less than 1, it means B→A, that is, the first data A depends on the second data B. And the distance of Dependency from 1 on the number line indicates the degree of that dependence. Here, the maximum value of Dependency when the relationship A→B holds is infinity, but the Dependency when in the relationship B→A is a value from 1 to 0. Therefore, when the relationship is B→A, the numerator and denominator may be swapped to obtain the positive-direction dependence degree, or the absolute value may be obtained from the logarithm of Dependency.
[0064] And finally, in step 60 (visualization of the dependency relationship), the dependency determination system 11 diagrams this based on the evaluation result by the asymmetry determination means 51 using the visualization means 91. FIG. 7 is a diagram showing the direction and degree of dependence between data sets diagrammed by the visualization means 91. "CRIM" in this diagram refers to the first data A, and "INDUS" refers to the second data B. In this example, the result is that the first data A depends on the second data B (B→A). Note that FIG. 7 shows the results of determining the dependency for not only "CRIM" and "INDUS" but also more items.
[0065] In this way, the dependency determination system 11 of this embodiment classifies data sets associated with each other into classes respectively, and determines the dependency between data from the bias of the class frequency of the other data in a certain class of one data set, thereby suppressing the exponential increase in the calculation amount with respect to the increase in the number of variables (items) and the number of data, and enabling the determination result to be obtained promptly. More specifically, according to the dependency determination system 11, even if the number of items to be determined increases, the increase in the calculation amount is limited to about the following formula. O(m×n 2 ) o: calculation amount n: number of variables (items) m: number of data
[0066] In the present embodiment, the first data value Av and the second data value Bv are each divided into 10 classes to determine the dependency between them. However, the number of classes An and Bn is not limited to 10 and can be arbitrarily adjusted. However, it should be at least larger than the sampling interval of the data set. For example, in the case of a data set with a range of 0.1 to 1.0 and an increment / decrement unit of 0.1, it should not be divided into 10 or more classes. If the number of classes An and Bn seems to be excessive, the following procedure may be considered to reduce the number of classes to a minimum of 5. (1) Check the number of classes with zero class frequency (2) If the number is 30% or less of the number of classes, there is no need to reduce the number of classes (end of consideration) (3) If the number is more than 40% of the number of classes, reduce the number of classes so that it becomes 30% or less (4) When the number of classes reaches the minimum value of 5, determine the number of classes to be 5 (end of consideration)
[0067] <Second Embodiment> Hereinafter, a second embodiment of the present invention will be described. FIG. 8 is a block diagram showing the functional configuration of the dependency determination system 12 according to this embodiment. Different from the dependency determination system 11 in the previous embodiment, the dependency determination system 12 can determine the dependency between data sets not only for numerical data sets but also for data sets that cannot be classified. Specifically, the dependency determination system 12 includes a grouping means 32 that supports not only numerical data but also other data, instead of the classification means 31 of the dependency determination system 11. Since the other functional configurations and the dependency determination method are the same as those of the dependency determination system 11, the following description of the duplicates will be omitted, and only the differences will be focused on for explanation.
[0068] FIG. 9 shows a part of a fictional data set, the occupation-salary data set, used in the description of the dependency determination system 12. The occupation-salary data set is data in CSV format that summarizes the occupation and salary of each worker and the attributes that can affect them.
[0069] In the above step 10 (input of dataset), the user of the dependency determination system 12 loads this occupation - salary dataset using the input means 21, and selects the item "Educational Level" which quantifies the highest educational attainment of the worker as the first data C, and the item "Salary" which indicates the salary of the worker as the second data D.
[0070] Figure 10(a) shows the states of the first data values Cv and the second data values Dv grouped by the grouping means 32. In the above step 20 (classification), the dependency determination system 12 groups each first data value Cv of the first data C, which is non - classifiable numerical data, into group Cg by the grouping means 32. For the second data D, which is classifiable numerical data, it is divided into 10 classes Dg, and each of the second data values Dv is assigned a class number (class #) from 0 to 9.
[0071] Figure 10(b) shows the process of obtaining the class frequency of the second data D in group Cg "1" by the distribution acquisition means 41. In step 31 (acquisition of class frequency), the distribution acquisition means 41 acquires the class frequency (not shown) of each class Dn to which these second data values Dv belong from the second data values Dv associated with the first data value Cv of group Cg "1" (similarly, the first data value Cv is also "1"). In the subsequent steps, the dependency can be determined in the same procedure as the previous embodiment by taking group Cg as class An of the previous embodiment and group Dg as class Bn.
[0072] Figure 11 shows a state where, in step 10 (input of a data set), a user of the dependency determination system 12 selects an item "occupation" indicating a worker's occupation as first data E and an item "race" indicating a worker's race as second data F. In this case, in step 20, the grouping means 32 groups each first data value Ev of the first data E, which is categorical data that cannot be classified hierarchically, into a group Eg, and also groups each second data value Fv of the first data F, which is also categorical data that cannot be classified hierarchically, into a group Fg. Also in this case, in subsequent steps, by setting group Eg as class An and group Fg as class Bn in the previous embodiment, the dependency can be determined in the same procedure as in the previous embodiment.
[0073] In this way, for a data set that cannot be classified hierarchically, the dependency determination system 12 enables the determination of the dependency between the data set that cannot be classified hierarchically and other data sets by treating the data values themselves in the same way as classes.
[0074] <Third Embodiment> Hereinafter, a third embodiment of the present invention will be described. FIG. 12 is a block diagram showing the functional configuration of a dependency determination system 13 according to this embodiment. The dependency determination system 13 includes a time lag setting means 71 in addition to the functions of the dependency determination system 12 in the previous embodiment. The time lag setting means 71 has a function of being able to associate pairs of time-series measurement data that are originally associated by having a common acquisition time by shifting their acquisition times. Since the other functional configurations and dependency determination methods are the same as those of the dependency determination system 12, the following description of the duplicates will be omitted, and only the differences will be focused on for explanation.
[0075] Figure 13(a) shows a part of an abnormal monitoring dataset, which is a fictional dataset used in the dependency determination system 13. The abnormal monitoring dataset is time-series measurement data in CSV format that combines, for example, the presence or absence of abnormalities (situation M) in facilities and equipment, etc., and the values S1 and S2 of the monitoring sensors at that time. As shown in Figure 13(a), the abnormal monitoring dataset is originally a dataset in which the situation M and the sensor values S1 and S2 are associated according to the acquisition time T. Note that the sensor values S1 and S2 are not data that directly indicate the presence or absence of abnormalities.
[0076] After the input of the abnormal monitoring dataset in the above step 10 (input of the dataset), the dependency determination system 13 of this embodiment can associate the situation M and the sensor values S1 and S2 with an arbitrary time width shift by the time lag setting means 71. Figure 13(b) is a diagram showing how the time lag setting means 71 associates the sensor values S1 and S2 acquired in the past with respect to each situation M from the acquisition time T. The items from the "Sensor 1-1" column to the "Sensor 2-3" column in the figure are items that associate the sensor values S1 and S2 with respect to each situation M while shifting them stepwise into the past. For example, in the "Sensor 1-1" column and the "Sensor 2-1" column, for the situation M "normal" acquired at time T "5", the sensor value S1 "Value41" and the sensor value S2 "Value42" acquired at time T "4" are associated.
[0077] The user can determine the dependency relationship between these datasets in the same manner as in the previous embodiment, using the situation M as the second data and any of the sensor value S1 or S2 items shifted into the past and associated as the first data. Thereby, data analysis such as specifying the sign of a certain event occurring becomes possible.
[0078] In this way, when detecting dependencies from data without direct association, the dependency determination system 13 can obtain the results with a significantly smaller amount of computation and in a significantly shorter time than using machine learning. For example, in a neural network, iterative calculations are used, so the required amount of computation becomes very large. Also, in many machine learning methods, although not as much as in neural networks, iterative calculations are still required. That is, one of the advantages of the dependency determination system 13 is that it can determine the dependencies between data with a small amount of computation without going through the process of "learning".
[0079] <Fourth Embodiment> Hereinafter, the fourth embodiment of the present invention will be described. FIG. 14 is a block diagram showing the functional configuration of the dependency determination system 14 according to this embodiment. The dependency determination system 14 includes first filter means 81 and second filter means 82 in addition to the functions of the dependency determination system 11 of the first embodiment. In this embodiment, the dependency possibility evaluation means 44 and the visualization means 91 of the dependency determination system 11 are omitted, but these may or may not be present. Since the other functional configurations and dependency determination methods are the same as those of the dependency determination system 11, the following duplicate explanations will be omitted, and only the differences will be focused on for explanation.
[0080] The first filter means 81 and the second filter means 82 are intended to detect explanatory variables with low explanatory power (whose influence on the target variable is considered weak) for the positioning item as the target variable included in the data set and remove them when there are items that can be multiple explanatory variables. That is, the purpose is to extract only the explanatory variables that are highly likely to be dependent on the target variable. By removing data with little influence on the target variable, various subsequent data analyses can be performed more efficiently and effectively. More specifically, for example, the dependency determination system 14 of this form is used as preprocessing for machine learning where the number of variable combination patterns becomes extremely large, or a case is assumed where the dependency determination system 14 is used for information selection in edge computing where a huge amount of sensor information is collected in real time.
[0081] Next, the functions of the first filter means 81 and the second filter means 82 of this form will be described using the example of FIG. 3. For example, assume that the "housing value index" item in FIG. 3 is the target variable and all other items are explanatory variables. At this time, the dependency determination system 14 determines the dependency possibility or dependency of each first data H of the second data I on the second data I, sequentially or in parallel, with the "housing value index" item as the second data I and the other items as the first data H. The method for determining the dependency possibility and dependency is the same as in the first embodiment.
[0082] Here, when removing unnecessary explanatory variables by the first filtering means 81, the first filtering means 81 extracts, from among a plurality of first data H, items in which the bias of the class frequency quantified by the bias evaluation means 42 or the weighted average value thereof by the averaging means 43 is equal to or greater than a predetermined magnitude, that is, items with a high degree of dependence. In this case, the processing after step 32 or 33 may be omitted. When estimating some output (objective variable) from some input (explanatory variable), this degree of dependence is strongly related to the estimation result. That is, an explanatory variable with a high degree of dependence has a large influence on the objective variable, and an explanatory variable with a low degree of dependence has a small influence on the objective variable. In order to reduce the scale of the estimation model of the objective variable, it is effective to reduce explanatory variables with low importance. For example, by removing explanatory variables in order of their low degree of dependence, the model scale can be reduced while minimizing a decrease in accuracy. Note that depending on the type and nature of the dataset to be determined, there are those whose degree of dependence is generally evaluated as low and those whose degree of dependence is generally evaluated as high. Therefore, it is desirable to appropriately adjust the "predetermined magnitude" of the above bias and weighted average in consideration of the balance between the reduction rate of the model scale and the accuracy.
[0083] On the other hand, when removing unnecessary explanatory variables by the second filtering means 82, the second filtering means 82 extracts, from among a plurality of first data, items for which it is specified by the asymmetry determination means 51 (Dependency) that the second data I depends or that the degree of dependence of the second data I is equal to or greater than a predetermined degree. Here, "equal to or greater than a predetermined degree" for Dependency assumes a case where all Dependencies are less than 1.0, for example, when guessing the number indicated by a digital image, and the specific degree varies depending on the determination target. Therefore, it is necessary to individually obtain and adjust the reference value of the "predetermined degree" here. Also, generally, there are many cases where it may be determined that there is "dependence" when Dependency is 1.5 or more. This reference value is also only a guideline value and a recommended value, and it is necessary to adjust it in consideration of the unique nature of each dataset.
[0084] In the above example, the "housing value index" item was used as the target variable and the other items were used as explanatory variables. However, it is possible to arbitrarily select which item to use as the target variable and which items to use as explanatory variables.
[0085] As described above, embodiments of the present invention have been described. However, the scope of the present invention is not limited thereto, and various modifications can be made without departing from the gist of the invention.
Explanation of Reference Numerals
[0086] 11, 12, 13, 14: Dependency determination system between data, 21: Input means, 31: Classification means, 32: Grouping means, 41: Distribution acquisition means, 42: Bias evaluation means, 43: Averaging means, 44: Dependency possibility evaluation means, 51: Asymmetry determination means, 71: Time lag setting means, 81: First filter means, 82: Second filter means, 91: Visualization means, A: First data (numerical data set), Av: First data value, Ac: Class of the first data, B: Second data (numerical data set), Bv: Second data value, Bc: Class of the second data, C: First data (numerical data set that cannot be classified), Cv: First data value, Cg: Group of the first data, D:: Second data (numerical data set that can be classified), Dv: Second data value, Dg: Group of the second data, E: First data (category data set), Ev: First data value, Eg: Group of the first data, F:: Second data (category data set), Fv: Second data value, Fg: Group of the second data, G: Measurement data set over time, T: Acquisition time, M: Presence or absence of abnormality, S1, S2: Sensor values
Claims
1. Input means for first data and second data, which are two types of numerical data sets linked to each other, Classification means for dividing each data value of the first data, which is a first data value, and each data value of the second data, which is a second data value, into a plurality of classes respectively, Distribution acquisition means for obtaining, in any class of the first data, the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in that class, Bias evaluation means for quantifying the bias in the number of occurrences between the classes of the second data, A system for determining the dependency between data.
2. The bias evaluation means quantifies the bias by the Gini Coefficient calculated from the histogram data obtained by accumulating the number of occurrences of each class of the second data in ascending order, The system for determining the dependency between data according to claim 1.
3. Further comprising dependency possibility evaluation means for evaluating that the greater the bias or its average value, the higher the possibility that the increase or decrease of the second data depends on the increase or decrease of the first data, and the smaller the bias, the lower the possibility, The system for determining the dependency between data according to claim 1.
4. There are a plurality of candidates for the data set that becomes the first data, Further comprising first filter means for extracting, from among the plurality of candidates, a candidate for which the bias in the number of occurrences between the classes of the second data quantified by the bias evaluation means or its average value is equal to or greater than a predetermined magnitude, The system for determining the dependency between data according to claim 1.
5. The distribution acquisition means obtains, for each of the plurality of classes of the first data, the number of occurrences of each class of the second data values associated with the first data values included in that class for each class, The bias evaluation means quantifies the bias in the number of occurrences between the classes of the second data for each class of the first data, Further comprising averaging means for obtaining a weighted average of the biases using, as a weight, the total number of occurrences of each class of the second data for each class of the first data, The system for determining the dependency between data according to claim 1.
6. The bias evaluation means quantifies the bias by the Gini Coefficient calculated from the histogram data obtained by accumulating the number of occurrences of each class of the second data in ascending order for each class of the first data, The averaging means obtains the weighted average by the following formula, The dependency determination system between data according to claim 5. 【Number 1】 An: Each class of the first data Gini B | An : Gini coefficient of the second data for each class of the first data numAn: The total number of the occurrences for each class of the first data totalGini A→B : weighted average of the bias of the second data
7. The distribution acquisition means further acquires, for any class in the whole of the second data, the number of occurrences of each class to which the first data values associated with the second data values included in the class belong, from the first data values associated with the second data values included in the class, The bias evaluation means further quantifies the bias in the number of occurrences between the classes of the first data, The dependency determination system between data according to claim 1.
8. The asymmetry determination means further includes, from the bias in the number of occurrences between the classes of the second data quantified by the bias evaluation means and the bias in the number of occurrences between the classes of the first data, evaluating the direction and degree of the dependency between the first data and the second data, The dependency determination system between data according to claim 7.
9. There are a plurality of candidates for the data set that becomes the first data, Among the plurality of candidates, the second filter means further extracts a candidate in which it is specified by the asymmetry determination means that the second data depends, or the degree of dependence of the second data is equal to or greater than a predetermined degree, The dependency determination system between data according to claim 8.
10. The distribution acquisition means acquires, for each of a plurality of classes in the whole of the first data, the number of occurrences of each class of the second data values associated with the first data values included in the class, for each class, The bias evaluation means quantifies the bias in the number of occurrences between the classes of the second data by the Gini coefficient calculated from the histogram data obtained by cumulatively arranging in ascending order the number of occurrences of each class of the second data for each of the plurality of classes of the first data, The averaging means further includes, using, as a weight, the total number of the applications of each class of the second data for each class in the plurality of classes of the first data, obtaining a weighted average of the bias of the second data, The distribution acquisition means further acquires, for each of a plurality of classes in the whole of the second data, the number of occurrences of each class of the first data values associated with the second data values included in the class, for each class, The bias evaluation means further quantifies the bias in the occurrences between the classes of the first data by using the Gini coefficient calculated from the histogram data obtained by cumulatively arranging in ascending order the occurrences of each class of the first data for each of the plurality of classes of the second data. The averaging means further obtains a weighted average of the bias of the first data, using, as weights, the total number of occurrences of each class of the first data for each class among the plurality of classes of the second data. The asymmetry determination means obtains the direction and degree of dependence between the first data and the second data by the following formula. The system for determining dependency between data according to claim 8. 【Number 2】 totalGini A→B : The weighted average of the bias of the second data totalGini B→A : weighted average of the bias of the first data Dependency: The direction and degree of dependence between the first data and the second data
11. The system for determining dependency between data according to claim 8 further includes visualization means for visualizing and displaying the direction of dependence evaluated by the asymmetry determination means in a diagram. The system for determining dependency between data according to claim 8.
12. Input means for the first data and the second data, which are two types of data sets associated with each other. Grouping means for classifying, into groups, each class obtained by classifying the first data value, which is each data value of the first data, and the second data value, which is each data value of the second data, into classes when the first data value or the second data value is numerical data that can be classified, or using each data value as a group when the first data value or the second data value is data that cannot be classified. Distribution acquisition means for acquiring, from the second data values associated with the first data values included in a group of the first data, the number of occurrences of each group to which these second data values belong. A system for determining dependency between data, comprising bias evaluation means for quantifying the bias in the number of occurrences between the groups of the second data. System for determining dependency between data.
13. Input means for the first data and the second data, which are two types of data sets associated with each other. Distribution acquisition means for acquiring the number of occurrences of each data value of the second data from the data values of the second data associated with the same data value of the first data. A system for determining dependency between data, comprising bias evaluation means for quantifying the bias in the number of occurrences between the data values of the second data. System for determining dependency between data.
14. The first data and the second data are time-series measurement data. The time lag setting means for associating the first data value and the second data value with each other by shifting their acquisition times instead of associating those having the same acquisition time. The dependency determination system between data according to claim 1.
15. A computer, Input means for the first data and the second data which are two types of numerical data sets associated with each other, Classification means for dividing each data value of the first data, which is a first data value, and each data value of the second data, which is a second data value, into a plurality of classes respectively, Distribution acquisition means for obtaining the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in the class in any one class of the first data, and Functioning as bias evaluation means for quantifying the bias of the number of occurrences between the classes of the second data, A program for determining the dependency between data.
16. An input step of inputting the first data and the second data which are two types of numerical data sets associated with each other, A classification step of dividing each data value of the first data, which is a first data value, and each data value of the second data, which is a second data value, into a plurality of classes respectively, A distribution acquisition step of obtaining the number of occurrences of each class to which these second data values belong from the second data values associated with the first data values included in the class in any one class of the first data, A bias evaluation step of quantifying the bias of the number of occurrences between the classes of the second data, and including, A method for determining the dependency between data.
17. In the distribution acquisition step, when the number of classes in which the number of occurrences is zero among all the classes of the second data exceeds a predetermined ratio of the total number, return to the classification step and divide the first data value and the second data into a smaller number of classes. The method for determining the dependency between data according to claim 16.
Citation Information
Patent Citations
Model creation device, information analyzer, model creation method, method for analyzing information, and program
JP2005107747A