Information processing program, information processing method, and information processing device
By using graph structure data to determine critical values for discretization, the proposed information processing program effectively aligns discretization boundaries with distribution characteristics, addressing the limitations of existing techniques and enhancing classification accuracy and explicability.
Patent Information
- Application Number
- PCT/JP2023/044885
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-19
AI Technical Summary
Existing techniques for discretizing continuous variables, such as Equal Width Discretization (EWD) and Gaussian Mixture Model (GMM) clustering, struggle to accurately divide sample sets at boundaries corresponding to the characteristics of the distribution, leading to misalignment between discretization thresholds and distribution boundaries.
An information processing program and apparatus that determines a threshold value for discretizing a sample set of continuous variables by identifying a critical value based on data indicating topological changes, using graph structure data such as Reeb graphs to ensure that boundaries align with distribution characteristics.
This approach allows for accurate division of continuous variables at boundaries corresponding to the distribution characteristics, improving classification accuracy and enhancing the explicability of discretization thresholds.
Smart Images

Figure JP2023044885_19062025_PF_FP_ABST
Abstract
Description
Information processing program, information processing method, and information processing device
[0001] The present invention relates to an information processing program, an information processing method, and an information processing device.
[0002] One known technique for discretizing continuous variables is discretization using equally spaced intervals, known as EWD (Equal Width Discretization).Other known techniques include clustering using a Gaussian Mixture Model (GMM).
[0003] Japanese Patent Application Laid-Open No. 2021-012531
[0004] However, the above techniques have an aspect in which it is difficult to divide a sample set of a continuous variable at a boundary that corresponds to the characteristics of the distribution contained in the sample set.
[0005] For example, in the above discretization using equally spaced intervals, the boundaries that divide the sample set of a continuous variable do not necessarily coincide with the boundaries between the distributions of the sample set, because the distributions of continuous variables with common characteristics are not necessarily spaced at equal intervals in the sample set of the continuous variable.
[0006] Furthermore, in the above clustering, the sample set of the continuous variable does not necessarily follow the distribution assumed by the GMM, and the number of clusters set in the GMM does not necessarily match the number of distributions included in the sample set of the continuous variable. Therefore, the boundaries between clusters obtained by the above clustering may not necessarily match the boundaries between distributions in the sample set.
[0007] In one aspect, an object of the present invention is to provide an information processing program, an information processing method, and an information processing device that can divide a sample set of a continuous variable at a boundary that corresponds to the characteristics of the distribution.
[0008] An information processing program according to one embodiment causes a computer to execute a process of determining, as a threshold value for discretizing a sample set, a value of a variable corresponding to a critical value identified based on data indicating a topology change for the sample set of variable values.
[0009] According to one embodiment, a sample set of a continuous variable can be divided at boundaries that correspond to characteristics of the distribution.
[0010] FIG. 1 is a block diagram showing an example of the functional configuration of a server device. FIG. 2 is a diagram showing a mathematical explanation of a Reeb graph. FIG. 3 is a diagram showing an example of mesh generation. FIG. 4 is a diagram showing an example of a Reeb graph. FIG. 5 is a diagram showing an example of a discretization result. FIG. 6 is a flowchart showing the procedure of discretization processing. FIG. 7 is a flowchart showing the procedure of calculation processing of a Reeb graph. FIG. 8 is a diagram showing an example of a hardware configuration.
[0011] Hereinafter, embodiments of the information processing program, information processing method, and information processing device according to the present application will be described with reference to the accompanying drawings. Each embodiment merely illustrates one example or aspect, and does not limit the range of values, functions, or usage scenarios. Furthermore, each embodiment can be appropriately combined within the scope that does not cause contradictions in the processing content.
[0012] <First Embodiment> <Overall Configuration> Fig. 1 is a block diagram showing an example of the functional configuration of a server device 10. The server device 10 shown in Fig. 1 provides a continuous variable discretization function that discretizes a sample set of continuous variables.
[0013] The server device 10 is an example of a computer that provides the continuous variable discretization function. In one embodiment, the server device 10 can provide the continuous variable discretization function by executing software that realizes the continuous variable discretization function. For example, the server device 10 can be realized as a server that provides the continuous variable discretization function on-premise. Alternatively, the server device 10 can be realized as a PaaS (Platform as a Service) or SaaS (Software as a Service) application, thereby providing the continuous variable discretization function as a cloud service.
[0014] The client terminal 30 corresponds to an example of a computer that receives the continuous variable discretization function described above. For example, the client terminal 30 may be realized by a desktop or laptop personal computer. This is merely an example, and the client terminal 30 may be realized by any computer, such as a mobile terminal device or a wearable terminal.
[0015] Note that while FIG. 1 shows an example in which the continuous variable discretization function is provided in a client-server system, this is merely an example, and the continuous variable discretization function may also be provided as a standalone function.
[0016] <Example of Usage Scenarios> The continuous variable discretization function described above may be applied in various situations. One example of a usage scenario is application to machine learning technologies that can realize explainable AI (artificial intelligence), so-called XAI (Explainable AI), such as Wide Learning (registered trademark).
[0017] In general, deep learning improves accuracy by training a single model using multiple layers of neural networks that mimic the structure of the neural circuits in the human brain, resulting in a complex model that humans cannot understand.
[0018] On the other hand, the above-mentioned machine learning technology generates a machine learning model that combines hypotheses and importance. That is, a large number of hypotheses are extracted by combining data items, and the importance of the hypotheses (knowledge chunks (hereinafter sometimes simply referred to as "chunks") is adjusted to build a machine learning model that is capable of highly accurate classification.
[0019] Knowledge chunks are simple rules that are easy for humans to understand, and they describe hypotheses that may hold for input-output relationships in a logical expression. Specifically, all combination patterns of data items in the input data are treated as hypotheses (chunks), and the importance of each hypothesis is adjusted based on the hit rate of the classification label for that hypothesis. For example, training is performed to lower the importance if there is a lot of overlap between the items that make up a knowledge chunk and the items that make up other knowledge chunks, and to raise the importance if there is little overlap.
[0020] Here, categorical data such as gender or marital status are not necessarily used as data items; continuous variables such as temperature and atmospheric pressure may also be incorporated into hypotheses. In such situations, discretizing continuous variables has a significant impact on classification accuracy. For example, if there is a discrepancy between the distribution trend of a continuous variable and the threshold (boundary) set during discretization, the accuracy of classification according to the values of the discretized continuous variable may decrease. From this perspective, the continuous variable discretization function, which can divide a sample set of continuous variables at boundaries corresponding to the distribution characteristics, is of great technical value.
[0021] <Configuration of Server Device 10> Next, an example of the functional configuration of the server device 10 according to this embodiment will be described. Fig. 1 schematically illustrates blocks related to the continuous variable discretization function of the server device 10. As shown in Fig. 1, the server device 10 includes a communication control unit 11, a storage unit 13, and a control unit 15. Note that Fig. 1 only illustrates a selection of functional units related to the continuous variable discretization function, and the server device 10 may include functional units other than those illustrated.
[0022] The communication control unit 11 is a functional unit that controls communication with other devices such as the client terminal 30. As just one example, the communication control unit 11 can be realized by a network interface card such as a LAN card. In one aspect, the communication control unit 11 receives a discretization request from the client terminal 30 requesting discretization of a sample set of continuous variables, or outputs a sample set of discrete variables obtained as a result of the discretization to the client terminal 30.
[0023] The storage unit 13 is a functional unit that stores various types of data. As an example, the storage unit 13 is realized by internal, external, or auxiliary storage of the server device 10. For example, the storage unit 13 stores sample set data 13A. Note that the sample set data 13A will be described together with a scene in which the sample set data 13A is referenced, generated, or registered.
[0024] The control unit 15 is a functional unit that performs overall control of the server device 10. For example, the control unit 15 may be realized by a hardware processor. As shown in FIG. 1 , the control unit 15 includes a receiving unit 15A, a calculating unit 15B, a determining unit 15C, and a discretizing unit 15D. Note that the control unit 15 may also be realized by hardwired logic or the like.
[0025] The reception unit 15A is a processing unit that receives various requests from the client terminal 30. As just one example, the reception unit 15A can receive a discretization request from the client terminal 30 to discretize a sample set of a continuous variable.
[0026] When receiving such a discretization request, the reception unit 15A can also receive a designation of a sample set of continuous variables to be discretized. In one aspect, the reception unit 15A can also receive a designation of one sample set from sample set data 13A in which multiple sample sets are stored in a database. In another aspect, the reception unit 15A can receive a sample set to be discretized from the client terminal 30 via the network NW. In a further aspect, the reception unit 15A can also receive a designation from among sample sets stored in a database server, file system, or the like on the network NW.
[0027] The calculation unit 15B is a processing unit that calculates data indicating a topology change from a sample set of continuous variables. An example of such data indicating a topology change is graph structure data. For example, the calculation unit 15B calculates a Reeb graph as an example of graph structure data indicating a topology change, but other graph structure data such as a Contour Tree may also be calculated.
[0028] Figure 2 shows a mathematical explanation of the Reeb graph. Given a continuous space M and a continuous function f:M→R, the equivalence relations defined on the continuous space M are illustrated. As shown in Figure 2, the equivalence relation v~u is given by f -1 The connected component of (f(v)) is f -1 For example, v in the continuous space M shown in Figure 2 is defined as the connected component of 1 , v 2 and v 3 In the example of 1 and v 2 is the equivalence relation v 1 ~v 2 On the other hand, v 1 and v 3 And, v 2 and v 3 is not an equivalence relation. The Reeb graph G1 is a quotient topological space with topology, and a topology change occurs near the value of f at the vertex.
[0029] The Reeb graph is a graph structure that represents how a geometric structure (isolines or isosurfaces) formed by a set of points with the same variable value changes topologically as the value of the variable changes. Therefore, it is possible to represent the topological changes related to the variable for which data discretization is desired. Therefore, the vertices of the Reeb graph can be considered as boundaries where the distribution characteristics change. In other words, the distribution trend does not change much unless the value of the Reeb graph vertex is exceeded. Therefore, by using the value of the Reeb graph vertex as an example of a critical value and discretizing a continuous variable based on the value of the Reeb graph vertex, it is possible to consider the characteristics of a distribution across multiple variables.
[0030] More specifically, the calculation unit 15B meshes the sample set of continuous variables received by the reception unit 15A. FIG. 3 is a diagram showing an example of mesh generation. FIG. 3 shows an example in which a set 20 of sample points having three continuous variables x, y, and z as elements is plotted on the YZ plane. As shown in FIG. 3, the set 20 of sample points is meshed according to an algorithm such as Delaunay triangulation. As a result, edges are set on the points included in the set 20 of sample points to generate triangular surfaces corresponding to the convex hull of the set 20 of sample points, and a mesh (geometric structure) 21 is generated from the set 20 of sample points.
[0031] The calculation unit 15B then calculates a Reeb graph for the variables to be discretized from the generated mesh. This Reeb graph can be calculated, for example, according to the Parsa algorithm described in Reference 1 below. For example, the calculation unit 15B sorts the vertices of the mesh in ascending order by the value of f. The calculation unit 15B then calculates contour lines near the larger and smaller values of each vertex. If the number of contour lines differs, the calculation unit 15B adds a node to the Reeb graph and inserts an edge into the node corresponding to the contour line.
[0032] Reference 1: Parsa, Salman. “A deterministic o (m log m) time algorithm for the reeb graph.” Proceedings of the twenty-eighth annual symposium on Computational geometry. 2012.
[0033] The determination unit 15C is a processing unit that determines critical values assigned to vertices of the Reeb graph as thresholds for discretizing a sample set. FIG. 4 is a diagram illustrating an example of a Reeb graph. FIG. 4 illustrates a Reeb graph G10 calculated for a variable x to be discretized from the mesh 21 illustrated in FIG. 3. As illustrated in FIG. 4, the values of the variable z of the vertices included in the Reeb graph G10 of the variable z, i.e., the nodes illustrated as spheres in the figure, are set as thresholds. For example, in the example of the Reeb graph G10 illustrated in FIG. 4, the set of {-54.7213, -51.011, ..., 55.3309, 76.3581} can be set as thresholds for discretizing the variable z. The set of sample points discretized using these thresholds corresponds to the edges of the Reeb graph G10. Note that FIG. 4 illustrates an example in which the minimum and maximum values of the vertices included in the Reeb graph G10 are also set as thresholds, but the minimum and maximum values can be excluded from the threshold setting targets.
[0034] It is possible to increase the threshold if the user has some intention other than to set the value of the variable corresponding to the vertex of the Reeb graph. In this case, different labels are assigned to samples with the same distribution characteristics.
[0035] The discretization unit 15D is a processing unit that discretizes the sample set of continuous variables accepted by the accepting unit 15A based on the threshold determined by the determining unit 15C. By such discretization, each continuous variable included in the sample set is converted into a discrete variable. In this case, the continuous variables do not necessarily have to be converted into discrete variables with consecutive numbers, and may be converted into categorical variables without logical order.
[0036] FIG. 5 is a diagram illustrating an example of a discretization result. FIG. 5 illustrates the correspondence between the variable z of each sample point in the set 20 of sample points shown in FIG. 3 , the labels of the discrete variables (categorical variables) when discretization using equally spaced intervals, i.e., EWD, is applied, and the labels of the discrete variables (categorical variables) when the continuous variable discretization function according to this embodiment is applied. Furthermore, FIG. 5 illustrates an example in which, in the discretization using equally spaced intervals, the label value is incremented by one at equal intervals of "3" from the minimum value of the variable z. Furthermore, FIG. 5 illustrates an example in which the discretization of the variable z is performed using thresholds set using the Reeb graph G10 shown in FIG. 4 , i.e., a set of {-54.7213, -51.011, ..., 55.3309, 76.3581}.
[0037] As shown in FIG. 5 , in EWD, the label "0" is assigned to a sample having a variable z value of "-54.7213" and a sample having a variable z value of "-51.7213". The label "1" is assigned to a sample having a variable z value of "-51.011". However, since the variable z value "-51.011" corresponds to a vertex of the Reeb graph, the topology does not change below -51.011. Despite this, in EWD, different labels are assigned to samples with the same topological characteristics. On the other hand, in this embodiment, the label "0" is assigned to a sample having a variable z value of "-54.7213", a sample having a variable z value of "-51.7213", and a sample having a variable z value of "-51.011". This makes it possible to assign a common label to samples with the same topological characteristics.
[0038] Furthermore, in EWD, the labels "36," "42," and "43" are assigned to the sample with variable z value "57.7213," the sample with variable z value "75.7213," and the sample with variable z value "76.3581," respectively. However, since the variable z value "55.3309" and the variable z value "76.3581" correspond to vertices in the Reeb graph, the topology does not change when the value is greater than 55.3309 but less than or equal to 76.3581. Nevertheless, in EWD, different labels are assigned to samples with the same topological characteristics. On the other hand, in this embodiment, the same label "0" is assigned to the sample with variable z value "57.7213," the sample with variable z value "75.7213," and the sample with variable z value "76.3581." This allows a common label to be assigned to samples with the same topological characteristics.
[0039] <Processing Flow> Fig. 6 is a flowchart showing the procedure of the discretization process. As shown in Fig. 6, when the receiving unit 15A receives a discretization request (step S101), the calculation unit 15B meshes a sample set of continuous variables specified in the discretization request of step S101 (step S102). As a result, a mesh is generated from the set of sample points.
[0040] Thereafter, the calculation unit 15B executes a loop process 1 that repeats the processes of step S103 and step S104 described below a number of times corresponding to the number M of variables to be discretized.
[0041] That is, the calculation unit 15B calculates a Reeb graph for the m-th variable from the mesh generated in step S102 (step S103). After that, the determination unit 15C determines the critical value corresponding to the vertex of the Reeb graph calculated in step S103 as the threshold to be used for discretizing the m-th variable.
[0042] By repeating the processes of steps S103 and S104, the threshold values used for discretization are determined for each of the M variables.
[0043] Thereafter, the discretization unit 15D executes loop process 2, which repeats the process of step S105 described below, the number of times corresponding to the number N of sample points. Furthermore, the discretization unit 15D executes loop process 3, which repeats the process of step S105 described below, the number of times corresponding to the number M of variables to be discretized.
[0044] That is, the discretization unit 15D discretizes the m-th variable of the n-th sample point based on the threshold value for the m-th variable (step S105).
[0045] By this loop process 3, M variables at the n-th sample point can be discretized. Furthermore, by loop process 2, M variables can be discretized for each of N sample points.
[0046] Fig. 7 is a flowchart showing the procedure of the calculation process of the Reeb graph. The process shown in Fig. 7 corresponds to the process of step S103 shown in Fig. 6. As shown in Fig. 7, the calculation unit 15B sorts the vertices of the mesh in ascending order by the value of f (step S301).
[0047] Next, the calculation unit 15B executes loop processing 1, which repeats the processing of step S302 and step S303 described below, until there are no more unprocessed vertices.
[0048] That is, the calculation unit 15B calculates contour lines near the larger and smaller values of each vertex (step S302). At this time, if the number of contour lines is different, the calculation unit 15B adds a node to the Reeb graph and inserts an edge to the node corresponding to the contour line (step S303).
[0049] The processes of steps S302 and S303 are repeated to calculate the Reeb graph.
[0050] <One Aspect of Effect> As described above, the server device 10 according to this embodiment acquires a variable value corresponding to a critical value identified based on graph structure data indicating a topology change for a sample set of variable values, and determines the acquired variable value as a threshold for discretizing the sample set. This allows continuous variables of samples in a distribution without topology change to be converted into the same discrete variable. In other words, continuous variables of samples in different distributions with topology change to be converted into different discrete variables. Therefore, the server device 10 according to this embodiment allows a sample set of a continuous variable to be divided at a boundary corresponding to the characteristics of the distribution. In the case of a continuous variable, samples in a distribution without topology change are considered to have the same properties. The characteristics of the samples are often related to the classification results in classification processes such as clustering. Therefore, dividing a sample set of a continuous variable at a boundary corresponding to the characteristics of the distribution contributes to improving classification accuracy. Furthermore, the discretization threshold described in this embodiment can be associated with a topology change of the continuous variable, thereby improving the interpretability of the determined threshold.
[0051] Although the embodiments of the disclosed device have been described above, the present invention may be embodied in various different forms other than the above-described embodiments. Therefore, other embodiments included in the present invention will be described below.
[0052] <Distribution and Integration> Furthermore, the components of each device shown in the figure do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. For example, the reception unit 15A, the calculation unit 15B, the determination unit 15C, or the discretization unit 15D may be connected via a network as an external device to the server device 10. Furthermore, the reception unit 15A, the calculation unit 15B, the determination unit 15C, or the discretization unit 15D may each be included in a separate device, and the functions of the server device 10 may be realized by the devices being connected to a network and operating together.
[0053] <Hardware Configuration> The various processes described in the above embodiments can be realized by executing a prepared program on a computer such as a personal computer, a workstation, etc. Therefore, an example of a computer that executes an information processing program having the same functions as those in the first and second embodiments will be described below with reference to FIG.
[0054] Fig. 8 is a diagram showing an example of a hardware configuration. As shown in Fig. 8, a computer 100 has an operation unit 110a, a speaker 110b, a camera 110c, a display 120, and a communication unit 130. The computer 100 also has a CPU 150, a ROM 160, a HDD 170, and a RAM 180. These units 110 to 180 are connected via a bus 140.
[0055] 8, an information processing program 170a that performs the same functions as the reception unit 15A, calculation unit 15B, determination unit 15C, and discretization unit 15D shown in the first embodiment is stored in the HDD 170. This information processing program 170a may be integrated or separated, similar to the components of the reception unit 15A, calculation unit 15B, determination unit 15C, and discretization unit 15D shown in FIG. 1. In other words, it is not necessary for all of the data shown in the first embodiment to be stored in the HDD 170, as long as the data used for processing is stored in the HDD 170.
[0056] Under such an environment, the CPU 150 reads the information processing program 170a from the HDD 170 and loads it into the RAM 180. As a result, the information processing program 170a functions as an information processing process 180a, as shown in FIG. 8. This information processing process 180a loads various data read from the HDD 170 into an area of the storage area of the RAM 180 allocated to the information processing process 180a, and executes various processes using the loaded data. For example, examples of processes executed by the information processing process 180a include the processes shown in FIGS. 6 and 7. Note that the CPU 150 does not necessarily need to operate all of the processing units shown in the first embodiment above; it is sufficient that the processing units corresponding to the processes to be executed are virtually implemented.
[0057] The information processing program 170a does not necessarily have to be stored in the HDD 170 or the ROM 160 from the beginning. For example, each program may be stored on a "portable physical medium" such as a flexible disk, a so-called FD, a CD-ROM, a DVD disk, a magneto-optical disk, or an IC card that is inserted into the computer 100. The computer 100 may then acquire and execute each program from such a portable physical medium. Alternatively, each program may be stored in another computer or server device connected to the computer 100 via a public line, the Internet, a LAN, a WAN, or the like, and the computer 100 may acquire and execute each program from such a computer or server device.
[0058] REFERENCE SIGNS LIST 10 Server device 11 Communication control unit 13 Storage unit 13A Sample set data 15 Control unit 15A Reception unit 15B Calculation unit 15C Determination unit 15D Discretization unit 30 Client terminal
Claims
1. A computer-executable information processing program, characterized in that it determines, as a threshold value when discretizing a sample set of variable values, a value of a variable corresponding to a critical value identified based on data indicating a topological change with respect to the sample set of variable values.
2. The information processing program according to claim 1, characterized in that the critical value corresponds to a vertex included in graph structure data, which is the data indicating the topological change.
3. The information processing program according to claim 1, characterized in that the determining process excludes the minimum value and the maximum value among the values of the variable corresponding to the critical value from the threshold value.
4. The information processing program according to claim 1, characterized in that the data indicating the topological change corresponds to a rave graph.
5. An information processing method, characterized in that a computer executes a process of determining, as a threshold value when discretizing a sample set of variable values, a value of a variable corresponding to a critical value identified based on data indicating a topological change with respect to the sample set of variable values.
6. The information processing method according to claim 5, characterized in that the critical value corresponds to a vertex included in graph structure data, which is the data indicating the topological change.
7. The information processing method according to claim 5, characterized in that the determining process excludes the minimum value and the maximum value among the values of the variable corresponding to the critical value from the threshold value.
8. The information processing method according to claim 5, characterized in that the data indicating the topological change corresponds to a rave graph.
9. An information processing apparatus including a control unit that executes a process of determining, as a threshold value when discretizing a sample set of variable values, a value of a variable corresponding to a critical value identified based on data indicating a topological change with respect to the sample set of variable values.
10. The information processing apparatus according to claim 9, characterized in that the critical value corresponds to a vertex included in graph structure data, which is the data indicating the topological change.
11. The information processing apparatus according to claim 9, wherein the determining process excludes the minimum value and the maximum value among the values of the variable corresponding to the threshold value from the threshold value.
12. The information processing apparatus according to claim 9, wherein the data indicating the topology change corresponds to a ray graph.
Citation Information
Patent Citations
Method for encoding and decoding object and device capable of utilizing the same
JP2001092991A
Automated discovery of causal relationships in mixed datasets
US20220156759A1