program

The program optimizes correlation network analysis by removing false positives and calculating element P-values, addressing computational inefficiencies in existing methods to handle large datasets effectively.

WO2026069723A1PCT designated stage Publication Date: 2026-04-02HIRATA CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for correlation network analysis of multivariate data, such as those described in Patent Document 1 and Non-Patent Document 1, require excessive computational resources, making their application to big data impractical.

Method used

A program that performs FPO analysis to remove false positive non-seed elements and incorporates FNI analysis to calculate element P-values, reducing computational complexity by optimizing module selection and integration processes.

Benefits of technology

Reduces computational complexity and improves the efficiency of correlation network analysis, enabling effective handling of large datasets while maintaining robustness and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024045259_02042026_PF_FP_ABST
    Figure JP2024045259_02042026_PF_FP_ABST
Patent Text Reader

Abstract

When configuring a seed element module from non-seed elements having a higher correlation with a seed element among elements of multivariate data to be subjected to correlation network analysis, a program causes a computer to: execute, for individual ranks while increasing the ranks, an FPO analysis process for removing false-positive non-seed elements from the non-seed elements having the higher correlation with the seed element; store, in a storage unit, configuration information pertaining to the module on which the FPO analysis process has been executed; and stop the FPO analysis process and terminate a search for a module indicating a maximum module F score of the ranks early when a module identical to the configuration information of the module stored in the storage unit is constructed. The program causes the computer to execute an FNI analysis process for capturing false-negative elements, and, simultaneously with the FNI analysis process, causes the computer to calculate an element P value using a recursive formula while computing, for an element of interest, an element specificity rate at individual correlation rankings of a candidate element that is an element outside of the module.
Need to check novelty before this filing date? Find Prior Art

Description

program

[0001] This invention relates to a program that causes a computer to analyze the relationships between multiple influencing components, and to identify modules that have a high influence on a particular effect, and to extract candidates.

[0002] Patent Document 1 describes a technique for performing correlation network analysis on multivariate data, which is comprehensive molecular information obtained through omics analysis. Non-Patent Document 1 describes a technique for performing correlation network analysis on multivariate data such as big data.

[0003] Patent No. 6318334

[0004] Itto Mannen, Yoshiyuki Ogata, and Hideyuki Suzuki, "A Revolutionary Approach to Multivariate Analysis: Development of ConfeitoGUIplus," Journal of the Japan Society for Biotechnology, Vol. 95, No. 7, July 2017.

[0005] However, implementing the technologies described in Patent Document 1 and Non-Patent Document 1 as they are would require an extremely large amount of computation, making their application to big data impractical.

[0006] The objective of this invention is to provide a program that can reduce the computational complexity of correlation network analysis of multivariate data.

[0007] One aspect of the present invention is a program that, when constructing a module of seed elements from non-seed elements with high correlation to seed elements among the elements of multivariate data to be analyzed for correlation network, causes a computer to perform an FPO analysis process to remove false positive non-seed elements from among the non-seed elements with high correlation to seed elements, increasing the rank from a lower limit to an upper limit, which is the number of elements other than seed elements to be incorporated into the module of seed elements, for each rank, stores the configuration information of the module on which the FPO analysis process has been performed in a storage unit, and if a module identical to the configuration information of the module stored in the storage unit is constructed, stops the FPO analysis process and terminates the search for the module showing the maximum module F score for that rank early.

[0008] One aspect of the present invention involves causing a computer to perform FNI analysis processing on multiple modules composed of elements of multivariate data to be analyzed for correlation network analysis, incorporating false negative elements, and simultaneously with the FNI analysis processing, calculating the element P-value for the element of interest using the following formula, while calculating the element specificity rate for each correlation rank of candidate elements which are elements outside the module. This is a program where N is the total number of elements in the multivariate data, n is the number of elements in the target module, edge is the number of edges of the candidate element within the target module, and degree is the degree of the candidate element.

[0009] According to the present invention, a program can be provided that can reduce the computational complexity of correlation network analysis for multivariate data.

[0010] This figure shows a schematic configuration example of an information processing device according to one embodiment. This flowchart shows an example of the procedure for an information processing method according to one embodiment. This figure shows an example of the configuration of multivariate data according to one embodiment. This figure illustrates an example of the correlation coefficient between elements according to one embodiment. This figure illustrates an example of the correlation coefficient between elements according to one embodiment. This figure illustrates an example of the FPO module selection method according to one embodiment. This figure illustrates an example of the FPO module selection method according to one embodiment. This figure illustrates an example of the FPO analysis according to one embodiment. This figure illustrates an example of the final element F score according to one embodiment. This figure illustrates an example of the final element F score according to one embodiment. This figure illustrates an example of the calculation method for the final element F score according to one embodiment. This figure illustrates an example of the integration process according to one embodiment. This figure illustrates an example of the integration result of the final FPO module according to one embodiment. This figure illustrates an example of the integration result of the FPO module corresponding to the final FPO module integration result in Figure 8. This figure illustrates an example of the calculation of the distance index of the final FPO module in Figure 8. This figure illustrates the FNI analysis process according to one embodiment. This figure shows an example of the results of comparing the analysis time of element P values ​​with and without a recursive formula.

[0011] Next, the program of this embodiment will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below. In all the figures used to describe the embodiments, components having the same function will be given the same reference numerals, and repeated explanations will be omitted.

[0012] Furthermore, "based on XX" as used in this application means "based on at least XX," and includes cases where it is based on another element in addition to XX. Also, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on something that has been calculated or processed. "XX" is any element (for example, any information).

[0013] The terms used in this embodiment are explained in Tables 1, 2, and 3 below.

[0014]

[0015]

[0016]

[0017] Figure 1 is a diagram showing a schematic configuration example of the information processing device 1 according to this embodiment. In Figure 1, the information processing device 1 comprises an input / output unit 11, a storage unit 12, and a control unit 13.

[0018] The input / output unit 11 receives data input and outputs data to and from an external device of the information processing device 1. The method of data input and output is not limited. The input / output unit 11 may send and receive data by communication, for example, via a communication line. The input / output unit 11 may also input and output data, for example, via a recording medium.

[0019] The storage unit 12 stores various types of data, such as programs executed by the control unit 13 and data used by the control unit 13. The storage unit 12 is composed of non-volatile memory such as a hard disk drive, magneto-optical disk drive, or flash memory, a read-only recording medium such as a CD-ROM, a volatile memory such as DRAM (Dynamic Random Access Memory), or a combination thereof.

[0020] The control unit 13 is equipped with a CPU (Central Processing Unit) and realizes various functions by executing programs stored in the storage unit 12. In this embodiment, the control unit 13 executes the correlation network analysis program 121 stored in the storage unit 12, thereby realizing the functions of the correlation coefficient calculation unit 131, FPO analysis unit 132, FPO module selection unit 133, FPO module integration unit 134, FNI analysis unit 135, FNI assignment processing unit 136, P-value calculation unit 137, and output data generation unit 138.

[0021] The information processing device 1 may be configured using a general-purpose computer device, or it may be configured as a dedicated hardware device. For example, the information processing device 1 may be configured using a server computer connected to a communication network such as the Internet. Furthermore, each function of the information processing device 1 may be realized by cloud computing. In addition, the information processing device 1 may be realized by a single computer, or its functions may be realized by distributing them among multiple computers. Furthermore, the information processing device 1 may be configured to launch a website using, for example, a WWW system.

[0022] The memory unit 12 has memory areas for storing a correlation coefficient list 122, module information 123 for each rank by element, module integration information 124 for each module F score (MF) and element F score (VF) threshold, final module integration information 125, evaluation value information 126 for each target step, final module information 127, and an FPO-analyzed module list 128. Furthermore, by calculating and adding a hash value to each piece of information, the retrieval of information stored in the memory unit 12 may be made faster.

[0023] Next, the information processing method according to this embodiment will be described with reference to Figure 2. Figure 2 is a flowchart showing an example of the procedure of the information processing method according to this embodiment.

[0024] (Step S1) The input / output unit 11 receives input data. The input data is multivariate data to be analyzed using correlation network analysis. The input data may be, for example, comprehensive molecular information obtained by omics analysis. The input data may be, for example, product information obtained in the manufacturing and inspection processes of a product for the purpose of analyzing the cause of defective products. The input data may be, for example, big data. The input data may be, for example, time series data.

[0025] Furthermore, there are no particular restrictions on the acquisition route, acquisition method, type, attributes, etc., of the multivariate data targeted for correlation network analysis according to this embodiment. For example, multivariate data obtained from omics analysis can be used. "Omics analysis" generally refers to the integrated analysis of comprehensive molecular information, and typical examples of comprehensive molecular information include transcriptome data, which is comprehensive information about gene transcripts, and metabolome data, which is comprehensive information about metabolites.

[0026] These comprehensive data can be obtained by any method or means, such as various gene analyses, gene expression analyses, and various mass spectrometry methods including liquid chromatography-mass spectrometry (LC-MS), gas chromatography-mass spectrometry (GC-MS), and capillary electrophoresis-mass spectrometry (CE-MS). Furthermore, there are no particular restrictions on the source of these comprehensive data, and they can include parts, organs, tissues, and cells from various types of plants, animals, microorganisms, and bacteria. In addition, multivariate data may be information obtained by any method from samples acquired from a certain environment and artificial products (e.g., processed foods).

[0027] The format of the input data may have a plurality of data sets each consisting of quantitative values corresponding to a plurality of elements. FIG. 3 is a diagram showing a configuration example of multivariate data (input data) according to the present embodiment. The input data may be a quantitative value table composed of rows corresponding to each element and columns corresponding to each data set as illustrated in FIG. 3. In the quantitative value table shown in FIG. 3, the first column is the element name (for example, element identifiers (ID001, ID002, etc. in FIG. 3)), the first row is the data set name (DATA_a, DATA_b, etc. in FIG. 3), and the quantitative value of data set y of element x is stored in the x-th row and y-th column (x and y are integers of 2 or more). For example, the data “data_b002” in the 3rd row and 3rd column is the quantitative value of data set “DATA_b” of element “ID002”.

[0028] (Step S2) The correlation coefficient calculation unit 131 calculates the correlation coefficient between elements based on the quantitative values included in the input data. The correlation coefficient calculation unit 131 generates a correlation coefficient list 122 sorted in descending order of the correlation coefficient for each element with respect to all elements included in the input data. Thereby, a correlation coefficient list 122 is generated for each of all elements included in the input data. The correlation coefficient list 122 of element x is a list in which pairs of the element name and the correlation coefficient are sorted in order from the one with the larger correlation coefficient with element x. The correlation coefficient lists 122 of each element generated by the correlation coefficient list 122 are stored in the storage unit 12.

[0029] The correlation coefficient between elements according to the present embodiment is, for example, any one of Pearson's product-moment correlation coefficient, Spearman's rank correlation coefficient, and cosine similarity. The correlation coefficient between elements is a real number of -1 or more and 1 or less.

[0030] Note that as the correlation coefficient between elements according to the present embodiment, the negative correlation coefficient may be replaced with “0”, or the negative correlation coefficient may not be replaced with “0”. When the negative correlation coefficient is not replaced with “0”, “negative correlation” and “no correlation” can be discriminated. Therefore, even when two elements having a strong positive correlation with respect to other elements and a negative correlation in the same module belong, they can be associated with each other. For this reason, the module can be made easier to interpret.

[0031] FIGS. 4A and 4B are virtual image diagrams for explaining an example of the correlation coefficient between elements according to the present embodiment. FIG. 4A shows an example of a module, and the elements are indicated by circles. Elements having a value equal to or greater than a specific correlation coefficient are connected by lines. The elements include seed elements and in-module elements. The symbols shown inside the circles are the identification information of the elements. In the example shown in FIG. 4A, the correlation coefficient ρ between element X and element Y XY is greater than zero, and consider the case where the correlation coefficient ρ between element X and element Z XZ is greater than zero. In this case, whether the correlation coefficient ρ between element Y and element Z YZ is greater than zero is unknown. FIG. 4B shows an image diagram of the values that the correlation coefficient ρ YZ can take. According to FIG. 4B, it can be seen that the correlation coefficient ρ YZ can take a value greater than zero, a value less than zero, or it may be unclear whether it is greater than or less than zero. Note that ρ represents the correlation coefficient. FIG. 4B shows the value range (-1 to 1) of the Pearson correlation value on the X-axis and Y-axis.

[0032] (Step S3) The FPO analysis unit 132 uses each of all the elements included in the input data as a seed element, and excludes (Out) false positive non-seed elements from the non-seed elements located at the upper positions (the ones with larger correlation coefficients) in the correlation coefficient list 122 of the seed elements (FPO analysis process). The FPO analysis process will be described below.

[0033] The FPO analysis unit 132 constructs a temporary module by connecting the seed element with edges the non-seed elements of a predetermined upper rank (up to rank M) in the correlation coefficient list 122 of the seed element. M is one of the values ​​within a predetermined range of module size (from the lower limit to the upper limit of module size). The number of elements (excluding the seed element) that constitute the module of the seed element is the rank (rank = M). The predetermined range of module size can be arbitrarily set by the operator by, for example, setting a MIN value and a MAX value in the information processing device 1 in advance. Here, the MIN value and MAX value are the module size excluding the seed element, and are the minimum and maximum values ​​that the rank can take. Therefore, the lower limit of the total number of elements (seed element and non-seed element) that constitute the module (module size) is "MIN value + 1", and the upper limit is "MAX value + 1".

[0034] Next, the FPO analysis unit 132 generates a list of correlation coefficients sorted in descending order for the provisional module, consisting of all elements included in the correlation coefficient list 122 of the seed element and each non-seed element. The FPO analysis unit 132 further connects elements in the generated list that have correlation coefficients greater than a predetermined correlation threshold with edges. Here, by setting multiple correlation thresholds, multiple provisional modules with different edge patterns are constructed. For example, by setting all possible values ​​as correlation thresholds, provisional modules with edge patterns for each of the said correlation thresholds are constructed. The lower the correlation threshold, the more edges are created. For example, instead of setting multiple correlation thresholds simultaneously, the correlation threshold may be initialized to the maximum value "1", and then provisional modules may be constructed for each correlation threshold while decreasing the correlation threshold. Next, the FPO analysis unit 132 calculates the F score and MF of each provisional module with different edge patterns using the following formula.

[0035]

[0036] n is the number of elements in the module. edge(i) is the number of edges of element i, and for the edge pattern corresponding to the current correlation coefficient, it is the number of elements connected to element i within the same module. i is an integer value between 1 and n, and is the index of the element. degree(i) is the degree of element i, and for the edge pattern corresponding to the current correlation coefficient, it is the total number of elements connected to element i both inside and outside the module.

[0037] MF is the harmonic mean of module density (MD) and module specificity (MS). MD represents the degree to which modules are interconnected internally as a whole. MS represents the degree to which modules are isolated (exclusively connected) from the outside. Therefore, modules with a large MF are isolated from the outside and have strong connections as a whole.

[0038] The FPO analysis unit 132 stores the maximum value of the MF calculated for each of the multiple provisional modules with different edge patterns, as well as the correlation threshold and edge pattern corresponding to the maximum value of MF, in the storage unit 12. The FPO analysis unit 132 calculates the F score, VF(i), of each non-seed element i for the correlation threshold and edge pattern corresponding to the maximum value of MF using the following formula.

[0039]

[0040] VF(i) is the harmonic mean of element density (VD(i)) and element singularity (VS(i)). VD(i) represents the proportion of element i that is connected to other elements within the module. VS(i) represents the proportion of the degree that element i has in the entire network that is connected to elements of the module. i is an integer value between 1 and n and is the index of the element. Therefore, an element i with a large VF(i) is exclusively and densely connected to the module and can be said to be a suitable element to be a component of that module.

[0041] The FPO analysis unit 132 determines that the non-seed element i that minimizes the VF(i) of the provisional module is a false positive, and removes the false positive non-seed element i from the provisional module. Next, the FPO analysis unit 132 recalculates the maximum value of MF and the VF(i) of each non-seed element i for the provisional module from which the false positive non-seed element i has been removed. This FPO analysis process to remove false positive non-seed elements i is repeated until the number of elements in the provisional module reaches the "lower limit of module size," i.e., "MIN value + 1." The FPO analysis unit 132 determines that the point in time when MF was at its maximum represents a state where all false positives have been removed without excess or deficiency, and stores the configuration information of the provisional module at the point in time when MF was at its maximum in the module information 123 as the maximum MF module in the "rank" of the seed elements.

[0042] The FPO analysis unit 132 performs FPO analysis processing for each rank within a predetermined range of ranks (from the lower limit (MIN value + 1) to the upper limit (MAX value + 1) of the module size at the time the provisional module was configured). As a result, for each seed element, the configuration information of the provisional module at the time when the MF was at its maximum is stored in the module information 123 as the maximum MF module for all ranks within the predetermined range of ranks.

[0043] Furthermore, the FPO analysis unit 132 performs the FPO analysis in step S3 while increasing the rank from the lower limit (MIN value) to the upper limit (MAX value), and stores the configuration information of the modules for which the FPO analysis in step S3 has been completed (FPO-analyzed modules) in the FPO-analyzed module list 128. If a module identical to an FPO-analyzed module already existing in the FPO-analyzed module list 128 is constructed, the FPO analysis process to remove false-positive non-seed elements may be stopped, and the search for the largest MF module of that rank may be terminated early.

[0044] The FPO analysis unit 132 uses the correlation threshold and edge pattern to determine whether the maximum value of MF is updated or not. If it determines that the maximum value of MF is not updated in a provisional module where MF has not yet been calculated among multiple provisional modules with different edge patterns, it may terminate the MF calculation and early terminate the search for the maximum MF module of that rank. By configuring it in this way, the FPO analysis unit 132 can terminate the MF calculation and early terminate the search for the maximum MF module of that rank if it determines that the maximum value of MF is not updated in a provisional module where MF has not yet been calculated among multiple provisional modules with different edge patterns. This reduces the computational load and shortens the FPO analysis processing time compared to a case where the MF calculation continues and the module optimization process continues even if the maximum value of MF is not updated.

[0045] (Step S4) The FPO module selection unit 133 selects the optimal module (referred to as the FPO module) for a given seed element based on the module information 123.

[0046] Figure 5A is a diagram illustrating the FPO module selection method according to this embodiment. The example in Figure 5A is time-series metabolome data of grated radish (2305 elements x 10 time-series data set). The Pearson product-moment correlation coefficient has a MIN value of "2" and a MAX value of "2305", and the element with identifier "c1487" was used as the seed element for FPO analysis. Figure 5A shows the maximum value of MF (vertical axis) corresponding to the rank (horizontal axis). The FPO module selection unit 133 determines the module M_PTmax of the rank that shows the maximum value PTmax (longest plateau length) during the period in which the maximum value of MF is not updated even if the rank value is increased. The FPO module selection unit 133 selects the determined module M_PTmax as the FPO module.

[0047] In the example shown in Figure 5A, the MF of the module constructed at "Rank = 623" was calculated to be 0.683, and no modules with an MF exceeding 0.683 were constructed until "Rank = 761". Therefore, the plateau length, which is the period during which the maximum MF is not updated even when the rank value is increased, is "138", which is obtained by subtracting "Rank = 623" from "Rank = 761". In this example, "138" became the longest plateau length (maximum plateau length) PTmax, so the module constructed at "Rank = 623" was selected as module M_PTmax of rank "623", which represents the maximum plateau length PTmax.

[0048] Traditionally, the module corresponding to the rank showing the highest MF across all ranks was selected as the FPO module. However, the rank showing the highest MF across all ranks is often at or near the upper limit (MAX value) of the rank. Therefore, traditionally, this often resulted in the creation of a large, biased module. Furthermore, because it depended on the setting of the rank's upper limit (MAX value), there was a possibility that the optimal FPO analysis results could not be obtained.

[0049] In contrast, according to this embodiment, a module with a large MF and an appropriate number of elements can be obtained, resulting in an improvement in the quality of the FPO analysis results. Furthermore, the FPO analysis results are less affected by the setting of the rank upper limit (MAX value). Therefore, a more robust FPO analysis can be performed.

[0050] Furthermore, the module information 123 for each rank can be used to verify the selection criteria for FPO modules. The output data generation unit 138 may use the module information 123 for each rank to generate data for a graph of the maximum value of MF, as exemplified in Figure 5A. This graph data can be displayed on a display screen or printed. This allows operators to easily verify the selection criteria for FPO modules.

[0051] Alternatively, as another method for selecting FPO modules, a criterion that maintains robustness regardless of the number of elements in the module or the distribution of correlation coefficients between elements may be adopted. For example, a criterion using the plateau length of the moving average graph of the maximum value of MF, as exemplified in Figure 5B, or a criterion where the growth rate of said graph falls below a certain value may be adopted. In addition, a criterion optimized for each dataset based on the distribution of correlation coefficients between elements may be adopted. In this case, the upper limit of the rank, i.e., the MAX value, does not need to be adjusted as a setting parameter, thus reducing the burden on the operator.

[0052] Figure 5B is a diagram illustrating the FPO module selection method according to this embodiment. The example in Figure 5B is time-series metabolome data of grated radish (2305 elements x 10-point time-series dataset). The result of FPO analysis processing was performed using an element with Pearson product-moment correlation coefficient, MIN value "2", MAX value "2305", and identifier "c1487" as the seed element. Figure 5B shows the 10-point moving average (vertical axis) of the maximum value of MF corresponding to the rank (horizontal axis). In the example shown in Figure 5B, the MF of the module constructed at "rank = 152" was calculated to be 0.590, and no modules with an MF exceeding 0.590 were constructed until "rank = 385". Therefore, the plateau length of the moving average, which is the period during which the maximum value of MF is not updated even when the rank value is increased, is "233", which is obtained by subtracting "rank = 152" from "rank = 385". In this example, "233" became the longest plateau length (maximum plateau length) of the moving average, so the module constructed with "rank = 152" was selected as the module M_MAPTmax with rank "152" that represents the maximum plateau length, MAPTmax.

[0053] The FPO module selection unit 133 may terminate the FPO analysis and FPO module selection if predetermined conditions are met. Figure 5C is a diagram illustrating an example of FPO analysis according to this embodiment. Figure 5C is time-series metabolome data of grated radish (2305 elements x 10 time-series data set). The result of FPO analysis processing was performed using an element with identifier "c1487" as a seed element, with Pearson product-moment correlation coefficient, MIN value "2", and MAX value "2305". Figure 5C shows the maximum value of MF (vertical axis) corresponding to the rank (horizontal axis). In this example, there are 1569 elements with a positive correlation coefficient to identifier "c1487", so the maximum rank of identifier "c1487" is 1569. Figure 5C shows the process when, in step S3, the FPO module selection unit 133 is executed each time the FPO analysis unit 132 finishes calculating the maximum MF for each rank, and determines whether or not the longest plateau length is updated. If the FPO module selection unit 133 determines that the longest plateau length is not updated, it may determine that the maximum value of the plateau length is not updated even if the rank changes (even if the rank increases), and may terminate the rank increase and terminate the FPO analysis and FPO module selection. In Figure 5C, when the rank exceeds 1432, the remaining number of ranks becomes "137", which is the maximum rank "1569" minus the rank "1432". Since the number of ranks "137" is less than the PTmax "138", it is determined that the PTmax will not be exceeded in the part where the rank is 1432 or higher, and the calculation is not performed in the part where it is indicated that no calculation is necessary (the part where the rank is 1432 or higher). By configuring it in this way, the FPO module selection unit 133 can terminate the FPO analysis and FPO module selection when the longest plateau length is not updated even if elements are added (even if the rank increases). Therefore, even if the longest plateau length is not updated, the computational load can be reduced compared to when the FPO analysis and FPO module selection for each rank are continued, and the time required for FPO analysis and FPO module selection can be shortened.

[0054] (Step S5) If the selection of FPO modules is completed using all elements of the input data as seed elements (Step S5, YES), proceed to Step S6-1. On the other hand, if there are still elements for which FPO modules have not yet been selected (Step S5, NO), return to Step S3 and perform FPO analysis and FPO module selection using the elements for which FPO modules have not yet been selected as seed elements.

[0055] (Step S6-1) The FPO analysis unit 132 calculates the maximum value of VF(i) for each element (i) within the FPO module (referred to as the final VF) and the MF based on the final VF (referred to as the final MF) for each seed element's FPO module. The final MF and final VF calculated for each seed element's FPO module are stored in the evaluation value information 126.

[0056] Figures 6A and 6B are hypothetical diagrams illustrating an example of the final VF according to this embodiment. Figure 6A shows an example of a module and elements. In Figure 6A, elements are represented by circles. Elements include the element of interest (the element for which the final VF is calculated), in-module elements, and out-of-module elements. The numbers shown inside the circles are the correlation ranks for the element of interest. As an example here, the higher the correlation, the higher the rank (smaller the value). In Figure 6A, lines connect elements with values ​​above a certain correlation coefficient between the element of interest and the in-module elements, and between the element of interest and the out-of-module elements. Solid lines indicate relationships between elements included in the edge, i.e., in-module elements. Dotted lines indicate relationships between elements not included in the edge, i.e., out-of-module elements.

[0057] The FPO analysis unit 132 calculates the final VF for each element (i) within an FPO module. Specifically, the FPO analysis unit 132 selects each element (i) within the FPO module as a focus element and calculates VF(i) for that focus element. Figure 6B shows the calculation result of VF(i) for one element (i) (focus element). In Figure 6B, the horizontal axis shows the correlation rank, and the vertical axis shows VF(i). The FPO analysis unit 132 obtains the maximum value of VF(i) (final VF) for one element (i) (focus element). In the example in Figure 6B, when the correlation rank is "5", there are 3 connections to elements within the module and 2 connections to elements outside the module. The degree is "5 = 2 + 3", the edge is 3, and the number of elements in the module (n) is 5. From equation (2) above, VF(i) is "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (5 + 5 - 1) = 0.67", which represents the maximum value of VF(i) (final VF).

[0058] Furthermore, the FPO analysis unit 132 may determine, in calculating the final VF, whether VF(i) will not exceed the maximum value even if the correlation rank increases. The FPO analysis unit 132 may terminate the process of calculating VF(i) if it determines that VF(i) will not exceed the maximum value even if the correlation rank increases.

[0059] Figure 6C is a hypothetical diagram illustrating an example of the method for calculating the final VF according to this embodiment. The FPO analysis unit 132 calculates the final VF for each element (i) within an FPO module. Specifically, the FPO analysis unit 132 designates each element (i) within the FPO module as a focus element and calculates VF(i) for that focus element. Figure 6C shows the calculation result of VF(i) for one element (i) (focus element). In Figure 6C, the horizontal axis represents the correlation rank, and the vertical axis represents VF(i). The FPO analysis unit 132 determines that VF(i) does not exceed the maximum value when the correlation rank is 8, and terminates the process of calculating VF(i). The reason why the calculation of VF(i) can be terminated is described below.

[0060]

[0061] In equation (3) above, VF is the element F score. n is the number of elements in the target module. edge is the number of edges in the target module. degree is the degree of the element. From equation (3) above, even if all the elements in a module are connected to the elements of other modules, edge = n-1, so we can say that edge ≤ n-1. At the time of calculating the final VF, the number of elements in the module n is a fixed value, so we can find the upper limit of the achievable final VF using only the degree.

[0062]

[0063] In equation (4) above, VF is the element F score. a The maximum VF value is defined as the point at correlation rank a. n is the number of elements in the target module. degree is the degree number of the element. In the example in Figure 6C, n=5, and at the point at correlation rank a=5, VF a Since this becomes = 0.67, VF aIn order to exceed this, d ≤ 8 is required by equation (4) above. Here, the value of degree depends on the correlation rank, and when the correlation rank is "9", the value of degree is 9. Therefore, after calculating the case for correlation rank "8", it is unnecessary to calculate for correlation ranks of "9" or higher. In the example in Figure 6C, when the correlation rank is "6", there are 3 connections to elements inside the module and 3 connections to elements outside the module, the degree number is "6 = 3 + 3", the edge number is 3, the number of elements in the module (n) is 5, and VF(i) is "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (6 + 5 - 1) = 0.60" from equation (2) above. In the case of a correlation rank of "7", there are 3 connections to elements within the module and 4 connections to elements outside the module, the degree number is "7 = 4 + 3", the edge number is 3, and the number of elements in the module (n) is 5, and VF(i) is given by equation (2) above as "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (7 + 5 - 1) = 0.54". In the case of a correlation rank of "8", there are 3 connections to elements within the module and 5 connections to elements outside the module. The degree count is "8 = 5 + 3", the edge count is 3, and the number of elements in the module (n) is 5. Therefore, VF(i) is calculated from equation (2) above as "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (8 + 5 - 1) = 0.50". Thus, as determined using equation (4), even when elements are added in order of highest correlation rank, if the element is an element outside the module, VF(i) will not exceed 0.67 (in the case of a correlation rank of "5"). By configuring it in this way, the amount of computation can be reduced and processing time can be shortened compared to calculating VF(i) for all correlation ranks.

[0064] The FPO analysis unit 132 calculates MF (referred to as the final MF) for one FPO module based on the final VF of each element (i) in the FPO module. Specifically, the FPO analysis unit 132 acquires edge(i) and degree(i) when the final VF is obtained for each element (i) in the FPO module, and calculates MF (final MF) from the above formula (1) using edge(i) and degree(i) of each acquired element (i).

[0065] Note that the method for calculating the final VF and the final MF may be any method for calculating VF(i) and MF as long as it does not depend on the seed element in the module for which the final VF and the final MF are calculated. For example, when the element P value described later is the smallest for a certain element (i) in the module with respect to other elements in the module, VF(i) may be used as the final VF.

[0066] (Step S6-2) The P value calculation unit 137 calculates the module P value for the FPO module of each seed element by the following formula.

[0067]

[0068] N is the total number of all elements of the multivariate data which is the input data. n is the number of elements of the module. edge avg is the integer closest to the average value of edge(i) of each element (i) used for calculating the final MF. degree avg is the integer closest to the average value of degree(i) of each element (i) used for calculating the final MF. q and r are positive integers, and r ≤ q. Note that instead of edge avg the number of edges in the module connecting elements having a correlation coefficient equal to or greater than the correlation threshold when the maximum value of MF of the FPO module is obtained may be used, and instead of degree avg the degree when the maximum value of MF of the FPO module is obtained may be used.

[0069] The module P-value calculated for each seed element's FPO module is stored in evaluation value information 126 as the module P-value after FPO analysis. The module P-value has the same meaning as probability used in statistical processing and represents the probability of an event occurring that is rarer than the observed event. The module P-value allows for testing whether the formed module is correct or not. The module P-value calculated by equation (5) above is obtained from the product of the probability of giving the same MD and the probability of giving the same MS when the modules are constructed randomly. The module P-value after FPO analysis allows the operator to evaluate the significance of the FPO module of each seed element. This adds credibility to the module.

[0070] (Step S7) The FPO module integration unit 134 performs a module integration process (modularization process) on the FPO module of each seed element. Each FPO module is configured separately for each seed element by steps S3-S5 described above. As a result, there are as many FPO modules as there are elements, and there is overlap in elements between the FPO modules. This overlap in elements is resolved by the module integration process. The module integration process is described below.

[0071] The FPO module integration unit 134 reconfigures the modules for each VF threshold within a predetermined range of the VF threshold (from the minimum value to the maximum value of the VF threshold). The VF threshold is set in advance with a constant step size from the minimum value to the maximum value of the VF threshold. As a module reconfiguration method, all existing edges are deleted for all FPO modules, and elements that are the final VFs above the VF threshold are reconnected with edges. At this time, if there are overlapping elements between FPO modules, the FPO modules are also integrated through the overlapping elements. This provides the integration result of the FPO modules for each VF threshold. The FPO module integration unit 134 stores the integration result of the FPO modules for each VF threshold in the module integration information 124. Note that the VF threshold is not limited to being set in a constant step size from the minimum value to the maximum value of the VF threshold. For example, all possible values ​​for MF and VF may be listed as the VF threshold, and integration may be performed for all MF and VF. In this case, the VF threshold does not need to be in fixed increments.

[0072] Furthermore, the FPO module integration unit 134 may choose not to perform the integration process on all FPO modules configured in steps S3-S5, but only on some of them. For example, based on the final VF, final MF, and module P value of each FPO module stored in the evaluation value information 126, FPO modules that do not meet predetermined conditions may be excluded from the integration process. For example, when the VF threshold is set from the minimum to the maximum value of the VF threshold, FPO modules whose final MF is less than the VF threshold, and FPO modules whose seed element's final VF is less than the VF threshold may be excluded from the integration process. In this way, by excluding FPO modules with low quality or significance from the integration process based on evaluation values, it is possible to improve the quality or significance of the modules being integrated.

[0073] Figure 7 is a hypothetical diagram illustrating an example of the integration process according to this embodiment. Figure 7A is a diagram illustrating the integration of modules in the modularization process, where the final VF is used to remove edges between elements. As an example, Figure 7A shows three FPO modules M7-1, M7-2, and M7-3, with elements indicated by circles, element identifiers "c001" to "c011", and the final MF of each FPO module and the final VF of each element calculated in step S6-1. In this example, element "c002" overlaps in FPO modules M7-1 and M7-2, and element "c008" overlaps in FPO modules M7-2 and M7-3. When the VF threshold is set to 0.5, the integration process constructs module M7-4 containing all elements. When the VF threshold is set to 0.75, the final VF of element "c002" in module M7-2 and element "c011" in module M7-3 are below the VF threshold and are therefore not connected by an edge, and are excluded by the integration process. As a result, modules M7-5 and M7-6 are formed by the integration process. When the VF threshold is set to 0.85, performing the integration process similarly results in the formation of modules M7-7, M7-8, and M7-9. Here, modules M7-8 and M7-9 have a small number of elements, 2 and 1 respectively, and are evaluated low. Furthermore, they do not show significant connections between multiple elements, which hinders the analysis of the FPO module integration results by the operator.

[0074] On the other hand, Figure 7B is an illustrative diagram of the modularization process, in which modules are integrated by excluding the components of each module using the final MF and the final VF of the seed elements, and by excluding the edges between elements using the final VF. Figure 7B shows the integration process when the VF threshold is set from the minimum to the maximum value of the VF threshold, and FPO modules whose final MF is less than the VF threshold, and FPO modules whose final VF of the seed elements is less than the VF threshold are excluded from the integration process. When the VF threshold is set to 0.85, modules M7-2 and M7-3, whose final MF is less than 0.85, are excluded from the integration process, and when the integration process is performed, only module M7-7 is formed. In this way, the integration process can prevent the formation of modules with low evaluation values, thereby improving the quality or significance of the formed modules.

[0075] The FPO module integration unit 134 stores the integration result (FPO module group) of FPO modules at the VF threshold, which is the maximum number of FPO modules that fit within a predetermined range of module size, in the final module integration information 125 as the final FPO module integration result.

[0076] Conventionally, the final integration result of the FPO module alone did not provide an overview of the positional relationships of the elements, making it difficult to analyze the internal structure of the FPO module. However, according to this embodiment, the integration result of the FPO module for each VF threshold is stored in the module integration information 124, so the operator can see how the integration result of the FPO module changes in accordance with changes in the VF threshold. This allows for an overview of the positional relationships of the elements in the final integration result of the FPO module, contributing to the analysis of the internal structure of the FPO module.

[0077] Figures 8 and 9 show examples of the final FPO module integration results according to this embodiment. Figures 8 and 9 show the results of FPO analysis performed using a dataset showing the social relationships of 34 members belonging to a karate club at a university in the United States, with Pearson's product-moment correlation coefficient, MIN value "5", and MAX value "34". Components are shown as circles, and each separated member is distinguished by the pattern inside the circle. The numbers are the identification numbers of the 34 members of the karate club. The solid lines show the relationship of connections obtained through modularization processing. Figure 8 shows the network depiction results for each obtained module. Figure 9A shows examples of FPO module integration results for each VF threshold corresponding to the final FPO module integration result. In the example of Figure 9A, the graph shows how the FPO module integration result (modularization of each element (vertical axis)) changes according to the change in the VF threshold (horizontal axis). For example, the VF threshold (VF th When the value is 0.81, the relationship between modules M8-1, M8-2, and M8-3 shown in Figure 8 can be understood from Figure 9A.

[0078] The output data generation unit 138 may use the module integration information 124 for each VF threshold to generate data (e.g., a graph) that shows the correspondence between the VF thresholds (horizontal axis) and the integration results of the FPO modules (modularization of each element (vertical axis)) as illustrated in Figure 9. This graph data can be displayed on a display screen or printed. This allows the operator to easily analyze the final integration results of the FPO modules.

[0079] The FPO module integration unit 134 may further calculate a distance index d, which indicates the final distance between the FPO modules after module integration, based on the integration results of the FPO modules for each VF threshold stored in the module integration information 124. The distance index is a positive real number and is calculated for all pairs of FPO modules. The lower the value of the distance index d, the closer the distance between the two modules is determined to be. By calculating the distance index d, the overall positional relationship between the FPO modules can be easily and quantitatively evaluated.

[0080] Figure 9B is a diagram illustrating an example of a method for calculating the distance index d of the final FPO module according to this embodiment. Figure 9B shows one method for calculating the distance index d based on the integration result of the final FPO module exemplified in Figure 9A. In Figure 9B, the VF shown in Figure 9A th For the three modules M8-1, M8-2, and M8-3 separated at a VF threshold of 0.81, we consider virtual modules M9-1, M9-2, and M9-3, which further include information about the structure of the modules when the VF threshold is increased. Point G in the graph is defined as the center point of the separated virtual modules M9-1, M9-2, and M9-3. Regarding the distance between the obtained FPO modules, by using the difference in VF thresholds and the number of components to calculate the distance to the center point of the virtual module, we can more accurately grasp the overall positional relationship of the elements in the integrated result of the FPO modules.

[0081] For example, when calculating the distance index between modules M9-1 and M9-2, the following steps (1) to (3) are performed: (1) The minimum VF threshold at which the components of modules M8-1 and M8-2 are divided by the module integration process, VF th,min Search. In this example, VF th,min = 0.76. (2) The components of virtual modules M9-1 and M9-2 set the VF threshold to VF th When set to the above, the average value of the minimum VF threshold that is excluded by the module integration process is VF. th, minavg The average of the minimum VF thresholds at which the 16 components are excluded is calculated. For example, in the case of virtual module M9-1, elements "23", "21", "19", "16", and "15" are excluded if the VF threshold is 0.95 or higher, elements "33", "31", "30", and "24" are excluded if the VF threshold is 0.92 or higher, elements "34", "29", "28", "27", and "10" are excluded if the VF threshold is 0.91 or higher, element "9" is excluded if the VF threshold is 0.90 or higher, and element "32" is excluded if the VF threshold is 0.86 or higher. Therefore, the average of the minimum VF thresholds at which the 16 components are excluded is calculated, VF th, minavg It is calculated using (M8-1). Specifically, VF th, minavg(M8-1) = (0.95 × 5 + 0.92 × 4 + 0.91 × 5 + 0.90 × 1 + 0.86 × 1) / 16 = 0.9212. Similarly, in the case of virtual module M9-2, elements "8", "4", "22", "2", and "18" are excluded if the VF threshold is 0.90 or higher, elements "1", "13", and "14" are excluded if the VF threshold is 0.88 or higher, element "20" is excluded if the VF threshold is 0.87 or higher, and element "3" is excluded if the VF threshold is 0.86 or higher. Thus, the average value of the minimum VF threshold at which 10 components are excluded is VF. th, minavg It is calculated using (M8-2). Specifically, VF th, minavg (M8-2) can be calculated as (0.9 × 5 + 0.88 × 3 + 0.87 × 1 + 0.86 × 1) / 10 = 0.8870. (3) The distance index d(M8-1, M8-2) is calculated by the following formula (6).

[0082]

[0083] In equation (6) above, d represents the distance index between modules. VF th This indicates the VF threshold. th,min This indicates the minimum VF threshold at which the components of modules M8-1 and M8-2 are separated by the module integration process. th, minavg This represents the average of the minimum VF thresholds at which components of both modules are excluded. In this example, d(M8-1, M8-2) = 2(0.81-0.76) + (0.8870-0.81) / 2 + (0.9212-0.81) / 2 = 0.1000 + 0.0385 + 0.0556 = 0.1941.

[0084] Applying the same calculation to virtual modules M9-2 and M9-3, we obtain d(M8-2, M8-3) = 0.0585 and d(M8-1, M8-3) = 0.1756. In the case of d(M8-2, M8-3), in the case of virtual module M9-3, elements "7", "6", "5", "12", and "11" are excluded when the VF threshold is 0.85 or higher, so the average value of the VF threshold at which these five components are excluded is VF. th, minavg It is calculated using (M8-3). Specifically, VF th, minavg(M8-3) = 0.85 × 5 / 5 = 0.8500 can be calculated. The minimum VF threshold at which the components of modules M8-2 and M8-3 are divided by the module integration process is VF. th,min Search. In this example, VF th,min = 0.81. In this example, d(M8-2, M8-3) = 2(0.81 - 0.81) + (0.8870 - 0.81) / 2 + (0.8500 - 0.81) / 2 = 0.0385 + 0.0200 = 0.0585 can be calculated. Also, in the case of d(M8-1, M8-3), the minimum VF threshold at which the components of modules M8-1 and M8-3 are separated by the module integration process, VF th,min Search. In this example, VF th,min = 0.76. In this example, d(M8-1, M8-3) = 2(0.81-0.76) + (0.9212-0.81) / 2 + (0.8870-0.81) / 2 = 0.1000 + 0.0556 + 0.0200 = 0.1756. From the above, it can be determined that M8-2 is closer to M8-1 than M8-3. In this way, the overall positional relationship of the elements in the final integrated result of the FPO module can be grasped more accurately.

[0085] Figure 9B shows the distance index calculated based on the final FPO module integration results illustrated in Figure 9A, expressed as a numerical value. In module M8-1, distance 0.1056 represents the average distance of each element from the minimum VF threshold at which module M8-2 is divided to the minimum VF threshold at which elements constituting module M8-1 are excluded. In module M8-1, distance 0.0556 represents the average distance of each element from the minimum VF threshold at which module M8-3 is divided to the minimum VF threshold at which elements constituting module M8-1 are excluded. In module M8-2, distance 0.0885 represents the average distance of each element from the minimum VF threshold at which module M8-1 is divided to the minimum VF threshold at which elements constituting module M8-2 are excluded. In module M8-2, distance 0.0385 represents the average distance of each element from the minimum VF threshold at which module M8-3 is divided to the minimum VF threshold at which elements constituting module M8-2 are excluded. In module M8-3, distance 0.0700 represents the average distance from the minimum VF threshold at which module M8-1 is divided to the minimum VF threshold at which the elements constituting module M8-3 are excluded. In module M8-3, distance 0.0200 represents the average distance of each element from the minimum VF threshold at which module M8-2 is divided to the minimum VF threshold at which the elements constituting module M8-3 are excluded.

[0086] (Step S8-1) The FPO module integration unit 134 calculates the final VF and final MF for each of the final FPO modules after module integration, similar to the FPO analysis unit 132 in step S6-1 above. The final VF and final MF calculated for each of the final FPO modules after module integration are stored in the evaluation value information 126.

[0087] (Step S8-2) The P-value calculation unit 137 calculates the module P-value for each of the final FPO modules after module integration using the above formula (5). The module P-values ​​calculated for each of the final FPO modules after module integration are stored in the evaluation value information 126 as module P-values ​​after module integration. By using the module P-values ​​after module integration, the operator can evaluate the significance of each of the final FPO modules after module integration. This makes it possible to add credibility to the modules.

[0088] (Step S9) The FNI analysis unit 135 targets any or all of the final FPO modules (final FPO modules) after module integration and performs an In process (FNI analysis process) to incorporate false-negative elements.

[0089] The final FPO modules to be subjected to FNI analysis processing may, for example, be specified by the user as multiple modules. Alternatively, the conditions for the final FPO modules to be subjected to FNI analysis processing (FNI analysis target conditions) may be set in advance in the information processing device 1, and the FNI analysis unit 135 may exclude final FPO modules that do not satisfy the FNI analysis target conditions and decide to subject the remaining final FPO modules to FNI analysis processing. The FNI analysis target conditions are, for example, modules that have a certain significance level (module P value) or higher. The FNI analysis processing will be described below.

[0090] Figure 10 is a hypothetical diagram illustrating the FNI analysis process according to this embodiment. In Figure 10, elements are represented by circles. Elements include elements of interest, candidate elements, elements within the target module, and elements outside the target module. Elements within the target module are those included in the final FPO module (final FPO module) after module integration, while elements outside the target module are those not included in the final FPO module (final FPO module) after module integration. The numbers shown inside the circles represent the correlation rank with respect to the candidate elements. Each element is connected by a line. Solid lines indicate the relationships between elements within the module, including the elements of interest. Dotted lines indicate the relationships between elements outside the module, including the candidate elements. The FNI analysis unit 135 repeatedly performs the FNI analysis on each element of interest, sequentially setting each of the elements included in all the final FPO modules targeted by the FNI analysis process as an element of interest. Therefore, the FNI analysis process is executed a number of times equal to the total number of elements included in all the final FPO modules targeted by the FNI analysis process. Here, we will refer to Figure 10 and explain using one of the elements of interest shown in Figure 10 as an example.

[0091] The FNI analysis unit 135 sets candidate elements to be included in the final FPO module (target module) to which the element of interest belongs. For example, the candidate elements to be included in the target module may be all elements that do not belong to the target module (i.e., all elements outside the target module). For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation rank with respect to the element of interest is up to a predetermined high rank. For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation coefficient with respect to the element of interest is above a predetermined correlation threshold. For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation rank with respect to the element of interest is up to a predetermined high rank AND whose correlation coefficient with respect to the element of interest is above a predetermined correlation threshold. For example, the candidate elements to be included in the target module may be limited to elements that do not belong to any final FPO module. The FNI analysis unit 135 performs a false negative determination for each candidate element with respect to the target module.

[0092] Specifically, the FNI analysis unit 135 refers to the candidate element i in order from the top (ranking from the highest correlation (correlation rank)) of the correlation coefficient list 122, and calculates the VS(i) of the above formula (2) for the target module of the candidate element i, "VS(i) = (number of edges in the target module of the candidate element i edge(i)) ÷ (number of degrees of the candidate element i degree(i))" for each rank. The FNI analysis unit 135 records information such as the correlation rank for all calculated VS(i) that exceed a predetermined VS threshold in the storage unit 12. The VS threshold can be arbitrarily set by the operator, for example. If the FNI analysis unit 135 calculates a VS(i) that exceeds a predetermined VS threshold, it determines that the candidate element is a false negative element with respect to the element of interest and incorporates the candidate element into the target module.

[0093] (Step S10) The P-value calculation unit 137 calculates the element P-value for the target module in the correlation rank corresponding to VS(i) that exceeds the VS threshold using the following formula.

[0094]

[0095] In equation (7) above, N is the total number of elements in the multivariate input data. n is the number of elements in the target module. edge is the number of edges (i) in the target module for candidate element i. degree is the degree (i) of candidate element i. q and r are positive integers such that r ≤ q.

[0096] Alternatively, the element P-value may be calculated using the recursive formula in equation (8) while calculating VS(i) for each correlation rank in step S9.

[0097]

[0098] In equation (8) above, N is the total number of elements in the multivariate data which is the input data. n is the number of elements in the module. edge is the number of edges (i) in the target module for candidate element i. degree is the degree number (i) of candidate element i. By calculating the element P value using the recursive formula in equation (8) while calculating VS (i) in step S9, the calculation of the number of combinations can be omitted, and the efficiency of processing can be improved.

[0099] Figure 11 shows the time required to perform FPO analysis using microarray data of Arabidopsis thaliana (dataset of 27127 x 9442 samples) with Pearson's product-moment correlation coefficient, MIN value "2", and MAX value "250", followed by FNI analysis (including steps S9 to S13-2 (described later)). As per the present invention, by executing steps S9 and S10 simultaneously and employing a recursive formula for calculating element P values, the analysis time was significantly reduced to 1 / 92.

[0100] The P-value calculation unit 137 stores information regarding the correlation rank with the smallest element P-value among the correlation ranks corresponding to VS(i) that exceed the VS threshold in the evaluation value information 126 for FNI analysis. The information regarding the correlation rank with the smallest element P-value is the information regarding the correlation rank that can be determined to be the most significant.

[0101] (Step S11) If FNI analysis processing is completed for all elements of the final FPO module targeted for FNI analysis processing (Step S11, YES), proceed to Step S12. On the other hand, if there are still elements for which FNI analysis processing has not been performed (Step S11, NO), return to Step S9 and perform FNI analysis processing on the elements for which FNI analysis processing has not been performed, with those elements as the elements of interest.

[0102] (Step S12) The FNI assignment processing unit 136 performs an assignment process to resolve duplicates for candidate elements that overlap in multiple final FPO modules generated by the FNI analysis process. The following assignment criteria are used in the assignment process.

[0103] Assignment Criteria: For multiple target modules where candidate element i overlaps, the following priority order is used: VS(i) > "Module P value of equation (5) above when candidate element i is incorporated into the target module" > Number of elements in the module. VS(i) has the highest priority, followed by module P value, and then the number of elements in the module.

[0104] First, candidate element i is retained only in the target module with the highest VS(i), and removed from all other target modules. On the other hand, if the VS(i) values ​​are equivalent, candidate element i is retained only in the target module with the smallest "module P value in equation (5) above when candidate element i is incorporated into the target module," and removed from all other target modules. If the VS(i) values ​​are equivalent and the "module P value in equation (5) above when candidate element i is incorporated into the target module" is equivalent among the target modules, candidate element i is retained only in the target module with the largest number of elements, and removed from all other target modules.

[0105] Furthermore, the assignment process is not limited to the assignment criteria described above, as long as it can resolve the duplication of elements between multiple modules that occurred during the FNI analysis process. For example, modules with duplicate elements may be merged through a module integration process (modularization process).

[0106] The FNI assignment processing unit 136 stores information about the final FPO module after the assignment process in the final module information 127. The final module information 127 may be individual module files, a file containing all modules, or a combination of individual module files and a file containing all modules. Information about the final FPO module before the assignment process may also be recorded in the storage unit 12.

[0107] According to this embodiment, if duplicate candidate elements occur in multiple final FPO modules as a result of performing FNI analysis on multiple final FPO modules, the duplication of candidate elements can be resolved. This eliminates the need for operators to individually specify the final FPO modules on which to perform FNI analysis, allowing them to perform FNI analysis on multiple final FPO modules simultaneously and obtain FNI analysis results without element duplication.

[0108] (Step S13-1) The FNI assignment processing unit 136 calculates the final VF and final MF for each of the final FPO modules after the assignment process, similar to the FPO analysis unit 132 in step S6-1 above. The final VF and final MF calculated for each of the final FPO modules after the assignment process are stored in the evaluation value information 126.

[0109] (Step S13-2) The P-value calculation unit 137 calculates a module P-value for each of the final FPO modules after the assignment process using the above formula (5). The module P-values ​​calculated for each of the final FPO modules after the assignment process are stored in the evaluation value information 126 as module P-values ​​after the assignment process. By using the module P-values ​​after the assignment process, the worker can evaluate the significance of each of the final FPO modules after the assignment process. This makes it possible to add credibility to the modules.

[0110] (Step S14) The output data generation unit 138 generates various types of output data. The output data is data that can be displayed on a display screen or printed. For example, the output data generation unit 138 generates drawing data that represents the configuration of each final FPO module after the assignment process. This drawing data visualizes the configuration of each final FPO module after the assignment process by drawing it.

[0111] For example, the output data generation unit 138 may generate graph data (such as the graph data exemplified in Figure 5A) showing the MF corresponding to each rank for each seed element obtained by the FPO analysis process. For example, the output data generation unit 138 may generate graph data (such as the graph data exemplified in Figure 9) showing the integration results by VF threshold obtained by the FPO module integration process.

[0112] The output data generated by the output data generation unit 138 may be output to an external device of the information processing device 1 by the input / output unit 11, or it may be stored in the storage unit 12. The output data stored in the storage unit 12 can be output to an external device by the input / output unit 11.

[0113] According to this embodiment, for example, it is possible to improve the accuracy of correlation network analysis on multivariate data such as comprehensive molecular information obtained by omics analysis or big data.

[0114] Although embodiments of the present invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments, and design modifications and the like are also included within the scope of the gist of the present invention.

[0115] Furthermore, computer programs for realizing the functions of each of the above-mentioned devices may be recorded on a computer-readable recording medium, and the programs recorded on this recording medium may be loaded into a computer system and executed. The term "computer system" here may include hardware such as an operating system and peripheral devices. Also, if a WWW system is used, the "computer system" shall also include the homepage provisioning environment (or display environment). Furthermore, "computer-readable recording medium" refers to writable non-volatile memory such as flexible disks, magneto-optical disks, ROMs, and flash memory, portable media such as DVDs (Digital Versatile Discs), and storage devices such as hard disks built into a computer system.

[0116] Furthermore, "computer-readable recording media" includes volatile memory (e.g., DRAM) within a computer system that acts as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, which retains the program for a certain period of time. The program may also be transmitted from the computer system storing it in a memory device to another computer system via a transmission medium or by transmission waves within the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be intended to implement only a part of the aforementioned functions. Furthermore, it may be a program that can implement the aforementioned functions in combination with a program already recorded in the computer system, a so-called differential file (differential program).

[0117] 1... Information processing device, 11... Input / output unit, 12... Storage unit, 121... Correlation network analysis program, 122... Correlation coefficient list, 123... Module information, 124... Module integration information, 125... Final module integration information, 126... Evaluation value information, 127... Final module information, 128... FPO analyzed module list, 13... Control unit, 131... Correlation coefficient calculation unit, 132... FPO analysis unit, 133... FPO module selection unit, 134... FPO module integration unit, 135... FNI analysis unit, 136... FNI assignment processing unit, 137... P-value calculation unit, 138... Output data generation unit

Claims

1. A program that, when constructing a module of seed elements from non-seed elements with high correlation to seed elements among the elements of multivariate data to be analyzed for correlation network, causes the computer to perform an FPO analysis process to remove false positive non-seed elements from among the non-seed elements with high correlation to seed elements, increasing the rank from a lower limit to an upper limit, which is the number of non-seed elements to be incorporated into the seed element module, for each rank, stores the configuration information of the module on which the FPO analysis process has been performed in a storage unit, and if a module identical to the module configuration information stored in the storage unit is constructed, stops the FPO analysis process and terminates the search for the module showing the highest module F score for that rank early.

2. The program according to claim 1, wherein the computer stores each of the following in its storage unit: a correlation coefficient list, module information for each rank by element, module F score and module integration information for each element F score threshold, final module integration information, evaluation value information for each target step, final module information, and FPO-analyzed module list; calculates and adds a hash value to each of the pieces of information stored in the storage unit; and identifies each piece of information by the hash value.

3. The program according to claim 1, wherein the computer is instructed to select the module of the rank that shows the maximum value during the period in which the best value of the module F score is not updated even when the rank is increased, with respect to the module F score in each rank from the lower limit to the upper limit of the rank, as the optimal module of the FPO analysis result, and the FPO analysis process and module selection are terminated when predetermined conditions are met.

4. The program according to claim 3, wherein the conditions for terminating the FPO analysis process and module selection are a comprehensive criterion based on the number of ranks for which the module F score has not been calculated, the maximum period during which the best value of the module F score has not been updated, and the period during which the module F score has not been updated at the current rank.

5. The program according to claim 1, wherein the computer is instructed to calculate an element F score, which is an index for evaluating the components of the module, for each of the plurality of elements constituting the module; to calculate a module F score, which is an index for evaluating the entire module, based on the maximum value of the element F scores; to determine whether the element F score will not exceed the maximum value even if the correlation rank is increased, and if it is determined that the element F score will not exceed the maximum value even if the correlation rank is increased, the program is instructed to terminate the process of calculating the element F score.

6. The program according to claim 5, wherein the criterion for determining whether or not the maximum value of the element F score is exceeded is a comprehensive criterion based on the number of elements in the module and the elements inside and outside the module.

7. The program according to claim 1, wherein, when the computer reconstructs modules for each seed element configured with each element of the multivariate data as a seed element, it associates elements whose element F scores are equal to or greater than the element F score threshold and integrates modules through overlapping elements; calculates an element F score, which is an index for evaluating the components of a module, for each of the multiple elements constituting the integrated module; calculates a module F score, which is an index for evaluating the entire module, based on the maximum value of the element F scores; determines whether the element F score will not exceed the maximum value even if the correlation rank is increased at the maximum value of the element F score; and terminates the process of calculating the element F score if it is determined that the element F score will not exceed the maximum value even if the correlation rank is increased.

8. The program according to claim 7, wherein the criterion for determining whether or not the maximum value of the element F score is exceeded is a comprehensive criterion based on the number of elements in the module and the elements inside and outside the module.

9. The computer is instructed to perform FNI analysis on multiple modules composed of elements from the multivariate data to be analyzed for correlation network analysis, incorporating false negative elements. Simultaneously with the FNI analysis, the computer is instructed to calculate the element p-value for the element of interest using the following formula, while calculating the element specificity rate for each correlation rank of candidate elements that are elements outside the module. In this program, N is the total number of elements in the multivariate data, n is the number of elements in the target module, edge is the number of edges of the candidate element within the target module, and degree is the degree of the candidate element.

10. The program according to claim 9, wherein the computer stores each of the following in its storage unit: a correlation coefficient list, module information for each rank by element, module F score and module integration information for each element F score threshold, final module integration information, evaluation value information for each target step, final module information, and FPO-analyzed module list; calculates and adds a hash value to each of the pieces of information stored in the storage unit; and identifies each piece of information by the hash value.

11. The program according to claim 9, wherein the computer is caused to perform an assignment process to resolve duplication of elements that overlap in multiple modules; to calculate an element F score, which is an index for evaluating the components of a module, for each of the multiple elements constituting the module on which the assignment process has been performed; to calculate a module F score, which is an index for evaluating the entire module, based on the maximum value of the element F scores; to determine whether the element F score will not exceed the maximum value even if the correlation rank is increased, and if it is determined that the element F score will not exceed the maximum value even if the correlation rank is increased, the process of calculating the element F score is terminated.

12. The program according to claim 11, wherein the criterion for determining whether or not the maximum value of the element F score is exceeded is a comprehensive criterion based on the number of elements in the module and the elements inside and outside the module.