Information processing device, information processing method, computer program, and computer-readable recording medium
The integration of false negative elements and duplication resolution in correlation network analysis enhances the accuracy and reliability of module identification in multivariate data analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-02
AI Technical Summary
Existing techniques for correlation network analysis of multivariate data suffer from insufficient analytical accuracy.
Incorporation of false negative elements into modules through FNI analysis and resolution of element duplication using FNI assignment processing to enhance the accuracy of correlation network analysis.
Improves the accuracy of correlation network analysis by reducing false positives and duplicates, leading to more robust and reliable module identification.
Smart Images

Figure JP2024034979_02042026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, computer program, and computer-readable recording medium.
[0001] The present invention relates to an information processing device, an information processing method, a computer program, and a computer-readable recording medium for analyzing the relationships between multiple influencing components, identifying modules that have a high influence on a particular effect, and extracting candidates.
[0002] Patent Document 1 describes a technique for performing correlation network analysis on multivariate data, which is comprehensive molecular information obtained through omics analysis. Non-Patent Document 1 describes a technique for performing correlation network analysis on multivariate data such as big data.
[0003] Patent No. 6318334
[0004] Itto Mannen, Yoshiyuki Ogata, and Hideyuki Suzuki, "A groundbreaking development for multivariate analysis: ConfeitoGUIplus," Journal of the Japan Society for Biotechnology, Vol. 95, No. 7, July 2017.
[0005] However, the techniques described in Patent Document 1 and Non-Patent Document 1 had insufficient analytical accuracy.
[0006] The objective of this invention is to improve the accuracy of correlation network analysis for multivariate data.
[0007] One aspect of the present invention is an information processing device comprising: an FNI analysis unit that performs FNI analysis processing to incorporate false negative elements into each of a plurality of modules composed of elements of multivariate data to be subjected to correlation network analysis; and an FNI assignment processing unit that performs assignment processing to resolve duplication for elements that overlap in the plurality of modules generated by the FNI analysis processing.
[0008] One aspect of the present invention is an information processing method performed by an information processing device, comprising: an FNI analysis step of performing an FNI analysis process to incorporate false negative elements into each of a plurality of modules composed of elements of multivariate data to be subjected to correlation network analysis; and an FNI assignment step of performing an assignment process to resolve duplication of elements that overlap in the plurality of modules generated by the FNI analysis process.
[0009] One aspect of the present invention is a computer program that causes a computer to perform an FNI analysis step, which involves performing an FNI analysis process to incorporate false negative elements into each of a plurality of modules composed of elements of multivariate data to be analyzed for correlation network analysis, and an FNI assignment step, which involves performing an assignment process to resolve duplicates in elements that overlap in the plurality of modules generated by the FNI analysis process.
[0010] One aspect of the present invention is a computer-readable recording medium that stores a computer program causing a computer to perform an FNI analysis step, which involves performing an FNI analysis process to incorporate false negative elements into each of a plurality of modules composed of elements of multivariate data to be analyzed for correlation network analysis, and an FNI assignment step, which involves performing an assignment process to resolve duplication for elements that overlap in the plurality of modules generated by the FNI analysis process.
[0011] According to the present invention, the effect is to improve the accuracy of correlation network analysis for multivariate data.
[0012] This figure shows a schematic configuration example of an information processing device according to one embodiment. This flowchart shows an example of the procedure for an information processing method according to one embodiment. This figure shows an example of the configuration of multivariate data according to one embodiment. This figure illustrates an example of the correlation coefficient between elements according to one embodiment. This figure illustrates an example of the correlation coefficient between elements according to one embodiment. This figure illustrates an example of the FPO module selection method according to one embodiment. This figure illustrates an example of the FPO module selection method according to one embodiment. This figure illustrates an example of the FPO analysis according to one embodiment. This figure illustrates an example of the final element F score according to one embodiment. This figure illustrates an example of the final element F score according to one embodiment. This figure illustrates an example of the calculation method for the final element F score according to one embodiment. This figure illustrates an example of the integration process according to one embodiment. This figure illustrates an example of the integration result of the final FPO module according to one embodiment. This figure illustrates an example of the integration result of the FPO module corresponding to the final FPO module integration result in Figure 8. This figure illustrates an example of the calculation of the distance index of the final FPO module in Figure 8. This figure illustrates the FNI analysis process according to one embodiment. This figure shows an example of the results of comparing the analysis time of element P values with and without a recursive formula.
[0013] Next, the information processing apparatus, information processing method, and computer program of this embodiment will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below. In all the drawings used to describe the embodiments, components having the same function will be given the same reference numerals, and repeated explanations will be omitted. Furthermore, in this application, "based on XX" means "based on at least XX," and includes cases where it is based on another element in addition to XX. Furthermore, "based on XX" is not limited to cases where XX is used directly, but also includes cases where it is based on XX after calculations or processing have been performed on it. "XX" is any element (for example, any information).
[0014] The terms used in this embodiment are explained in Tables 1, 2, and 3 below.
[0015]
[0016]
[0017]
[0018] Figure 1 is a diagram showing a schematic configuration example of the information processing device 1 according to this embodiment. In Figure 1, the information processing device 1 comprises an input / output unit 11, a storage unit 12, and a control unit 13.
[0019] The input / output unit 11 receives data input and outputs data to and from an external device of the information processing device 1. The method of data input and output is not limited. The input / output unit 11 may send and receive data by communication, for example, via a communication line. The input / output unit 11 may also input and output data, for example, via a recording medium.
[0020] The storage unit 12 stores various types of data, such as programs executed by the control unit 13 and data used by the control unit 13. The storage unit 12 is composed of non-volatile memory such as a hard disk drive, magneto-optical disk drive, or flash memory, a read-only recording medium such as a CD-ROM, a volatile memory such as DRAM (Dynamic Random Access Memory), or a combination thereof.
[0021] The control unit 13 is equipped with a CPU (Central Processing Unit) and realizes various functions by executing programs stored in the storage unit 12. In this embodiment, the control unit 13 executes the correlation network analysis program 121 stored in the storage unit 12, thereby realizing the functions of the correlation coefficient calculation unit 131, FPO analysis unit 132, FPO module selection unit 133, FPO module integration unit 134, FNI analysis unit 135, FNI assignment processing unit 136, P-value calculation unit 137, and output data generation unit 138.
[0022] The information processing device 1 may be configured using a general-purpose computer device, or it may be configured as a dedicated hardware device. For example, the information processing device 1 may be configured using a server computer connected to a communication network such as the Internet. Furthermore, each function of the information processing device 1 may be realized by cloud computing. In addition, the information processing device 1 may be realized by a single computer, or its functions may be realized by distributing them among multiple computers. Furthermore, the information processing device 1 may be configured to launch a website using, for example, a WWW system.
[0023] The memory unit 12 has memory areas for storing a correlation coefficient list 122, module information 123 for each rank by element, module integration information 124 for each module F score (MF) and element F score (VF) threshold, final module integration information 125, evaluation value information 126 for each target step, final module information 127, and an FPO-analyzed module list 128. Furthermore, by calculating and adding a hash value to each piece of information, the retrieval of information stored in the memory unit 12 may be made faster.
[0024] Next, the information processing method according to this embodiment will be described with reference to Figure 2. Figure 2 is a flowchart showing an example of the procedure of the information processing method according to this embodiment.
[0025] (Step S1) The input / output unit 11 receives input data. The input data is multivariate data to be analyzed using correlation network analysis. The input data may be, for example, comprehensive molecular information obtained by omics analysis. The input data may be, for example, product information obtained in the manufacturing and inspection processes of a product for the purpose of analyzing the cause of defective products. The input data may be, for example, big data. The input data may be, for example, time series data.
[0026] Note that the multivariate data to be analyzed by the correlation network analysis according to this embodiment is not particularly limited in terms of its acquisition route, acquisition method, type, attributes, etc. For example, multivariate data obtained by omics analysis can be cited. "Omics analysis" generally means integrating and analyzing individual comprehensive molecular information. Representative examples of comprehensive molecular information include transcriptome data, which is comprehensive information on gene transcripts, and metabolome data, which is comprehensive information on metabolites. These comprehensive data can be acquired by any method or means, such as various gene analyses, gene expression analyses, and various mass analyses such as liquid chromatography-mass spectrometry (LC-MS), gas chromatography-mass spectrometry (GC-MS), and capillary electrophoresis-mass spectrometry (CE-MS). Furthermore, there is no particular limitation on the acquisition source of these comprehensive data, and examples include sites, organs, tissues, and cells derived from various types of animals, plants, microorganisms, and bacteria. Furthermore, the multivariate data may be information obtained by any method from samples and artificial products (e.g., processed foods, etc.) obtained from a certain environment.
[0027] The format of the input data only needs to have a plurality of data sets consisting of quantitative values corresponding to each of the plurality of elements. FIG. 3 is a diagram showing a configuration example of the multivariate data (input data) according to this embodiment. The input data may be a quantitative value table composed of rows corresponding to each element and columns corresponding to each data set, as exemplified in FIG. 3. In the quantitative value table shown in FIG. 3, the first column is the element name (e.g., element identifier (ID001, ID002, etc. in FIG. 3)), the first row is the data set name (DATA_a, DATA_b, etc. in FIG. 3), and the quantitative value of the data set y of the element x is stored in the x-th row and y-th column (x and y are integers of 2 or more). For example, the data "data_b002" in the 3rd row and 3rd column is the quantitative value of the data set "DATA_b" of the element "ID002".
[0028] (Step S2) The correlation coefficient calculation unit 131 calculates the correlation coefficients between elements based on the quantitative values of the input data. The correlation coefficient calculation unit 131 generates a correlation coefficient list 122 for each element of the input data, sorted in descending order of correlation coefficient. As a result, a correlation coefficient list 122 is generated for each element of the input data. The correlation coefficient list 122 for element x is a list in which pairs of element name and correlation coefficient are sorted in descending order of the correlation coefficient with element x. The correlation coefficient lists 122 for each element generated by the correlation coefficient list 122 are stored in the storage unit 12.
[0029] The correlation coefficient between elements in this embodiment is, for example, one of the following indices: Pearson's product-moment correlation coefficient, Spearman's rank correlation coefficient, or cosine similarity. The correlation coefficient between elements is a real number between -1 and 1.
[0030] In this embodiment, the correlation coefficient between elements may be set to "0" for negative correlation coefficients, or it may not be set to "0". If negative correlation coefficients are not set to "0", it is possible to distinguish between "negative correlation" and "no correlation". Therefore, even if two elements that have a strong positive correlation with other elements and also have a negative correlation belong to the same module, they can be associated with each other. This makes it easier to interpret the modules.
[0031] Figures 4A and 4B are hypothetical diagrams illustrating an example of the correlation coefficient between elements according to this embodiment. Figure 4A shows an example of a module, with elements represented by circles. Elements with a correlation coefficient greater than or equal to a certain value are connected by lines. Elements include seed elements and elements within the module. The symbols shown inside the circles are the identification information of the elements. In the example shown in Figure 4A, the correlation coefficient ρ between element X and element Y is shown. XY ρ is greater than zero, and is the correlation coefficient between element X and element Z. XZ Let's consider the case where is greater than zero. In this case, the correlation coefficient ρ between element Y and element Z is ρ YZ It is unclear whether it is greater than zero or not. Figure 4B shows the correlation coefficient ρ YZ Figure 4B shows an illustrative diagram of the possible values of the correlation coefficient ρ. YZIt can be understood that it can take a value greater than zero, a value less than zero, or the prediction of whether it is greater than zero or less than zero is unclear. Note that ρ represents the correlation coefficient. FIG. 4B shows the value range (-1 to 1) of the Pearson correlation value on the X-axis and Y-axis.
[0032] (Step S3) The FPO analysis unit 132 uses each element of the input data as a seed element, and excludes (Out) false-positive non-seed elements from the non-seed elements located at the top (the ones with larger correlation coefficients) in the correlation coefficient list 122 of the seed elements (FPO analysis process). The FPO analysis process will be described below.
[0033] The FPO analysis unit 132 constructs a temporary module that connects the non-seed elements located at a predetermined top (up to the Mth place) in the correlation coefficient list 122 of the seed elements to the seed elements with edges. M is any value within a predetermined range of the module size (from the lower limit value to the upper limit value of the module size). The number of elements (excluding the seed elements) that make up the module of the seed elements is the rank (rank = M). The predetermined range of the module size can be arbitrarily set by the operator, for example, by setting the MIN value and the MAX value in the information processing apparatus 1 in advance. Here, the MIN value and the MAX value are the module sizes excluding the seed elements, and are the minimum and maximum values that the rank can take. Therefore, the lower limit value of the total number of elements (seed elements and non-seed elements) that make up the module (module size) is "MIN value + 1", and the upper limit value is "MAX value + 1".
[0034] Next, the FPO analysis unit 132 generates a list of correlation coefficients sorted in descending order for the provisional module, consisting of all elements included in the correlation coefficient list 122 of the seed element and each non-seed element. The FPO analysis unit 132 further connects elements in the generated list that have correlation coefficients greater than a predetermined correlation threshold with edges. Here, by setting multiple correlation thresholds, multiple provisional modules with different edge patterns are constructed. For example, by setting all possible values as correlation thresholds, provisional modules with edge patterns for each of the said correlation thresholds are constructed. The lower the correlation threshold, the more edges are created. For example, instead of setting multiple correlation thresholds simultaneously, the correlation threshold may be initialized to the maximum value "1", and then provisional modules may be constructed for each correlation threshold while decreasing the correlation threshold. Next, the FPO analysis unit 132 calculates the F score and MF of each provisional module with different edge patterns using the following formula.
[0035]
[0036] n is the number of elements in the module. edge(i) is the number of edges of element i, and for the edge pattern corresponding to the current correlation coefficient, it is the number of elements connected to element i within the same module. i is an integer value between 1 and n, and is the index of the element. degree(i) is the degree of element i, and for the edge pattern corresponding to the current correlation coefficient, it is the total number of elements connected to element i both inside and outside the module.
[0037] MF is the harmonic mean of module density (MD) and module specificity (MS). MD represents the degree to which modules are interconnected internally as a whole. MS represents the degree to which modules are isolated (exclusively connected) from the outside. Therefore, modules with a large MF are isolated from the outside and have strong connections as a whole.
[0038] The FPO analysis unit 132 stores the maximum value of the MF calculated for each of the multiple provisional modules with different edge patterns, as well as the correlation threshold and edge pattern corresponding to the maximum value of MF, in the storage unit 12. The FPO analysis unit 132 calculates the F score, VF(i), of each non-seed element i for the correlation threshold and edge pattern corresponding to the maximum value of MF using the following formula.
[0039]
[0040] VF(i) is the harmonic mean of element density (VD(i)) and element singularity (VS(i)). VD(i) represents the proportion of element i that is connected to other elements within the module. VS(i) represents the proportion of the degree that element i has in the entire network that is connected to elements of the module. i is an integer value between 1 and n and is the index of the element. Therefore, an element i with a large VF(i) is exclusively and densely connected to the module and can be said to be a suitable element to be a component of that module.
[0041] The FPO analysis unit 132 determines that the non-seed element i that minimizes the VF(i) of the provisional module is a false positive, and removes the false positive non-seed element i from the provisional module. Next, the FPO analysis unit 132 recalculates the maximum value of MF and the VF(i) of each non-seed element i for the provisional module from which the false positive non-seed element i has been removed. This FPO analysis process to remove false positive non-seed elements i is repeated until the number of elements in the provisional module reaches the "lower limit of module size," i.e., "MIN value + 1." The FPO analysis unit 132 determines that the point in time when MF was at its maximum represents a state where all false positives have been removed without excess or deficiency, and stores the configuration information of the provisional module at the point in time when MF was at its maximum in the module information 123 as the maximum MF module in the "rank" of the seed elements.
[0042] The FPO analysis unit 132 performs FPO analysis processing for each rank within a predetermined range of ranks (from the lower limit (MIN value + 1) to the upper limit (MAX value + 1) of the module size at the time the provisional module was configured). As a result, for each seed element, the configuration information of the provisional module at the time when the MF was at its maximum is stored in the module information 123 as the maximum MF module for all ranks within the predetermined range of ranks.
[0043] Furthermore, the FPO analysis unit 132 performs the FPO analysis in step S3 while increasing the rank from the lower limit (MIN value) to the upper limit (MAX value), and stores the configuration information of the modules for which the FPO analysis in step S3 has been completed (FPO-analyzed modules) in the FPO-analyzed module list 128. If a module identical to an FPO-analyzed module already existing in the FPO-analyzed module list 128 is constructed, the FPO analysis process to remove false-positive non-seed elements may be stopped, and the search for the largest MF module of that rank may be terminated early.
[0044] The FPO analysis unit 132 uses the correlation threshold and edge pattern to determine whether the maximum value of MF is updated or not. If it determines that the maximum value of MF is not updated in a provisional module where MF has not yet been calculated among multiple provisional modules with different edge patterns, it may terminate the MF calculation and early terminate the search for the maximum MF module of that rank. By configuring it in this way, the FPO analysis unit 132 can terminate the MF calculation and early terminate the search for the maximum MF module of that rank if it determines that the maximum value of MF is not updated in a provisional module where MF has not yet been calculated among multiple provisional modules with different edge patterns. This reduces the computational load and shortens the FPO analysis processing time compared to a case where the MF calculation continues and the module optimization process continues even if the maximum value of MF is not updated.
[0045] (Step S4) The FPO module selection unit 133 selects the optimal module (referred to as the FPO module) for a given seed element based on the module information 123.
[0046] Figure 5A is a diagram illustrating the FPO module selection method according to this embodiment. The example in Figure 5A is time-series metabolome data of grated radish (2305 elements x 10 time-series data set). The Pearson product-moment correlation coefficient has a MIN value of "2" and a MAX value of "2305", and the element with identifier "c1487" was used as the seed element for FPO analysis. Figure 5A shows the maximum value of MF (vertical axis) corresponding to the rank (horizontal axis). The FPO module selection unit 133 determines the module M_PTmax of the rank that shows the maximum value PTmax (longest plateau length) during the period in which the maximum value of MF is not updated even if the rank value is increased. The FPO module selection unit 133 selects the determined module M_PTmax as the FPO module.
[0047] In the example shown in Figure 5A, the MF of the module constructed at "Rank = 623" was calculated to be 0.683, and no modules with an MF exceeding 0.683 were constructed until "Rank = 761". Therefore, the plateau length, which is the period during which the maximum MF is not updated even when the rank value is increased, is "138", which is obtained by subtracting "Rank = 623" from "Rank = 761". In this example, "138" became the longest plateau length (maximum plateau length) PTmax, so the module constructed at "Rank = 623" was selected as module M_PTmax of rank "623", which represents the maximum plateau length PTmax.
[0048] Traditionally, the module corresponding to the rank showing the highest MF across all ranks was selected as the FPO module. However, the rank showing the highest MF across all ranks is often at or near the upper limit (MAX value) of the rank. Therefore, traditionally, this often resulted in the creation of a large, biased module. Furthermore, because it depended on the setting of the rank's upper limit (MAX value), there was a possibility that the optimal FPO analysis results could not be obtained.
[0049] In contrast, according to this embodiment, a module with a large MF and an appropriate number of elements can be obtained, resulting in an improvement in the quality of the FPO analysis results. Furthermore, the FPO analysis results are less affected by the setting of the rank upper limit (MAX value). Therefore, a more robust FPO analysis can be performed.
[0050] Furthermore, the module information 123 for each rank can be used to verify the selection criteria for FPO modules. The output data generation unit 138 may use the module information 123 for each rank to generate data for a graph of the maximum value of MF, as exemplified in Figure 5A. This graph data can be displayed on a display screen or printed. This allows operators to easily verify the selection criteria for FPO modules.
[0051] Alternatively, as another method for selecting FPO modules, a criterion that maintains robustness regardless of the number of elements in the module or the distribution of correlation coefficients between elements may be adopted. For example, a criterion using the plateau length of the moving average graph of the maximum value of MF, as exemplified in Figure 5B, or a criterion where the growth rate of said graph falls below a certain value may be adopted. In addition, a criterion optimized for each dataset based on the distribution of correlation coefficients between elements may be adopted. In this case, the upper limit of the rank, i.e., the MAX value, does not need to be adjusted as a setting parameter, thus reducing the burden on the operator.
[0052] Figure 5B is a diagram illustrating the FPO module selection method according to this embodiment. The example in Figure 5B is time-series metabolome data of grated radish (2305 elements x 10-point time-series dataset). The result of FPO analysis processing was performed using an element with Pearson product-moment correlation coefficient, MIN value "2", MAX value "2305", and identifier "c1487" as the seed element. Figure 5B shows the 10-point moving average (vertical axis) of the maximum value of MF corresponding to the rank (horizontal axis). In the example shown in Figure 5B, the MF of the module constructed at "rank = 152" was calculated to be 0.590, and no modules with an MF exceeding 0.590 were constructed until "rank = 385". Therefore, the plateau length of the moving average, which is the period during which the maximum value of MF is not updated even when the rank value is increased, is "233", which is obtained by subtracting "rank = 152" from "rank = 385". In this example, "233" became the longest plateau length (maximum plateau length) of the moving average, so the module constructed with "rank = 152" was selected as the module M_MAPTmax with rank "152" that represents the maximum plateau length, MAPTmax.
[0053] The FPO module selection unit 133 may terminate the FPO analysis and FPO module selection if predetermined conditions are met. Figure 5C is a diagram illustrating an example of FPO analysis according to this embodiment. Figure 5C is time-series metabolome data of grated radish (2305 elements x 10 time-series data set). The result of FPO analysis processing was performed using an element with identifier "c1487" as a seed element, with Pearson product-moment correlation coefficient, MIN value "2", and MAX value "2305". Figure 5C shows the maximum value of MF (vertical axis) corresponding to the rank (horizontal axis). In this example, there are 1569 elements with a positive correlation coefficient to identifier "c1487", so the maximum rank of identifier "c1487" is 1569. Figure 5C shows the process when, in step S3, the FPO module selection unit 133 is executed each time the FPO analysis unit 132 finishes calculating the maximum MF for each rank, and determines whether or not the longest plateau length is updated. If the FPO module selection unit 133 determines that the longest plateau length is not updated, it may determine that the maximum value of the plateau length is not updated even if the rank changes (even if the rank increases), and may terminate the rank increase and terminate the FPO analysis and FPO module selection. In Figure 5C, when the rank exceeds 1432, the remaining number of ranks becomes "137", which is the maximum rank "1569" minus the rank "1432". Since the number of ranks "137" is less than the PTmax "138", it is determined that the PTmax will not be exceeded in the part where the rank is 1432 or higher, and the calculation is not performed in the part where it is indicated that no calculation is necessary (the part where the rank is 1432 or higher). By configuring it in this way, the FPO module selection unit 133 can terminate the FPO analysis and FPO module selection when the longest plateau length is not updated even if elements are added (even if the rank increases). Therefore, even if the longest plateau length is not updated, the computational load can be reduced compared to when the FPO analysis and FPO module selection for each rank are continued, and the time required for FPO analysis and FPO module selection can be shortened.
[0054] (Step S5) If the selection of FPO modules is completed using all elements of the input data as seed elements (Step S5, YES), proceed to Step S6-1. On the other hand, if there are still elements for which FPO modules have not yet been selected (Step S5, NO), return to Step S3 and perform FPO analysis and FPO module selection using the elements for which FPO modules have not yet been selected as seed elements.
[0055] (Step S6-1) The FPO analysis unit 132 calculates the maximum value of VF(i) for each element (i) within the FPO module (referred to as the final VF) and the MF based on the final VF (referred to as the final MF) for each seed element's FPO module. The final MF and final VF calculated for each seed element's FPO module are stored in the evaluation value information 126.
[0056] Figures 6A and 6B are hypothetical diagrams illustrating an example of the final VF according to this embodiment. Figure 6A shows an example of a module and elements. In Figure 6A, elements are represented by circles. Elements include the element of interest (the element for which the final VF is calculated), in-module elements, and out-of-module elements. The numbers shown inside the circles are the correlation ranks for the element of interest. As an example here, the higher the correlation, the higher the rank (smaller the value). In Figure 6A, lines connect elements with values above a certain correlation coefficient between the element of interest and the in-module elements, and between the element of interest and the out-of-module elements. Solid lines indicate relationships between elements included in the edge, i.e., in-module elements. Dotted lines indicate relationships between elements not included in the edge, i.e., out-of-module elements.
[0057] The FPO analysis unit 132 calculates the final VF for each element (i) within an FPO module. Specifically, the FPO analysis unit 132 selects each element (i) within the FPO module as a focus element and calculates VF(i) for that focus element. Figure 6B shows the calculation result of VF(i) for one element (i) (focus element). In Figure 6B, the horizontal axis shows the correlation rank, and the vertical axis shows VF(i). The FPO analysis unit 132 obtains the maximum value of VF(i) (final VF) for one element (i) (focus element). In the example in Figure 6B, when the correlation rank is "5", there are 3 connections to elements within the module and 2 connections to elements outside the module. The degree is "5 = 2 + 3", the edge is 3, and the number of elements in the module (n) is 5. From equation (2) above, VF(i) is "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (5 + 5 - 1) = 0.67", which represents the maximum value of VF(i) (final VF).
[0058] Furthermore, the FPO analysis unit 132 may determine, in calculating the final VF, whether VF(i) will not exceed the maximum value even if the correlation rank increases. The FPO analysis unit 132 may terminate the process of calculating VF(i) if it determines that VF(i) will not exceed the maximum value even if the correlation rank increases.
[0059] Figure 6C is a hypothetical diagram illustrating an example of the method for calculating the final VF according to this embodiment. The FPO analysis unit 132 calculates the final VF for each element (i) within an FPO module. Specifically, the FPO analysis unit 132 designates each element (i) within the FPO module as a focus element and calculates VF(i) for that focus element. Figure 6C shows the calculation result of VF(i) for one element (i) (focus element). In Figure 6C, the horizontal axis represents the correlation rank, and the vertical axis represents VF(i). The FPO analysis unit 132 determines that VF(i) does not exceed the maximum value when the correlation rank is 8, and terminates the process of calculating VF(i). The reason why the calculation of VF(i) can be terminated is described below.
[0060]
[0061] In equation (3) above, VF is the element F score. n is the number of elements in the target module. edge is the number of edges in the target module. degree is the degree of the element. From equation (3) above, even if all the elements in a module are connected to the elements of other modules, edge = n-1, so we can say that edge ≤ n-1. At the time of calculating the final VF, the number of elements in the module n is a fixed value, so we can find the upper limit of the achievable final VF using only the degree.
[0062]
[0063] In equation (4) above, VF is the element F score. a The maximum VF value is defined as the point at correlation rank a. n is the number of elements in the target module. degree is the degree number of the element. In the example in Figure 6C, n=5, and at the point at correlation rank a=5, VF a Since this becomes = 0.67, VF aIn order to exceed this, d ≤ 8 is required by equation (4) above. Here, the value of degree depends on the correlation rank, and when the correlation rank is "9", the value of degree is 9. Therefore, after calculating the case for correlation rank "8", it is unnecessary to calculate for correlation ranks of "9" or higher. In the example in Figure 6C, when the correlation rank is "6", there are 3 connections to elements inside the module and 3 connections to elements outside the module, the degree number is "6 = 3 + 3", the edge number is 3, the number of elements in the module (n) is 5, and VF(i) is "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (6 + 5 - 1) = 0.60" from equation (2) above. In the case of a correlation rank of "7", there are 3 connections to elements within the module and 4 connections to elements outside the module, the degree number is "7 = 4 + 3", the edge number is 3, and the number of elements in the module (n) is 5, and VF(i) is given by equation (2) above as "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (7 + 5 - 1) = 0.54". In the case of a correlation rank of "8", there are 3 connections to elements within the module and 5 connections to elements outside the module. The degree count is "8 = 5 + 3", the edge count is 3, and the number of elements in the module (n) is 5. Therefore, VF(i) is calculated from equation (2) above as "VF(i) = (2 × edge(i)) / (degree(i) + n - 1) = (2 × 3) / (8 + 5 - 1) = 0.50". Thus, as determined using equation (4), even when elements are added in order of highest correlation rank, if the element is an element outside the module, VF(i) will not exceed 0.67 (in the case of a correlation rank of "5"). By configuring it in this way, the amount of computation can be reduced and processing time can be shortened compared to calculating VF(i) for all correlation ranks.
[0064] The FPO analysis unit 132 calculates MF (referred to as the final MF) for one FPO module based on the final VF of each element (i) in the FPO module. Specifically, the FPO analysis unit 132 acquires edge(i) and degree(i) when the final VF is obtained for each element (i) in the FPO module, and calculates MF (final MF) from the above formula (1) using edge(i) and degree(i) of each acquired element (i).
[0065] Note that the method for calculating the final VF and the final MF may be any method for calculating VF(i) and MF as long as it does not depend on the seed elements in the module for which the final VF and the final MF are calculated. For example, when the element P value described later is the minimum for a certain element (i) in the module with respect to other elements in the module, the VF(i) may be used as the final VF.
[0066] (Step S6-2) The P value calculation unit 137 calculates the module P value for each FPO module of the seed elements according to the following formula.
[0067]
[0068] N is the total number of all elements of the multivariate data which is the input data. n is the number of elements of the module. edge avg is the integer closest to the average value of edge(i) of each element (i) used for calculating the final MF. degree avg is the integer closest to the average value of degree(i) of each element (i) used for calculating the final MF. q and r are positive integers, and r ≤ q. Note that instead of edge avg the number of edges in the module connecting elements having a correlation coefficient equal to or greater than the correlation threshold when the maximum value of MF of the FPO module is obtained may be used, and instead of degree avg the degree when the maximum value of MF of the FPO module is obtained may be used.
[0069] The module P-value calculated for each seed element's FPO module is stored in evaluation value information 126 as the module P-value after FPO analysis. The module P-value has the same meaning as probability used in statistical processing and represents the probability of an event occurring that is rarer than the observed event. The module P-value allows for testing whether the formed module is correct or not. The module P-value calculated by equation (5) above is obtained from the product of the probability of giving the same MD and the probability of giving the same MS when the modules are constructed randomly. The module P-value after FPO analysis allows the operator to evaluate the significance of the FPO module of each seed element. This adds credibility to the module.
[0070] (Step S7) The FPO module integration unit 134 performs a module integration process (modularization process) on the FPO module of each seed element. Each FPO module is configured separately for each seed element by steps S3-S5 described above. As a result, there are as many FPO modules as there are elements, and there is overlap in elements between the FPO modules. This overlap in elements is resolved by the module integration process. The module integration process is described below.
[0071] The FPO module integration unit 134 reconfigures the modules for each VF threshold within a predetermined range of the VF threshold (from the minimum value to the maximum value of the VF threshold). The VF threshold is set in advance with a constant step size from the minimum value to the maximum value of the VF threshold. As a module reconfiguration method, all existing edges are deleted for all FPO modules, and elements that are the final VFs above the VF threshold are reconnected with edges. At this time, if there are overlapping elements between FPO modules, the FPO modules are also integrated through the overlapping elements. This provides the integration result of the FPO modules for each VF threshold. The FPO module integration unit 134 stores the integration result of the FPO modules for each VF threshold in the module integration information 124. Note that the VF threshold is not limited to being set in a constant step size from the minimum value to the maximum value of the VF threshold. For example, all possible values for MF and VF may be listed as the VF threshold, and integration may be performed for all MF and VF. In this case, the VF threshold does not need to be in fixed increments.
[0072] Furthermore, the FPO module integration unit 134 may choose not to perform the integration process on all FPO modules configured in steps S3-S5, but only on some of them. For example, based on the final VF, final MF, and module P value of each FPO module stored in the evaluation value information 126, FPO modules that do not meet predetermined conditions may be excluded from the integration process. For example, when the VF threshold is set from the minimum to the maximum value of the VF threshold, FPO modules whose final MF is less than the VF threshold, and FPO modules whose seed element's final VF is less than the VF threshold may be excluded from the integration process. In this way, by excluding FPO modules with low quality or significance from the integration process based on evaluation values, it is possible to improve the quality or significance of the modules being integrated.
[0073] Figure 7 is a hypothetical diagram illustrating an example of the integration process according to this embodiment. Figure 7A is a diagram illustrating the integration of modules in the modularization process, where the final VF is used to remove edges between elements. As an example, Figure 7A shows three FPO modules M7-1, M7-2, and M7-3, with elements indicated by circles, element identifiers "c001" to "c011", and the final MF of each FPO module and the final VF of each element calculated in step S6-1. In this example, element "c002" overlaps in FPO modules M7-1 and M7-2, and element "c008" overlaps in FPO modules M7-2 and M7-3. When the VF threshold is set to 0.5, the integration process constructs module M7-4 containing all elements. When the VF threshold is set to 0.75, the final VF of element "c002" in module M7-2 and element "c011" in module M7-3 are below the VF threshold and are therefore not connected by an edge, and are excluded by the integration process. As a result, modules M7-5 and M7-6 are formed by the integration process. When the VF threshold is set to 0.85, performing the integration process similarly results in the formation of modules M7-7, M7-8, and M7-9. Here, modules M7-8 and M7-9 have a small number of elements, 2 and 1 respectively, and are evaluated low. Furthermore, they do not show significant connections between multiple elements, which hinders the analysis of the FPO module integration results by the operator.
[0074] On the other hand, Figure 7B is an illustrative diagram of the modularization process, in which modules are integrated by excluding the components of each module using the final MF and the final VF of the seed elements, and by excluding the edges between elements using the final VF. Figure 7B shows the integration process when the VF threshold is set from the minimum to the maximum value of the VF threshold, and FPO modules whose final MF is less than the VF threshold, and FPO modules whose final VF of the seed elements is less than the VF threshold are excluded from the integration process. When the VF threshold is set to 0.85, modules M7-2 and M7-3, whose final MF is less than 0.85, are excluded from the integration process, and when the integration process is performed, only module M7-7 is formed. In this way, the integration process can prevent the formation of modules with low evaluation values, thereby improving the quality or significance of the formed modules.
[0075] The FPO module integration unit 134 stores the integration result (FPO module group) of FPO modules at the VF threshold, which is the maximum number of FPO modules that fit within a predetermined range of module size, in the final module integration information 125 as the final FPO module integration result.
[0076] Conventionally, the final integration result of the FPO module alone did not provide an overview of the positional relationships of the elements, making it difficult to analyze the internal structure of the FPO module. However, according to this embodiment, the integration result of the FPO module for each VF threshold is stored in the module integration information 124, so the operator can see how the integration result of the FPO module changes in accordance with changes in the VF threshold. This allows for an overview of the positional relationships of the elements in the final integration result of the FPO module, contributing to the analysis of the internal structure of the FPO module.
[0077] Figures 8 and 9 show examples of the final FPO module integration results according to this embodiment. Figures 8 and 9 show the results of FPO analysis performed using a dataset showing the social relationships of 34 members belonging to a karate club at a university in the United States, with Pearson's product-moment correlation coefficient, MIN value "5", and MAX value "34". Components are shown as circles, and each separated member is distinguished by the pattern inside the circle. The numbers are the identification numbers of the 34 members of the karate club. The solid lines show the relationship of connections obtained through modularization processing. Figure 8 shows the network depiction results for each obtained module. Figure 9A shows examples of FPO module integration results for each VF threshold corresponding to the final FPO module integration result. In the example of Figure 9A, the graph shows how the FPO module integration result (modularization of each element (vertical axis)) changes according to the change in the VF threshold (horizontal axis). For example, the VF threshold (VF th When the value is 0.81, the relationship between modules M8-1, M8-2, and M8-3 shown in Figure 8 can be understood from Figure 9A.
[0078] The output data generation unit 138 may use the module integration information 124 for each VF threshold to generate data (e.g., a graph) that shows the correspondence between the VF thresholds (horizontal axis) and the integration results of the FPO modules (modularization of each element (vertical axis)) as illustrated in Figure 9. This graph data can be displayed on a display screen or printed. This allows the operator to easily analyze the final integration results of the FPO modules.
[0079] The FPO module integration unit 134 may further calculate a distance index d, which indicates the final distance between the FPO modules after module integration, based on the integration results of the FPO modules for each VF threshold stored in the module integration information 124. The distance index is a positive real number and is calculated for all pairs of FPO modules. The lower the value of the distance index d, the closer the distance between the two modules is determined to be. By calculating the distance index d, the overall positional relationship between the FPO modules can be easily and quantitatively evaluated.
[0080] Figure 9B is a diagram illustrating an example of a method for calculating the distance index d of the final FPO module according to this embodiment. Figure 9B shows one method for calculating the distance index d based on the integration result of the final FPO module exemplified in Figure 9A. In Figure 9B, the VF shown in Figure 9A th For the three modules M8-1, M8-2, and M8-3 separated at a VF threshold of 0.81, we consider virtual modules M9-1, M9-2, and M9-3, which further include information about the structure of the modules when the VF threshold is increased. Point G in the graph is defined as the center point of the separated virtual modules M9-1, M9-2, and M9-3. Regarding the distance between the obtained FPO modules, by using the difference in VF thresholds and the number of components to calculate the distance to the center point of the virtual module, we can more accurately grasp the overall positional relationship of the elements in the integrated result of the FPO modules.
[0081] For example, when calculating the distance index between modules M9-1 and M9-2, the following steps (1) to (3) are performed: (1) The minimum VF threshold at which the components of modules M8-1 and M8-2 are divided by the module integration process, VF th,min Search. In this example, VF th,min = 0.76. (2) The components of virtual modules M9-1 and M9-2 set the VF threshold to VF th When set to the above, the average value of the minimum VF threshold that is excluded by the module integration process is VF. th, minavg The average of the minimum VF thresholds at which the 16 components are excluded is calculated. For example, in the case of virtual module M9-1, elements "23", "21", "19", "16", and "15" are excluded if the VF threshold is 0.95 or higher, elements "33", "31", "30", and "24" are excluded if the VF threshold is 0.92 or higher, elements "34", "29", "28", "27", and "10" are excluded if the VF threshold is 0.91 or higher, element "9" is excluded if the VF threshold is 0.90 or higher, and element "32" is excluded if the VF threshold is 0.86 or higher. Therefore, the average of the minimum VF thresholds at which the 16 components are excluded is calculated, VF th, minavg It is calculated using (M8-1). Specifically, VF th, minavg(M8-1) = (0.95 × 5 + 0.92 × 4 + 0.91 × 5 + 0.90 × 1 + 0.86 × 1) / 16 = 0.9212. Similarly, in the case of virtual module M9-2, elements "8", "4", "22", "2", and "18" are excluded if the VF threshold is 0.90 or higher, elements "1", "13", and "14" are excluded if the VF threshold is 0.88 or higher, element "20" is excluded if the VF threshold is 0.87 or higher, and element "3" is excluded if the VF threshold is 0.86 or higher. Thus, the average value of the minimum VF threshold at which 10 components are excluded is VF. th, minavg It is calculated using (M8-2). Specifically, VF th, minavg (M8-2) can be calculated as (0.9 × 5 + 0.88 × 3 + 0.87 × 1 + 0.86 × 1) / 10 = 0.8870. (3) The distance index d(M8-1, M8-2) is calculated by the following formula (6).
[0082]
[0083] In equation (6) above, d represents the distance index between modules. VF th This indicates the VF threshold. th,min This indicates the minimum VF threshold at which the components of modules M8-1 and M8-2 are separated by the module integration process. th, minavg This represents the average of the minimum VF thresholds at which components of both modules are excluded. In this example, d(M8-1, M8-2) = 2(0.81-0.76) + (0.8870-0.81) / 2 + (0.9212-0.81) / 2 = 0.1000 + 0.0385 + 0.0556 = 0.1941.
[0084] Applying the same calculation to virtual modules M9-2 and M9-3, we obtain d(M8-2, M8-3) = 0.0585 and d(M8-1, M8-3) = 0.1756. In the case of d(M8-2, M8-3), in the case of virtual module M9-3, elements "7", "6", "5", "12", and "11" are excluded when the VF threshold is 0.85 or higher, so the average value of the VF threshold at which these five components are excluded is VF. th, minavg It is calculated using (M8-3). Specifically, VF th, minavg(M8-3) = 0.85 × 5 / 5 = 0.8500 can be calculated. The minimum VF threshold at which the components of modules M8-2 and M8-3 are divided by the module integration process is VF. th,min Search. In this example, VF th,min = 0.81. In this example, d(M8-2, M8-3) = 2(0.81 - 0.81) + (0.8870 - 0.81) / 2 + (0.8500 - 0.81) / 2 = 0.0385 + 0.0200 = 0.0585 can be calculated. Also, in the case of d(M8-1, M8-3), the minimum VF threshold at which the components of modules M8-1 and M8-3 are separated by the module integration process, VF th,min Search. In this example, VF th,min = 0.76. In this example, d(M8-1, M8-3) = 2(0.81-0.76) + (0.9212-0.81) / 2 + (0.8870-0.81) / 2 = 0.1000 + 0.0556 + 0.0200 = 0.1756. From the above, it can be determined that M8-2 is closer to M8-1 than M8-3. In this way, the overall positional relationship of the elements in the final integrated result of the FPO module can be grasped more accurately.
[0085] Figure 9B shows the distance index calculated based on the final FPO module integration results illustrated in Figure 9A, expressed as a numerical value. In module M8-1, distance 0.1056 represents the average distance of each element from the minimum VF threshold at which module M8-2 is divided to the minimum VF threshold at which elements constituting module M8-1 are excluded. In module M8-1, distance 0.0556 represents the average distance of each element from the minimum VF threshold at which module M8-3 is divided to the minimum VF threshold at which elements constituting module M8-1 are excluded. In module M8-2, distance 0.0885 represents the average distance of each element from the minimum VF threshold at which module M8-1 is divided to the minimum VF threshold at which elements constituting module M8-2 are excluded. In module M8-2, distance 0.0385 represents the average distance of each element from the minimum VF threshold at which module M8-3 is divided to the minimum VF threshold at which elements constituting module M8-2 are excluded. In module M8-3, distance 0.0700 represents the average distance from the minimum VF threshold at which module M8-1 is divided to the minimum VF threshold at which the elements constituting module M8-3 are excluded. In module M8-3, distance 0.0200 represents the average distance of each element from the minimum VF threshold at which module M8-2 is divided to the minimum VF threshold at which the elements constituting module M8-3 are excluded.
[0086] (Step S8-1) The FPO module integration unit 134 calculates the final VF and final MF for each of the final FPO modules after module integration, similar to the FPO analysis unit 132 in step S6-1 above. The final VF and final MF calculated for each of the final FPO modules after module integration are stored in the evaluation value information 126.
[0087] (Step S8-2) The P-value calculation unit 137 calculates the module P-value for each of the final FPO modules after module integration using the above formula (5). The module P-values calculated for each of the final FPO modules after module integration are stored in the evaluation value information 126 as module P-values after module integration. By using the module P-values after module integration, the operator can evaluate the significance of each of the final FPO modules after module integration. This makes it possible to add credibility to the modules.
[0088] (Step S9) The FNI analysis unit 135 targets any or all of the final FPO modules (final FPO modules) after module integration and performs an In process (FNI analysis process) to incorporate false-negative elements.
[0089] The final FPO modules to be subjected to FNI analysis processing may, for example, be specified by the user as multiple modules. Alternatively, the conditions for the final FPO modules to be subjected to FNI analysis processing (FNI analysis target conditions) may be set in advance in the information processing device 1, and the FNI analysis unit 135 may exclude final FPO modules that do not satisfy the FNI analysis target conditions and decide to subject the remaining final FPO modules to FNI analysis processing. The FNI analysis target conditions are, for example, modules that have a certain significance level (module P value) or higher. The FNI analysis processing will be described below.
[0090] Figure 10 is a hypothetical diagram illustrating the FNI analysis process according to this embodiment. In Figure 10, elements are represented by circles. Elements include elements of interest, candidate elements, elements within the target module, and elements outside the target module. Elements within the target module are those included in the final FPO module (final FPO module) after module integration, while elements outside the target module are those not included in the final FPO module (final FPO module) after module integration. The numbers shown inside the circles represent the correlation rank with respect to the candidate elements. Each element is connected by a line. Solid lines indicate the relationships between elements within the module, including the elements of interest. Dotted lines indicate the relationships between elements outside the module, including the candidate elements. The FNI analysis unit 135 repeatedly performs the FNI analysis on each element of interest, sequentially setting each of the elements included in all the final FPO modules targeted by the FNI analysis process as an element of interest. Therefore, the FNI analysis process is executed a number of times equal to the total number of elements included in all the final FPO modules targeted by the FNI analysis process. Here, we will refer to Figure 10 and explain using one of the elements of interest shown in Figure 10 as an example.
[0091] The FNI analysis unit 135 sets candidate elements to be included in the final FPO module (target module) to which the element of interest belongs. For example, the candidate elements to be included in the target module may be all elements that do not belong to the target module (i.e., all elements outside the target module). For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation rank with respect to the element of interest is up to a predetermined high rank. For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation coefficient with respect to the element of interest is above a predetermined correlation threshold. For example, the candidate elements to be included in the target module may be limited to all elements outside the target module whose correlation rank with respect to the element of interest is up to a predetermined high rank AND whose correlation coefficient with respect to the element of interest is above a predetermined correlation threshold. For example, the candidate elements to be included in the target module may be limited to elements that do not belong to any final FPO module. The FNI analysis unit 135 performs a false negative determination for each candidate element with respect to the target module.
[0092] Specifically, the FNI analysis unit 135 refers to the candidate element i in order from the top (ranking from the highest correlation (correlation rank)) of the correlation coefficient list 122, and calculates the VS(i) of the above formula (2) for the target module of the candidate element i, "VS(i) = (number of edges in the target module of the candidate element i edge(i)) ÷ (number of degrees of the candidate element i degree(i))" for each rank. The FNI analysis unit 135 records information such as the correlation rank for all calculated VS(i) that exceed a predetermined VS threshold in the storage unit 12. The VS threshold can be arbitrarily set by the operator, for example. If the FNI analysis unit 135 calculates a VS(i) that exceeds a predetermined VS threshold, it determines that the candidate element is a false negative element with respect to the element of interest and incorporates the candidate element into the target module.
[0093] (Step S10) The P-value calculation unit 137 calculates the element P-value for the target module in the correlation rank corresponding to VS(i) that exceeds the VS threshold using the following formula.
[0094]
[0095] In equation (7) above, N is the total number of elements in the multivariate input data. n is the number of elements in the target module. edge is the number of edges (i) in the target module for candidate element i. degree is the degree (i) of candidate element i. q and r are positive integers such that r ≤ q.
[0096] Alternatively, the element P-value may be calculated using the recursive formula in equation (8) while calculating VS(i) for each correlation rank in step S9.
[0097]
[0098] In equation (8) above, N is the total number of elements in the multivariate data which is the input data. n is the number of elements in the module. edge is the number of edges (i) in the target module for candidate element i. degree is the degree number (i) of candidate element i. By calculating the element P value using the recursive formula in equation (8) while calculating VS (i) in step S9, the calculation of the number of combinations can be omitted, and the efficiency of processing can be improved.
[0099] Figure 11 shows the time required to perform FPO analysis using microarray data of Arabidopsis thaliana (dataset of 27127 x 9442 samples) with Pearson's product-moment correlation coefficient, MIN value "2", and MAX value "250", followed by FNI analysis (including steps S9 to S13-2 (described later)). As per the present invention, by executing steps S9 and S10 simultaneously and employing a recursive formula for calculating element P values, the analysis time was significantly reduced to 1 / 92.
[0100] The P-value calculation unit 137 stores information regarding the correlation rank with the smallest element P-value among the correlation ranks corresponding to VS(i) that exceed the VS threshold in the evaluation value information 126 for FNI analysis. The information regarding the correlation rank with the smallest element P-value is the information regarding the correlation rank that can be determined to be the most significant.
[0101] (Step S11) If FNI analysis processing is completed for all elements of the final FPO module targeted for FNI analysis processing (Step S11, YES), proceed to Step S12. On the other hand, if there are still elements for which FNI analysis processing has not been performed (Step S11, NO), return to Step S9 and perform FNI analysis processing on the elements for which FNI analysis processing has not been performed, with those elements as the elements of interest.
[0102] (Step S12) The FNI assignment processing unit 136 performs an assignment process to resolve duplicates for candidate elements that overlap in multiple final FPO modules generated by the FNI analysis process. The following assignment criteria are used in the assignment process.
[0103] Assignment Criteria: For multiple target modules where candidate element i overlaps, the following priority order is used: VS(i) > "Module P value of equation (5) above when candidate element i is incorporated into the target module" > Number of elements in the module. VS(i) has the highest priority, followed by module P value, and then the number of elements in the module.
[0104] First, candidate element i is retained only in the target module with the highest VS(i), and removed from all other target modules. On the other hand, if the VS(i) values are equivalent, candidate element i is retained only in the target module with the smallest "module P value in equation (5) above when candidate element i is incorporated into the target module," and removed from all other target modules. If the VS(i) values are equivalent and the "module P value in equation (5) above when candidate element i is incorporated into the target module" is equivalent among the target modules, candidate element i is retained only in the target module with the largest number of elements, and removed from all other target modules.
[0105] Furthermore, the assignment process is not limited to the assignment criteria described above, as long as it can resolve the duplication of elements between multiple modules that occurred during the FNI analysis process. For example, modules with duplicate elements may be merged through a module integration process (modularization process).
[0106] The FNI assignment processing unit 136 stores information about the final FPO module after the assignment process in the final module information 127. The final module information 127 may be individual module files, a file containing all modules, or a combination of individual module files and a file containing all modules. Information about the final FPO module before the assignment process may also be recorded in the storage unit 12.
[0107] According to this embodiment, if duplicate candidate elements occur in multiple final FPO modules as a result of performing FNI analysis on multiple final FPO modules, the duplication of candidate elements can be resolved. This eliminates the need for operators to individually specify the final FPO modules on which to perform FNI analysis, allowing them to perform FNI analysis on multiple final FPO modules simultaneously and obtain FNI analysis results without element duplication.
[0108] (Step S13-1) The FNI assignment processing unit 136 calculates the final VF and final MF for each of the final FPO modules after the assignment process, similar to the FPO analysis unit 132 in step S6-1 above. The final VF and final MF calculated for each of the final FPO modules after the assignment process are stored in the evaluation value information 126.
[0109] (Step S13-2) The P-value calculation unit 137 calculates a module P-value for each of the final FPO modules after the assignment process using the above formula (5). The module P-values calculated for each of the final FPO modules after the assignment process are stored in the evaluation value information 126 as module P-values after the assignment process. By using the module P-values after the assignment process, the worker can evaluate the significance of each of the final FPO modules after the assignment process. This makes it possible to add credibility to the modules.
[0110] (Step S14) The output data generation unit 138 generates various types of output data. The output data is data that can be displayed on a display screen or printed. For example, the output data generation unit 138 generates drawing data that represents the configuration of each final FPO module after the assignment process. This drawing data visualizes the configuration of each final FPO module after the assignment process by drawing it.
[0111] For example, the output data generation unit 138 may generate graph data (such as the graph data exemplified in Figure 5A) showing the MF corresponding to each rank for each seed element obtained by the FPO analysis process. For example, the output data generation unit 138 may generate graph data (such as the graph data exemplified in Figure 9) showing the integration results by VF threshold obtained by the FPO module integration process.
[0112] The output data generated by the output data generation unit 138 may be output to an external device of the information processing device 1 by the input / output unit 11, or it may be stored in the storage unit 12. The output data stored in the storage unit 12 can be output to an external device by the input / output unit 11.
[0113] According to this embodiment, for example, it is possible to improve the accuracy of correlation network analysis on multivariate data such as comprehensive molecular information obtained by omics analysis or big data.
[0114] Although embodiments of the present invention have been described in detail above with reference to the drawings, the specific configuration is not limited to these embodiments, and design modifications and the like are also included within the scope of the gist of the present invention.
[0115] Furthermore, computer programs for realizing the functions of each of the above-mentioned devices may be recorded on a computer-readable recording medium, and the programs recorded on this recording medium may be loaded into a computer system and executed. The term "computer system" here may include hardware such as an operating system and peripheral devices. Also, if a WWW system is used, the "computer system" shall also include the homepage provisioning environment (or display environment). Furthermore, "computer-readable recording medium" refers to writable non-volatile memory such as flexible disks, magneto-optical disks, ROMs, and flash memory, portable media such as DVDs (Digital Versatile Discs), and storage devices such as hard disks built into a computer system.
[0116] Furthermore, "computer-readable recording media" includes volatile memory (e.g., DRAM) within a computer system that acts as a server or client when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, which retains the program for a certain period of time. The program may also be transmitted from the computer system storing it in a memory device to another computer system via a transmission medium or by transmission waves within the transmission medium. Here, the "transmission medium" for transmitting the program refers to a medium that has the function of transmitting information, such as a network (communication network) such as the Internet or a communication line (communication line) such as a telephone line. The program may also be intended to implement only a part of the aforementioned functions. Furthermore, it may be a program that can implement the aforementioned functions in combination with a program already recorded in the computer system, a so-called differential file (differential program).
[0117] 1... Information processing device, 11... Input / output unit, 12... Storage unit, 121... Correlation network analysis program, 122... Correlation coefficient list, 123... Module information, 124... Module integration information, 125... Final module integration information, 126... Evaluation value information, 127... Final module information, 128... FPO analyzed module list, 13... Control unit, 131... Correlation coefficient calculation unit, 132... FPO analysis unit, 133... FPO module selection unit, 134... FPO module integration unit, 135... FNI analysis unit, 136... FNI assignment processing unit, 137... P-value calculation unit, 138... Output data generation unit
Claims
1. An information processing device comprising: an FNI analysis unit that performs FNI analysis processing to incorporate false negative elements into each of several modules composed of elements of the multivariate data to be analyzed for correlation network analysis; and an FNI assignment processing unit that performs assignment processing to resolve duplication for elements that overlap in the multiple modules generated by the FNI analysis processing.
2. The FNI assignment processing unit, for modules containing duplicate elements, retains the duplicate elements only in the module with the highest VS(i) calculated by the following formula, and removes the duplicate elements from all other target modules. The information processing apparatus according to claim 1, wherein edge(i) is the number of edges of the overlapping element i, degree(i) is the degree of the overlapping element i, and i is an integer value between 1 and n, and is the index of the element.
3. The FNI assignment processing unit, when a module contains duplicate elements and the VS(i) values are equivalent across modules, retains the duplicate elements only in the module with the smallest module P value calculated by the following formula, and removes the duplicate elements from the other modules. N is the total number of elements in the multivariate data, n is the number of elements in the module, and edge avg This is the integer closest to the average number of edges for each element used to calculate the final module F score, and degree avg The information processing apparatus according to claim 2, wherein is the integer closest to the average value of the order of each element used in calculating the final module F score, and q and r are positive integers such that r ≤ q.
4. The information processing apparatus according to claim 3, wherein the FNI assignment processing unit, when the VS(i) values are the same among the modules and the module P values are the same among the modules, leaves the duplicate elements only in the module with the largest number of elements and removes the duplicate elements from the other modules.
5. An information processing method executed by an information processing device, comprising: an FNI analysis step of performing an FNI analysis process to incorporate false negative elements into each of a plurality of modules composed of elements of multivariate data subject to correlation network analysis; and an FNI assignment step of performing an assignment process to resolve duplication for elements that overlap in the plurality of modules generated by the FNI analysis process.
6. A computer program for causing a computer to perform the following steps: an FNI analysis step which performs an FNI analysis process to incorporate false negative elements into each of the multiple modules composed of elements of the multivariate data to be analyzed for correlation network analysis; and an FNI assignment step which performs an assignment process to resolve duplicates of elements that overlap in the multiple modules generated by the FNI analysis process.
7. A computer-readable recording medium containing a computer program for executing the following steps: an FNI analysis step which performs an FNI analysis process to incorporate false negative elements into each of several modules composed of elements of the multivariate data subject to correlation network analysis; and an FNI assignment step which performs an assignment process to resolve duplicates in the multiple modules resulting from the FNI analysis process.