Evaluation method, visualization method, evaluation device, and visualization device
Patent Information
- Application Number
- JP2024548172
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2023-08-31
- Filing Date
- 2023-08-31
- Publication Date
- 2025-09-12
AI Technical Summary
Current methods for evaluating clustering results in omics data, particularly in unsupervised settings, face challenges in determining biological robustness due to sensitivity to setting values and noise, and lack consideration of biological information, leading to suboptimal cluster validation and selection.
A method and device that calculate a quantitative index of biological robustness by evaluating the concordance between cluster assignments across multiple omics data, using weighted statistics and similarity measures to output an index indicating the biological validity of clusters, and visualize these results for better understanding.
This approach provides a robust evaluation of clustering results by quantifying biological robustness, reducing noise susceptibility and improving cluster validation, enabling more accurate selection of algorithms and settings, and offering a visual representation for clearer insights.
Smart Images

Figure 2024062895000001 
Figure 2024062895000002
Abstract
Description
Evaluation method, visualization method, evaluation device, and visualization device
[0001] The present invention relates to an evaluation method, a visualization method, an evaluation device, and a visualization device, and more particularly to a technique for evaluating clusters or clustering of omics data and a technique for visualizing the evaluation results.
[0002] [Challenges in clustering evaluation] Clustering technology is a method for forming sets (called clusters) of sample groups classified based on arbitrarily defined similarities, and is a widely used method as one of the search methods in data analysis.
[0003] Clustering is considered a type of unsupervised machine learning, and is generally performed in an unsupervised setting. Here, "supervised" refers to a situation where input data and correct answer information, such as "the input data for sample i is X and belongs to cluster K," are given in advance. "Unsupervised," which is the primary method for clustering analysis, is the opposite, with the problem setting being "only the input data X for sample i is given, and classification is performed in a situation where it is unknown which cluster is correct." Hereafter, in this text, we will assume that clustering is performed in an unsupervised setting.
[0004] The challenge with unsupervised clustering is evaluating the execution results. Evaluation makes it possible to verify whether the clusters obtained are valid. Based on the evaluation results, the user can (1) adjust the setting values (values that the user generally needs to set; hyperparameters in machine learning) or (2) select the appropriate clustering method or algorithm from among the multiple available methods.
[0005] Regarding (1) above, it is known that even with the same clustering algorithm, results can vary significantly depending on the settings. Regarding (2), various clustering methods have been developed, and the method with the best performance often differs depending on the application, the scale and / or the nature of the data. Therefore, evaluation is important as one criterion for appropriately selecting the method and its settings to be used.
[0006] Evaluation in supervised machine learning is typically performed by calculating an index (called an external criterion) based on the degree of agreement between the correct label and the predicted label. For example, for five samples, if the correct labels {sample 1, sample 2, sample 3, sample 4, sample 5} = {A, B, A, C, C} are given, the agreement (accuracy rate) of the predicted labels {A, B, B, A, C} is 3 / 5 = 0.6. However, because clustering analysis is performed in an unsupervised setting, i.e., no correct labels are given, it is usually difficult to calculate an index based on the degree of agreement with the correct answer.
[0007] Therefore, various evaluation indices (called internal criteria) for clustering results that do not use correct labels have been considered. One example is an index based on the cohesion within the same cluster and the divergence between different clusters. Applying this to the above case, a clustering in which "sample 1 belonging to predicted cluster A has a small distance from sample 4 belonging to the same predicted cluster A, and conversely, has a large distance from sample 2 belonging to predicted cluster B" is considered to be an appropriate clustering. Here, the distance between samples is a value calculated from the observed value.
[0008] Among these internal criteria, it has been reported that there are indicators that show good correlation with external criteria in experiments using artificially generated data. However, there are problems such as the results being obtained using artificial data under limited conditions, and the optimal internal criteria differing depending on the target data, and therefore evaluation of clustering results using internal criteria remains a challenge.
[0009] [Clustering Omics Data] First of all, the collective term for genes present in a living organism is called the genome, and data measuring the genome is called genomic data. Furthermore, by relating this to a series of phenomena occurring within the organism, it is possible to expand this concept, and in addition to the genome, other terms have been proposed, such as the epigenome, which is a collective term for acquired regulatory factors that do not involve changes in gene sequence, the transcriptome, which is a collective term for gene transcription products, and the proteome, which is a collective term for proteins. In recent years, the concept of omics / omics data has emerged, generalizing these "collective terms for biological substances that exist at the same level in life phenomena."
[0010] In the biomedical field, clustering analysis of various omics data is widely conducted. For example, by clustering the transcriptome data of cancer patients, it is possible to classify the same cancer into subtype groups with different characteristics, and research is being conducted to select highly accurate treatments and / or medications appropriate for each subtype. Furthermore, the concept of multi-omics has emerged in recent years, and attempts are being made to more precisely model complex biological phenomena by integrating and analyzing information obtained from multiple omics.
[0011] As such clustering techniques, the techniques described in Non-Patent Documents 1 to 3 are known.
[0012] [Technology of Non-Patent Document 1] In Non-Patent Document 1, Lu et al. propose an evaluation index called "Robustness" for the robustness of a clustering algorithm to a set value. Here, "robust" refers to the insensitivity of the response of the clustering algorithm results to changes in the set value.
[0013] "Robustness" is the ability to perform clustering using a single algorithm under different settings (e.g., changing the number of clusters), and the more common elements the results have, the more robust the result. Here's an example: (i) Clustering is performed 10 times using algorithm X on three samples under different settings. (ii) When looking at sample a and sample b, the number of times a and b belong to the same cluster is 7. In this case, for (a, b), the result is 7 / 10 = 0.7. (iii) Similarly, for (b, c), the result is 5 / 10 = 0.5, and for (c, a), the result is 0.3. (iv) In this case, the "Robustness" of X is (0.7 + 0.5 + 0.3) / 3 = 0.5.
[0014] Among internal criteria, "Robustness" focuses on the selection of algorithms for omics data clustering analysis. In particular, for unsupervised datasets that are sensitive to parameters, such as omics data clustering analysis, it can be difficult to determine appropriate parameters. In such situations, a method that is insensitive to parameters is preferable, which can be achieved by selecting an algorithm with a high "Robustness."
[0015] On the other hand, "Robustness" is an index that gives one score to one algorithm, and therefore cannot be used for the above-mentioned clustering task "(1) Adjustment of setting values."
[0016] [Technology of Non-Patent Document 2] Another issue in clustering analysis methods is the issue of clustering stability. Some clustering algorithms use random numbers to generate initial states. Such algorithms do not guarantee that the cluster assignments obtained will be the same each time they are run. Since solutions obtained using different initial states are likely to have similar evaluation values, it is difficult to determine which execution result is best, even if an appropriate evaluation index is used.
[0017] Therefore, there is a method called "consensus clustering" that derives the final cluster assignment by extracting the common parts of multiple cluster assignments performed with randomly determined initial values. The above-mentioned Non-Patent Document 2 is a representative method, and its purpose is to obtain reliable clusters.
[0018]
[0019]
[0020]
[0021]
[0022] [Technology of Non-Patent Document 3] In Non-Patent Document 3, the "Consensus Clustering" of Non-Patent Document 2 is applied to multi-omics data.
[0023] First, we perform clustering for each omics (using the clustering algorithm based on previous studies) and convert the results into binary vectors. Then, we combine these to generate a matrix with dimensions (sum of the number of clusters for each omics) × number of samples.
[0024] Specifically, the samples belonging to the first cluster of omics 1 are samples 1, 2, and 5, which we will express as "omics1-c1 = [1, 1, 0, 0, 1]." Now, if there are three omics, their cluster assignments can be expressed as follows: omics1-c1=[1, 1, 0, 0, 1] omics1-c2=[0, 1, 0, 1, 0] omics1-c3=[1, 0, 0, 1, 0] omics2-c1=[1, 0, 1, 1, 1] omics2-c2=[0, 1, 1, 0, 0] omics3-c1=[1, 0, 0, 1, 1] In the matrix combining these, each row can be considered a "cluster feature." Using this matrix as input, we perform the "Consensus Clustering" described in Non-Patent Document 2. The final result is a "cluster of clusters" in which similar clusters extracted from each omics are merged together.
[0025] [Clustering Methods in Multi-Omics Data Analysis] In addition to the above Non-Patent Documents 1 to 3, clustering analysis is also used in the analysis of multi-omics data. To date, several clustering methods specialized for multi-omics data analysis have been developed. Although these are not techniques based on internal criteria, they are equally effective in obtaining more valid clusters (especially from a biological perspective).
[0026] [Generalizing the concept of multi-omics] The concept of multi-omics is generally referred to as "multi-view" or "multi-modal" in the field of machine learning (omics = view / modal). Multi-view data, including multi-omics data, is characterized by a data structure that contains signals from multiple different sources for a single sample. For example, a set of data containing a face photo and a person's profile has two views, "image" and "text," and when the image and text are processed together, it becomes multi-view. The same is true for omics data; suppose each sample contains data measuring gene expression data, methylation data, and protein expression data. In this case, each omics corresponds to a view, and when all three are handled simultaneously, it becomes multi-omics.
[0027] "A robustness metric for biological data clustering algorithms," Yuping Lu et al., [Retrieved September 9, 2022], Internet (https: / / pubmed.ncbi.nlm.nih.gov / 31874625 / ); "Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data," Stefano Monti et al., [Retrieved September 9, 2022], Internet (https: / / link.springer.com / article / 10.1023 / A:1023949509487); "Multiplatform analysis of 12 cancer types reveals molecular classification within and across tissues of origin," Katherine A Hoadley et al., [Retrieved September 9, 2022], Internet (https: / / pubmed.ncbi.nlm.nih.gov / 25109877 / ).
[0028] [Problems with the Prior Art] The technology described in Non-Patent Document 1 uses results from various settings for one algorithm, making it impossible to compare settings between them. For example, it cannot be used to determine an appropriate number of clusters (one of the settings). Furthermore, the technology described in Non-Patent Document 1 does not allow evaluation of the properties of the clusters. Because the evaluation is based on changes in multiple results, information about the obtained clusters cannot be obtained. Furthermore, the technology described in Non-Patent Document 1 does not take biological information into consideration. Obtaining results that are common to various settings does not lead to biological robustness.
[0029] Furthermore, the technology described in Non-Patent Document 2 outputs multiple cluster assignments by repeating subsamples, resulting in the presence of samples that are not used in clustering. Furthermore, the technology described in Non-Patent Document 2 repeatedly performs the same sample population, so the resulting clusters can only take into account one omics. Therefore, if there is a bias in a specific omics, the resulting clusters will be affected by that bias. Furthermore, the technology described in Non-Patent Document 2 does not take biological information into account. Therefore, the commonality of multiple subsamples obtained from a single omics does not translate to biological robustness.
[0030] Furthermore, in the technology described in Non-Patent Document 3, the results of clustering performed on each omics are stored in a binary vector and directly input into "Consensus Cluster," and this method does not provide an internal criterion for the output cluster. Therefore, although a score for "stability" proposed by "Consensus Clustering" can be obtained, this is a viewpoint on the non-determinism of the algorithm and has no biological background.
[0031] [Challenges in Clustering Omics Data] Clustering analysis of omics data, like general clustering analysis, is performed in an unsupervised setting, and there are challenges with evaluation based on internal criteria, etc. In addition, a characteristic of omics data is that, since it is a measurement of living organisms, it has the property of including noise derived from biological reactions (cellular heterogeneity, temporal changes due to dynamics) in addition to measurement-related noise. A challenge in clustering omics data is to develop a method that is biologically valid (hereinafter referred to as "biologically robust") and less susceptible to such noise and bias. However, conventional techniques such as those described in Non-Patent Documents 1 to 3 have not been able to adequately address these challenges.
[0032] The present invention has been made in consideration of the above circumstances, and one embodiment of the present invention provides an evaluation method and evaluation device capable of outputting an index indicating the biological robustness of clusters of multi-omics data. Another embodiment of the present invention provides an evaluation method and evaluation device capable of outputting an index indicating the biological robustness of clustering results of multi-omics data. Yet another embodiment of the present invention provides a visualization method and visualization device capable of visualizing the results of evaluation by these evaluation methods and evaluation devices.
[0033] An evaluation method according to a first aspect of the present invention is a method for evaluating clusters, which is executed by an evaluation device having a processor. The processor acquires cluster assignment information obtained from clustering performed independently and by any method for each omics on multi-omics data consisting of two or more omics, calculates the degree of agreement between the cluster assignments of different omics for each cluster from the cluster assignment information, calculates a first quantitative index for each cluster based on the degree of agreement, which indicates the biological robustness of the cluster, and outputs the calculated first quantitative index.
[0034] In the evaluation method according to the second aspect, in the first aspect, the processor, in calculating the degree of similarity, counts the number of times that each pair of samples belongs to the same cluster in the cluster assignment information for all clusters, for each omics, for all pairs of samples belonging to each cluster, weights the counted number for each omics by a specified weight across all omics, and calculates the degree of similarity for each cluster from the weighted results; and in calculating the first quantitative index, calculates a first statistic for the degree of similarity for each cluster as the first quantitative index.
[0035] The evaluation method according to the third aspect is the second aspect, in which the processor calculates one or more of the sum, average, median, minimum, maximum, and mode of the weighted counts as the first statistical quantity.
[0036] The evaluation method according to the fourth aspect is the second or third aspect, wherein the processor weights each omics by multiplying the number of times it has been counted by a designated constant.
[0037] The evaluation method according to the fifth aspect is the fourth aspect, wherein the processor performs weighting by changing the constant by which some of the two or more omics are multiplied and the constant by which the remaining omics are multiplied.
[0038] The evaluation method according to the sixth aspect is any one of the second to fifth aspects, wherein the processor weights the results of counting for any combination of omics when the results match specified conditions, and when the results of counting for any combination of samples when the results match specified conditions.
[0039] In the evaluation method of the seventh aspect, in the first aspect, the processor calculates the similarity for each cluster of each omics with all clusters belonging to different omics in calculating the degree of similarity, and calculates a first quantitative index using a second statistical quantity for the similarity as the degree of similarity.
[0040] The evaluation method according to the eighth aspect is the seventh aspect, wherein the processor calculates one or more of the sum, average, median, minimum, maximum, and mode of the similarities as the second statistical quantity.
[0041] In the evaluation method of the ninth aspect, in the seventh or eighth aspect, the processor calculates one of an inner product between clusters, a cosine similarity, a Jaccard coefficient, a Hamming distance, a Dice coefficient, and a correlation function as the similarity.
[0042] The evaluation method of the tenth aspect is any one of the seventh to ninth aspects, in which the processor assigns a weight to each omics in the calculation of similarity when any combination of omics meets specified conditions and when any combination of samples meets specified conditions.
[0043] An evaluation method according to an eleventh aspect is a method for evaluating clustering results, which is executed by an evaluation device having a processor, in which the processor calculates a first quantitative index for each cluster using the evaluation method according to any one of the first to tenth aspects, calculates a third statistic of the first quantitative index, and outputs the third statistic as a second quantitative index indicating the biological robustness of the clustering result.
[0044] The evaluation method according to the twelfth aspect is the eleventh aspect, in which the processor calculates one of the sum, average, median, minimum, maximum, and mode of the first quantitative index as the third statistical quantity.
[0045] In addition, a program that causes a computer to execute the evaluation method according to any one of the first to twelfth aspects, and a non-transitory, tangible recording medium on which computer-readable code of such a program is recorded, can also be cited as aspects of the present invention.
[0046] A visualization method according to a thirteenth aspect is a visualization method for cluster evaluation, which is executed by an evaluation device having a processor. The processor acquires cluster assignment information obtained from clustering performed independently for each omics by an arbitrary method for multi-omics data consisting of two or more omics, expresses each cluster extracted for each omics as a vector having samples belonging to each cluster as components, compresses the vector to three or less dimensions by an arbitrary dimensionality reduction method, plots each cluster on coordinates in the compressed dimensions, and outputs a first quantitative index calculated by the evaluation method according to any one of the first to twelfth aspects in association with each cluster.
[0047] A visualization method according to a fourteenth aspect is a method for visualizing cluster evaluation, which is executed by an evaluation device having a processor, in which the processor constructs a graph structure in which each cluster is a node and the similarity is an edge, based on the similarity calculated by the evaluation method according to the seventh aspect, and outputs the constructed graph structure and a first quantitative index calculated by the evaluation method according to any one of the first to tenth aspects, associating each cluster in the graph structure with the first quantitative index.
[0048] In addition, a program that causes a computer to execute the visualization method according to the thirteenth or fourteenth aspect, and a non-transitory, tangible recording medium on which computer-readable code of such a program is recorded, can also be cited as aspects of the present invention.
[0049] The evaluation device according to a fifteenth aspect is a cluster evaluation device comprising a processor, which acquires cluster assignment information obtained from clustering performed independently and by any method for each omics on multi-omics data consisting of two or more omics, calculates the degree of agreement between the cluster assignments of different omics from the cluster assignment information, calculates a first quantitative index for each cluster based on the degree of agreement, which indicates the biological robustness of the cluster, and outputs the calculated first quantitative index.
[0050] The evaluation device according to the fifteenth aspect may have a configuration for executing the same processes as those of the second to tenth aspects.
[0051] In the evaluation device of the 16th aspect, in the 15th aspect, the processor, in calculating the degree of similarity, counts the number of times that each pair of samples belongs to the same cluster in the cluster assignment information for all clusters, for each omics, for all pairs of samples belonging to each cluster, weights the number of times counted for each omics by a specified weight across all omics, and calculates the degree of similarity for each cluster from the weighted results; and in calculating the first quantitative index, calculates a first statistic for the degree of similarity for each cluster as the first quantitative index.
[0052] In the evaluation device of the seventeenth aspect, in the fifteenth aspect, the processor calculates the similarity between each cluster of each omics and all clusters belonging to different omics in calculating the degree of similarity, and calculates a first quantitative index using a second statistical quantity of the similarity as the degree of similarity.
[0053] The evaluation device of the 18th aspect is an evaluation device for evaluating clustering results, and is equipped with a processor. The processor calculates a first quantitative index for each cluster using the evaluation device of any one of the 15th to 17th aspects, calculates a third statistic of the first quantitative index, and outputs the third statistic as a second quantitative index that indicates the biological robustness of the clustering result.
[0054] The evaluation device according to the eighteenth aspect may have a configuration for executing the same processing as that of the twelfth aspect.
[0055] A visualization device according to a 19th aspect is a visualization device for cluster evaluation, and includes a processor. The processor acquires cluster assignment information obtained from clustering performed independently for each omics by an arbitrary method for multi-omics data consisting of two or more omics, expresses each cluster extracted for each omics as a vector whose components are samples belonging to each cluster, compresses the vector to three or less dimensions using an arbitrary dimensionality reduction method, plots each cluster on coordinates in the compressed dimensions, and outputs a first quantitative index calculated by an evaluation device according to any one of the 15th to 17th aspects in association with each cluster.
[0056] A visualization device according to the twentieth aspect is a visualization device for cluster evaluation, and includes a processor. The processor constructs a graph structure in which each cluster is a node and the similarity is an edge based on the similarity calculated by the evaluation device according to the seventeenth aspect, and outputs the constructed graph structure and a first quantitative index calculated by the evaluation device according to any one of the fifteenth to seventeenth aspects, associating each cluster in the graph structure with the first quantitative index.
[0057] FIG. 1 is a conceptual diagram showing clustering of multi-omics data. FIG. 2 is a diagram showing the configuration of an evaluation device according to a first embodiment. FIG. 3 is a diagram showing the configuration of a processing unit. FIG. 4 is a flowchart showing processing of a cluster evaluation method. FIG. 5 is a diagram showing calculation of a first quantitative index. FIG. 6 is a diagram showing calculation of a first quantitative index by weighting the counting number. FIG. 7 is a diagram showing an example of index calculation taking into account the similarity between clusters. FIG. 8 is a flowchart showing processing of a clustering result evaluation method. FIG. 9 is a diagram showing the functional configuration of a processing unit according to a second embodiment. FIG. 10 is a flowchart showing processing of a cluster evaluation visualization method. FIG. 11 is a diagram showing an example of cluster evaluation visualization. FIG. 12 is another flowchart showing processing of a cluster evaluation visualization method. FIG. 13 is a diagram showing another example of cluster evaluation visualization.
[0058] Hereinafter, embodiments of an evaluation method and evaluation device, and a visualization method and visualization device according to the present invention will be described. In the description, the accompanying drawings will be referred to as necessary.
[0059] [First Embodiment] [Concept of Clustering Multi-omics Data] First, clustering of multi-omics data will be described with reference to the conceptual diagram in Fig. 1. In the example of Fig. 1, a sample population is composed of data on cells from, for example, breast cancer patients, and this data is measured for three omics (see part (a) in the figure). Each omics data is high-dimensional data (e.g., each sample has thousands to hundreds of thousands of features).
[0060] 1 assumes that visualization using dimensionality reduction is applied after clustering (similar to the visualization example described below). In this case, the samples circled in part (a) of FIG. 1 (e.g., subtypes 600A to 600D) indicate breast cancer subtypes. Since the subtypes are used in subsequent steps, it is preferable to determine them appropriately.
[0061] However, for example, if subtypes are determined based solely on the genome (omics 1), four subtypes will be obtained in the example of Figure 1, but if clustering is performed using a different setting or method, the number of subtypes may be reduced to, for example, two (subtypes 602A to 602B in part (b) of Figure 1).
[0062] Here, we turn our attention to other omics. Even if the omics are different, they are the measurement results of the same sample population 600, so we assume that common properties potentially exist across the omics. Therefore, by identifying clusters commonly obtained across different omics (omics 2: epigenome, omics 3: proteome), it is expected that "biologically plausible subtypes" can be defined.
[0063] [Configuration of Evaluation Apparatus] FIG. 2 is a diagram showing the schematic configuration of the evaluation apparatus 10 (evaluation apparatus) according to the first embodiment of the present invention. As shown in FIG. 2, the evaluation apparatus 10 includes a processing unit 100 (processor, computer), a storage unit 200, a display unit 300, and an operation unit 400. These components are interconnected to transmit and receive necessary information. Various installation configurations can be adopted for these components. Each component may be installed in a single location (e.g., within a single enclosure or room) or may be installed in remote locations and connected via a network. The evaluation apparatus 10 can also connect to an external server 500 and / or an external database 510 via a network NW such as the Internet, and can acquire multi-omics data, receive cluster assignment information, and the like as needed. Furthermore, the evaluation apparatus 10 can store processing results and the like in the external server 500 and / or the external database 510.
[0064] [Configuration of Processing Unit] FIG. 3 is a diagram showing the configuration of the processing unit 100. As shown in the figure, the processing unit 100 includes a processor 110 (processor, computer), a ROM 130 (Read Only Memory, ROM), and a RAM 140 (Random Access Memory, RAM). The processor 110 performs overall control of the processing performed by each unit of the processing unit 100 and has the functions of an information acquisition unit 112, a coincidence calculation unit 114, a quantitative index calculation unit 116, and an output control unit 118. The functions of these units are summarized as follows: the information acquisition unit 112 (processor) acquires (receives or calculates) cluster assignment information; the coincidence calculation unit 114 (processor) calculates the coincidence between the cluster assignments of different omics for each cluster based on the cluster assignment information; the quantitative index calculation unit 116 (processor) calculates a first quantitative index indicating the biological robustness of the cluster for each cluster based on the coincidence; and the output control unit 118 (processor) outputs the calculated first quantitative index and the like to the storage unit 200 and / or monitor 310.
[0065] The information acquisition unit 112 can receive multi-omics data and cluster assignment information from an external server 500, an external database 510, etc., via the network NW, and / or from a recording medium such as the storage unit 200. The output control unit 118 can display the received or generated information and calculation results (first quantitative index, second quantitative index, etc.) on the monitor 310 and / or store them in the storage unit 200.
[0066] The functions of each unit of the processing unit 100 and the processor 110 described above can be realized using various processors and recording media. The various processors include, for example, a central processing unit (CPU), which is a general-purpose processor that executes software (programs) to realize various functions. The various processors also include a graphics processing unit (GPU), which is a processor specialized for image processing, and a programmable logic device (PLD), such as a field programmable gate array (FPGA), whose circuit configuration can be changed after manufacturing. A configuration using a GPU is effective for image learning and recognition. Furthermore, the various processors described above also include dedicated electrical circuits, such as an application-specific integrated circuit (ASIC), which is a processor with a circuit configuration designed specifically to execute specific processing.
[0067] The functions of each unit may be realized by a single processor, or by multiple processors of the same or different types (e.g., multiple FPGAs, a combination of a CPU and an FPGA, or a combination of a CPU and a GPU). Furthermore, multiple functions may be realized by a single processor. Examples of multiple functions configured by a single processor include: a first configuration, as typified by a computer, in which a single processor is configured by combining one or more CPUs and software, and this processor realizes multiple functions; a second configuration, as typified by a system-on-chip (SoC), in which a processor is used to realize the functions of the entire system on a single IC (Integrated Circuit) chip; and various functions are thus configured as hardware structures using one or more of the various processors described above. Furthermore, the hardware structures of these various processors are, more specifically, electrical circuits combining circuit elements such as semiconductor devices. These electrical circuits may realize the above-mentioned functions using logical operations such as logical sum, logical product, logical negation, exclusive OR, and combinations of these.
[0068] When the processor or electrical circuit executes the software (program), the computer-readable code for the software to be executed (e.g., various processors and electrical circuits constituting the processing unit 100, and / or a combination thereof) is stored in a non-transitory and tangible recording medium such as ROM 130, and the computer references the software. The software stored in the non-transitory and tangible recording medium includes programs (e.g., evaluation programs, visualization programs) for executing the evaluation method and visualization method of the present invention and data used during execution (e.g., multi-omics data, cluster assignment information, setting values of hyperparameters required for program execution, etc.). The code may be recorded in various recording media such as magneto-optical recording devices and semiconductor memories instead of ROM 130. During processing using the software, for example, RAM 140 is used as a temporary storage area, and data stored in a recording medium such as an EEPROM (Electronically Erasable and Programmable Read Only Memory) or flash memory (not shown) may also be referenced. The storage unit 200 may also be used as a "non-transitory and tangible recording medium."
[0069] The processing using the processing unit 100 having the above-described configuration will be described in detail later.
[0070] [Configuration of the storage unit] The storage unit 200 is composed of various storage devices such as a magneto-optical recording device, a hard disk, and a semiconductor memory, and their control units, and can store the above-mentioned multi-omics data, cluster assignment information, the first quantitative index, the second quantitative index, the degree of agreement between cluster assignments, the similarity between clusters, etc.
[0071] [Configuration of Display Unit] The display unit 300 includes a monitor 310 (display device) configured with a display such as a liquid crystal display, and can display input data, execution results of the evaluation method (evaluation program) and visualization method (visualization program), etc. The monitor 310 may be configured with a touch panel display to accept user instruction input.
[0072] [Configuration of Operation Unit] The operation unit 400 includes a keyboard 410 and a mouse 420, and a user can perform operations related to the execution of the evaluation method (evaluation program) and visualization method (visualization program) according to the present invention, the display of results, etc. via the operation unit 400. The operation unit 400 may also include other operation devices.
[0073] [Cluster Evaluation Method] A cluster evaluation method executed by the evaluation device 10 (processor) will be specifically described.
[0074] [Acquisition of Cluster Assignment Information] Figure 4 is a flowchart showing the processing of the cluster evaluation method. As shown in part (a) of Figure 5, the information acquisition unit 112 (processor) acquires cluster assignment information obtained from clustering performed independently for each omics using an arbitrary method for multi-omics data consisting of two or more omics (step S100: information acquisition step). Note that in the present invention, the number of omics may be two or more, and there is no upper limit.
[0075] The "cluster assignment information" acquired in step S100 is the result of clustering performed on the multi-omics data using an arbitrary method independent for each omics, and holds information on "to which cluster each sample is assigned for each omics" as shown in part (b) of Figure 5. Note that the information acquisition unit 112 may receive pre-calculated cluster assignment information from a recording device such as the external server 500 or the external database 510, or may receive multi-omics data from these recording devices, assign clusters, and thereby generate cluster assignment information.
[0076] [Calculation of coincidence] The coincidence calculation unit 114 (processor) calculates the coincidence between the cluster assignments of different omics for each cluster from the acquired cluster assignment information (step S110: coincidence calculation step). The coincidence calculation unit 114 counts the number of times all pairs of samples belonging to each sample belong to the same cluster in the cluster assignment information for all clusters, weights the counted number of times by a specified weight, and can calculate the coincidence for each cluster from the weighted result.
[0077]
[0078]
[0079]
[0080]
[0081] Thus, the coincidence calculation unit 114 can calculate the coincidence of cluster assignments in all omics using the following formula (7) (step S110).
[0082]
[0083]
[0084] The quantitative index calculation unit 116 can determine which statistical quantity to use to calculate the first quantitative index through a user operation via the keyboard 410 and / or the mouse 420 (operation unit 400). The quantitative index calculation unit 116 may also determine which statistical quantity to use without relying on a user operation.
[0085] The output control unit 118 (processor) outputs the calculated first quantitative index (step S130: index output process). The first quantitative index is an index indicating the biological robustness of the cluster, and the larger the index value, the higher the biological robustness of the cluster. The output control unit 118 may store the first quantitative index in the memory unit 200 or may display it on the monitor 310 (display device). The display on the monitor 310 may be, for example, letters, numbers, figures, symbols, or a combination thereof indicating the index, and the color of the letters, numbers, figures, symbols, etc. may be changed depending on the value of the index. In addition, the output control unit 118 preferably outputs the cluster and the first quantitative index for that cluster in association with each other.
[0086] In the first embodiment, the biological robustness of a cluster of multi-omics data can be evaluated using such a first quantitative index. The same applies to the following aspects.
[0087]
[0088] Figure 6 shows how the first quantitative index is calculated by weighting the counted number of times. The information acquisition unit 112 (processor) acquires cluster assignment information for each omics (step S100: information acquisition step). Part (a) of Figure 6 shows examples of cluster assignment information for omics 1 to 3 (Tables 720 to 724). Note that in Tables 720 to 724 and Tables 726 and 728, the upper triangle is symmetrical to the lower triangle (similar to Table 716 in Figure 5).
[0089] [Weighting Mode 1] In the above-mentioned formula (7), the weights for each omics are all set to 1, but the weights for each omics are changed in mode 1. Specifically, the coincidence calculation unit 114 can assign weights using the following formula (9), for example.
[0090]
[0091]
[0092] The quantitative index calculation unit 116 can calculate the first quantitative index from the weighted result using the above-mentioned formula (8) (step S120: index calculation process), and the output control unit 118 can output the calculated first quantitative index (step S130: index output process). As described above for formula (8), the statistic for the weighted result (first statistic) can also use one or more arbitrary statistics such as sum, mean, median, minimum, maximum, mode, etc. The type of statistic to be used corresponds to the type of function F to be used in formula (8) (the same applies to the following aspects).
[0093] [Weighting Mode 2] In mode 2, the coincidence calculation unit 114 multiplies the weight by 3 only when the results of omics 1 and omics 2 coincide, as shown in the following formula (11).
[0094]
[0095] The condition "when the results of omics 1 and omics 2 match" and the value "three times" are examples of weighting, and other conditions and values may be used. The results of weighting using formula (11) are shown in table 728 in part (b2) of Figure 6. In embodiment 2, the calculation and output of the first quantitative index can be performed using formula (8), as in embodiment 1.
[0096] [Other Weighting Aspects] The above-described aspects 1 and 2 are examples of weighting, and the coincidence calculation unit 114 may also perform weighting using other methods, such as "conditionally matching the results of three or more omics" (one aspect of "weighting is performed when the counting results for any combination of omics match the specified condition") or "weighting is performed only when the results of a specific sample pair match specific omics" (one aspect of "weighting is performed when the counting results for any combination of samples match the specified condition"). Furthermore, the weighting method may be "addition of weights" in addition to multiplication of weights. Calculating a quotient or difference is possible by setting a weight less than 1 or less than 0.
[0097] [Aspect of Calculating Similarity Between Clusters as an Index] In the present invention, the coincidence calculation unit 114 (processor) calculates the similarity between each cluster of each omics and all clusters belonging to different omics in calculating the coincidence, and the quantitative index calculation unit 116 (processor) calculates the first quantitative index using the second statistical value of the similarity as the coincidence. Figure 7 is a diagram showing an example of index calculation taking into account the similarity between clusters.
[0098] The information acquisition unit 112 can acquire cluster assignment information in the same manner as in step S100 of Fig. 4 (parts (a) and (b) of Fig. 7), and the similarity calculation unit 114 calculates the similarity between each cluster of each omics and all clusters belonging to different omics. The similarity calculation unit 114 can calculate, for example, any of the inner product between clusters, cosine similarity, Jaccard coefficient, Hamming distance, Dice coefficient (Dice), and correlation function as the similarity, but the similarity is not limited to these examples.
[0099]
[0100] A large value of this first quantitative index indicates a high biological robustness of the cluster. In the example above, the robustness of Cluster 1 of Omics 1 (1.0) is higher than the robustness of Cluster 1 of Omics 3 (0.67).
[0101] [Weighting Each Omic in Similarity Calculation] The coincidence calculation unit 114 may weight each omics in the similarity calculation described above. For example, as shown in part (d) of FIG. 7, if common noise or the like is known in a specific omics pair (the pair of omics 1 and omics 3 in the example of the same figure), applying a weight of less than 1 can prevent overfitting to the noise. Conversely, as shown in part (b2) of FIG. 6, if it is determined that a match in a specific omics pair (the pair of omics 1 and omics 2 in the example of the same figure) increases reliability, a weight greater than 1 may be applied or multiplied.
[0102] [Method for Evaluating Clustering Results] In the above-described method for evaluating clusters, one index (first quantitative index) is assigned to each cluster. In the present invention, as described below, a quantitative index of the biological robustness of the clustering results can be obtained by calculating one or more statistics from the first quantitative index for each cluster.
[0103] 8 is a flowchart showing the process of the clustering result evaluation method. The components of the processing unit 100 (processor) calculate a first quantitative index for each cluster in the same manner as described above for the cluster evaluation method (step S200: information acquisition step, degree of coincidence calculation step, index calculation step, index output step). The process of step S200 is the same as the process described above for steps S100 to S120 in FIG. 4, and a detailed description thereof will be omitted. Note that, in the process of step S200, any of the variations described above for steps S100 to S120 can be used.
[0104] The quantitative index calculation unit 116 (processor) calculates one of the sum, mean, median, minimum, maximum, and mode of the first quantitative indexes as a third statistical quantity (step S210: statistical quantity calculation step). The quantitative index calculation unit 116 can determine which statistical quantity to calculate in response to a user operation via the operation unit 400 or automatically without relying on a user operation. The output control unit 118 (processor) outputs the third statistical quantity as a second quantitative index indicating the biological robustness of the clustering result (i.e., with respect to the clustering method or clustering setting value set) (step S220: index output step).
[0105] [Second Embodiment] [Visualization of Cluster Evaluation (Aspect 1)] In the present invention, the index described above in the first embodiment can be used to visualize cluster evaluation (evaluation of the biological robustness of clusters or cluster assignment). FIG. 9 is a diagram showing the configuration of a processing unit 100A (processor) in an evaluation device 10 (evaluation device) according to the second embodiment. The processing unit 100A further includes a visualization unit 120 (processor) in addition to the processing unit 100 (see FIG. 3) according to the first embodiment. The other configuration of the evaluation device 10 according to the second embodiment is the same as that of the first embodiment.
[0106] Fig. 10 is a flowchart showing the process of the cluster evaluation visualization method (aspect 1), and Fig. 11 is a diagram showing the visualization of the cluster evaluation. In the flowchart of Fig. 10, the calculation of the first quantitative index is the same as in the first embodiment (see steps S100 to S120 in Fig. 4), and the information acquisition unit 112 acquires cluster assignment information.
[0107]
[0108] The visualization unit 120 compresses the obtained vectors to three or less dimensions using an arbitrary dimension reduction method (step S310: dimension reduction process) and plots each cluster on coordinates in the compressed dimensions (step S320: plotting process). The output control unit 118 further outputs a first quantitative index calculated using the same method as in the first embodiment in association with each plotted cluster (step S330: output process).
[0109] Part (b) of Figure 11 shows the state in which the above-mentioned vectors have been compressed into two dimensions using principal component analysis. In this figure, each point represents a cluster, and it can be seen that similar clusters are plotted at similar coordinates. Principal component analysis is a method for compressing multidimensional data (here, vectors) into lower dimensions while preserving as much information as possible. In addition to principal component analysis, other dimensionality reduction methods include the UMAP (Uniform Manifold Approximation and Projection) method and the t-SNE (t-Distributed Stochastic Neighbor Embedding) method. It is preferable that the visualization unit 120 select a dimensionality reduction method appropriate for the properties of the data. The visualization unit 120 may also select a dimensionality reduction method in response to a user operation via the operation unit 400.
[0110] Part (c) of FIG. 11 is a table (table 722) showing the first quantitative index for each cluster of each omics, and part (b) of FIG. 11 displays the first quantitative index included in this table near the plot of each cluster (one example of association output). The output control unit 118 may display the plot results on the monitor 310 (display device) or store them in the storage unit 200. A diagram of the plot results may also be output by a printer (not shown). This visualization allows the user to easily grasp the biological robustness of the clusters.
[0111] [Visualization of Cluster Evaluation (Aspect 2)] Fig. 12 is a flowchart showing the processing of the visualization method of cluster evaluation (Aspect 2), and Fig. 13 is a diagram showing the visualization of cluster evaluation. In the flowchart of Fig. 12, each part of the processing unit 100 (processor) can calculate the first quantitative index in the same way as in the first embodiment (see steps S100 to S120 in Fig. 4; step S410). Furthermore, each part of the processing unit 100 (processor) can calculate the similarity (step S400) in the same way as described above with reference to Fig. 7.
[0112] Based on the similarities calculated in step S400, the visualization unit 120 constructs a graph structure in which each cluster is a node and the similarities are edges (step S420: graph structure construction process), and the output control unit 118 outputs the constructed graph structure and the first quantitative index by associating each cluster in the graph structure with the first quantitative index (step S430: output process).
[0113] FIG. 13 illustrates the visualization process according to aspect 2. Part (a) of FIG. 13 is a table showing similarities (table 718; identical to part (e) of FIG. 7 ), and part (b) of FIG. 13 is a diagram illustrating a graph structure. In this diagram, the first quantitative index (same as part (b) of FIG. 11 ) shown in part (c) of FIG. 13 is output near each cluster. Note that in the graph structure diagram (part (b) of FIG. 13 ), the thickness of the edges indicates the magnitude of similarity. The output control unit 118 may perform emphasis processing, such as changing the size of the symbols (ellipses in FIG. 13 ) indicating clusters according to the magnitude of the first quantitative index, or changing the line type or color of the symbols, or changing the line type or color of the edges according to the magnitude of similarity. The output control unit 118 may display the constructed graph structure diagram on the monitor 310 (display device) or store it in the storage unit 200. The diagram may also be output by a printer (not shown). This type of visualization also allows users to easily grasp the biological robustness of the clusters.
[0114] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described aspects, and various modifications are possible without departing from the spirit of the present invention.
[0115] 10 Evaluation device 100 Processing unit 100A Processing unit 110 Processor 112 Information acquisition unit 114 Matching degree calculation unit 116 Quantitative index calculation unit 118 Output control unit 120 Visualization unit 130 ROM 140 RAM 200 Storage unit 300 Display unit 310 Monitor 400 Operation unit 410 Keyboard 420 Mouse 500 External server 510 External database 600 Sample population 600A Subtype 600B Subtype 600C Subtype 600D Subtype 602A Subtype 602B Subtype NW Network S100 to S220 Each step of the evaluation method S300 to S330 Each step of the visualization method S400 to S430 Each step of the visualization method
Claims
1. A method for evaluating clusters, executed by an evaluation device equipped with a processor, wherein the processor: acquires cluster assignment information obtained from clustering performed independently for each omics using an arbitrary method on multi-omics data consisting of two or more omics; calculates, for each cluster, a degree of agreement between the cluster assignments of different omics from the cluster assignment information; calculates, for each cluster, a first quantitative index indicating the biological robustness of the cluster based on the degree of agreement; and outputs the calculated first quantitative index.
2. The evaluation method of claim 1, wherein the processor, in calculating the degree of similarity, counts the number of times that each pair of samples belongs to the same cluster for each omics in the cluster assignment information for all clusters, for all pairs of samples belonging to each cluster, weights the number of times counted for each omics by a specified weight across all omics, and calculates the degree of similarity for each cluster from the weighted results; and, in calculating the first quantitative index, calculates a first statistical quantity for the degree of similarity for each cluster as the first quantitative index.
3. The evaluation method according to claim 2, wherein the processor calculates one or more of the sum, average, median, minimum, maximum, and mode of the weighted counts as the first statistical quantity.
4. The evaluation method according to claim 2 or 3, wherein the processor assigns the weight to each omics by multiplying the counted number by a designated constant.
5. The evaluation method according to claim 4, wherein the processor performs the weighting by changing the constant by which some of the two or more omics are multiplied and the constant by which the remaining omics are multiplied.
6. The evaluation method of claim 2 or 3, wherein the processor performs the weighting when the counting result for any combination of omics matches a specified condition, and when the counting result for any combination of samples matches a specified condition.
7. The evaluation method of claim 1, wherein the processor, in calculating the degree of similarity, calculates the similarity for each cluster of each omics with all clusters belonging to different omics, and calculates the first quantitative index using a second statistical quantity for the similarity as the degree of similarity.
8. The evaluation method according to claim 7, wherein the processor calculates one or more of the sum, average, median, minimum, maximum, and mode of the similarities as the second statistical quantity.
9. The evaluation method according to claim 7 or 8, wherein the processor calculates, as the similarity, one of an inner product between the clusters, a cosine similarity, a Jaccard coefficient, a Hamming distance, a Dice coefficient, and a correlation function.
10. The evaluation method described in claim 7 or 8, wherein the processor, in calculating the similarity, assigns a weight to each omics when any combination of omics meets specified conditions and when any combination of samples meets specified conditions.
11. A method for evaluating clustering results, executed by an evaluation device having a processor, wherein the processor calculates the first quantitative index for each cluster using the evaluation method described in any one of claims 1, 2, 3, 7, and 8, calculates a third statistic of the first quantitative index, and outputs the third statistic as a second quantitative index indicating the biological robustness of the clustering result.
12. The evaluation method according to claim 11, wherein the processor calculates one of the sum, mean, median, minimum, maximum, and mode of the first quantitative index as the third statistical quantity.
13. A visualization method for cluster evaluation executed by an evaluation device equipped with a processor, wherein the processor: acquires cluster assignment information obtained from clustering performed independently for each omics using an arbitrary method for multi-omics data consisting of two or more omics; expresses each cluster extracted for each omics as a vector having samples belonging to each cluster as components; compresses the vector to three or less dimensions using an arbitrary dimensionality compression method; plots each cluster on coordinates in the compressed dimensions; and outputs the first quantitative index calculated by the evaluation method described in any one of claims 1, 2, 3, 7, and 8 in association with each cluster.
14. A visualization method for cluster evaluation executed by an evaluation device equipped with a processor, wherein the processor constructs a graph structure in which each cluster is a node and the similarity is an edge, based on the similarity calculated by the evaluation method described in claim 7, and outputs the constructed graph structure and the first quantitative index calculated by the evaluation method described in any one of claims 1, 2, 3, 7, and 8, associating each cluster in the graph structure with the first quantitative index.
15. A cluster evaluation device comprising a processor, which: acquires cluster assignment information obtained from clustering performed independently for each omics using an arbitrary method on multi-omics data consisting of two or more omics; calculates a degree of agreement between cluster assignments of different omics from the cluster assignment information; calculates a first quantitative index for each cluster based on the degree of agreement, which indicates the biological robustness of the cluster; and outputs the calculated first quantitative index.
16. The evaluation device described in claim 15, wherein the processor, in calculating the degree of similarity, counts the number of times that each pair of samples belongs to the same cluster for each omics in the cluster assignment information for all clusters, for all pairs of samples belonging to each cluster, weights the number of times counted for each omics by a specified weight across all omics, and calculates the degree of similarity for each cluster from the weighted results; and in calculating the first quantitative index, calculates a first statistic for the degree of similarity for each cluster as the first quantitative index.
17. The evaluation device described in claim 15, wherein the processor, in calculating the degree of coincidence, calculates the similarity between each cluster of each omics and all clusters belonging to different omics, and calculates the first quantitative index using a second statistical quantity for the similarity as the degree of coincidence.
18. An evaluation device for clustering results, comprising a processor, wherein the processor calculates the first quantitative index for each cluster using the evaluation device described in any one of claims 15 to 17, calculates a third statistic of the first quantitative index, and outputs the third statistic as a second quantitative index indicating the biological robustness of the clustering result.
19. A visualization device for cluster evaluation, comprising a processor, which acquires cluster assignment information obtained from clustering performed independently for each omics using an arbitrary method for multi-omics data consisting of two or more omics, expresses each cluster extracted for each omics as a vector having samples belonging to each cluster as components, compresses the vector to three or less dimensions using an arbitrary dimensionality compression method, plots each cluster on coordinates in the compressed dimensions, and outputs the first quantitative index calculated by the evaluation device described in any one of claims 15 to 17 in association with each cluster.
20. A visualization device for cluster evaluation, comprising a processor, which constructs a graph structure in which each cluster is a node and the similarity is an edge based on the similarity calculated by the evaluation device described in claim 17, and outputs the constructed graph structure and the first quantitative index calculated by the evaluation device described in any one of claims 15 to 17, associating each cluster in the graph structure with the first quantitative index.