Feature selection methods, multi-class classification methods, feature selection devices, multi-class classification devices, and feature sets

By employing feature selection and multi-class classification methods, and utilizing pairwise coupling and binary class classifiers, the problems of sample discrimination accuracy and cost in multi-class classification are solved, achieving high-precision and low-cost sample discrimination in cancer diagnosis.

CN115104028BActive Publication Date: 2025-11-14FUJIFILM CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180014238.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-13
Filing Date
2021-02-05
Publication Date
2025-11-14
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve robust and high-precision sample classification across multiple categories, especially in biological fields such as cancer diagnosis, where sample classification accuracy declines and costs are high, making it impossible to effectively utilize omics data for discrimination.

Method used

A feature selection method is adopted, which uses pairwise coupling and quantitative evaluation of the discriminability between features to construct a multi-class classifier by combining a binary class classifier and a elimination hierarchy method, selects robust discriminative positions, and uses a simple combination of classifiers for multi-class discrimination.

Benefits of technology

It achieves high-precision and low-cost sample discrimination in complex multi-class classification problems, and can screen out a small number of DNA methylation sites from a large number of DNA methylation sites for robust discrimination, which is applicable to fields such as cancer diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104028B_ABST
    Figure CN115104028B_ABST
Patent Text Reader

Abstract

The object of this invention is to provide a multi-class classification method, a multi-class classification apparatus, and a feature selection method, a feature selection apparatus, and a feature set for such multi-class classification, all of which involve selecting feature quantities and classifying samples into any one of multiple classes based on the values ​​of the selected feature quantities. In this invention, the multi-class classification problem is addressed in conjunction with feature quantity selection. Feature quantity selection is a method of pre-selecting, literally, the feature quantities required for subsequent processing (especially multi-class classification in this invention) from a large number of feature quantities possessed by the sample. Multi-class classification is a discrimination problem that determines which of multiple classes a given unknown sample belongs to.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-class classification method, a multi-class classification apparatus, and a feature selection method, a feature selection apparatus, and a feature set for such multi-class classification, which select feature quantities and classify samples into any one of a plurality of classes based on the values ​​of the selected feature quantities. Background Technology

[0002] In recent years, machine learning has made progress in its application and development in industry, but feature selection and multi-class classification remain major challenges. Various feature selection methods exist, but one approach focuses on pairwise coupling of classes (see “Non-Patent Document 1” below). Specifically, the technique described in Non-Patent Document 1 focuses on “binary class classification” with two basic classes, performs pairwise coupling of classes, and focuses on and selects the discriminative power of feature quantities.

[0003] Furthermore, as a method for multi-class classification, there is the known OVO method (One-Versus-One: one-to-one) which repeatedly performs two-class discrimination.

[0004] Furthermore, in fields such as biology, feature selection and multi-class classification methods have been actively researched, particularly for cancer. These generally involve the application of common machine learning methods, such as feature selection methods based on t-tests or information gain, and classification methods based on SVM (Support Vector Machine), random forests, and Naive Bayes. Such techniques are described, for example, in Patent Document 1.

[0005] Previous technical documents

[0006] Non-patent literature

[0007] Non-patent document 1: "Feature selection for multi-class classification using pairwise class discriminatory measure and covering concept", Hyeon Ji et al., ELECTRONICS LETTERS, 16th March 2000, vol.36, No.6, p.524-525

[0008] Patent documents

[0009] Patent Document 1: Japanese Patent Publication No. 2012-505453 Summary of the Invention

[0010] The technical problem to be solved by the invention

[0011] The research described in Non-Patent Document 1 only focuses on feature selection, directly using existing methods in subsequent multi-class classification. Furthermore, the present invention does not explicitly describe any extensions to the set-covering problem as described later. It also fails to verify the independence between feature quantities used for robustness selection, and only assumes basic multi-class classification, without introducing classes that do not require discrimination. Therefore, it is difficult to directly apply to extended multi-class classification. Similarly, the technology described in Patent Document 1 does not consider examining the genome required for discrimination as a set-covering problem in detail.

[0012] Furthermore, in methods that repeatedly perform binary discrimination for multi-class classification, the voting method has already pointed out the problem of "the order of higher-level classifications being unreliable." And the elimination hierarchy method has already pointed out the problem of "difficulty in determining the comparison order."

[0013] In the field of biological feature selection and multi-class classification, in many reported cases based on mRNA expression levels, there is a problem of "accuracy decreasing when the number of classes reaches around 10." For example, in one report of a multi-class cancer classifier developed based on mutation information, the result exceeded an F-value of 0.70, and it could identify 5 types of cancer. Feature selection and multi-class classification based on DNA methylation have also been studied. However, the applicability remains limited to a small number of small-scale experiments.

[0014] In recent years, research has also emerged that applies deep learning. However, due to the inherent uncertainty of omics data, learning cannot proceed smoothly (the sample size is small relative to the number of parameters; even with open data, the number of tumor records that can be obtained is less than 10,000, relative to the existence of hundreds of thousands of methylation sites). Even if it is successful, for example in diagnostic applications, there are unacceptable issues due to the lack of clear reasons for the judgment.

[0015] Thus, in the prior art, it is impossible to robustly and accurately classify a sample with multiple feature values ​​into any one of multiple classes based on the values ​​of a selected subset of feature values.

[0016] This invention was made in view of the following situation, and its object is to provide a multi-class classification method and apparatus capable of robustly and accurately classifying samples having multiple feature values ​​into any one of multiple classes based on the values ​​of a selected subset of feature values. Furthermore, the object of this invention is to provide a feature value selection method, feature value selection apparatus, and feature value set for such multi-class classification.

[0017] means for solving technical problems

[0018] The first aspect of the present invention relates to a feature selection method that selects a set of feature quantities for determining which of two or more N classes a sample belongs to. The feature selection method includes: an input step, inputting a learning dataset consisting of a known set of samples belonging to a given class and a set of feature quantities of the known sample set; and a selection step, selecting, based on the learning dataset, a set of feature quantities required for class determination of an unknown sample belonging to an unknown class from the set of feature quantities. The selection step includes: a quantification step, quantifying the discriminability between two classes based on each feature quantity of the selected set of feature quantities by pairwise coupling of two of the N classes according to the learning dataset; and an optimization step, statistically analyzing the quantified discriminability for all pairwise couplings and selecting a combination of feature quantities that optimizes the statistical results.

[0019] The feature selection method involved in the second method, in the first method, further includes: a first marking step, marking a portion of the given classes as a first group of classes that do not need to be distinguished from each other; and a first exclusion step, excluding the pairwise couplings between the marked first groups of classes that do not need to be distinguished from each other from the expanded pairwise couplings.

[0020] The feature selection method involved in the third method is in the first or second method, and the selection process includes: a similarity evaluation process, which evaluates the similarity between feature quantities based on the discriminability of each feature quantity for each pair of couplings; and a priority setting process, which sets the priority of the feature quantities to be selected based on the similarity evaluation results.

[0021] The feature selection method involved in the fourth approach is similarity in the third approach, which is the repeatability and / or inclusion relationship of the discriminability of each pair of couplings.

[0022] The feature selection method involved in the fifth approach is the same as that in the third or fourth approach, where similarity is the distance or a distance-based metric between each pair of coupled discriminative vectors.

[0023] The feature selection method involved in the sixth method, in any of the methods from the first to the fifth method, also has a selection number input step, wherein the selection number input step inputs the selection number M of the feature quantity in the selection step, and the optimization is based on maximizing the minimum value of the statistical value in all paired couplings of the M selected feature quantities.

[0024] The feature selection method involved in Method 7, in any of Methods 1 to 6, further includes the following optimization steps: an importance input step, which inputs the importance of the class or pairwise discrimination; and a weighting assignment step, which assigns weights based on importance during statistics.

[0025] The feature selection method involved in Method 8 involves selecting more than 25 features in any of Methods 1 to 7 during the selection process.

[0026] The feature quantity selection method involved in the 9th method is the same as that in the 8th method, where the number of feature quantities selected in the selection process is more than 50.

[0027] The feature quantity selection method involved in Method 10 is the same as that in Method 9, where the number of feature quantities selected in the selection process is more than 100.

[0028] The feature selection procedure involved in the 11th aspect of the present invention causes a computer to execute the feature selection method involved in any of the 1st to 10th aspects.

[0029] The multi-class classification method according to the 12th aspect of the present invention, when N is an integer of 2 or more, determines which of N classes a sample belongs to based on the characteristic quantity of the sample. The multi-class classification method includes: an input step and a selection step, which are performed using the characteristic quantity selection method according to any of the 1st to 10th aspects; and a determination step, which performs class determination for an unknown sample based on the selected characteristic quantity group. The determination step includes an acquisition step for obtaining the characteristic quantity value of the selected characteristic quantity group and a class determination step for performing class determination based on the obtained characteristic quantity value. In the determination step, the class determination for the unknown sample is performed by constructing a multi-class discriminator that uses the selected characteristic quantity group in a pairwise coupled manner.

[0030] Figure 1 This is a schematic diagram of a multi-class classification problem with accompanying feature selection, processed in the 12th manner of the present invention. Feature selection (step 1) is a method of pre-selecting features from a large number of features possessed by the sample for subsequent processing (especially multi-class classification in the present invention), as the name suggests (the feature selection method involved in any of the 1st to 10th manners). That is, a large number of features are pre-obtained from a specified dataset (the so-called training dataset), and features (feature sets) required for subsequent processing are selected based on this information. Furthermore, when a (unknown) sample is actually given, multi-class classification is performed only with reference to the small number of pre-selected features (feature sets). In addition, since the unknown sample is classified based on features selected only in the training dataset, robust feature selection is naturally preferred.

[0031] Feature selection is particularly useful when costs (including time and expense) are incurred in order to obtain feature quantities from a reference sample (including acquiring and storing them). Therefore, for example, the mechanism for obtaining feature quantities from reference learning data can differ from the mechanism for obtaining feature quantities from reference unknown samples, or a suitable feature quantity acquisition mechanism can be developed and prepared based on the selection of a small number of feature quantities.

[0032] On the other hand, multi-class classification (step 2) is a discrimination problem that determines which of several classes a given unknown sample belongs to, a general problem in machine learning. However, many real-world multi-class classification problems are not simply about choosing one of N classes. For example, even if multiple classes actually exist, there are cases where the discrimination itself is unnecessary. Conversely, for example, there are multiple groups of samples with different states mixed together in a sample set labeled as class 1. A method that can tolerate such complex and extended multi-class classification is preferred.

[0033] As the simplest feature selection method, it is also possible to evaluate all selection methods that choose a small number of features from a large number of candidates using a training dataset. However, due to the risk of overlearning the training dataset and the large number of candidates that cannot be fully evaluated, a framework is needed.

[0034] This section illustrates an example of applying the first aspect of the present invention (multi-class classification accompanied by feature selection) to the biological field. Inherent DNA methylation patterns exist in cancerous tissues and other body tissues. Furthermore, human blood contains cell-free DNA (cfDNA), and cfDNA originating from cancer is also detected. Therefore, by analyzing the methylation pattern of cfDNA, the presence or absence of cancer can be determined, and if cancer is present, the primary lesion can be identified. In other words, early cancer screening via blood sampling can be achieved, guiding patients to appropriate, more detailed examinations.

[0035] Therefore, determining whether a disease is cancerous or non-cancerous and identifying the tissue of origin from DNA methylation patterns is extremely important. This can be defined as a multi-class classification problem of identifying cancer from blood or normal tissue. However, the human body involves many organs (e.g., 8 major types of cancer and more than 20 types of normal tissue), cancer has subtypes, and even cancers in the same organ can have different states, making it a very difficult classification problem.

[0036] Furthermore, based on the assumptions provided for screening, the aim is to minimize measurement costs, thus prohibiting the direct use of expensive arrays that comprehensively measure methylation sites. Therefore, it is necessary to pre-screen from hundreds of thousands of DNA methylation sites to identify the required few sites; that is, feature selection is required in the preceding stage.

[0037] Therefore, a technique (the method proposed in this invention) that screens a small number of DNA methylation sites from a vast number of sites, and that can distinguish cancer from normal tissue based on these sites, while also determining the characteristic selection and multi-classification of the source tissue, is useful. Furthermore, the number of sites selected from, for example, 300 out of 300,000 DNA methylation sites exceeds 10 to the power of 1,000, thus indicating that a comprehensive exploratory approach is not feasible.

[0038] Therefore, the inventors of this application list DNA methylation sites that act as switches that contribute to robustness in discrimination, and propose a feature selection method based on combinatorial exploration that fully covers pairwise discrimination of the required classes. Furthermore, they propose a method that uses only the robustness-discriminating parts of the selected sites to construct a multi-class classifier by combining a simple binary class classifier with a hierarchical elimination method.

[0039] Therefore, a multi-class classification system is capable of handling feature selections that address various characteristics associated with real-world problems. In fact, it is applicable to multi-class classifications, such as those used in the cancer diagnosis example mentioned above, where the combined categories of cancer and normal cells far exceed 10. The feature selection and multi-class classification method proposed by the inventors of this application is extremely useful in industry.

[0040] Furthermore, this description is merely one specific example, and the 12th aspect of the present invention is not limited to the biological field. In fact, just as many common machine learning techniques can be applied to the biological field, it is permissible to apply techniques developed in the biological field to common machine learning problems.

[0041] The multi-class classification method involved in Method 13, in Method 12, utilizes the statistically significant difference in feature quantities in the learning dataset between paired coupled classes during the quantification process.

[0042] The multi-class classification method involved in Method 14, in Method 12 or 13, in the quantification process, when a threshold set based on the reference learning dataset is given as a feature quantity of an unknown sample belonging to any of the paired coupled classes, utilizes the probability that the class to which the unknown sample belongs can be correctly determined based on the given feature quantity.

[0043] In the multi-class classification method involved in Method 15, in any of Methods 12 to 14, the quantification value of discriminability in the quantification process is the value after multiple tests and corrections on the statistical probability value based on the number of feature quantities.

[0044] The multi-class classification method involved in the 16th method, in any of the 12th to 15th methods, further includes: a subclass setting step, which clusters more than one sample belonging to a class from the learning dataset according to a given feature quantity, thereby forming clusters, and setting each cluster as a subclass in a class; a second labeling step, which labels each subclass in a class as a second non-discriminable class group that does not need to be distinguished from each other in a class; and a second exclusion step, which excludes the pairwise coupling between the labeled second non-discriminable class groups from the unfolded pairwise coupling.

[0045] In the multi-class classification method involved in Method 17, in any of Methods 12 to 16, statistics are the calculation of the sum or average of quantitative values ​​of discriminability.

[0046] The multi-class classification method involved in Method 18, in any of Methods 12 to 17, also has a target threshold input step, wherein the target threshold input step inputs a target threshold T representing the statistical value of the statistical result, and the optimization is to set the minimum value of the statistical value in all paired couplings based on the selected feature quantity as above the target threshold T.

[0047] In the multi-class classification method involved in Method 19, in any of Methods 12 to 18, during the determination process, binary class discriminators that are coupled with each pair and associated with each other by using selected feature sets are constructed, and the binary class discriminators are combined to construct a multi-class discriminator.

[0048] The multi-class classification method involved in Method 20, in any of Methods 12 to 19, further includes: a step of evaluating the similarity between a sample and each class using a binary class discriminator; and a step of constructing a multi-class discriminator based on the similarity.

[0049] The multi-class classification method involved in Method 21, in any of Methods 12 to 20, further includes: a step of evaluating the similarity between a sample and each class using a binary class discriminator; and a step of constructing a multi-class discriminator by reapplying the binary class discriminator used for similarity evaluation between classes to classes with higher similarity.

[0050] In any of the methods from method 12 to method 21, the multi-class classification method involved in method 22 constructs a decision tree that is coupled with each pair of selected feature groups in the decision process, and combines one or more decision trees to form a multi-class discriminator.

[0051] The multi-class classification method involved in Method 23 is the same as that in Method 22, where in the decision-making process, a multi-class discriminator is constructed from decision trees and combinations of decision trees as a random forest.

[0052] The multi-class classification method involved in Method 24, in any of Methods 12 to 23, determines the class to which the live tissue slice belongs from N classes by measuring the omics information of the live tissue slice.

[0053] The multi-class classification method involved in Method 25, in any of Methods 12 to 24, determines the class to which the live tissue slide belongs from N classes by measuring the on / off state information of the omics of the live tissue slide.

[0054] The multi-class classification method involved in Method 26 requires that the number of classes to be identified in any of Methods 12 to 25 be 10 or more.

[0055] The multi-class classification method involved in Method 27 requires that the number of classes to be identified in Method 26 be 25 or more.

[0056] The multi-class classification program according to the 28th aspect of the present invention causes a computer to execute the multi-class classification method according to any one of the 12th to 27th aspects. Furthermore, a non-transitory recording medium containing computer-readable code of the program according to the 28th aspect can also be cited as an aspect of the present invention.

[0057] The feature selection device according to the 29th aspect of the present invention selects a set of feature quantities for determining which of two or more N classes a sample belongs to. The feature selection device includes a first processor that performs the following processes: input processing, inputting a learning dataset consisting of a known set of samples belonging to a given class and a set of feature quantities of the known sample set; and selection processing, selecting, based on the learning dataset, a set of feature quantities required for class determination of an unknown sample belonging to an unknown class from the set of feature quantities. The selection processing includes: quantification processing, quantifying the discriminability between two classes based on each feature quantity of the selected set of feature quantities by pairwise coupling of two of the N classes according to the learning dataset; and optimization processing, statistically analyzing the quantified discriminability for all pairwise couplings and selecting a combination of feature quantities that optimizes the statistical results.

[0058] The multi-class classification apparatus according to the 30th aspect of the present invention, when N is an integer of 2 or more, determines which of N classes a sample belongs to based on the characteristic quantity of the sample. The multi-class classification apparatus includes: a feature quantity selection device according to the 29th aspect; and a second processor, the second processor performing the following processes: input processing and selection processing, using the feature quantity selection device; and determination processing, performing class determination for an unknown sample based on the selected feature quantity group. The determination processing includes an acquisition processing for obtaining the feature quantity values ​​of the selected feature quantity group and a class determination processing for performing class determination based on the acquired feature quantity values. In the determination processing, the class determination for the unknown sample is performed by constructing a multi-class discriminator that uses the selected feature quantity group in a pairwise coupled manner.

[0059] The feature set involved in the 31st aspect of the present invention is used by a multi-class classification device to determine which of two or more N classes a sample belongs to. The feature set has a feature set dataset of samples belonging to each class that is the object. When the discriminability between two classes based on each feature set of the selected feature set is quantified by referring to the feature set dataset through pairwise coupling of two of the N classes, the pairwise coupling is marked as being discriminable with at least one feature.

[0060] In Method 32, the feature set involved in Method 31, when quantifying the discriminability between two classes based on each feature of the selected feature set by referring to the feature set in pairwise coupling by combining two of the N classes, is marked as being discriminable with at least five features in all pairwise couplings.

[0061] In Method 33, the feature set involved in Method 31, when quantifying the discriminability between two classes based on each feature of the selected feature set by combining two of the N classes in pairwise coupling, is marked as being discriminable with at least 10 features in all pairwise couplings.

[0062] In Method 34, the feature set involved in Method 31, when quantifying the discriminability between two classes based on each feature of the selected feature set by combining two of the N classes in pairwise coupling, is marked as being discriminable with at least 60 features in all pairwise couplings.

[0063] The feature set involved in Method 35 is in any of Methods 31 to 34, where the number of selected features is less than 5 times the minimum coverage suggested.

[0064] The feature set involved in method 36 should have more than 10 classes to be identified in any of methods 31 to 35.

[0065] The feature set involved in method 37 requires that the number of classes to be identified in method 36 be 25 or more.

[0066] The feature set involved in method 38 is 25 or more in any of methods 31 to 37.

[0067] The feature set involved in method 39 is more than 50 selected features in method 38.

[0068] The feature set involved in method 40 is more than 100 features selected in method 39. Attached Figure Description

[0069] Figure 1 This is a schematic diagram representing a multi-class classification problem accompanied by feature selection.

[0070] Figure 2 It is a diagram showing the structure of a multi-class classification device.

[0071] Figure 3 This is a diagram showing the structure of the processing unit.

[0072] Figure 4 This is a flowchart illustrating the processing of a multi-class classification method.

[0073] Figure 5 This is a diagram representing the classification based on switching characteristics.

[0074] Figure 6 It is a diagram representing the matrix used to determine the switch values.

[0075] Figure 7 It is a diagram showing the determination of switch values / state values.

[0076] Figure 8 It is a graph that represents the pairwise expansion of classes that do not need to be judged.

[0077] Figure 9 This is a diagram illustrating the case of subclass imports.

[0078] Figure 10 This is a diagram illustrating the creation of a cyclic sort.

[0079] Figure 11 This is a diagram representing the final elimination match.

[0080] Figure 12 It is a graph that represents the detailed contents of the dataset.

[0081] Figure 13 This is a graph showing the comparison results of the discrimination accuracy of the present invention and the existing method.

[0082] Figure 14 This is a graph showing the comparison results of the robustness of the present invention with existing methods.

[0083] Figure 15 This is a graph showing the relationship between the number of selected features and the discrimination accuracy (F-value).

[0084] Figure 16 This is a table showing examples of diagrams illustrating the criteria for judgment.

[0085] Figure 17 It is a graph representing the relationship between the number of selected features and the minimum coverage.

[0086] Figure 18 It is a table showing the relationship between the minimum coverage number and the minimum F value. Detailed Implementation

[0087] Hereinafter, with reference to the accompanying drawings, the embodiments of the feature selection method, feature selection procedure, multi-class classification method, multi-class classification procedure, feature selection device, multi-class classification device, and feature set involved in the present invention will be described in detail.

[0088] <First Embodiment>

[0089] <Simplified Structure of a Multi-Classification Sorting Device>

[0090] Figure 2 This is a diagram showing the schematic structure of the multi-class classification device according to the first embodiment. For example... Figure 2 As shown, the multi-class classification device 10 (feature selection device, multi-class classification device) according to the first embodiment includes a processing unit 100 (first processor, second processor), a storage unit 200, a display unit 300, and an operation unit 400, which are interconnected to transmit and receive the required information. These components can be arranged in various ways; each component can be installed in one location (within a frame, in an interior, etc.) or in separate locations connected via a network. Furthermore, the multi-class classification device 10 (input processing unit 102; see reference...) Figure 3 It connects to external server 500 and external database 510 via the Internet or other network NW, and can obtain information such as multi-class classification samples, learning datasets, and feature sets as needed.

[0091] <Structure of the Processing Unit>

[0092] like Figure 3As shown, the processing unit 100 includes an input processing unit 102, a selection processing unit 104, a determination processing unit 110, a CPU 116 (Central Processing Unit), a ROM 118 (Read Only Memory), and a RAM 120 (Random Access Memory). The input processing unit 102 performs input processing to input a learning dataset consisting of known sample groups of known classes and feature sets of known sample groups from a storage device on the storage unit 200 or a network. The selection processing unit 104 performs selection processing to select the feature sets required for class determination of unknown samples of unknown classes from the feature sets based on the input learning dataset, and includes a quantification processing unit 106 and an optimization processing unit 108. The determination processing unit 110 performs class determination (determination processing) for unknown samples based on the selected feature sets, and includes an acquisition processing unit 112 and a class determination processing unit 114. The output processing unit 115 outputs processing conditions or processing results through display, storage, printing, etc. Furthermore, the processing of these components is carried out under the control of CPU116 (the first processor and the second processor).

[0093] The functions of each part of the processing unit 100 described above can be implemented using various processors and recording media. These processors include, for example, a CPU (Central Processing Unit), a general-purpose processor that executes software (programs) to implement various functions. Furthermore, these processors also include GPUs (Graphics Processing Units), which are specialized processors for image processing, and programmable logic devices (PLDs), such as FPGAs (Field-Programmable Gate Arrays), whose circuit structures can be modified after manufacturing. Using a GPU is effective when performing image learning or recognition. Moreover, dedicated circuits, such as ASICs (Application Specific Integrated Circuits), which have circuit structures specifically designed for performing specific processes, are also included in these processors.

[0094] Each function can be implemented by a single processor, or by multiple processors of the same or different types (e.g., multiple FPGAs, a combination of CPU and FPGA, or a combination of CPU and GPU). Furthermore, a single processor can implement multiple functions. As examples of a single processor performing multiple functions, firstly, as in the case of a computer, a processor is composed of a combination of one or more CPUs and software, and this processor implements multiple functions. Secondly, as in the case of a System-on-Chip (SoC), a processor is used to implement the overall system functions using a single IC (Integrated Circuit) chip. Thus, for various functions, one or more of the aforementioned processors are used as the hardware structure. More specifically, the hardware structure of these various processors is a circuit composed of circuit elements such as semiconductor components. These circuits can also be circuits that implement the aforementioned functions using logical OR, logical AND, logical negation, XOR, and logical operations combining them.

[0095] When the aforementioned processor or circuit executes the software (program), code that can be read by the computer executing the software (e.g., various processors or circuits constituting the processing unit 100, and / or combinations thereof) is stored in a non-transitory recording medium such as ROM 118, and the computer refers to the software. The software stored in the non-transitory recording medium includes programs (feature selection program, multi-class classification program) for executing the feature selection method and / or multi-class classification method involved in this invention, and data used during execution (data related to the acquisition of learning data, data for feature selection and class determination, etc.). The code may not be recorded in ROM 118, but rather in a non-transitory recording medium such as various optical-magnetic recording devices or semiconductor memories. When processing the software, for example, RAM 120 is used as a temporary storage area, and data stored in EEPROM (Electrically Erasable and Programmable Read Only Memory, not shown) may also be referenced. The storage unit 200 may also be used as a "non-transitory recording medium".

[0096] The details of the processing of the processing unit 100 of the above structure will be described later.

[0097] <Structure of the storage section>

[0098] The storage unit 200 comprises various storage devices such as hard disks and semiconductor memories, as well as their control units, and is capable of storing the aforementioned learning dataset, execution conditions and results of selection processing or class determination processing, feature sets, etc. The feature set is used by the multi-class classification device 10 to determine which of two or more N (N being an integer of two or more) classes a sample belongs to. The feature set contains a dataset of feature values ​​for samples belonging to each of the classes to which the sample is to be classified. When quantifying the discriminability between two classes based on each feature value of the selected feature set by combining two of the N classes in a pairwise coupling, the feature set is marked as being discriminable using at least one feature value among all pairwise couplings. This feature set can be generated through the input process (input processing) and selection process (selection processing) in the feature selection method (feature selection device) of the present invention. Furthermore, this feature set is preferably marked as being discriminable using at least five or more feature values, more preferably as being discriminable using at least ten or more feature values, and even more preferably as being discriminable using at least 60 or more feature values. Furthermore, this feature set is effective when there are 10 or more classes to be identified, and even more effective when there are 25 or more. It is also effective when there are 50 or more selected features, and even more effective when there are 100 or more.

[0099] <Structure of the Display Section>

[0100] The display unit 300 includes a monitor 310 (display device) composed of a display such as a liquid crystal display, which can display the acquired learning data, or the results of selection processing and / or class determination processing. The monitor 310 may also be composed of a touch panel type display, which can accept user input.

[0101] <Structure of the operating unit>

[0102] The operation unit 400 includes a keyboard 410 and a mouse 420, and the user can perform operations related to the execution and result display of the multi-class classification method involved in the present invention through the operation unit 400.

[0103] <1. Feature selection methods and processing of multi-class classification methods>

[0104] Figure 4This is a flowchart illustrating the basic processing of the feature selection method and multi-class classification method of the present invention. The feature selection method of the present invention selects a set of feature values ​​used to determine which of two or more N classes a sample belongs to. Furthermore, the multi-class classification method of the present invention, when N is an integer of 2 or more, determines which of the N classes a sample belongs to based on the sample's feature values. The multi-class classification method includes: an input step (step S100), inputting a learning dataset consisting of a known set of samples belonging to a given class and a set of feature values ​​of the known sample sets; a selection step (step S110), selecting a set of feature values ​​required for class determination of an unknown sample with an unknown class from the set of feature values ​​based on the learning dataset; and a determination step (step S120), performing class determination for the unknown sample based on the selected set of feature values. The determination step includes an acquisition step (step S122) for obtaining the feature value of the selected set of feature values ​​and a class determination step (step S124) for performing class determination based on the obtained feature value.

[0105] The selection process includes: a quantification step (step S112), which quantifies the discriminability between two classes based on each feature quantity of the selected feature quantity group by pairwise coupling of two of N classes according to the learning dataset; and an optimization step (step S114), which statistically analyzes the quantified discriminability for all pairwise couplings and selects combinations of feature quantity groups that optimize the statistical results. Furthermore, in the determination step, a multi-class discriminator using the selected feature quantity group is constructed in association with the pairwise couplings to determine the class of the unknown sample.

[0106] <2. Basic Principles of the Invention>

[0107] The present invention is particularly preferred in cases where features with properties close to binary values ​​are selected and the class is determined by combining such features like "switches". That is, cases where the features are not quantitatively combined linearly or nonlinearly, but this is not necessarily simple and becomes a sufficiently complex problem when there are a large number of switches. Therefore, the present invention is based on the principle of "exploring and selecting a large number of combinations of features with switching functions, and constructing a multi-class classifier from a simple classifier".

[0108] Figure 5 This is a diagram illustrating the aforementioned "feature quantity with switching function". Figure 5 Part (a) represents the case of class classification based on feature X' and feature Y', which becomes a complex and non-linear classification. In contrast, Figure 5Part (b) represents the case of class classification based on feature X and feature Y, which is a simple and linear classification. From the viewpoint of high-precision and high-robustness class classification, it is preferable to select feature quantities that have a switching function as shown in part (b) of the figure.

[0109] In addition, given the learning dataset, any sample is assigned values ​​for multiple common feature quantities (e.g., methylation location) (in addition, as values, some may also contain "missing values": hereinafter denoted as NA) and one correct class label (e.g., cancer or non-cancer, and tissue classification) (the learning dataset is input by the input processing unit 102 (input process, input processing: step S100)).

[0110] Furthermore, for the sake of simplicity, the above premise is set here, but so-called semi-supervised learning can also be introduced even when a portion of the samples are not assigned the correct class label. Since it is a combination with a well-known method, two representative processing examples are simply presented. A method that can simultaneously use (1) as preprocessing, and assign a certain class label to the samples that are not assigned the correct class label based on the data of the samples assigned the correct class label, and (2) iteratively inferring the class of other unknown samples based on the data that has been temporarily assigned class labels, and taking the one with the higher accuracy as the "correct label", and adding new learning data to learn, etc.

[0111] <2.1 Methods for Selecting Feature Quantities>

[0112] In this section, the selection of feature quantities (step S110: selection process) of the selection processing unit 104 (quantification processing unit 106, optimization processing unit 108) will be explained. First, the principle of feature quantity selection (selection process, selection procedure) in this invention will be explained in a simplified manner. The method of sequential expansion will then be explained. Finally, the steps for introducing all expanded feature quantity selections will be summarized. Furthermore, all feature quantities mentioned in this section refer to the feature quantities of the learning data.

[0113] <2.2 The principle of feature selection: reduced to the set covering problem>

[0114] First, the principle of feature selection (selection process) for multi-class classification will be explained. In this section, for simplicity, it is assumed that all feature values ​​of samples belonging to the same class are completely identical, and the feature value takes a definite binary (0 or 1) value.

[0115] When the value of feature i of class s is set to X i (s) When “being able to distinguish between class s and t by selecting feature set f” means that any feature quantity is different, that is, satisfying the following equation (1).

[0116] [Formula 1]

[0117]

[0118] Therefore, the necessary and sufficient condition for being able to mutually distinguish all given classes C = {1,2,…,N} satisfies the following equation (2).

[0119] [Formula 2]

[0120]

[0121] Here, the binary relation is expanded in pairs, and the XOR Y of the binary feature quantity i of classes s and t is introduced into the pair k = {s, t} ∈ P2(C) of the binary combination. i (k) (Refer to the following formula (3)), which is called the "discrimination switch" ( Figure 5 ).

[0122] [Formula 3]

[0123]

[0124] Figure 6 This is a diagram showing the calculation of the switch. Figure 6 Part (a) represents a table of binary feature values ​​#1 to #5 (values ​​are 0 or 1; binary feature values) for classes A, B, and C. Part (b) represents the cases where classes A, B, and C are expanded in pairs to form pairs {A,B}, {A,C}, and {B,C}. Figure 6 Part (c) represents the XOR of the binary feature values ​​for each pair (a value of 0 or 1; the discrimination switch value). For example, for the pair {A, B}, the discrimination switch value of feature #1 is 0, which means "it is impossible to distinguish the pair {A, B} based on feature #1 (it is impossible to determine which class A or B the sample belongs to)". In contrast, for the pair {A, B}, the discrimination switch value of feature #2 is 1, so it can be known that "the pair {A, B} can be distinguished based on the value of feature #2".

[0125] In summary, the necessary and sufficient condition for being able to mutually distinguish all given classes C can be rewritten as the following equation (4).

[0126] [Formula 4]

[0127]

[0128] That is, if all feature sets are set as F, then the selection of features for multi-class classification can be reduced to selecting a subset that satisfies the above formula. The set covering problem.

[0129] Alternatively, the "set covering problem" can be defined, for example, as "given a set U and a subset S of the power set of U, the problem of selecting a subset of S in such a way that it contains (= covers) all elements of U at least once" (or other definitions are also possible).

[0130] Here, for the switch set I of characteristic quantity i i ={k|Y i (k) =1} is a subset of the binary combination P2(C). Therefore, I = {Ii|i∈F} corresponding to all feature sets F is a subset of its set family, the power set of P2(C). That is, this problem is "given a subset I (corresponding to F) of the power set of P2(C), choose a subset of I (corresponding to f) that contains all elements of P2(C) at least once", which can be regarded as a set covering problem. Specifically, for all pairs expanded in pairs, it is necessary to select the feature quantity (and / or combination thereof) whose discrimination switch value is at least one "1". Figure 6 In the example, you can select "Feature #2, #4", "Feature #3, #4", or "Feature #2, #3, #4". Additionally, when the feature value is NA, the paired discrimination switch values ​​are automatically zero.

[0131] <2.3 Use a quantitative value of discriminability to replace XOR>

[0132] Here, if the feature quantity is originally a binary value, the feature quantity or its representative value (median value, etc.) can be directly regarded as discriminability. However, feature quantities are usually not limited to binary values, and even samples belonging to the same class can fluctuate to various values. Therefore, it is preferable that the quantification processing unit 106 (selection processing unit 104) replaces the discrimination switch value (XOR) with the quantitative value (quantified value) of discriminability based on the feature quantity of the learning dataset.

[0133] First, the quantitative processing unit 106 calculates the distribution parameter θ of the characteristic quantity i belonging to class s based on the measurement value group of the characteristic quantity i. i (s) and distribution D(θ) i (s) (Step S112: Quantification process). It is particularly preferable to quantify discriminability based on the distribution or distribution parameters. Furthermore, samples with a characteristic value of NA can be excluded from the quantification process. Of course, if all samples have a value of NA, then their characteristic values ​​cannot be used.

[0134] For example, the quantification processing unit 106 can process paired parameters θ i (s) With θ i (t)To determine if there is a significant difference between the two values, a statistical test is performed to obtain the p-value. Specifically, Welch's t-test can be used. Welch's t-test works as follows: assuming a normal distribution, a generally applicable method (as a graph, based on the approximate distributions of the characteristic values ​​of s and t) is used. Figure 7 (Which part, (a) or (b), should be used to determine the significance of the difference?) Of course, appropriate distributions and corresponding statistical tests can also be used in a timely manner based on the statistical properties of the characteristic quantity, or the observation and analysis results.

[0135] Figure 7 It is a diagram representing the determined image of the switch value and the state value. Figure 7 Part (a) involves using feature quantities in the discrimination of pairs {A,B}. The quantification processing unit 106 pre-sets a threshold based on the learning data (the values ​​of the positions of the two vertical lines in the figure), and determines the discrimination switch state value based on the measured value of the target sample (step S112: quantification process). If the measured value belongs to distribution A, the state value is +1; if it belongs to distribution B, the state value is -1; and if it belongs to the retention region, the state value is 0. On the other hand, Figure 7 Part (b) is the case where features are not originally used for the discrimination of pairs {A,B} (Y). i ({A,B}) =0).

[0136] However, especially when there are a large number of candidate features, if the determination is repeated across all feature sets F, it will fall into multiple comparison tests. Therefore, it is preferable that the quantification processing unit 106 corrects the p-value group obtained for the same pair k = {s,t} to a so-called q-value group (step S112: quantification process). Methods for correcting multiple tests include, for example, the Bonferroni method or the BH method [Benjamini, Y., and Y. Hochberg, 1995], and more preferably the so-called FDR (False Discovery Rate) method of correction to the latter, but it is not limited to these.

[0137] As shown in the following formula (5), the quantification processing unit 106 compares the obtained q value with the preset reference value α and assigns 0 or 1 to the discrimination switch (in particular, the case where the discrimination switch is 1 is called "marking").

[0138] [Formula 5]

[0139]

[0140] Furthermore, from the perspective of the extended set coverage problem, the discrimination switch is decentralized and binary in the above, but it can also be set to 1-q, for example, to handle continuous variables.

[0141] Furthermore, the p-value or q-value is a statistical difference, not a probability that can distinguish a sample. Therefore, the quantification processing unit 106 can also quantify the probability of correctly identifying the class of an unknown sample by a given feature quantity belonging to either of the paired coupled classes, based on an appropriate threshold set according to the reference learning dataset. Moreover, the quantification processing unit 106 can also perform multiple test corrections on such statistical probability values ​​according to the number of feature quantities.

[0142] Furthermore, it can be used not only as a benchmark related to statistical tests, but also as a benchmark value that has a certain difference from the mean, or even as a substitute for the mean. Of course, various statistical measures other than the mean or standard deviation can also be used as benchmarks.

[0143] <2.4 Extending the set covering problem to optimization problems such as maximizing the minimum pairwise covering number>

[0144] When the features are probabilistic variables, even with a discrimination switch marked, it is not always possible to accurately identify paired features. Therefore, the extended set covering problem is preferred.

[0145] Therefore, as shown in equation (6), the quantification processing unit 106 (selection processing unit 104) determines redundancy as the paired coverage number Z. f (k) Calculate the quantitative values ​​of each discriminability (calculate the total value as the statistical value; step S112: quantification process).

[0146] [Formula 6]

[0147]

[0148] Z f (k) The definition is not limited to that shown in equation (6). For example, for the continuous variable version of -Y i (k) , which can be used as the probability of failure in all discriminations, is defined as (1-Y i (k) The product of Y can also be expressed using a suitable threshold U, based on Y. i (k) Calculate the probability of success among at least U discriminations. Furthermore, the average of each discriminability can also be calculated. Thus, various statistical methods can be considered.

[0149] Next, from the perspective of "optimizing to minimize the bottleneck of discrimination", the optimization processing unit 108 (selection processing unit 104) can set the number of features to be selected to m. For example, by the following formula (7), the feature selection problem is reduced to the problem of maximizing the minimum pair coverage (step S114: optimization process, optimization processing).

[0150] [Formula 7]

[0151]

[0152] The above is a summary example when determining the number of features to be selected (the case where the number of features to be selected M is input, i.e., the case where the selection number input process / process has been performed). Conversely, the optimization processing unit 108 (selection processing unit 104) can set a threshold (target threshold T) in the minimum pairwise coverage number (the minimum value of the statistical value of discriminability) (target threshold input process / process) and select features in a way that satisfies this threshold (step S114: optimization process / process, selection process / process). In this case, it is of course preferable to select fewer features, and especially preferable to select the minimum.

[0153] Alternatively, combining the two can be considered as a way to optimize the process.

[0154] Because the set covering problem is an area of ​​active research, various solutions exist. The problem of maximizing the minimum cover number, an extension of it, can also be addressed with roughly the same steps. However, since it is typically an NP-complete problem, rigorous solutions are not easily found.

[0155] Therefore, it is of course preferable to find a rigorous solution, which solves the problem of maximizing the minimum paired coverage number literally or the problem of achieving the set coverage number with the fewest features. However, the optimization processing unit 108 (selection processing unit 104) may also use a method to find a local minimum by increasing the coverage number as much as possible through heuristic methods or by minimizing the number of selected features.

[0156] Specifically, for example, the optimization processing unit 108 (selection processing unit 104) can employ a simple greedy exploration step. In addition to the minimum pairwise coverage number of the currently selected feature set, a method such as "sequentially defining the i-th smallest i-th bit pairwise coverage number and sequentially selecting the feature quantity that maximizes the i-th bit pairwise coverage number of the smaller i" can also be considered.

[0157] Furthermore, the importance of input classes or pairs (step S112: quantification process, importance input process / process) can also be assigned a weight based on this importance during optimization (weighting process / process). For example, the above equation (7) can be modified into the following equation (8).

[0158] [Formula 8]

[0159] argmax min{Zk / wk}…(8)

[0160] Here, w kThis indicates the importance of pairwise comparisons. Alternatively, the importance of the class can be specified, set to w. k =w s w t And so on, and determine the importance of pairs based on the importance of the class. In addition, of course, reflecting the importance of the class in the pairwise calculation based on the product is just one example, and the specific calculation of the weighted average can also be other methods with the same theme.

[0161] Specifically, for example, in the identification of pathological tissues, where the distinction between disease A and disease B is particularly important, while the distinction between disease B and disease C is not important, it is preferable to set a large value for wk = {A, B} and a small value for wk = {B, C}. This allows for appropriate feature selection or classification (diagnosis) methods for cases where early detection of disease A is particularly important but the symptoms are similar to those of disease B, and cases where early detection of diseases B and C is not important but the symptoms differ significantly.

[0162] <2.5 Exclusion of Similar Features>

[0163] Generally, features with high similarity (similarity) values ​​among similar values ​​in the overall class of objects are highly correlated. Therefore, considering the robustness of the discrimination, it is preferable to avoid repeated selection. Furthermore, in the optimization exploration described above, efficiency can be improved if |F| can be reduced. Therefore, the optimization processing unit 108 (selection processing unit 104) preferably pre-screens the features to be considered based on the similarity evaluation results (step S110: selection process / process, similarity evaluation process / process, priority setting process / process). In fact, for example, there are hundreds of thousands or more methylation sites.

[0164] Here, we will define Y as the feature quantity i. i (k) The set I of k = 1 i ={k|Y i (k) =1} is called the "switch set". Based on this switch set, the similarity (or similarity degree) of feature quantities can be considered, that is, the same value relationship (repetition relationship) and the inclusion relationship of feature quantities.

[0165] For feature i, the collection becomes I. i =I l All l are used to create a set of identical features U as shown in equation (9). i Furthermore, collection becomes All l, as shown in equation (10), create a feature set H. i .

[0166] [Formula 9]

[0167]

[0168] [Formula 10]

[0169]

[0170] The same-value feature set is a set obtained by grouping repetitive features, and the included feature set is a set obtained by grouping attribute features. If the selection is based on a single representative feature, then highly similar features can be excluded. Therefore, for example, the similarity exclusion feature set can be used to replace all feature sets F as shown in the following equation (11).

[0171] [Formula 11]

[0172]

[0173] Of course, the selection processing unit 104 can consider only the set of features with the same value or include one of the features in the set as similarity, or it can create other metrics. For example, it can also consider calculating the vector distance between features (the distance between discriminability vectors), or a method that considers distances below a certain threshold as similar features. In addition to simple distances, it can also import arbitrary distances or metrics based on the distance calculated after normalizing the discriminability of multiple features.

[0174] Furthermore, while screening is implemented as described above, the selection processing unit 104 can also use a method to determine the ease of selection by lowering the selection preference order (priority) of features with similar characteristics that have already been selected (priority setting process) when conducting optimization exploration. Of course, it is also possible to use a method that increases the selection preference order (priority) of features with low similarity to the already selected features (priority setting process).

[0175] <2.6 Importing Pairs (Collections of Classes) That Do Not Require Mutual Authentication>

[0176] A class-based binary relation involves |P2(C)| = ... N C2. This simply takes all binary relations of the class, but in practice, there are sometimes pairs that do not need to be evaluated.

[0177] For example, in the case of a hypothetical cancer diagnosis problem (see the examples described later), it is necessary to distinguish between cancerous tissues and between cancerous tissues and normal tissues, but it is not necessary to distinguish between normal tissues.

[0178] Therefore, the selection processing unit 104 can suppress the pairwise expansion of some class binary relations. That is, based on the set C of classes that must be determined... T and the set of classes C that do not need to be determined N(The first step does not require class grouping), segment the given class C = {c|c∈C} T C N}, consider C T With C T Between, and C T With C N Between (paired expansion), on the other hand, excluding C from the class binary relation. N The process is compared with each other (step S110: process selection, first marked process / process, first excluded process / process). That is, the selection processing unit 104 calculates P2(C)' by the following formula (12) and replaces the previous P2(C) with P2(C)'.

[0179] [Formula 12]

[0180] P2(C)=P2(C)\{{s,t}|s≠t∈C N}…(12)

[0181] In addition, there can be more than two such divisions or markers.

[0182] Figure 8 This is a diagram representing the suppression of a portion of paired expansions. In Figure 8 In the example, classes T1, T2, ..., Tm are groups of classes that need to be distinguished between classes (e.g., cancerous tissue), while classes N1, N2, ..., Nn are groups of classes that need to be distinguished as "not T (not cancerous tissue)" but do not need to be distinguished from each other (e.g., normal tissue).

[0183] In this case, the selection processing unit 104 performs pairwise expansion between classes T (e.g., classes T1 and T2, classes T1 and T3, etc.) and between classes T and classes N (e.g., classes T1 and N1, classes T1 and N2, etc.), but does not perform pairwise expansion between classes N.

[0184] <2.7 Importing Subclasses from Sample Clustering>

[0185] Even when samples are correctly labeled, multiple groups with different states may sometimes coexist within samples of the same name. Even if class discrimination by name is sufficient, the discrimination switch cannot be correctly assigned because the features do not necessarily follow the same distribution parameters.

[0186] For example, cancers also have subtypes, and even cancers in the same organ can have different states [Holm, Karolina, et al., 2010]. However, if the premise is suitable for screening (used in conjunction with detailed examination), it is not necessary to distinguish subtypes.

[0187] Therefore, in order to correspond to the subtype, special class units called subclasses that do not need to be distinguished from each other can also be imported (step S110: select process, subclass set process / process, second mark process / process).

[0188] Subclasses can be automatically formed from samples. However, since identification from a single feature is difficult, it is considered that the processing unit 104 clusters the samples according to all features (given features) for each class, with an appropriate number of clusters L (or a minimum cluster size n). C This involves dividing the data into subclasses and clusters. For example, such as... Figure 9 As shown in part (a), samples belonging to a certain class (here, class B) are clustered using all features, and based on the results, they are segmented into subclasses X and Y, as shown in part (b) of the figure. In this example, if class B is segmented into subclasses X and Y, then feature i can be used to distinguish between class A and subclass Y of class B. However, there are also cases where a class is accidentally divided into multiple subclasses, in which case it is meaningless to forcibly consider them as "subclasses".

[0189] Furthermore, since various clustering methods exist, clustering can be performed using other methods, and the criterion for clusters can also be set in various ways.

[0190] For example, if class J is partitioned into {J1, J2, ..., J...} L If the second group does not need to be distinguished, then the given class C = {1,2,…,J,…,N} can be extended as shown in the following equation (13).

[0191] [Formula 13]

[0192] C +J ={1,2,…,J1,J2,…,J L ,…,N}…(13)

[0193] Similar to the previous one, the binary relation excludes pairs between subclasses that do not need to be judged, and replaces them as shown in the following formula (14) (second exclusion step).

[0194] [Formula 14]

[0195] P2(C +J ) ′-J =P2(C +J ) ′ \{{s,t}|s≠t∈J *}…(14)

[0196] Additionally, it will include the preceding term C. N The final class binary relation, applied sequentially within this group, is denoted as P2(C). +C ) '-C .

[0197] <2.8 Summary of the steps of the feature selection method>

[0198] This paper summarizes the steps of the feature selection method (selection process and selection treatment of selection processing unit 104) proposed by the inventors of this application.

[0199] (i) Define a set of classes C that do not need to be determined from a given set of classes C. N .

[0200] (ii) Cluster the samples by all features for each class, so that each cluster corresponds to a subclass (a subclass is a special class that does not need to be distinguished from another class).

[0201] (iii) Define the pairwise expansion P2(C) of all binary relations that are the objects of discrimination, except for those binary relations that do not need to be discriminated. +C ) '-C .

[0202] (iv) Based on the samples belonging to each class, the distribution parameters are estimated, and statistical tests are used to determine the significance of the characteristic quantity between class pairs k = {s, t}, for the discriminant switch Y. i (k={s,t}) Allocate 0 / 1.

[0203] (v) Create a set of similar exclusion features F' by constructing a set of features with the same value and a set of features containing the same value using the discrimination switch.

[0204] (vi) For the pairwise expansion P2(C) of the object class to be determined +C ) '-C Overall, select from F' such that the number of paired coverages Z is determined based on the discriminant switch and the calculated number of coverages. f (k) The feature set f (feature set) that maximizes the minimum value.

[0205] However, steps i to vi described above are only one example covering all of them, and it is not necessary to implement all of them; some steps may be omitted. Furthermore, alternative methods described or suggested in each section can also be used. Additionally, the multi-class classification apparatus 10 may only perform the feature selection method steps (feature selection method, feature selection processing) to obtain the feature set for multi-class classification.

[0206] <3. Multi-class classification methods>

[0207] In this section, the processing performed by the class determination processing unit 114 (determination processing unit 110) (step S120: determination process, determination process) will be explained. First, an example of the structure of a binary class classifier (binary class discriminator) based on the selected feature quantity (selected feature quantity group, feature quantity set) will be explained (class determination process, determination process). Next, an example of a method (class determination process, determination process) in which a multi-class classifier (multi-class discriminator) is constructed by the binary class classifier through two stages of (1) cyclic matching and sorting and (2) decisive elimination matching (constructing a multi-class discriminator that establishes a relationship with the selected feature quantity group by pair coupling) will be explained.

[0208] <3.1 Structure of Binary Class Classifier>

[0209] It is desirable to utilize the structure of selected feature quantities that facilitate pairwise discrimination. Therefore, a binary classifier can be constructed solely based on the combination of pairs and feature quantities marked with discrimination switches (a binary class discriminator is constructed by selecting a set of feature quantities and establishing a connection with each pair). Furthermore, during class classification, the acquisition processing unit 112 acquires the feature quantity values ​​of the selected feature quantity set (step S122: acquisition process, acquisition processing).

[0210] For example, the class determination processing unit 114 (determination processing unit 110) can compare with the learning distribution to determine the discrimination switch state y for class pair {s,t} of a given sample j (whose class is unknown) and select feature quantity i. i (k=(s,t),j) (Step S124: Class determination process, see reference) Figure 7 First, the distribution is extrapolated based on the learning data, and the significance level is determined. Figure 7 Whether it is the state shown in part (a) or the state shown in part (b), a threshold is preset in the case of "significant difference". Furthermore, the class determination processing unit 114 only calculates the distribution to which a given sample belongs (or whether a distribution to which it belongs) based on the value of the feature quantity when classifying a given sample if "significant difference" is selected, and determines the discrimination switch state value as shown in the following formula (15) (step S124: class determination process).

[0211] [Formula 15]

[0212]

[0213] Additionally, the "?" in the above formula indicates that the class of sample x is unknown. Furthermore, when the value of the sample's feature quantity is NA, y is set to 0.

[0214] The class determination processing unit 114 (determination processing unit 110) performs statistical analysis on the data to calculate the discrimination score r. j(s,t), and as shown in equations (16) and (17), a binary classifier B is constructed. j (s,t)(Step S124: Class determination process).

[0215] [Formula 16]

[0216]

[0217] [Formula 17]

[0218]

[0219] <3.2 Steps of Multi-Class Classification (1): Cyclic Matching and Sort>

[0220] The class determination processing unit 114 (determination processing unit 110) can further sum the above-mentioned discrimination scores (however, the number of discrimination switches is normalized, so it is preferable to take their sign value) and calculate the class score (pair score) as shown in the following formula (18) (step S124: class determination process).

[0221] [Formula 18]

[0222]

[0223] This score represents the degree of similarity between the unknown sample j and class s. Furthermore, the class determination processing unit 114 (determination processing unit 110) generates candidate classes based on the descending order of this score and creates a cyclic matching sorting G (step S124: class determination process). During this process, a replacement process can also be performed (if the class score is positive, it is replaced with +1; if it is zero, it remains ±0; if it is negative, it is replaced with -1).

[0224] Figure 10 This is a diagram illustrating the creation of a cyclic matching sort. First, as... Figure 10 As shown in part (a), the class determination processing unit 114 calculates the sign value of the discrimination score for each class pair ({A,B}, {A,C}, ...) (sgn(r) of equation (17)). j(s,t))). For example, regarding class pair {A,B}, it becomes "Regarding the sample, when considering the value of feature #1, it is similar to class A (sign value = +1), and when considering the value of feature #2, it cannot be said to be either class A or B (sign value = 0)...", with a subtotal of 24. Therefore, it can be said that "the sample is similar to A in class A and B" (the subtotal is positive and the larger the absolute value, the higher the similarity). Furthermore, regarding class pair {A,C}, it becomes "Regarding the sample, when considering the value of feature #3, it is similar to class C (sign value = -1), and when considering the value of feature #4, it is similar to class A (sign value = +1)...", with a subtotal of -2. Therefore, it can be said that "the sample is not similar to either class A or C (or, is slightly similar to class C)".

[0225] When calculating the time for all class pairs in this way, it is possible to obtain Figure 10 The results shown in part (b) are as follows. For example, {A,*} represents the comparison result of class A with all other classes, and the total score after the above substitution is 7. Similarly, the total for class D is 10. Furthermore, the class determination processing unit 114 determines the score based on this total, as follows: Figure 10 Section (c) shows the (sorted) candidate classes for discrimination. In this example, the totals for classes D, N, and A are 10, 8, and 7 respectively, with class D being number 1, class N being number 2, and class A being number 3.

[0226] <3.3 Steps of Multi-Class Classification (2): Final Elimination Matching>

[0227] In multi-class classification systems that address this problem, the discrimination between similar classes often becomes a performance bottleneck. Therefore, in this invention, features that can distinguish all pairs of feature sets (feature sets), including those between similar classes, are selected.

[0228] In contrast, in the aforementioned cyclic matching sorting G, it is expected that highly similar classes will cluster near the top position, but the class scores are largely determined by comparison with the sorted lower-ranking classes. That is, the sorted uppermost class (in...) Figure 10 In the example, the ordering of classes D, N, and A is not necessarily reliable.

[0229] Therefore, as shown in equation (19), the class determination processing unit 114 (determination processing unit 110) can eliminate the irregular match T based on the superior class g of the cyclic matching sort. j To determine the final classification class (step S124: classification process).

[0230] [Formula 19]

[0231] T j (G1, G2, ..., G g ) = T j (G1, ..., G)g-2 B j (G g-1 G g ))=…=B j (G1, B) j (G2, ..., B) j (G g-1 G g )…))…(19)

[0232] That is, the class determination processing unit 114 reapplies the binary class classifier to the pairs of the next two classes from the top g classes in the list to determine the winning remainder, and reduces the number of lists one by one, taking the same steps in sequence (finally, the top g class is compared with the winning remainder class).

[0233] For example, such as Figure 11 As shown, from the top three classes (classes D, N, A) in the list, class scores are calculated for classes N and A, which are the next two classes, to determine the winning remainder (class N or A). The same method is used to calculate class scores for class D, which is the top class in the cyclic sort, and the winning remainder class. Additionally, "the position up to the specified point in the cyclic sort is used as the target for the final elimination match" (in... Figure 11 In the example, "up to number 3" is not specifically limited.

[0234] <3.4 Structure of Other Multi-Class Classifiers>

[0235] Furthermore, the above is one example of a classifier structure; various machine learning methods can also be used. For example, it can be a structure that is basically a random forest, or a structure in which only the decision tree using the discriminative switch of the selected feature quantity is effective (decision process) in the intermediate decision tree. Specifically, the class decision processing unit 114 (decision processing unit 110) can construct decision trees that are coupled with each pair to establish an association using the selected feature quantity group, and combine more than one decision tree to construct a multi-class discriminator (step S124: class decision process). At this time, the class decision processing unit 114 can also construct a multi-class discriminator as a random forest based on the decision tree and the combination of decision trees (step S124: class decision process).

[0236] <4. Output>

[0237] The output processing unit 115 can output the input data or the conditions and results of the aforementioned processing, either based on the user's operation via the operation unit 400 or independently of the user's operation. For example, it can output the input learning dataset, the selected feature set, the results of round-robin matching or elimination matching, etc., by displaying them on a display device such as a monitor 310, storing them in a storage device such as a storage unit 200, or printing them using a printer (not shown). (Output process, output handling; regarding...) Figure 16 (To be described later).

[0238] <5. Test Data and Examples>

[0239] The inventors of this application selected eight types of cancer (colorectal cancer, stomach cancer, lung cancer, breast cancer, prostate cancer, pancreatic cancer, liver cancer, and cervical cancer) as diagnostic targets. These cancers account for approximately 70% of cancer cases in Japan [Hori M, Matsuda T, et al., 2015], and therefore considered them suitable for early screening.

[0240] Furthermore, since normal tissue needs to encompass all tissues that can flow into the bloodstream, in addition to the organs corresponding to the eight types of cancer mentioned above, a total of 24 other conceivable organs such as blood, kidneys, and thyroid are listed.

[0241] As a feasibility study, assuming the identification of extracted cell blocks (living tissue slices), open data containing measurements of methylation sites were collected from a total of 5,110 samples. Figure 12 ).

[0242] For cancerous tumors and normal organs (excluding blood), 4,378 samples were collected from the registry data of "The Cancer Genome Atlas" (TCGA) [Tomczak, Katarzyna, et al., 2015]. In addition, 732 blood samples were collected [Johansson, Asa, Stefan Enroth, and Ulf Gyllensten, 2013].

[0243] The sample classification (including the source tissue for distinguishing between cancer and non-cancer samples) is assigned according to the registration annotation information.

[0244] Furthermore, a total of methylation measurements were taken at 485,512 sites, but excluding sites where all sample values ​​(NA) could not be measured, the total number of sites was 291,847. Additionally, the data obtained after normalization and other post-processing were directly used in the above registration data.

[0245] Furthermore, the entire dataset is mechanically divided equally, with one set used as the learning dataset and the other as the test dataset.

[0246] The experimental topics set in this embodiment are as follows.

[0247] i. Prepare a dataset of approximately 5,000 samples.

[0248] Allocation categories (32 in total): Cancer (8 types) or normal tissue (24 types)

[0249] Characteristic quantity (methylation position): Approximately 300,000 items

[0250] ii. From the above half of the learning dataset, pre-select up to 10 to 300 methylation sites (omics information, omics on / off state information) of items that can be used for discrimination (at the same time, learn parameters such as subclass segmentation or distribution parameters).

[0251] iii. (Especially from the remaining half of the test dataset) (independently, sample by sample) answer the discrimination question for a given sample.

[0252] Input: Selected methylation site measurements of the sample (up to 300 items corresponding to selection ii).

[0253] Output: Estimated class = "Cancer + Source tissue (choose from 8 types)" or choose from 9 types of "Non-cancer (only 1 type)".

[0254] In addition, in the embodiments, the following method was used as a conventional method for comparison with the proposed method (the method of the present invention).

[0255] • Feature selection method: Shannon entropy benchmark with methylation site studies [Kadota, Koji, et al., 2006; Zhang, Yan, et al., 2011]

[0256] • Multi-class classification: Naive Bayes classifier (simple but known for its high performance [Zhang, Harry, 2004])

[0257] <5.1 Comparison Results of the Proposed Method and Existing Methods>

[0258] <5.1.1 Discrimination accuracy of test data>

[0259] Learning data was used to select 277 locations (omics information, omics on / off state information) to confirm the discrimination accuracy of the test data. The proposed method (the multi-class classification method of this invention) was compared with existing methods. Figure 13 The result indicates that the proposed method has high discrimination accuracy across all projects.

[0260] Compared to the average F-value of 0.809 for existing methods, the proposed method achieves an average F-value of 0.953. Furthermore, in existing methods, there are instances where the F-value / sensitivity / fitness remains below 0.8 in lung cancer, pancreatic cancer, and gastric cancer, but in the proposed method, it achieves above 0.8 in all categories.

[0261] <5.1.2 Robustness of the Judgment>

[0262] The robustness of the judgment is confirmed by the average F-score difference between learning and testing in the preceding term, and the proposed method is compared with existing methods. Figure 14The results show that the proposed method has excellent robustness (F-value decreased by 0.008).

[0263] In the existing method, the average F-value of 0.993 is roughly perfect relative to the learning data, but the accuracy drops significantly (difference of 0.185) on the test data, indicating that it has fallen into overlearning.

[0264] On the other hand, in the proposal method, the average decrease in F-value remained at 0.008. Furthermore, the discrimination ability for pancreatic cancer was relatively low in the proposal method (F-value 0.883), but also relatively low during the learning phase (F-value 0.901). This proposal method suggests that, to some extent, the discrimination accuracy and tendency in the test data can be predicted at the completion of the learning phase.

[0265] <5.1.3 Relationship between the number of selected features and the discrimination accuracy>

[0266] The relationship between the number of selected features and the discrimination accuracy (F-value) was confirmed. Figure 15 The results show that the discrimination accuracy is significantly improved when 50 to 100 selections are used, while there is a tendency for saturation when 150 to 300 selections are used.

[0267] Therefore, especially in the problem of determining "whether it is cancer or non-cancer" based on the methylation pattern of cfDNA and in cancer diagnosis of the tissue of origin, it is indicated that the discriminative power is insufficient when selecting 10 features, and at least 25 to 100 items of multi-item measurement are required (therefore, in such multi-class classification problems with a large number of classes, the number of features (selection feature group) selected in the selection step (selection process) is preferably 25 or more, more preferably 50 or more, and most preferably 100 or more).

[0268] <5.1.4 Exclusion of similar features and import of pairs that do not require discrimination>

[0269] In the proposed method, similarity features (similarity evaluation process, similarity evaluation processing) are not selected. Furthermore, pairs that do not require discrimination are introduced.

[0270] There are 291,847 valid methylation sites (the feature quantity of this problem), of which 59,052 similar features (iso-value relationships, inclusion relationships) were identified and can be eliminated as out-of-object pairs (reduction of 20.2%). Furthermore, by dividing the original 32 classes into 89 classes based on sample clustering, the total number of simple pairs increases to 4,005. Of these, 551 out-of-object pairs between normal tissues and cancer subclasses can be eliminated (reduction of 13.8%).

[0271] At the same time, it can reduce the exploration space by 31.2%. By excluding similar features and importing pairs that do not require discrimination, it can be confirmed that the exploration of discriminant switch combinations is made more efficient.

[0272] <5.1.5 Subclass Segmentation>

[0273] In the proposed method, sample clustering is introduced to divide a given cluster into subclasses. The pairwise combinations that do not require discrimination are also important, thus confirming the effectiveness of merging both.

[0274] For comparison, no subclassing was performed, and no feature selection was imported for pairs of tissues that did not require discrimination. The same steps were performed on other experiments. As a result, even when limited to cancerous tissue, the discrimination accuracy decreased from 95.9% to 85.6% (normal tissue increased to 24 types without segmentation, so the comparison was limited to cancerous tissue, especially to confirm the effectiveness of subclassing).

[0275] It can confirm that high-precision discrimination is achieved by using subclass segmentation and paired imports that do not require discrimination.

[0276] <5.1.6 Simultaneous Use of Winner-Winning Matchmaking>

[0277] In the proposed method, both cyclic matching sorting (in this case, the first class is called the "pre-selected top-level class") and elimination matching are used in the multi-class classification.

[0278] Of the 2,555 test cases, 278 instances showed discrepancies between the pre-selected top-level class and the correct class. Of these, 162 were corrected to the correct classification through a final elimination matching process. Conversely, 19 instances showed the opposite (the pre-selected top-level class matched the correct class, but the final elimination matching changed the classification to incorrect).

[0279] That is, by using the elimination matching method simultaneously and subtracting the discrimination error of the pre-selected top-level class, a 51.4% correction can be achieved, improving the overall accuracy by 5.6%. This confirms that this method effectively leverages the performance of a pairwise binary class classifier.

[0280] In this proposed method, the steps for discrimination, the types of comparative studies, and the characteristic quantities on which the basis is based are clearly defined. Therefore, it is possible to trace the discrimination results and easily identify and explain the differences from the characteristic quantities or thresholds used as the basis. It can be said to be an "explanatory AI" that is particularly advantageous for medical diagnoses that require discriminative basis.

[0281] Figure 16 This is a table showing examples of graphs illustrating the decision criteria (extracting examples of actual decision shifts from the test data). Figure 16Part (a) shows the superclass and result of the classification, as well as the score. In the example in this figure, the sample is classified as "cancer tissue 1" with a score of 79, and the next most similar sample is "cancer tissue 3" with a score of 76.

[0282] Similarly, in the seven rows from "cancer tissue 1" to "normal tissue 1", various scores R can be identified. i (s). Moreover, in the three rows from the row "<cancer tissue 1|cancer tissue 3>" to the row "<cancer tissue 1|cancer tissue 5>", it is possible to identify the various paired discrimination scores r. j (s,t).

[0283] Furthermore, in Figure 16 The table shown in section (b) provides a summary of how the selected features (recorded as markers in the table) contribute to the various discrimination scores. Of course, besides... Figure 7 In addition to the distribution plot of the learning data shown in section (a), visualizations such as plotting the values ​​of each sample on the graph can also be added.

[0284] Thus, according to the proposed method (this invention), after classification (selection), by reversing the processing steps and illustrating each score, the criteria for judgment can be confirmed and visualized. Therefore, the reliability of the final judgment result can be inferred based on other candidate similarity class scores or judgment scores. Furthermore, by identifying the feature quantities that serve as the criteria, their interpretation can be used for post-classification examination.

[0285] <Relationship between the number of selected features and the minimum coverage>

[0286] The relationship between the number of selected features and the minimum coverage in the above embodiments is shown in... Figure 17 The curve graph.

[0287] [Formula 20]

[0288]

[0289] Here, a linear relationship with a slope of approximately 1 / 5 is obtained. This indicates that for a multi-class classification problem with a high degree of internal subclass segmentation (8 cancer classes / 24 normal classes), approximately every 5 features selected can cover the feature set for all class discriminations. That is, this demonstrates the significant effect of the method disclosed in this invention, which reduces feature selection to a set coverage problem and expands upon it, effectively increasing the minimum coverage number in multi-class classification problems. Furthermore, by... Figure 17It can be seen that by fine-tuning the obtained feature set, it is possible to create a feature set with a very small portion of the overall feature set, specifically, a feature set with a high discriminative power that is less than 5 times the minimum coverage required. Such a small number of feature sets with sufficient minimum coverage is of great value.

[0290] <Relationship between minimum coverage number and minimum F-value>

[0291] The relationship between the minimum coverage number in the selected feature set and the minimum F-value (the minimum discriminative ability F-value in the test data among the discriminant object classes) is shown in the figure. Figure 18 The curve graph.

[0292] [Formula 21]

[0293]

[0294] Therefore, it can be seen that performance is almost non-existent when the minimum coverage is 0. Around the minimum coverage of 5, the minimum F-value becomes 0.8; around the minimum coverage of 10, it becomes 0.85; and around the minimum coverage of 60, it becomes 0.9. That is, firstly, it can be seen that if a feature set with a minimum coverage of at least 1 is not selected, performance is almost non-existent. Furthermore, the specific benchmark for the required F-value varies depending on the problem. Since 0.80, 0.85, and 0.90 are easily understood benchmarks, feature sets with a minimum coverage of 5 or more, or 10 or more, or 60 or more are valuable. Combined with the previous point (the relationship between the number of selected features and the minimum coverage), the ability of this invention to "achieve coverage with a relatively small number of selected features (less than 5 times the suggested minimum coverage)" is particularly valuable.

[0295] Furthermore, the above-described embodiment regarding "methylation location and live tissue classification" is merely one specific example. The method of this invention has been sufficiently generalized and can be applied to any feature selection and multi-class classification outside of the biological field. For example, when classifying people in captured images (e.g., Asia, Oceania, North America, South America, Eastern Europe, Western Europe, Middle East, Africa), the method of this invention can select features based on a large number of features such as facial size or shape, skin color, hair color, and / or the position, size, and shape of the eyes, nose, and mouth, and then use the selected features for multi-class classification. Furthermore, the method of this invention can also be applied to agricultural, forestry, and fishery products or industrial products, or to feature selection and classification for various statistical data.

[0296] The embodiments and other examples of the present invention have been described above, but the present invention is not limited to the above-described manner, and various modifications can be made without departing from the spirit of the present invention.

[0297] Symbol Explanation

[0298] 10-Multi-class classification device, 100-Processing unit, 102-Input processing unit, 104-Selection processing unit, 106-Quantification processing unit, 108-Optimization processing unit, 110-Decision processing unit, 112-Acquisition processing unit, 114-Class determination processing unit, 115-Output processing unit, 116-CPU, 118-ROM, 120-RAM, 200-Storage unit, 300-Display unit, 310-Monitor, 400-Operation unit, 410-Keyboard, 420-Mouse, NW-Network, S100~S124-Each processing step of the multi-class classification method.

Claims

1. A feature selection method, wherein the feature selection method selects a set of features used to determine which of two or more N classes a sample belongs to, the feature selection method having: The input process involves inputting a learning dataset consisting of a known set of samples belonging to a given class that becomes an object and a set of features of the known sample set. and In the selection process, based on the learning dataset, a set of features is chosen from the set of features required for class determination of unknown samples belonging to an unknown class. The selection process has the following characteristics: The quantification process involves pairwise coupling of two of the N classes and quantifying the discriminability between the two classes based on each feature quantity of the selected feature quantity group according to the learning dataset. The optimization process involves statistically analyzing the quantified discriminability for all the paired couplings and selecting a combination of feature sets that optimize the results of the statistics. The first marking step involves marking a portion of the given classes as a first group of classes that do not require mutual discrimination; and The first exclusion step is to exclude the pairwise couplings between the marked first non-discriminable class groups from the unfolded pairwise couplings.

2. The feature selection method according to claim 1, wherein, The selection process has the following characteristics: The similarity evaluation process assesses the similarity between feature quantities based on the discriminability of each feature quantity for each paired coupling; and The priority setting process sets the priority of the feature quantities to be selected based on the evaluation results of the similarity.

3. The feature selection method according to claim 2, wherein, The similarity refers to the repeating and / or inclusion relationships of discriminative pairs of couplings.

4. The feature selection method according to claim 2, wherein, The similarity is the distance between each pair of coupled discriminative vectors or a metric based on the distance.

5. The feature selection method according to claim 1, further comprising: The selection process involves inputting the number M of feature quantities to be selected. The optimization is based on maximizing the minimum of the statistical values ​​in all pairwise couplings of M selected features.

6. The feature selection method according to claim 1, wherein, The optimization process has the following characteristics: Importance input process, importance of input class or pairwise judgment; and The weighting process assigns weights based on the importance during the statistics.

7. The feature selection method according to claim 1, wherein, The number of features selected in the selection process is 25 or more.

8. The feature selection method according to claim 7, wherein, The number of features selected in the selection process is 50 or more.

9. The feature selection method according to claim 8, wherein, The number of features selected in the selection process is 100 or more.

10. A recording medium that is non-transitory and computer-readable, and which records a program that causes a computer to execute the feature selection method according to any one of claims 1 to 9.

11. A multi-class classification method, wherein when N is an integer greater than or equal to 2, the method determines which of N classes a sample belongs to based on the sample's feature values, the multi-class classification method having: The input process and the selection process are performed using the feature quantity selection method according to any one of claims 1 to 9; and The determination process involves classifying the unknown sample based on a selected set of feature values. This determination process includes an acquisition step for obtaining feature values ​​from the selected set of feature values ​​and a class determination step for classifying the sample based on the acquired feature values. In the determination process, the class determination for the unknown sample is performed by constructing a multi-class discriminator that is associated with the selected set of feature quantities through pairwise coupling.

12. The multi-class classification method according to claim 11, wherein, The selection process has the following characteristics: The similarity evaluation process assesses the similarity between feature quantities based on the discriminability of each feature quantity for each paired coupling; and The priority setting process sets the priority of the feature quantities to be selected based on the evaluation results of the similarity.

13. The multi-class classification method according to claim 11, wherein, In the quantification process, the statistically significant difference in the feature quantities in the learning dataset between the paired classes is utilized.

14. The multi-class classification method according to any one of claims 11 to 13, wherein, In the quantification process, when a threshold set based on the learning dataset is given, indicating the feature quantity of an unknown sample belonging to any of the paired classes, the probability of correctly determining the class to which the unknown sample belongs is utilized based on the given feature quantity.

15. The multi-class classification method according to any one of claims 11 to 13, wherein, In the quantification process, the quantified value of discriminability is the value after multiple tests and corrections on the statistical probability value based on the number of characteristic quantities.

16. The multi-class classification method according to any one of claims 11 to 13, further comprising: The subclass setting process involves clustering samples belonging to one or more classes from the learning dataset based on given feature values ​​to form clusters, and setting each cluster as a subclass of the class. The second marking step involves marking each subclass within a class as a second group of classes that do not require mutual discrimination within that class; and The second exclusion process excludes the pairwise couplings between the marked second non-discriminable class groups from the unfolded pairwise couplings.

17. The multi-class classification method according to any one of claims 11 to 13, wherein, The statistics are the calculation of the total or average of the quantitative values ​​of the discriminability.

18. The multi-class classification method according to any one of claims 11 to 13, further comprising: The target threshold input step involves inputting a target threshold T, representing the statistical value of the statistical result. The optimization involves setting the minimum of the statistical values ​​in all paired couplings based on the selected feature quantity to be above the target threshold T.

19. The multi-class classification method according to any one of claims 11 to 13, wherein, In the determination process, Each binary class discriminator is constructed and associated with each paired set of selected features. The binary class discriminator is combined to form the multi-class discriminator.

20. The multi-class classification method according to any one of claims 11 to 13, further comprising: The process of evaluating the similarity of a sample to each class using a binary class discriminator; and The process of constructing the multi-class discriminator based on the similarity.

21. The multi-class classification method according to any one of claims 11 to 13, further comprising: The process of evaluating the similarity of a sample to each class using a binary class discriminator; and The process of constructing the multi-class discriminator involves reapplying the binary class discriminator used for evaluating the similarity between classes with higher similarity levels.

22. The multi-class classification method according to any one of claims 11 to 13, wherein, In the determination process, Decision trees are constructed by using selected feature sets to establish relationships between paired couplings. Combining one or more of the aforementioned decision trees to construct a multi-class discriminator.

23. The multi-class classification method according to claim 22, wherein, In the decision-making process, the decision tree and combinations of decision trees constitute the multi-class discriminator as a random forest.

24. The multi-class classification method according to any one of claims 11 to 13, wherein, By measuring the omics information of a live tissue slice, the class to which the live tissue slice belongs is determined from the N classes.

25. The multi-class classification method according to any one of claims 11 to 13, wherein, By measuring the on / off state information of the omics of the living tissue slice, the class to which the living tissue slice belongs is determined from the N classes.

26. The multi-class classification method according to any one of claims 11 to 13, wherein, The number of classes to be identified is 10 or more.

27. The multi-class classification method according to claim 26, wherein, The number of classes to be identified is 25 or more.

28. A recording medium that is non-transitory and computer-readable, and which records a program that causes a computer to execute the multi-class classification method according to any one of claims 11 to 27.

29. A feature selection device for selecting a set of features used to determine which of two or more N classes a sample belongs to, said feature selection device comprising a first processor, The first processor performs the following processing: Input processing, wherein the input consists of a learning dataset comprising a known set of samples belonging to a given class of objects and a set of features of the known sample sets; and The selection process involves choosing, based on the learning dataset, a set of features from the set of features required for class determination of unknown samples with unknown class affiliation. The selection process has the following characteristics: The quantitative processing involves pairwise coupling of two of the N classes, and quantifying the discriminability between the two classes based on each feature quantity of the selected feature quantity group according to the learning dataset. The optimization process involves statistically analyzing the quantified discriminability for all paired couplings and selecting a combination of feature sets that optimize the statistical results. The first marking process involves marking a portion of the given classes as a first group of classes that do not require mutual discrimination; and The first exclusion process excludes the pairwise couplings between the marked first non-discriminable class groups from the expanded pairwise couplings.

30. A multi-class classification device, wherein when N is an integer greater than or equal to 2, determines which of N classes a sample belongs to based on the sample's characteristic values, the multi-class classification device comprising: The feature quantity selection device according to claim 29; and Second processor, The second processor performs the following processing: The input processing and the selection processing utilize the feature selection device; and The determination process involves classifying the unknown sample based on a selected set of feature values. This determination process includes acquiring feature values ​​from the selected set of feature values ​​and class determining the sample based on those acquired feature values. In the determination process, the class determination for the unknown sample is performed by constructing a multi-class discriminator that is associated with the selected set of feature quantities through pairwise coupling.

31. A set of features used by a multi-class classification device to determine which of two or more N classes a sample belongs to, said set of features comprising: The feature dataset of samples belonging to various categories that become objects. When quantifying the discriminability between two classes based on each feature quantity of the selected feature quantity group by pairwise coupling of two of the N classes and referring to the feature quantity dataset, all pairwise couplings are marked as discriminable with at least one feature quantity. A subset of the aforementioned categories are designated as the first non-discriminable category group, which does not require mutual discrimination. Exclude the pairwise couplings between the first group of groups that do not require discrimination from the expanded pairwise couplings.

32. The feature set according to claim 31, wherein, When quantifying the discriminability between the two classes based on each feature quantity of the selected feature quantity group by pairwise coupling of two of the N classes with reference to the feature quantity dataset, all pairwise couplings are marked as being discriminable with at least 5 features.

33. The feature set according to claim 32, wherein, When quantifying the discriminability between the two classes based on each feature quantity of the selected feature quantity group by pairwise coupling of two of the N classes with reference to the feature quantity dataset, all pairwise couplings are marked as being discriminable with at least 10 features.

34. The feature set according to claim 33, wherein, When quantifying the discriminability between the two classes based on each feature quantity of the selected feature quantity group by pairwise coupling of two of the N classes with reference to the feature quantity dataset, all pairwise couplings are marked as being discriminable with at least 60 features.

35. The set of features according to any one of claims 31 to 34, wherein, The number of selected features should be less than 5 times the minimum coverage suggested.

36. The set of features according to any one of claims 31 to 34, wherein, The number of classes to be identified is 10 or more.

37. The feature set according to claim 36, wherein, The number of classes to be identified is 25 or more.

38. The set of features according to any one of claims 31 to 34, wherein, The number of selected features is 25 or more.

39. The feature set according to claim 38, wherein, The number of selected features is more than 50.

40. The feature set according to claim 39, wherein, The number of selected features is more than 100.

Citation Information

Patent Citations

  • Algorithm for sub-disease type and prognosis using gene expression profiling

    JP2012505453A

  • Image processing device and image processing program

    WO2008139825A1