Information processing device, information processing method, and computer program
Patent Information
- Application Number
- JP2025510629
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-10
AI Technical Summary
Mass spectrometry results for disease diagnosis and classification are prone to variations due to various factors, leading to suboptimal diagnostic accuracy, as they typically only consider specific substances and ignore the relative intensity relationships between multiple substances in the sample.
An information processing device and method that utilizes comparison data from mass spectrometry results to classify samples by establishing substance group pairs based on peak intensity relationships, reducing variation influence and improving accuracy through trained models, including data on ratios and average values, and setting pairs based on mutual correlations.
Enhances the accuracy of sample classification and disease diagnosis by reducing the impact of variations in mass spectrometry results, allowing for more precise identification of diseases like cancer using biological samples.
Abstract
Description
Information processing device, information processing method, and computer program
[0001] The technology disclosed in this specification relates to information processing for classifying samples.
[0002] Various markers are used for diagnosing diseases. For example, CA19-9 is used as a tumor marker for pancreatic duct cancer and bile duct cancer (see, for example, Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2023-37003
[0004] It is conceivable to classify whether a subject suffers from a specific disease by referring to the peak area of a specific substance obtained by mass spectrometry of a sample (blood, urine, etc.) collected from the subject. However, mass spectrometry results can vary due to various factors. Therefore, while standard substances are used in quantitative mass spectrometry of specific substances, such quantitative mass spectrometry does not utilize the contents of substances other than the specific substance being measured simultaneously for diagnosis, leaving room for improvement in diagnostic accuracy. Note that this issue is not limited to cases where the purpose is to diagnose a specific disease, but is a common issue when performing some kind of classification by referring to the results of mass spectrometry of a sample.
[0005] This specification discloses a technique that can solve the above-mentioned problems.
[0006] The technology disclosed in this specification can be realized, for example, in the following forms.
[0007] (1) The information processing device disclosed in this specification includes a model acquisition unit, a target data acquisition unit, and a classification execution unit. The model acquisition unit acquires a trained model that receives comparison data as input and outputs a classification result for the sample. The comparison data is data that includes data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups for at least one substance group pair, each of which is a pair of two substance groups consisting of at least one substance, in a mass spectrometry result for a plurality of substances contained in the sample. The target data acquisition unit acquires the comparison data for the target sample. The classification execution unit inputs the comparison data for the target sample into the trained model, thereby classifying the target sample and outputting the classification result.
[0008] In this manner, the information processing device classifies target samples using a trained model that inputs comparison data and outputs sample classification results. The comparison data includes data indicating the magnitude relationship between the peak intensities of substances belonging to one substance group and those of substances belonging to the other substance group for at least one substance group pair, each of which is a pair of two substance groups consisting of at least one substance, in the mass spectrometry results for multiple substances contained in the sample. Therefore, by using the comparison data, the influence of variations in values between samples due to various factors can be relatively reduced, thereby improving the accuracy of sample classification using mass spectrometry results. Furthermore, because the comparison data is data created using mass spectrometry results for multiple substances contained in the sample, the accuracy of sample classification using mass spectrometry results can be improved compared to a configuration that uses mass spectrometry results for only a specific substance.
[0009] (2) In the information processing device, the comparison data may include data indicating the ratio of the peak intensity of a substance belonging to one of the substance groups to the peak intensity of a substance belonging to the other of the substance groups. This configuration allows for the formation of substance group pairs using any combination of substances, not just pairs of substances whose peak intensities in the samples are similar and whose magnitude relationship is reversed depending on the classification to which they belong. This increases the degree of freedom in setting substance group pairs, making it possible to obtain more desirable comparison data and effectively improve the accuracy of sample classification based on mass spectrometry results.
[0010] (3) In the information processing device, the comparison data may include data indicating the ratio of the peak intensity of each substance belonging to one substance group to the average value of the peak intensities of all substances belonging to the other substance group. By adopting this configuration, the value of the comparison data is determined for each substance, thereby preventing a decrease in the amount of information. This allows comparison data with a larger amount of information to be obtained, effectively improving the accuracy of sample classification based on mass spectrometry results.
[0011] (4) In the information processing device, the comparison data may include data on a plurality of the substance group pairs. By adopting this configuration, comparison data with a larger amount of information can be obtained, and the accuracy of sample classification based on the mass spectrometry results can be effectively improved.
[0012] (5) The information processing device may further include a pair setting unit that sets the at least one substance group pair based on the correlation between the plurality of substances. By adopting this configuration, a substance group pair that contributes to more appropriate classification can be set, and the accuracy of sample classification based on mass spectrometry results can be effectively improved.
[0013] (6) In the information processing device, the sample may be a biological sample collected from a human. By adopting this configuration, it is possible to improve the accuracy of classification of the biological sample based on the results of mass spectrometry of the biological sample.
[0014] (7) In the information processing device, the biological sample may be at least one of urine, bile, sweat, saliva, blood, cerebrospinal fluid, ascites, pleural effusion, pancreatic juice, gastric juice, and intestinal fluid. By adopting this configuration, it is possible to improve the accuracy of classification of the biological sample by referring to the mass spectrometry results of the biological sample.
[0015] (8) In the information processing device, the classification execution unit may be configured to perform the classification for a specific disease. By adopting this configuration, it is possible to improve the accuracy of diagnosing a specific disease by referring to a mass spectrometry result.
[0016] (9) In the information processing device, the specific disease may be cancer. By adopting this configuration, it is possible to improve the accuracy of cancer diagnosis with reference to the mass spectrometry results.
[0017] The technology disclosed in this specification can be realized in various forms, such as an information processing device, an information processing method, a computer program that realizes those methods, a non-transitory recording medium on which that computer program is recorded, etc.
[0018] FIG. 1 is an explanatory diagram showing the diagnostic model MO of this embodiment; FIG. 1 is an explanatory diagram showing the diagnostic model MO of this embodiment; FIG. 2 is an explanatory diagram showing the mass analysis data MD and the comparison data CD; FIG. 3 is an explanatory diagram showing the general configuration of the information processing device 100;
[0019] A. Embodiments: A-1. Overview of Diagnostic Model MO: FIGS. 1 and 2 are explanatory diagrams that schematically show the diagnostic model MO in this embodiment. The diagnostic model MO is a trained model for diagnosing a specific disease, and is used, for example, to classify whether or not a subject suffers from a specific disease. Examples of specific diseases include cancer (biliary tract cancer, pancreatic cancer, stomach cancer, colon cancer, etc.) and diseases other than cancer (inflammatory bowel disease, etc.). In the following description, the specific disease is assumed to be cancer (more specifically, biliary tract cancer).
[0020] As shown in FIGS. 1 and 2 , the diagnostic model MO receives input of comparison data CD identified based on mass spectrometry data MD of a biological sample collected from a subject. Examples of biological samples that can be used include urine, bile, sweat, saliva, blood, cerebrospinal fluid, ascites, pleural effusion, pancreatic juice, gastric juice, and intestinal fluid. In the following description, the biological sample is assumed to be urine, and the target of mass spectrometry is assumed to be multiple substances contained in urine (metabolites (metabolomes)). In mass spectrometry of urine, the peak areas (peak intensities) of a large number of substances, for example, approximately 1,500 to 2,000, are measured.
[0021] In mass spectrometry, various factors can cause variability in results, such as sample collection conditions, sample processing conditions before analysis, daily setup of the analytical equipment (batch-to-batch error), and substances present in the sample that interfere with ionization (matric effect). Furthermore, biological samples such as urine also vary in concentration, and mass spectrometry of biological samples tends to produce relatively large variability in analysis results. Therefore, in this embodiment, comparison data CD is used, as described below, which is determined based on mass analysis data MD indicating the mass analysis results.
[0022] As shown in Figures 1 and 2, in order to generate comparison data CD, a plurality of substance group pairs (hereinafter simply referred to as "pairs") are set in advance for a plurality of substances contained in a sample. A substance group pair is a pair of two substance groups (a first substance group and a second substance group), each of which is composed of at least one substance. In the example shown in Figure 1, the two substance groups constituting each substance group pair are both composed of one substance. For example, the first substance group pair P1 is a pair of a first substance group composed of one substance A1 and a second substance group composed of another substance B1. In this manner, in this specification, a unit composed of a single element is also referred to as a "group."
[0023] 2, each substance group pair is made up of two substance groups, each of which is made up of one or more substances. For example, the first substance group pair P1 is a pair of a first substance group made up of four substances A1, A2, A3, and A4, and a second substance group made up of two substances B1 and B2.
[0024] The substance group pairs are set based on the correlation between the substances contained in the sample, as will be described in detail later.
[0025] The comparison data CD for each substance group pair includes data indicating the magnitude relationship between the peak area of a substance belonging to one substance group (the first substance group or the second substance group) and the peak area of a substance belonging to the other substance group (the second substance group or the first substance group). For example, in the example of Fig. 1, the comparison data CD for each substance group pair is binary data of "0 / 1" that takes the value "0" if the peak area of the substance belonging to the first substance group is smaller than the peak area of the substance belonging to the second substance group, and conversely, takes the value "1" if the peak area of the substance belonging to the first substance group is equal to or greater than the peak area of the substance belonging to the second substance group. For example, for the first substance group pair P1, the peak area a1 of substance A1 constituting the first substance group is smaller than the peak area b1 of substance B1 constituting the second substance group, so the comparison data CD takes the value "0", and for the second substance group pair P2, the peak area c1 of substance C1 constituting the first substance group is greater than or equal to the peak area d1 of substance D1 constituting the second substance group, so the comparison data CD takes the value "1".
[0026] Furthermore, the comparison data CD for each substance group pair may include data indicating the ratio of the peak area of a substance belonging to one substance group to the peak area of a substance belonging to the other substance group. For example, in the example of FIG. 2 , the comparison data CD for each substance belonging to one substance group constituting each substance group pair is the value obtained by dividing the peak area of each substance belonging to that one substance group by the average peak area of all substances belonging to the other substance group. For example, for substance A1 belonging to the first substance group of the first substance group pair P1, the value obtained by dividing the peak area a1 of substance A1 by the average peak area bm (=(b1+b2) / 2) of all substances (substances B1, B2) belonging to the second substance group of the pair. Furthermore, for substance B1 belonging to the second substance group of the pair, the value obtained by dividing the peak area b1 of substance B1 by the average peak area am (=(a1+a2+a3+a4) / 4) of all substances (substances A1, A2, A3, A4) belonging to the first substance group of the pair. In each substance group pair, if the value of the comparison data CD calculated for a substance belonging to one substance group is less than 1, it can be said that the peak area of the substance belonging to that substance group is smaller than the peak area of the substance belonging to the other substance group, and conversely, if the value of the comparison data CD calculated for a substance belonging to one substance group is 1 or more, it can be said that the peak area of the substance belonging to that substance group is equal to or larger than the peak area of the substance belonging to the other substance group. Therefore, in the example of Figure 2, the comparison data CD for each substance group pair can be said to be data that indicates the magnitude relationship between the peak area of a substance belonging to one substance group and the peak area of a substance belonging to the other substance group.
[0027] In this way, the comparison data CD for each substance group pair is information indicating which of the substances belonging to one substance group has a larger peak area than the substance belonging to the other substance group. This comparison data CD may also include information indicating which of the substances belonging to one substance group has a larger peak area than the substance belonging to the other substance group. By using such comparison data CD, the influence of variations in values between samples due to various factors can be relatively reduced.
[0028] 3 is an explanatory diagram conceptually showing mass spectrometry data MD and comparison data CD. In FIG. 3, "Batch 1," "Batch 2," and "Batch 3" represent sample collection units, and sample collection conditions may differ depending on the batch. Of the two "Batch 2" data, the data on the right is for samples collected from healthy subjects, and the remaining data is for samples collected from biliary tract cancer patients.
[0029] The upper part of Figure 3 conceptually illustrates an example of mass spectrometry data MD. The mass spectrometry data MD is data indicating the peak areas of multiple substances contained in a sample (e.g., guanine, dopamine_3_o_sulfate, glycochenodeoxycholate_3_sulfate, etc.) identified by mass spectrometry of the sample. In Figure 3, the peak area of each substance is expressed as concentration. As shown in the figure, even in samples collected from biliary tract cancer patients, the trends in the peak areas of each substance in urine can vary significantly depending on the date of mass spectrometry measurement.
[0030] The bottom section of Figure 3 conceptually illustrates an example of the comparison data CD. In the example shown in Figure 3, similar to the example shown in Figure 1, each of the two substance groups constituting each substance group pair consists of a single substance. The comparison data CD for each substance group pair is binary data (0 / 1), taking a value of "0" if the peak area of the substance belonging to the first substance group is smaller than the peak area of the substance belonging to the second substance group, and a value of "1" if the peak area of the substance belonging to the first substance group is equal to or greater than the peak area of the substance belonging to the second substance group. For biliary tract cancer patients, the trends in the comparison data CD for each substance group pair are similar regardless of the batch, and significantly different from the trends in the comparison data CD for healthy individuals. In this way, by using the comparison data CD for each substance group pair, the influence of variations in values between samples due to various factors can be relatively reduced.
[0031] 1 and 2, the output of the diagnostic model MO is diagnostic result data RD indicating the diagnostic result. The diagnostic result is a classification result of the target sample, more specifically, a classification result of whether or not the subject from whom the target sample, i.e., urine, was collected has biliary tract cancer.
[0032] In this way, for a substance group pair, which is a pair of two substance groups each consisting of at least one substance, in the mass spectrometry results for a plurality of substances contained in a sample, by using a diagnostic model MO, which is a trained model that inputs comparison data CD including data indicating the magnitude relationship between the peak area of a substance belonging to one substance group and the peak area of a substance belonging to the other substance group and outputs the classification result of the sample, it is possible to relatively reduce the influence of variations in values between samples due to various factors, thereby realizing more accurate sample classification, i.e., more accurate disease diagnosis. The classification method in this embodiment is also called Inverse Pair Boosting (abbreviated as "IPB").
[0033] A-2. Configuration of Information Processing Device 3: Next, the configuration of the information processing device 100 for creating the diagnostic model MO and performing a diagnosis using the diagnostic model MO will be described. Figure 4 is an explanatory diagram showing a schematic configuration of the information processing device 100. The information processing device 100 is configured by a computer (PC, server, etc.).
[0034] The information processing device 100 includes a control unit 110, a storage unit 120, a display unit 130, an operation input unit 140, and an interface unit 150. These units are connected to each other so as to be able to communicate with each other via a bus 190. The information processing device 100 may also include a speaker as an output means.
[0035] The display unit 130 of the information processing device 100 is configured, for example, by a liquid crystal display or the like, and displays various images and information. The operation input unit 140 is configured, for example, by a keyboard, mouse, buttons, a microphone, a trackpad, or the like, and accepts operations and instructions from an administrator. The display unit 130 may also function as the operation input unit 140 by being equipped with a touch panel. The interface unit 150 is configured, for example, by a LAN interface, a USB interface, or the like, and communicates with other devices via wired or wireless connections.
[0036] The storage unit 120 of the information processing device 100 is configured, for example, with a ROM, RAM, a hard disk drive (HDD), etc., and is used to store various programs and data, and as a work area when executing various programs, and as a temporary storage area for data. For example, the storage unit 120 stores a diagnostic program CP, which is a computer program for executing the diagnostic model acquisition process and diagnostic process described below. The diagnostic program CP is provided in a state stored on a computer-readable recording medium (not shown), such as a CD-ROM, DVD-ROM, or USB memory, or is provided in a state that can be obtained from an external device (a server or other terminal device on a network) via the interface unit 150, and is stored in the storage unit 120 in a state that is operable on the information processing device 100.
[0037] Furthermore, learning data LD, diagnostic models MO, and diagnostic result data RD are stored in advance or during various processes described below in the storage unit 120 of the information processing device 100. The contents of this information and data will be explained in conjunction with the explanations of the various processes described below.
[0038] The control unit 110 of the information processing device 100 is configured with, for example, a CPU, and controls the operation of the information processing device 100 by executing a computer program read from the storage unit 120. For example, the control unit 110 reads and executes a diagnostic program CP from the storage unit 120, thereby performing various processes described below. In this case, the control unit 110 functions as a raw data acquisition unit 111, a pair setting unit 112, a learning data acquisition unit 113, a model acquisition unit 114, a target data acquisition unit 115, and a classification execution unit 116. The functions of each of these units will be described in conjunction with the explanation of the various processes described below.
[0039] A-3. Diagnostic Model Acquisition Processing: Next, the diagnostic model acquisition processing executed by the information processing device 100 of this embodiment will be described. FIG. 5 is a flowchart showing the diagnostic model acquisition processing in this embodiment. The diagnostic model acquisition processing is processing for acquiring the diagnostic model MO described above. In this embodiment, the information processing device 100 acquires the diagnostic model MO by creating the diagnostic model MO by itself using predetermined machine learning. The diagnostic model acquisition processing is started in response to a start instruction being input by the user operating the operation input unit 140 of the information processing device 100.
[0040] First, the raw data acquisition unit 111 ( FIG. 4 ) of the information processing device 100 acquires raw data to be used for creating the diagnostic model MO (S110). The raw data is mass analysis data MD (see FIGS. 1 and 2 ) showing the results of mass analysis of urine as biological samples collected from multiple biliary tract cancer patients and healthy individuals (controls). The raw data specifies the peak areas of numerous substances (metabolites) contained in the urine. The raw data is acquired via the interface unit 150 or the operation input unit 140.
[0041] Next, the pair setting unit 112 (FIG. 4) of the information processing device 100 sets a plurality of substance group pairs (S120). As described above, a substance group pair is a pair of two substance groups, each of which is composed of at least one substance (see FIGS. 1 and 2). The pair setting unit 112 sets a plurality of substance group pairs based on the correlation between the plurality of substances contained in the sample.
[0042] A specific method for setting multiple substance group pairs using the pair setting unit 112 is as follows, for example. First, for substances (metabolites) that are present in more than 95% of healthy subjects, a Wilcoxon rank-sum test, for example, is performed to select substances with a sufficiently large statistical significance. The selected substances can be said to be cancer-specific substances. Of the selected substances, the peak area data of substances that are decreased relative to healthy subjects (satisfying the relationship Log2 fold change (L2FC) < 0) is multiplied by -1, the sign is inverted, and the Pearson correlation coefficient is calculated. Unsupervised clustering (for example, using the R package M3C) using this correlation coefficient as the distance is performed up to maxK = 10, and multiple substances are classified into the number of clusters that results in the lowest p-value. By mixing the data with the inverted sign and clustering as described above, substance groups showing an inverse correlation with the inverted sign are classified into the same cluster. Pairs are created between these inversely correlated substance groups within the same cluster. A cluster is checked to see whether it contains enough substances (e.g., five or more) showing an inverse correlation. If a cluster does not contain enough substances, it is paired with another cluster with a similar Pearson correlation coefficient. Each cluster established in this way becomes a substance group pair. In cancer, metabolic reprogramming causes abnormalities in various metabolic pathways. When metabolism stalls at a certain point in a metabolic pathway, the pre-metabolic substance group increases, and conversely, the post-metabolic substance group decreases. By comprehensively extracting the inverse correlations between multiple substance groups on metabolic pathways that arise in this way, it is possible to establish substance group pairs consisting of two substance groups that highlight metabolic differences between biliary tract cancer patients and healthy individuals.
[0043] Next, the training data acquisition unit 113 ( FIG. 4 ) of the information processing device 100 acquires the above-described comparison data CD for the plurality of substance group pairs that have been set (S130). For example, as shown in FIG. 1 , the training data acquisition unit 113 acquires binary data of "0 / 1" as the comparison data CD for the plurality of substance group pairs. The value is "0" if the peak area of the substance belonging to the first substance group is smaller than the peak area of the substance belonging to the second substance group, and "1" if the peak area of the substance belonging to the first substance group is equal to or greater than the peak area of the substance belonging to the second substance group. Alternatively, as shown in FIG. 2 , the training data acquisition unit 113 calculates the comparison data CD for the plurality of substance group pairs by dividing the peak area of each substance belonging to one substance group by the average value of the peak areas of all substances belonging to the other substance group. The training data acquisition unit 113 creates training data LD using the acquired comparison data CD. The learning data LD is data in which the comparison data CD is associated with whether the subject is a biliary tract cancer patient or a healthy subject.
[0044] Next, the model acquisition unit 114 ( FIG. 4 ) of the information processing device 100 creates a diagnostic model MO through machine learning using the learning data LD (S140). Various known machine learning algorithms can be used for the machine learning used to create the diagnostic model MO. For example, a random forest or a support vector machine may be used to create the diagnostic model MO. The diagnostic model MO created through machine learning is stored in the storage unit 120 of the information processing device 100. This completes the diagnostic model MO acquisition process ( FIG. 5 ).
[0045] A-4. Diagnostic Processing: Next, the diagnostic processing executed by the information processing device 100 of this embodiment will be described. FIG. 6 is a flowchart showing the diagnostic processing in this embodiment. The diagnostic processing is a processing for classifying a target sample using the diagnostic model MO, more specifically, for classifying whether or not the subject from whom the target sample was collected is suffering from biliary tract cancer. The diagnostic processing is started in response to a start instruction being input by the user operating the operation input unit 140 of the information processing device 100.
[0046] First, the target data acquisition unit 115 ( FIG. 4 ) of the information processing device 100 acquires comparison data CD for a target sample (S210). The target sample is urine collected from a subject. As described above, the comparison data CD includes, for each substance group pair, data indicating the magnitude relationship between the peak area of a substance belonging to one substance group and the peak area of a substance belonging to the other substance group, as well as data indicating the ratio of the peak area of a substance belonging to one substance group to the peak area of a substance belonging to the other substance group. The target data acquisition unit 115 acquires mass analysis data MD for the target sample via the interface unit 150 and creates comparison data CD based on the mass analysis data MD. Alternatively, the target data acquisition unit 115 acquires comparison data CD previously created based on the mass analysis data MD via the interface unit 150.
[0047] Next, the classification execution unit 116 ( FIG. 4 ) of the information processing device 100 performs a diagnosis using the comparison data CD for the subject sample and the diagnostic model MO (S220). Specifically, the classification execution unit 116 inputs the comparison data CD for the subject sample into the diagnostic model MO, and obtains a classification result (a classification result indicating whether the subject is suffering from biliary tract cancer) output from the diagnostic model MO. The classification execution unit 116 generates diagnostic result data RD indicating the diagnostic result and stores it in the storage unit 120 of the information processing device 100.
[0048] Next, the classification execution unit 116 outputs the diagnosis result based on the diagnosis result data RD (S230). For example, the classification execution unit 116 displays the diagnosis result on the display unit 130. This completes the diagnostic process.
[0049] A-5. Example: An example of the above-mentioned diagnostic model MO will be described below. Figure 7 is an explanatory diagram showing the results of principal component analysis of mass analysis data MD. In Figure 7, the X-axis represents the first principal component (PC1) and the Y-axis represents the second principal component (PC2). Each axis represents the contribution rate of each component. Also, in Figure 7, the hatching of each plot indicates the classification of the target data. In the figure, "CCA" represents a sample from a biliary tract cancer patient, "No_PDAC" represents a sample from a patient who underwent surgery due to suspected pancreatic cancer but whose pathological diagnosis revealed that the cancer was not present, "CTR" represents a sample from a healthy subject (control), "IPMC" represents a sample from a type of precancerous lesion (post-cancerous) called intraductal papillary mucinous carcinoma, "IPMN" represents a sample from a type of precancerous lesion (pre-cancerous) called intraductal papillary mucinous neoplasm, and "PDAC" represents a sample from a pancreatic cancer patient. Also, in Figure 7, the shape of each plot indicates the batch to which the target data belongs. These points also apply to Figures 8 and 9, which will be described later.
[0050] As shown in Figure 7, the results of the principal component analysis of the mass spectrometry data MD show that the biliary tract cancer patient samples (CCA, black) and healthy control samples (CTR, open) are clustered on the right side of the graph area, while the pancreatic cancer patient samples (PDAC, vertical hatching) are clustered on the left side of the graph area, with the two samples separated horizontally. In this example, the biliary tract cancer patient samples and healthy control samples were collected first thing in the morning, while the pancreatic cancer patient samples were collected after anesthesia was administered. Therefore, it is believed that the variability due to differences in sample collection conditions became the largest variance vector (first principal component). Furthermore, in the results of this principal component analysis, the samples from batch 4 (diamonds) are clustered on the top side of the graph area, while the samples from other batches are clustered on the bottom side of the graph area, with the two samples separated vertically. Therefore, the second principal component is believed to represent differences between batches. On the other hand, the distribution of biliary tract cancer patient samples (CCA, black) and the distribution of healthy subject samples (CTR, open) completely overlap in the graph area, and the differences between the biliary tract cancer patient samples and the healthy subject samples are below the third principal component.
[0051] FIG. 8 is an explanatory diagram showing the results of principal component analysis of mass spectrometry data MD, narrowed down to cancer-specific substances (features), as described above. As shown in FIG. 8, by selecting "cancer-specific substances," substances unrelated to the effects of anesthesia can be collected. As a result, the distance between the biliary tract cancer patient samples (CCA, black) and healthy control samples (CTR, open), which were separated along the first principal component (variation due to differences in sample collection conditions) in the principal component analysis results shown in FIG. 7, and the pancreatic cancer patient samples (PDAC, vertical hatching) has been significantly reduced. Furthermore, the biliary tract cancer patient samples (CCA, black) and healthy control samples (CTR, open) are separated from each other along the second principal component. However, in the principal component analysis results shown in FIG. 8, the biliary tract cancer patient samples (CCA, black) and healthy control samples (CTR, open) are not separated along the first principal component. In other words, the first principal component can still be said to be the variation caused by differences in sample collection conditions.
[0052] FIG. 9 is an explanatory diagram showing the results of principal component analysis of the comparison data CD generated based on the mass analysis data MD. The results of the principal component analysis shown in FIG. 9 clearly separate biliary tract cancer patient samples (CCA, black) from healthy control samples (CTR, white). Specifically, even if a positive value of the first principal component is classified as a biliary tract cancer patient, and a negative value of the first principal component is classified as a healthy control, diagnosis can be performed with high accuracy. In other words, the results of this principal component analysis indicate that the presence or absence of biliary tract cancer is the largest factor contributing to the variability, contributing to the diagnosis of biliary tract cancer. Furthermore, since the healthy control samples (CTR, white) and the pancreatic cancer patient samples (PDAC, vertical hatching) are significantly separated vertically, the second principal component can be said to represent the variability of whether or not the sample is pancreatic cancer. Variability due to differences between batches is represented by the third principal component and below. The contribution rate of the first principal component was 22%, and that of the second principal component was 18.9%, boosting 40.9% of the generated data to contribute to the diagnosis of biliary tract cancer and pancreatic cancer, while the contribution rate of the third principal component decreased to 6%, resulting in a relative reduction in variability due to differences between batches.
[0053] From the results of the principal component analysis shown in Figures 7 to 9, it can be said that by using the diagnostic model MO created by learning using the comparison data CD, the influence of variation in values between samples caused by various factors can be relatively reduced, and more accurate diagnosis of disease can be achieved.
[0054] FIG. 10 is an explanatory diagram showing the accuracy of the diagnostic model MO of this embodiment. FIG. 10 shows the specificity and sensitivity when the diagnostic model MO was created by learning using data submitted in 2021 and then verified using data submitted in 2022 (i.e., when a prospective study was conducted). As shown in FIG. 10, in the diagnosis of biliary tract cancer using the diagnostic model MO (classification of whether or not a patient has biliary tract cancer), sensitivity was 82.8%, specificity was 100%, and ROC-AUC was 0.994, indicating that highly accurate diagnosis was achieved.
[0055] Fig. 11 is an explanatory diagram showing the accuracy of the diagnostic model MO of this embodiment. Fig. 11 shows the specificity and sensitivity for each stage of pancreatic cancer (stages 1 to 4) when the diagnostic model MO of this embodiment is applied to pancreatic cancer data. As shown in Fig. 11, in diagnosing pancreatic cancer using the diagnostic model MO of this embodiment, highly accurate diagnoses are achieved not only for stages 2 to 4, but also for stage 1, which is relatively difficult to diagnose because it is an early stage.
[0056] 12 and 13 are explanatory diagrams showing the results of a comparison between creatinine correction as a comparative example and the diagnostic model MO of this embodiment. FIGS. 12 and 13 respectively show the results of principal component analysis and classification accuracy when two data sets (batch 1 and batch 2) are classified using creatinine correction and the diagnostic model MO of this embodiment. As shown in FIG. 12 , with creatinine correction, batch 1 and batch 2 are separated, and the variability due to differences between the batches is relatively large, resulting in relatively low classification accuracy. On the other hand, as shown in FIG. 13 , with classification using the diagnostic model MO of this embodiment (IPB), batch 1 and batch 2 are not separated, and the variability due to differences between the batches is relatively small, resulting in relatively high classification accuracy. Thus, classification using the diagnostic model MO of this embodiment can improve classification accuracy compared to classification using creatinine correction.
[0057] A-6. Effects of the Present Embodiment: As described above, the information processing device 100 of the present embodiment is an apparatus for classifying samples and includes a model acquisition unit 114, a target data acquisition unit 115, and a classification execution unit 116. The model acquisition unit 114 acquires a diagnostic model MO, which is a trained model that receives comparison data CD as input and outputs a sample classification result. The comparison data CD includes data indicating the magnitude relationship between the peak area of a substance belonging to one substance group and the peak area of a substance belonging to the other substance group for at least one substance group pair, each of which is a pair of two substance groups consisting of at least one substance, in a mass spectrometry result for a plurality of substances contained in the sample. The target data acquisition unit 115 acquires the comparison data CD for the target sample. The classification execution unit 116 inputs the comparison data CD for the target sample into the diagnostic model MO, thereby classifying the target sample and outputting the classification result.
[0058] As described above, the information processing device 100 of this embodiment classifies target samples using a diagnostic model MO, which is a trained model that receives comparison data CD as input and outputs sample classification results. The comparison data CD includes data indicating the magnitude relationship between the peak area of a substance belonging to one substance group and the peak area of a substance belonging to the other substance group for at least one substance group pair, each of which is a pair of two substance groups consisting of at least one substance, in the mass analysis results for multiple substances contained in the sample. Therefore, by using the comparison data CD, the influence of variations in values between samples due to various factors can be relatively reduced, thereby improving the accuracy of sample classification using mass analysis results. Furthermore, because the comparison data CD is data created using mass analysis results for multiple substances contained in the sample, the accuracy of sample classification using mass analysis results can be improved compared to a configuration that uses mass analysis results for only specific substances.
[0059] Furthermore, in the information processing device 100 of this embodiment, the comparison data CD may include data indicating the ratio of the peak area of a substance belonging to one substance group to the peak area of a substance belonging to the other substance group. With this configuration, substance group pairs can be formed using any combination of substances, not just pairs of substances whose peak areas in the samples are similar and whose magnitude relationship is reversed depending on the classification to which they belong. This increases the degree of freedom in setting substance group pairs, making it possible to obtain more preferable comparison data CD and effectively improve the accuracy of sample classification using mass spectrometry results.
[0060] Furthermore, in the information processing device 100 of this embodiment, the comparison data CD may include data indicating the ratio of the peak area of each substance belonging to one substance group to the average value of the peak areas of all substances belonging to the other substance group. With this configuration, the value of the comparison data CD is determined for each substance, thereby preventing a decrease in the amount of information. This allows the comparison data CD to be obtained with a larger amount of information, effectively improving the accuracy of sample classification based on mass spectrometry results.
[0061] Furthermore, in the information processing device 100 of this embodiment, the comparison data CD includes data on a plurality of substance group pairs, which makes it possible to obtain comparison data CD with a larger amount of information, thereby effectively improving the accuracy of sample classification based on the mass spectrometry results.
[0062] The information processing device 100 of this embodiment further includes a pair setting unit 112 that sets at least one substance group pair based on the correlation between a plurality of substances. This makes it possible to set a substance group pair that contributes to more appropriate classification, thereby effectively improving the accuracy of sample classification based on mass spectrometry results.
[0063] Furthermore, in the information processing device 100 of this embodiment, the target sample may be a biological sample collected from a human. This configuration can improve the accuracy of biological sample classification with reference to the results of mass spectrometry of the biological sample. Furthermore, the biological sample may be at least one of urine, bile, sweat, saliva, blood, cerebrospinal fluid, ascites, pleural effusion, pancreatic juice, gastric juice, and intestinal fluid. This configuration can improve the accuracy of biological sample classification with reference to the results of mass spectrometry of the biological sample.
[0064] Furthermore, in the information processing device 100 of this embodiment, the classification execution unit 116 may perform classification for a specific disease. With this configuration, it is possible to improve the accuracy of diagnosing a specific disease with reference to the mass spectrometry results. Furthermore, the specific disease may be cancer. With this configuration, it is possible to improve the accuracy of diagnosing a cancer with reference to the mass spectrometry results.
[0065] B. Modifications: The technology disclosed in this specification is not limited to the above-described embodiment, and can be modified in various forms without departing from the spirit of the invention. For example, the following modifications are also possible.
[0066] The configuration of the information processing device 100 in the above embodiment is merely an example and can be modified in various ways. Furthermore, the contents of the diagnostic model acquisition process and the diagnostic process in the above embodiment are merely an example and can be modified in various ways. For example, in the above embodiment, the information processing device 100 acquires the diagnostic model MO by generating it itself, but the information processing device 100 may also acquire a diagnostic model MO generated by another device.
[0067] The data structure and machine learning algorithm used to create the diagnostic model MO in the above embodiment are merely examples and can be modified in various ways. For example, the comparison data CD used to create the diagnostic model MO may be composed of data for a single substance group pair, or may be composed of data for multiple substance group pairs. Furthermore, in the example of FIG. 1 , the comparison data CD for each substance group pair may be the value obtained by dividing the peak area of a substance constituting one substance group by the peak area of a substance constituting the other substance group. Furthermore, in the example of FIG. 2 , the comparison data CD for each substance group pair may be the value obtained by dividing the average value of the peak areas of all substances belonging to one substance group by the average value of the peak areas of all substances belonging to the other substance group. Furthermore, in the example of FIG. 2 , other representative values (e.g., medians) may be used instead of the average values. Furthermore, the mass analysis data MD and the comparison data CD do not need to be data identifying specific substance names, but may be data identifying the positions (e.g., m / z values) of peaks in a mass spectrum obtained by mass analysis.
[0068] Before inputting the comparison data CD into the diagnostic model MO, a principal component analysis may be performed, and the first and second principal components may be used as inputs to the diagnostic model MO.
[0069] In the above embodiment, information processing for classifying whether or not a subject from whom a sample was collected has biliary tract cancer is exemplified, but the technology disclosed in this specification is not limited to biliary tract cancer and can be similarly applied to classifying whether or not a subject has a specific disease. Furthermore, the technology disclosed in this specification is not limited to diagnosis and can be similarly applied to classifying samples from some perspective.
[0070] In the above embodiments, at least one of the functional units included in the control unit 110 of the information processing device 100 may be included in another device rather than the control unit 110 of the information processing device 100. The diagnostic model acquisition process and the diagnostic process in the above embodiments do not necessarily have to be executed by a single device, and may each be executed by a different device. The steps of the diagnostic model acquisition process in the above embodiments do not necessarily have to be executed by a single device, and may each be executed by a different device. Similarly, the steps of the diagnostic process in the above embodiments do not necessarily have to be executed by a single device, and may each be executed by a different device. In the above embodiments, part of the configuration realized by hardware may be replaced with software, and conversely, part of the configuration realized by software may be replaced with hardware.
[0071] The machine learning method may be a single method such as a random forest or a support vector machine, or an ensemble model (a stack model or a blend model) that combines several of these methods.
[0072] The results of mass spectrometry are not limited to the peak area of each substance, but may be any peak intensity value.
[0073] 100: Information processing device 110: Control unit 111: Raw data acquisition unit 112: Pair setting unit 113: Learning data acquisition unit 114: Model acquisition unit 115: Target data acquisition unit 116: Classification execution unit 120: Storage unit 130: Display unit 140: Operation input unit 150: Interface unit 190: Bus CP: Diagnostic program LD: Learning data MD: Mass spectrometry data MO: Diagnostic model RD: Diagnostic result data
Claims
1. a model acquisition unit that receives comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, in a mass analysis result for a plurality of substances contained in a sample, and acquires a trained model that outputs a classification result for the sample; a target data acquisition unit that acquires the comparison data for the target sample; a classification execution unit that inputs the comparison data about the target sample into the trained model, thereby classifying the target sample, and outputs the classification result; An information processing device comprising:
2. 2. The information processing device according to claim 1, An information processing device, wherein the comparison data includes data indicating a ratio of a peak intensity of a substance belonging to one of the substance groups to a peak intensity of a substance belonging to the other of the substance groups.
3. 3. The information processing device according to claim 2, An information processing device, wherein the comparison data includes data indicating a ratio of peak intensity of each substance belonging to the one substance group to an average value of peak intensities of all substances belonging to the other substance group.
4. 4. The information processing device according to claim 1, The information processing device, wherein the comparison data includes data for a plurality of the substance group pairs.
5. The information processing device according to any one of claims 1 to 3, further comprising: an information processing device comprising: a pair setting unit that sets the at least one substance group pair based on correlations between the plurality of substances;
6. 4. The information processing device according to claim 1, The information processing device, wherein the sample is a biological sample collected from a human.
7. 7. The information processing device according to claim 6, The information processing device, wherein the biological sample is at least one of urine, bile, sweat, saliva, blood, cerebrospinal fluid, ascites, pleural effusion, pancreatic juice, gastric juice, and intestinal fluid.
8. 7. The information processing device according to claim 6, The classification execution unit performs the classification for a specific disease.
9. 9. The information processing device according to claim 8, The information processing device, wherein the specific disease is cancer.
10. An information processing device comprising: a model acquisition unit that, for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, receives comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups in a mass analysis result for a plurality of substances contained in a sample, and generates a trained model by executing training that outputs a classification result for the sample.
11. a step of acquiring a trained model that inputs comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, in a mass spectrometry result for a plurality of substances contained in a sample, and outputs a classification result for the sample; obtaining said comparison data for a target sample; performing classification of the target sample by inputting the comparison data for the target sample into the trained model, and outputting the classification results; An information processing method comprising:
12. An information processing method comprising: a step of performing learning to generate a trained model that inputs comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, in mass spectrometry results for a plurality of substances contained in a sample; and outputs the classification result of the sample.
13. On the computer, a process of acquiring a trained model that inputs comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, in a mass analysis result for a plurality of substances contained in a sample, and outputs a classification result for the sample; obtaining the comparative data for the target sample; a process of inputting the comparison data for the target sample into the trained model, thereby classifying the target sample, and outputting the classification result; A computer program that executes
14. On the computer, A computer program that executes a process to generate, by performing learning, a trained model that receives as input comparison data including data indicating the magnitude relationship between the peak intensity of a substance belonging to one of the substance groups and the peak intensity of a substance belonging to the other of the substance groups, for at least one substance group pair, which is a pair of two substance groups each consisting of at least one substance, in the mass analysis results of a plurality of substances contained in a sample, and outputs the classification result of the sample.
15. 11. The information processing device according to claim 1, The substance group pair is a pair of two of the substance groups that exhibit an anti-correlation.
16. 13. An information processing method according to claim 11 or claim 12, An information processing method, wherein the substance group pair is a pair of two substance groups that exhibit an anti-correlation.
17. 15. A computer program according to claim 13 or claim 14, comprising: The substance group pair is a pair of two of the substance groups that exhibit an inverse correlation.