Information Systems

The information system uses prediction models and machine learning to identify gene functions related to disease progression, overcoming limitations of WGCNA by accurately extracting relevant biological functions and reducing analysis costs.

JP7844387B2Active Publication Date: 2026-04-13HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2023-05-29
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing methods, such as WGCNA, struggle to sufficiently narrow down gene functions and genes associated with disease progression and exacerbation, making it difficult to identify biological functions related to disease progression and worsening.

Method used

An information system comprising a processor and storage device that generates prediction models using omics data to evaluate and extract biological functions related to disease progression, utilizing a target variable table, metadata table, and machine learning techniques to assess feature importance.

Benefits of technology

Enables the extraction of biological functions, such as gene functions, that are highly associated with disease progression and severity, allowing for more efficient identification and reduction of analysis costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007844387000001
    Figure 0007844387000001
  • Figure 0007844387000002
    Figure 0007844387000002
  • Figure 0007844387000003
    Figure 0007844387000003
Patent Text Reader

Abstract

To provide an information system capable of extracting biological functions related to disease progression and aggravation.SOLUTION: The information system includes a processor and a storage. The storage stores an objective variable table related to objective variables and a meta-information table of biological functions and features related to those biological functions. The processor is configured to generate a prediction model that predicts the objective variable using the features of each biological function as explanatory variables, evaluate the accuracy of the prediction model, and based on the evaluation results of the prediction model, extract objective variables and related biological functions.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information system.

Background Art

[0002] In recent years, drug discovery and treatment development have been carried out through omics data analysis that comprehensively analyzes biomolecules such as genes, proteins, or metabolites. Methods for extracting genes related to disease progression and exacerbation have been developed in the past. In Non-Patent Document 1, using a method called WGCNA (Weighted gene co-expression network analysis), gene functions and gene identification related to the pathological progression of adrenocortical carcinoma were performed.

Prior Art Documents

Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] As described above, gene functions and gene extraction related to disease progression and exacerbation have been performed using co-expression network analysis such as WGCNA. However, it was considered that it is not easy to sufficiently narrow down gene functions and genes related to disease progression and exacerbation with only co-expression network analysis such as WGCNA.

[0005] Non-patent document 1 indicated that 110 gene functions were suggested to be associated with the progression of adrenocortical carcinoma by WGCNA, but no further analysis was conducted to narrow down which of these 110 gene functions were important. This suggests that the above analysis alone is insufficient to narrow down the gene functions.

[0006] Therefore, there is a challenge in identifying the biological functions associated with disease progression and worsening of the disease. [Means for solving the problem]

[0007] The present invention provides the following information system. This information system comprises a processor and a storage device. The storage device stores a target variable table for the target variable and a metadata table for biological functions and features related to said biological functions. The processor generates a prediction model for each biological function, using features as explanatory variables to predict the target variable, evaluates the accuracy of the prediction model, and extracts biological functions related to the target variable based on the evaluation results of the prediction model. [Effects of the Invention]

[0008] According to the present invention, it is possible to extract biological functions related to disease progression and worsening of the disease. Other problems, configurations, and effects not mentioned above will be clarified by the following description of embodiments for carrying out the invention. [Brief explanation of the drawing]

[0009] [Figure 1] This figure shows an example of the configuration of an important biological function presentation system. [Figure 2] This figure shows an example of the hardware configuration of a system for presenting important biological functions. [Figure 3] This figure shows an example of the structure of the target variable table held by the important biological function presentation system. [Figure 4] This figure shows an example of the structure of a feature table held by a system that displays important biological functions. [Figure 5]This figure shows an example of the structure of the metadata table held by the important biological function presentation system. [Figure 6] This figure shows an example of the configuration of a predictive model management table held by a critical biological function presentation system. [Figure 7] This figure shows an example of the configuration of the predictive model performance information management table held by the important biological function presentation system. [Figure 8] This figure shows an example of the structure of the feature importance table held by the Important Biological Function Presentation System. [Figure 9] This diagram shows an example of the processing flow of an important biological function presentation system. [Figure 10] This figure shows an example of the processing flow for predictive model construction and performance evaluation within the processing flow of a critical biological function presentation system. [Figure 11] This figure shows an example of the processing flow for calculating feature importance within the processing flow of the Important Biological Function Presentation System. [Modes for carrying out the invention]

[0010] Embodiments of the present invention will be described below with reference to the drawings. The embodiments are illustrative examples for explaining the present invention, and have been omitted and simplified as appropriate for clarity of explanation. The present invention can also be implemented in various other forms. Unless otherwise specified, each component may be singular or plural. The positions, sizes, shapes, and ranges of the components shown in the drawings may not represent their actual positions, sizes, shapes, and ranges in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the positions, sizes, shapes, and ranges disclosed in the drawings. Examples of various types of information may be described using terms such as "table," "list," and "queue," but these types of information may also be represented by other data structures. For example, various types of information such as "XX table," "XX list," and "XX queue" may be referred to as "XX information." When describing identification information, terms such as "identification information," "identifier," "name," "ID," and "number" are used, and these terms are interchangeable. When there are multiple components with the same or similar function, they may be described using the same symbol but with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description. In embodiments, processing performed by executing a program may be described. Here, the computer executes the program using a processor (e.g., CPU, GPU) and performs processing defined by the program using memory resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the main entity performing the processing by executing the program may be the processor. Similarly, the main entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The main entity performing the processing by executing the program may be an arithmetic unit, and may include dedicated circuits that perform specific processing. Here, dedicated circuits include, for example, FPGAs (Field Programmable Gate Arrays), ASICs (Application Specific Integrated Circuits), CPLDs (Complex Programmable Logic Devices), etc. The program may be installed on the computer from the program source. The program source may be, for example, a program distribution server or a storage medium readable by the computer. If the program source is a program distribution server, the program distribution server includes a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to other computers. In addition, in the embodiment, two or more programs may be implemented as one program, or one program may be implemented as two or more programs.

[0011] In this embodiment, an information system that supports the analysis of omics data such as genes and proteins is described. The technology of this information system, as an example, relates to extracting gene functions and other biological information, such as genes, that contribute to predicting disease onset and severity from omics data. This information system can, for example, reduce the costs of analyzing omics data, thus contributing from an economic standpoint.

[0012] First, an example of the configuration of the Important Biological Function Presentation System will be explained with reference to Figure 1. Figure 1 is a diagram showing an example of the configuration of the Important Biological Function Presentation System. The Important Biological Function Presentation System 100 (information system) is a system that reads data, generates models using machine learning technology, calculates feature importance from the generated machine learning model, and plots the results. The Important Biological Function Presentation System 101 exemplified here comprises a data acquisition unit 102, a prediction model generation and evaluation unit 103, a feature importance calculation unit 104, and a result plotting unit 105.

[0013] The data acquisition unit 102 shown in FIG. 1 has a function of capturing, as an example, the data shown in Data 106. The gene expression level data 107 is data obtained by quantitatively measuring the amount of RNA (Ribonucleic acid), which is the expression product of a gene. The protein expression level data 108 is data obtained by quantitatively measuring the amount of protein, which is the expression product of RNA. The patient information data 109 is data that describes background information and test value information of patients, such as age, gender, and clinical test values, used as explanatory variables in a machine learning model, and numerical information used as a target variable in the machine learning model. The meta-information data 110 is data that defines features used as explanatory variables when constructing a model using machine learning techniques. The data captured by the data acquisition unit 102 is managed as a target variable table 111, a feature table 112, and a meta-information table 113 within the important biological function presentation system 101.

[0014] The prediction model generation / evaluation unit 103 has a function of generating a prediction model using machine learning techniques, a function of evaluating the performance of the generated prediction model, and a function of saving the generated prediction model. The generated prediction model is held within the important biological function presentation system 101 as a prediction model file 117. Also, information regarding the generated prediction model is stored in a prediction model management table 114 and a prediction model performance information management table 115.

[0015] The feature importance calculation unit 104 has a function of calculating the importance of features from the prediction model generated by the prediction model generation / evaluation unit 102. The calculated feature importance is stored in a feature importance table 116. The result drawing unit 105 has a function of displaying the performance information of the model saved by the prediction model generation / evaluation unit 102 and the feature importance output by the feature importance calculation unit 104. It is assumed that each function of the important biological function presentation system 101 can be controlled from the input / output terminal 118. Details of the configuration of each table held in the important biological function presentation system 101 will be described later.

[0016] Next, an example of the hardware configuration of the important biological function presentation system 101 will be described with reference to FIG. 2. The important biological function presentation system 101 includes a host CPU (Central Processing Unit) 201 that is an example of a processor and performs various processes, a host memory 202, a peripheral IF 203, a storage device 204, a communication IF 205, and a bus 206. The host CPU 201, the host memory 202, the peripheral I / F 203, the storage device 204, and the communication IF 205 are connected via the bus 206 and can exchange information with each other.

[0017] Among these, the host CPU 201 is an arithmetic unit that executes the program held in the storage device 204. The host memory 202 is a volatile memory device used as a working memory and a temporary buffer for input / output data when the host CPU 201 executes the above program. The peripheral IF 203 is an interface that connects various peripheral devices such as input / output devices such as a mouse, a keyboard, a monitor, and external storage such as a USB (Universal Serial Bus) memory to the important biological function presentation system 101.

[0018] The storage device 204 is composed of a magnetic disk device, a flash ROM (Read Only Memory), etc., and stores an OS, various drivers, various application programs, and various information used in the programs (for example, information set by an administrator or a maintenance person, etc.).

[0019] The communication IF 205 provides an interface when the important biological function presentation system 101 performs communication. There may be two or more of these communication IFs 205.

[0020] In the Important Biological Function Presentation System 101, the host CPU 201 reads the program from the storage device 204 into the host memory 202 and executes it, thereby implementing the aforementioned functions necessary for the Important Biological Function Presentation System 101, namely the data acquisition unit 102, the prediction model generation and evaluation unit 103, the feature importance calculation unit 104, and the result plotting unit 105.

[0021] Next, an example of the information stored in the target variable table 111, which is provided in the storage device 204 of the important biological function presentation system 101, will be explained using Figure 3. The target variable table 111 stores the values ​​of patient ID 301 and target variable value 302. The column for patient ID 301 is the area for storing ID information representing the patient's identification information. The column for target variable value 302 is the area for storing the value of the target variable corresponding to each patient ID. For example, row 303 indicates that the value of the target variable for a patient with patient ID "#1" is "1". Note that Figure 3 shows the case where the target variable value 302 is binary (two-class classification), but it may also be a continuous quantity or have multiple values ​​(multi-class classification).

[0022] Next, an example of the information stored in the feature table 112 provided in the storage device 204 of the important biological function presentation system 101 will be explained using Figure 4. The feature table 112 stores the values ​​of patient ID 301, patient information 401 obtained from patient information data 109, and expression level information 402 obtained from gene expression level data 107 or protein expression level data 108. The patient information 401 and expression level information 402 stored in the feature table 112 are used as features (explanatory variables) when the prediction model generation and evaluation unit 103 generates a prediction model. Each column of patient information 401 is an area that stores patient background information such as gender and age, or information on clinical test values ​​such as albumin and creatinine. The expression level information 402 is an area that stores information such as gene expression level and protein expression level. For example, in row 403, the feature data for patient ID "#1" indicates that the gender value is "0", the age value is "43", gene expression level A is "0.54", gene expression level B is "6.54", and gene expression level C is "2.76".

[0023] Next, an example of the information stored in the metadata table 113 provided in the memory device 204 of the important biological function presentation system 101 will be explained using Figure 5. The metadata table 113 stores the values ​​of function ID 501, function name 502, and feature information 503. The function ID 501 column stored in the metadata table 113 is the area for storing IDs that represent the identification information of functions (biological functions). The function name 502 column stored in the metadata table 113 is the area for storing function names corresponding to function IDs. Each column of feature information 503 stored in the metadata table 113 is the area for storing feature names belonging to the function corresponding to function name 502. Here, the number of columns in feature information 503 differs depending on the function corresponding to function name 502.

[0024] The feature information 503 corresponding to function ID 501 can be specified from the input / output terminal 118. Furthermore, it can be automatically entered based on previous analysis such as WGCNA, rather than being specified manually. Here, the feature information 503 corresponding to function name ID 501 is the gene corresponding to that function if it is a gene function. The relationship between gene function and gene may use information defined in a known GO (Gene Ontology), and if there is a relationship (annotation) between a certain GO and a gene, the GO can be the function name and the feature information the annotated gene. For example, if the function name ID is "GO:0090248", it will be associated with the gene "cdh2". Such gene functions and genes can be entered by the user from the input / output terminal 118, or the combination of gene function and gene identified using co-expression network analysis such as WGCNA can be automatically entered. Also, the combination of function (biological function) and the feature belonging to that function is not limited to gene function and gene; it may also be a combination of function (biological function) and feature related to other biological information, such as protein function and the protein belonging to that protein function.

[0025] For example, it is possible to manually add other factors such as age and sex to the metadata table 113. This allows for the extraction of gene functions and genes that are important for disease development, while also considering confounding factors between gene expression levels and disease onset.

[0026] Each row in the metadata table 113 represents a feature (explanatory variable) used when the prediction model generation and evaluation unit 103 generates a prediction model. For example, row 504 indicates that for a function ID of "K001" and a function name of "DNA damage repair," the features are "age" for feature 1, "gene AA" for feature 2, and "gene AR" for feature 3. This means that a prediction model consisting of the features "age," "gene AA," and "gene AR" is generated.

[0027] Next, an example of the information stored in the prediction model management table 114, which is provided in the storage device 204 of the important biological function presentation system 101, will be explained using Figure 6. The prediction model management table 114 stores the values ​​of model ID 601, storage directory 602, model file name 603, target variable table name 604, and function ID 605. The column for model ID 601 is the area that stores ID information representing the model's identification information. The columns for storage directory 602 and model file name 603 are the areas that store the directory name and file name, respectively, where the prediction model file 117 is stored. The area for target variable table name 604 is the area that stores the target variable table name used when generating the prediction model. The area for function ID 605 is the area that stores the function ID corresponding to the list of features used when generating the prediction model.

[0028] For example, row 606 indicates that the prediction model file corresponding to model ID "M001" is stored in " / home / user / model / " as listed in the storage directory column. It also indicates that the filename of the prediction model file with model ID "M001" is "Target01.K001.model". Furthermore, it indicates that the prediction model with model ID "M001" uses the target variable table "Target01" as its target variable and is generated from the feature list corresponding to feature ID "K001".

[0029] Next, an example of the information stored in the prediction model performance information management table 115, which is provided in the storage device 204 of the important biological function presentation system 101, will be explained using Figure 7. The prediction model performance information management table 115 stores the values ​​of performance ID 701, model ID 601, function ID 605, performance index 702, and performance value 703. The column for performance ID 701 is the area for storing the ID representing the identification information of the performance information. The column for performance index 702 is the area for storing the index name used when the prediction model generation and evaluation unit 103 evaluates the performance of the prediction model. The column for performance value 703 is the area for storing the value when evaluated using the index of performance index 702.

[0030] For example, row 704 indicates that performance ID "P001" stores performance information related to model ID "M001". Row 704 also indicates that the prediction model for model ID "M001" was generated from the feature list corresponding to function ID "K001", and that the performance value when evaluated using the performance metric "AUC" is "0.876".

[0031] Here, we allow each model to be evaluated using multiple performance metrics. In other words, the prediction model performance information management table 115 allows for multiple rows to be stored for each model ID.

[0032] Next, an example of the information stored in the feature importance table 116, which is provided in the memory device 204 of the important biological function presentation system 101, will be explained using Figure 8. The feature importance table 116 stores the values ​​of importance ID 801, model ID 601, function ID 605, feature name 802, importance index 803, and importance 804. The column for importance ID 801 is the area that stores the ID representing the identification information of importance. The column for feature name 802 is the area that stores the name of the feature used when evaluating the importance of the feature in the feature importance calculation unit 104. The column for importance 804 is the area that stores the value when evaluated using the index of importance index 803.

[0033] For example, row 805 indicates that the importance ID "I001" stores the importance related to model ID "M001". Also, row 805 indicates that the prediction model for model ID "M001" is generated from the feature list corresponding to function ID "K001", and that the importance of the feature name "Age" in model ID "M001" when evaluated using "PI" is "0.006".

[0034] Next, an example of the processing flow of the important biological function presentation system 101 will be explained using Figure 9.

[0035] First, the Important Biological Function Presentation System 101 reads table data from the table data acquired by the Data Acquisition Unit 102 and stored in the Storage Device 204 to be used as input for the Predictive Model Generation and Evaluation Unit 103. Initially, the target variable table to be used by the Predictive Model Generation and Evaluation Unit 103 is specified by the user via the Input / Output Terminal 118, and the Important Biological Function Presentation System 101 reads the specified table data (Step S901). Next, the feature quantity table to be used by the Predictive Model Generation and Evaluation Unit 103 is specified by the user via the Input / Output Terminal 118, and the Important Biological Function Presentation System 101 reads the specified table data (Step S902). Then, the metadata table to be used by the Predictive Model Generation and Evaluation Unit 103 is specified by the user via the Input / Output Terminal 118, and the Important Biological Function Presentation System 101 reads the specified table data (Step S903).

[0036] Next, the important biological function presentation system 101 receives the table data read in steps S901, S902, and S903, and the prediction model generation and evaluation unit 103 performs prediction model construction and performance evaluation (step S904). Details of the processing in step S904 will be described later.

[0037] After step S904 is executed, the important biological function presentation system 101 calculates feature importance from the predictive model constructed in step S904 (step S905). Details of the processing in step S905 will be described later.

[0038] Finally, the important biological function presentation system 101 plots the performance information of the prediction model generated in step S904 and the information regarding the importance of the features calculated in step S905 (step S906), and outputs the plotted content to the input / output terminal 118 (step S907). Here, the plotting method in step S906 includes not only outputting table information in tabular format, but also, for example, plotting in the form of a network diagram or plotting in graph format.

[0039] Next, we will explain an example of the processing flow for step S904 in the processing flow described in Figure 9, namely the predictive model construction and performance evaluation by the predictive model generation and evaluation unit 103, using Figure 10.

[0040] First, the important biological function presentation system 101 specifies which machine learning algorithm to use to construct the prediction model (step S1001). The machine learning algorithms selectable in step S1001 are explainable machine learning algorithms. That is, the selection is limited to machine learning algorithms in which the importance of features can be calculated.

[0041] Next, the important biological function presentation system 101 sets a hyperparameter search range to search for the optimal hyperparameters of the machine learning algorithm, and the setting is read by the prediction model generation and evaluation unit 103 (step S1002). The hyperparameters for which the search range is set differ for each machine learning algorithm.

[0042] Next, the important biological function presentation system 101 sets the performance evaluation method for the prediction model, and the prediction model generation and evaluation unit 103 reads the setting (step S1003). Here, multiple indicators may be selected to be used for performance evaluation. Also, if a validation method such as k-fold cross validation is to be executed during performance evaluation, it is specified in step S1003.

[0043] The important biological function presentation system 101 generates a prediction model using the values ​​set in steps S1001, S1002, and S1003. The prediction model generates as many models as there are function IDs in the metadata table read in step S903, and determines whether there are any unprocessed function IDs (step S1004). If there are function IDs for which a prediction model has not been generated (step S1004: unprocessed), the process proceeds to step S1005. On the other hand, if there are no function IDs for which a prediction model has not been generated (step S1004: all completed), the process ends.

[0044] In step S1005, the important biological function presentation system 101 obtains the function ID for generating the prediction model and the corresponding feature data from the feature table read in step S902. For example, if the features corresponding to the function ID are "Age" and "Gene AR", the system reads the data for "Age" and "Gene AR" from the feature table.

[0045] Next, in step S1006, the important biological function presentation system 101 searches for the optimal hyperparameters for generating a predictive model using the target variable table data read in step S901 and the feature table data read in step S1005. The hyperparameters are searched from the range set in step S1002. Whether the hyperparameters are optimal is determined by the performance evaluation method set in step S1003.

[0046] After the hyperparameter search is completed in step S1006, the important biological function presentation system 101 constructs a predictive model using the optimal hyperparameters (step S1007). In addition, the performance of the predictive model constructed in step S1007 is evaluated based on the method set in step S1003 (step S1008).

[0047] Next, the important biological function presentation system 101 stores the prediction model constructed in step S1007 as a prediction model file 117 in the storage device 204 (step S1009).

[0048] The Important Biological Function Presentation System 101 outputs information necessary for management, such as the storage location and file name of the prediction model file, to the Prediction Model Management Table 114. Then, it outputs the performance evaluation results from step S1008 to the Prediction Model Performance Information Management Table 115 (step S1010). Here, high performance of the prediction model means that the feature table data read in step S1005 is strongly related to the prediction of the target variable table data read in step S901. In other words, the function to which the feature table data read in step S1005 belongs is strongly related to the prediction of the target variable representing disease onset, etc., so gene functions strongly associated with disease onset can be extracted using only the performance information of the prediction model.

[0049] Furthermore, a high contribution from explanatory variables such as genes in the prediction model indicates a strong correlation with the prediction of the target variable. In other words, genes with high importance in the feature importance table are strongly correlated with the prediction of the target variable representing disease onset, and therefore, genes strongly associated with disease onset can be extracted based solely on the contribution of the prediction model.

[0050] In this way, for example, by generating a predictive model using genes associated with gene functions, which are automatically entered based on the WGCNA analysis results or entered by the user, it becomes possible to identify important gene functions and genes from the performance and contribution of the predictive model. This allows for more efficient narrowing down of candidates.

[0051] Furthermore, according to this embodiment, as an example, it is possible to output both the importance of gene function and the importance of the gene using only the generated predictive model. This makes it possible to extract important genes without omission (without needing to further narrow the perspective) compared to analyses that refer to protein expression level information and are applicable to genes that contribute to protein expression but cannot target genes that do not express proteins.

[0052] Next, we will explain an example of the processing flow when calculating feature importance in step S905 of the processing flow described in Figure 9, i.e., in the feature importance calculation unit 104, using Figure 11.

[0053] First, the Important Biological Function Presentation System 101 specifies how to calculate the importance (step S1101). The importance calculation methods selectable in step S1101 are those that can be calculated by the machine learning algorithm of the prediction model created in step S904. Next, a threshold for the performance of the model whose importance is calculated and output to the feature importance table 116 is set, and the setting is read by the feature importance calculation unit 104 (step S1102). If the feature importance is to be calculated from all the prediction models constructed in step S904, the lower limit of the threshold that the indicator used for performance evaluation can take is specified.

[0054] The Important Biological Function Presentation System 101 calculates feature importance by reading the number of prediction model files corresponding to the number of prediction models constructed in step S904. That is, it reads the number of prediction model files corresponding to the number of function IDs present in the metadata table read in step S903 and determines whether there are any unprocessed function IDs (step S1103). If there are function IDs that have not been loaded as prediction models (step S1103: unprocessed), the process proceeds to step S1104. On the other hand, if there are no function IDs that have not been loaded as prediction models (step S1103: all completed), the process ends.

[0055] Next, in step S1104, the important biological function presentation system 101 reads the performance information of the prediction model corresponding to the function ID from the prediction model performance information management table 115. The read performance information is compared with the threshold set in step S1102. If the performance of the read prediction model is above the threshold (step S1105: Yes), the process proceeds to step S1106. If the performance of the read prediction model is below the threshold set in step S1102 (step S1105: No), the process returns to step S1103 and executes the process with the prediction model corresponding to the next function ID.

[0056] In step S1106, the important biological function presentation system 101 calculates feature importance from the prediction model based on the calculation method set in step S1101. After the processing in step S1106 is completed, the calculated feature importance is stored in the feature importance table 116 (step S1107).

[0057] According to the embodiment described above, the following information system is provided. This information system comprises a processor (for example, a host CPU 201) and a storage device 204. The storage device 204 stores an objective variable table 111 for the objective variable and a metadata table 113 for biological functions and features related to said biological functions. The processor generates a prediction model for each biological function, using the features as explanatory variables to predict the objective variable. The processor then evaluates the accuracy of the prediction model and extracts the biological functions related to the objective variable based on the evaluation results of the prediction model.

[0058] According to this, by identifying highly accurate models among predictive models based on features related to biological functions, it is possible to extract biological functions that are highly associated with disease.

[0059] Furthermore, the processor may extract characteristic features of biological functions in conjunction with the extraction of biological functions.

[0060] According to this method, related features can also be extracted.

[0061] The processor may also calculate the importance of the features of the prediction model and extract the features of the biological functions in conjunction with the extraction of biological functions.

[0062] According to this method, it is possible to identify features that influence the target variable from among numerous features related to biological functions.

[0063] Furthermore, the metadata table 113 may be generated by a co-expression network.

[0064] According to this method, it is possible to identify features related to biological functions and to automatically input combinations of biological functions and features related to those biological functions.

[0065] Furthermore, the metadata table 113 may allow users to add biological functions and features related to those biological functions.

[0066] According to this, the user can manually edit the contents of the metadata table 113.

[0067] Furthermore, the metadata table 113 may not only be automatically generated (not only by a co-expression network such as WGCNA), but may also allow users to add biological functions and features related to those biological functions.

[0068] According to this, in addition to automatically generated content, users can, for example, manually add features and other elements, taking into account the influence of other factors.

[0069] Furthermore, as an example, biological functions can be gene functions, and characteristic features can be genes. Also, as an example, biological functions can be protein functions, and characteristic features can be proteins.

[0070] Although embodiments have been described above, the present invention is not limited to the embodiments described above, and includes various modifications and equivalent configurations within the spirit of the attached claims. For example, the embodiments described above are described in detail for the purpose of explaining the present invention in an easy-to-understand manner, and the present invention is not necessarily limited to having all the configurations described. Also, for example, some of the configurations of the embodiments may be added, deleted, or replaced with other configurations.

[0071] Although an example using the host CPU 201 as the processor was described, any device that performs the specified processing will suffice, and other semiconductor devices may be used.

[0072] The input / output terminal 118 can be configured, for example, as a suitable computer device. The input / output terminal 118 may be configured to include, for example, a processor, a storage device, a communication device, an input device, and an output device.

[0073] The important biological function display system 101 and the input / output terminal 118 may be connected by wire or wirelessly. Furthermore, the important biological function display system 101 may be connected to a single input / output terminal 118 or to multiple input / output terminals 118.

[0074] Communication between the important biological function display system 101 and the input / output terminal 118 can use, for example, a fifth-generation mobile communication system, so-called 5G (5th Generation), which enables "multiple simultaneous connections" and "ultra-low latency." By taking advantage of the features of new systems such as 5G and later, even when, for example, many input / output terminals 118 are connected simultaneously and the amount of data communication increases, the amount of data stored in the storage device 204 becomes large, and the amount of data communication between the input / output terminal 118 and the storage device becomes large, communication delay can be suppressed.

[0075] The important biological function display system 101 may be connected to a remote input / output terminal 118 and located in the cloud. Furthermore, a newer system, such as 5G or later, may be used for wireless communication between the important biological function display system 101 in the cloud and the input / output terminal 118. [Explanation of symbols]

[0076] 101 Important Biological Function Presentation System 102 Data Acquisition Unit 103 Predictive Model Generation and Evaluation Department 104 Feature Importance Calculation Unit 105 Result drawing section 106 data 107 Gene expression data 108 Protein expression data 109 Patient Information Data 110 Metadata 111 Dependent Variable Table 112 Feature Table 113 Metadata Table 114 Predictive Model Management Table 115 Predictive Model Performance Information Management Table 116 Feature Importance Table 117 Predictive Model Files 118 Input / Output Terminals 201 CPU 202 memory 203 Peripheral IF 204 Storage device 205 Communication IF Bus 206

Claims

1. Processor and Memory device and Equipped with, The aforementioned storage device is It stores a table of target variables for a target variable representing disease progression, a target variable representing disease severity, and / or a target variable representing disease onset, and a metadata table for biological functions and features related to said biological functions. The aforementioned processor, A predictive model is generated that predicts the target variable using the feature quantities as explanatory variables for each of the aforementioned biological functions. The processor calculates the performance value of the prediction model by evaluating the accuracy of the prediction model using a set index. The processor extracts the biological functions associated with the predictive model whose performance value is higher than a predetermined value as the biological functions associated with the target variable. An information system characterized by the following features.

2. The information system according to claim 1, The aforementioned processor, Along with extracting the aforementioned biological functions, characteristic quantities of the aforementioned biological functions are extracted. An information system characterized by the following features.

3. The information system according to claim 1, The aforementioned processor, Using a calculation method that calculates a value relating to the importance of the set feature quantities, the aforementioned value for the feature quantities of the prediction model is calculated from the prediction model whose performance value is equal to or greater than the set threshold. Along with extracting the aforementioned biological functions, characteristic quantities of the aforementioned biological functions are extracted, The aforementioned processor, Using the aforementioned values, we identify the features that influence the target variable from the features related to the biological function. An information system characterized by the following features.

4. The information system according to claim 1, The aforementioned metadata table is generated by the co-expression network. An information system characterized by the following features.

5. The information system according to claim 1, The aforementioned metadata table allows users to add biological functions and features related to those biological functions. An information system characterized by the following features.

6. The information system according to claim 4, The aforementioned metadata table allows users to add biological functions and features related to those biological functions. An information system characterized by the following features.

7. The information system according to claim 1, The aforementioned biological function is a gene function, The aforementioned feature is a gene. An information system characterized by the following features.

8. The information system according to claim 1, The aforementioned biological function is a protein function. The aforementioned feature is a protein. An information system characterized by the following features.

9. The information system according to claim 1, The aforementioned information system is It is connected to an input / output terminal that performs data input and output, A fifth-generation mobile communication system is used for communication with the aforementioned input / output terminal. An information system characterized by the following features.

10. The information system according to claim 1, The aforementioned information system is It is connected to an input / output terminal that performs data input and output, and is located in the cloud. An information system characterized by the following features.

Citation Information

Patent Citations

  • Medical analysis system

    JP2011520206A

  • Parallel processing apparatus, parallel processing method, and parallelization processing program

    JP2018077547A

  • Medical analysis system

    US20200395128A1