Apparatus and method for predicting homologous recombination deficiency

The machine learning-based HRD prediction device addresses the limitations of existing methods by using transcriptome data from multiple cancer types to accurately predict HRD, enabling personalized treatment and reducing analytical costs and complexity.

WO2025135480A1PCT designated stage expired Publication Date: 2025-06-26SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/017105
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-11-04
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing methods for predicting homologous recombination deficiency (HRD) in cancer cells require extensive patient-derived materials and genetic experiments, limiting their applicability and accuracy, especially when applied to individual patient data or different cancer types.

Method used

A machine learning-based prediction device and method that utilizes transcriptome data from multiple cancer types to predict HRD, enabling individual sample gene-set analysis and reducing the time and cost associated with traditional sequencing methods.

Benefits of technology

The method effectively identifies key genes associated with HRD, allowing for accurate prediction of HRD in various cancer types, thereby facilitating personalized treatment approaches and reducing the complexity and cost of genetic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024017105_26062025_PF_FP_ABST
    Figure KR2024017105_26062025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus for predicting homologous recombination deficiency (HRD) using artificial intelligence, according to the present invention, comprises: a data collection unit that collects transcript data, gene expression level data, and homologous recombination deficiency score data of various types of cancer patients; a first candidate gene selection unit; a first prediction model generation unit; a second candidate gene selection unit; a third candidate gene selection unit; a second prediction model generation unit; and a homologous recombination deficiency prediction value calculation unit. Accordingly, homologous recombination deficiency for individual samples can be more accurately predicted.
Need to check novelty before this filing date? Find Prior Art

Description

Homologous recombination deficiency prediction device and method thereof

[0001] The present invention relates to a technique for predicting homologous recombination deficiency, and more particularly, to a technique for predicting homologous recombination deficiency based on machine learning.

[0002] Homologous recombination deficiency (HRD) is a phenomenon caused by mutations in genes that constitute the chromosome repair mechanism and the resulting problems with homologous recombination function, and has been observed in various cancer cells. Genes that directly cause homologous recombination deficiency include BRCA1 and BRCA2, and mutations in these genes have been observed in certain cancers. Studies have reported that the expression of solid cancers such as breast cancer, ovarian cancer, and pancreatic cancer is correlated with the phenotype of homologous recombination deficiency. In particular, cancer patients with the phenotype of homologous recombination deficiency show a high survival rate and sensitivity to PARP target inhibitors. Therefore, active research is being conducted to predict the phenotype and use it for precise diagnosis and personalized medicine.

[0003] Previous studies attempting to predict such homologous recombination deficiency involved whole genome sequencing and whole exome sequencing based on next-generation sequencing of specific cancer cells. In addition, a study was reported that reduced the cost and time required for whole genome analysis by applying a single nucleotide polymorphism array. The analysis using single nucleotide mutations applied the concept of 'HRD score' to interpret homologous recombination deficiency by dividing it into three types of detailed phenotypic traits, which are composed of loss of heterozygosity (LOH), telomeric allelic imbalance (TAI), and large-scale transitions (LST), respectively.

[0004] In the case of the general homologous recombination deficiency prediction method (next-generation sequencing-based whole-genome / single-nucleotide mutation array analysis method), there is a problem that various patient-derived materials (cancer cells and additional blood cells for genomic mutation control) and additional genomic experiments are required to predict homologous recombination deficiency.

[0005] To address this, a method using multiple transcriptomes (mRNAs) rather than the entire genome is based on the expression profiles (genomic signatures) of 230 mRNAs. However, because this method predicts homologous recombination deficiency using markers derived from gene expression levels in a limited sample (breast cancer, some cell lines), it is only applicable to the same cancer type and has limitations, such as low predictive rates when applied to datasets other than the cohort used.

[0006] Additionally, these methods are commonly optimized for cohort-based data applications, making them difficult to apply to individual patient data.

[0007] The inventors of the present invention have been researching and working to solve the problems of such conventional homologous recombination deficiency prediction techniques. Through machine learning utilizing big data derived from multiple cancers, the inventors of the present invention have successfully completed a homologous recombination deficiency prediction device and method that predicts homologous recombination deficiency using multiple transcriptome (mRNA) expression patterns (genomic signatures), and enables prediction of homologous recombination deficiency for individual patients through machine learning-based single sample gene-set analysis.

[0008]

[0009] [National Research and Development Project Supporting This Invention]

[0010] [Project ID]1465039638

[0011] [Assignment Number] HI22C0645000023

[0012] [Ministry Name] Ministry of Health and Welfare

[0013] [Name of Project Management (Specialist) Institution] Korea Health Industry Development Institute

[0014] [Research Project Name] Healthcare Talent Development Support Project (National Health Promotion Fund)

[0015] [Research Project Name] Korean Ovarian Cancer Anti-Cancer Based on Homologous Recombination Deficiency Multifactorial Analysis

[0016] Development of response prediction technology

[0017] [Contribution rate] 1 / 6

[0018] [Name of the project performing organization] Seoul National University Industry-Academic Cooperation Foundation

[0019] [Research Period] January 1, 2023 - December 31, 2023

[0020]

[0021] [National Research and Development Project Supporting This Invention]

[0022] [Project ID]1711194244

[0023] [Assignment Number] KD000153

[0024] [Ministry Name] Ministry of Science and ICT

[0025] [Name of Project Management (Specialist) Institution] (Foundation) Inter-Ministry Full-cycle Medical Device Research and Development Project Group

[0026] [Research Project Name] Inter-Ministry Full-cycle Medical Device Research and Development (Ministry of Science and Technology)

[0027] [Research Project Name] Anticancer Drug Sensitivity Through Multifactor Analysis of Ovarian Cancer Organoids

[0028] Diagnostic kit development

[0029] [Contribution rate] 1 / 6

[0030] [Name of Project Performing Organization] MBD Co., Ltd.

[0031] [Research Period] January 1, 2023 - December 31, 2023

[0032]

[0033] [National Research and Development Project Supporting This Invention]

[0034] [Project ID]1711194085

[0035] [Assignment Number] 00208519

[0036] [Ministry Name] Ministry of Science and ICT

[0037] [Name of Project Management (Specialist) Institution] National Research Foundation of Korea

[0038] [Research Project Name] Individual Basic Research (Ministry of Science and ICT)

[0039] [Research Project Name] Development of Tumor Microenvironment Heating Technology Based on Tumor-Immune Mimics

[0040] [Contribution rate] 1 / 6

[0041] [Name of the project performing organization] Seoul National University Industry-Academic Cooperation Foundation

[0042] [Research Period] March 1, 2023 - February 29, 2024

[0043]

[0044] [National Research and Development Project Supporting This Invention]

[0045] [Project ID]1711184184

[0046] [Assignment Number] 2021R1F1A1062369

[0047] [Ministry Name] Ministry of Science and ICT

[0048] [Name of Project Management (Specialist) Institution] National Research Foundation of Korea

[0049] [Research Project Name] Individual Basic Research (Ministry of Science and ICT)

[0050] [Research Project Name] Development of Anticancer Immuno-Intensifying Technology Based on a Tumor Microenvironment Simulation Model

[0051] [Contribution rate] 1 / 6

[0052] [Name of Project Performing Organization] Bundang Seoul National University Hospital

[0053] [Research Period] July 1, 2023 - June 30, 2024

[0054]

[0055] [National Research and Development Project Supporting This Invention]

[0056] [Project ID]1711200033

[0057] [Assignment Number] 00272547

[0058] [Ministry Name] Ministry of Science and ICT

[0059] [Name of Project Management (Specialist) Institution] National Research Foundation of Korea

[0060] [Research Project Name] Individual Basic Research (Ministry of Science and ICT)

[0061] [Research Project Title] Development of a Multipurpose Platform for Analysis of Tumor-Immunotherapy Responses

[0062] [Contribution rate] 1 / 6

[0063] [Name of the project performing organization] Seoul National University Industry-Academic Cooperation Foundation

[0064] [Research Period] September 1, 2023 - February 29, 2024

[0065]

[0066] [National Research and Development Project Supporting This Invention]

[0067] [Project ID]2460000153

[0068] [Assignment Number] RS-2022-KH129481

[0069] [Ministry Name] Ministry of Health and Welfare

[0070] [Name of Project Management (Specialist) Institution] Korea Health Industry Development Institute

[0071] [Research Project Name] Research-Oriented Hospital Development (R&D)

[0072] [Research Project Name] Establishment of a Multi-Omics and Organoid-Based Gene Therapy Efficacy Evaluation System

[0073] [Contribution rate] 1 / 6

[0074] [Name of Project Performing Organization] Bundang Seoul National University Hospital

[0075] [Research Period] July 1, 2022 - December 31, 2030

[0076]

[0077] The purpose of the present invention is to provide an artificial intelligence-based homologous recombination deficiency prediction device and method.

[0078] In addition, another purpose of the present invention is to enable prediction of homologous recombination deficiency for individual samples by applying machine learning using big data to multiple cancer types.

[0079] Additionally, it aims to reduce the time and cost required to predict homologous recombination deficiency in biological / chemical samples in the laboratory by using computer-based artificial intelligence.

[0080] Meanwhile, other unspecified purposes of the present invention will be additionally considered within a range that can be easily inferred from the detailed description and effects thereof below.

[0081] The homologous recombination deficiency prediction device according to the present invention is:

[0082] The present invention is characterized by comprising: a data collection unit that collects transcriptome data, gene expression level data, and homologous recombination deficiency (HRD) score data of various types of cancer patients; a first candidate gene selection unit that selects a first candidate gene from the collected data using a differentially expressed gene (DEG) analysis technique; a first prediction model generation unit that trains a machine learning model using the expression level data, transcriptome data, and HRD score data of the candidate gene among the collected data as learning data to generate a first prediction model for predicting key genes; a second candidate gene selection unit that selects a second candidate gene through statistical analysis based on gene importance for the first candidate gene; a third candidate gene selection unit that selects a third candidate gene, which is a final biomarker gene, among the second candidate genes; a second prediction model generation unit that generates a second prediction model trained by the third candidate gene; and a homologous recombination deficiency prediction value calculation unit that calculates a prediction value for homologous recombination deficiency by applying the second prediction model to the third candidate gene selected by the third candidate gene selection unit among an individual sample gene group.

[0083] The third candidate gene selection unit is characterized in that it randomly selects the second candidate genes and performs bootstrap analysis to select the third candidate genes.

[0084] The third candidate gene selected above is characterized by having an influence value of positive or negative depending on its importance in homologous recombination deficiency.

[0085] The transcriptome data of the above-mentioned various types of cancer patients is characterized by being data standardized to TPM (transcripts per million).

[0086] The above homologous recombination deficiency prediction value calculation unit is characterized in that it performs individual sample gene group analysis by the ssGSEA (single sample Gene set Enrichment Analysis) method.

[0087] A method for predicting homologous recombination deficiency according to another embodiment of the present invention,

[0088] The method is characterized by comprising the steps of collecting transcriptome data, gene expression level data, and homologous recombination deficiency (HRD) score data of various types of cancer patients; selecting a first candidate gene from the collected data using a differentially expressed gene (DEG) analysis technique; training a machine learning model using the expression level data, transcriptome data, and HRD score data of the candidate gene among the collected data as learning data to generate a first prediction model for predicting key genes; selecting a second candidate gene through statistical analysis based on gene importance for the first candidate gene; selecting a third candidate gene, which is a final biomarker gene, among the second candidate genes; and calculating a prediction value for homologous recombination deficiency by applying the second prediction model to the selected third candidate gene among a group of individual sample genes.

[0089] The step of selecting the third candidate gene is characterized in that the second candidate gene is randomly selected and a bootstrap analysis is performed to select the third candidate gene.

[0090] The third candidate gene selected above is characterized by having an influence value of positive or negative depending on its importance in homologous recombination deficiency.

[0091] The transcriptome data of the above-mentioned various types of cancer patients is characterized by being data standardized to TPM (transcripts per million).

[0092] The step of calculating a prediction value for homologous recombination deficiency is characterized by performing individual sample gene set analysis by the ssGSEA (single sample Gene set Enrichment Analysis) method.

[0093] According to the present invention, there is an effect of being able to derive major genes related to homologous recombination deficiency through machine learning based on transcriptome big data acquired from various cancer types.

[0094] In addition, by using the derived genes to create a machine learning-based homologous recombination deficiency judgment model, there is an advantage in being able to judge the homologous recombination deficiency association with various cancer types with high accuracy.

[0095] Additionally, it has the effect of enabling prediction of homologous recombination deficiency for each patient through a machine learning-based individual sample gene group analysis process.

[0096] Meanwhile, even if the effect is not explicitly mentioned herein, it is added that the effect and its provisional effect described in the following specification expected by the technical features of the present invention are treated as described in the specification of the present invention.

[0097] Figure 1 is a schematic structural diagram of a homologous recombination deficiency prediction device according to a preferred embodiment of the present invention.

[0098] FIG. 2 shows the correlation between a homologous recombination deficiency prediction value predicted by a homologous recombination deficiency prediction device according to a preferred embodiment of the present invention and an actual homologous recombination deficiency value.

[0099] Figure 3 is a schematic flowchart of a method for predicting homologous recombination deficiency according to another preferred embodiment of the present invention.

[0100] FIG. 4 is a schematic flowchart of a bootstrap analysis method for predicting homologous recombination deficiency according to another preferred embodiment of the present invention.

[0101] FIG. 5 illustrates an example of a computer program that implements a method for predicting homologous recombination deficiency according to another preferred embodiment of the present invention.

[0102] ※ It is to be noted that the attached drawings are provided for reference only to help understand the technical concept of the present invention, and the scope of the rights of the present invention is not limited thereby.

[0103] Hereinafter, with reference to the drawings, the configuration of the present invention, guided by various embodiments thereof, and the effects resulting from such configurations will be examined. In describing the present invention, detailed descriptions of related, well-known functions that are obvious to those skilled in the art and that may unnecessarily obscure the gist of the present invention will be omitted.

[0104] Terms such as "first" and "second" may be used to describe various components, but the components should not be limited by these terms. These terms may only be used to distinguish one component from another. For example, without departing from the scope of the present invention, a "first component" may be referred to as a "second component," and similarly, a "second component" may also be referred to as a "first component." Furthermore, singular expressions include plural expressions unless the context clearly dictates otherwise. Terms used in the embodiments of the present invention may be interpreted as having meanings commonly known to those of ordinary skill in the art, unless otherwise defined.

[0105] Hereinafter, with reference to the drawings, the configuration of the present invention guided by various embodiments of the present invention and the effects resulting from the configuration will be examined.

[0106]

[0107] Figure 1 is a schematic structural diagram of a homologous recombination deficiency prediction device using artificial intelligence according to a preferred embodiment of the present invention.

[0108] The homologous recombination deficiency prediction device (100) according to the present invention may include a control unit (not shown) including one or more processors and memory.

[0109] The control unit may include a data collection unit (110), a first candidate gene selection unit (120), a first prediction model generation unit (130), a second candidate gene selection unit (140), a third candidate gene selection unit (150), a second prediction model generation unit (160), and a homologous recombination deficiency prediction value generation unit (170).

[0110] The memory of the control unit stores various information necessary for the operation of the operating homologous recombination deficiency prediction device (100). The information stored in the memory may include, but is not limited to, gene-related data received from the data collection unit (110), information for the control operation of the control unit, information processed and analyzed by the control unit, and program information related to the homologous recombination deficiency prediction method described below.

[0111] For example, the memory may include, but is not limited to, a hard disk type, a magnetic media type, a compact disc read only memory (CD-ROM), an optical media type, a magneto-optical media type, a multimedia card micro type, a flash memory type, a read only memory type, or a random access memory type, depending on its type. In addition, the memory may include, but is not limited to, a cache, a buffer, a main memory, an auxiliary memory, or a separately provided storage system depending on its use / location.

[0112] The control unit can perform various control operations of the homologous recombination deficiency prediction device (100). That is, the control unit can perform control on the homologous recombination deficiency prediction device (100), control analysis of genetic information processed by the homologous recombination deficiency prediction device (100), etc., and control execution of the homologous recombination deficiency prediction method described below. In addition, the control unit can control the operations of the remaining components of the homologous recombination deficiency prediction device (100). For example, the control unit may include, but is not limited to, a processor, which is hardware, or a process, which is software executed by the processor.

[0113] The data collection unit (110) collects transcriptome (mRNA) data, gene expression level data, and homologous recombination deficiency (HRD) score data of various types of cancer patients. The transcriptome data may be transcriptome sequence data.

[0114] For example, the data collection unit (110) can select cancer types for which transcriptome data is reported in the Pan-Cancer database and use them to train a machine learning model.

[0115] To this end, the data collection unit (110) can receive necessary data from the database using wired / wireless communication technologies. For example, the wired / wireless communication technologies may perform wireless communication such as 5G (5th generation communication), LTE-A (long term evolution-advanced), LTE (long term evolution), Bluetooth, BLE (Bluetooth low energy), NFC (near field communication), and WiFi communication, or may use wired communication technologies such as cable communication, but are not limited thereto.

[0116] The data collection unit (110) can also receive single nucleotide mutation array data and calculate an HRD score based on the data. Calculating an HRD score using single nucleotide mutation data is a well-known technique, and a detailed description thereof is omitted here.

[0117] The data collection unit (110) can normalize the collected transcriptome data and gene expression data to standardize transcript expression levels according to various cancer types. For example, the transcriptome data of all cancer types can be grouped together and standardized to TPM (transcripts per million) data, or the data can be normalized by converting the data into a normal distribution by taking the log of each gene expression data. In addition, various normalization methods such as RPKM (reads per kilobase per millions mapped reads) and TMM (trimmed means of M values) can be used.

[0118] The first candidate gene selection unit (120) uses the normalized data collected by the data collection unit (110) to select first candidate genes likely to be associated with homologous recombination deficiency. For example, differentially expressed gene (DEG) analysis techniques may be used for gene selection. For example, among the training samples, the distribution of HRD scores is normal, and 4,436 genes associated with homologous recombination deficiency can be initially selected by performing DEG (Diferentially expressed gene) analysis on the training sample of an ovarian cancer patient with the highest average.

[0119] The first prediction model generation unit (130) generates and trains the first prediction model, which is a homologous recombination deficiency prediction model, using the gene expression level data, transcript data, and HRD score of the gene sample selected by the first candidate gene selection unit (120).

[0120] The variable used to train the homologous recombination deficiency prediction model is the normalized gene expression level, and the predicted value of the prediction model is the HRD score calculated for each patient.

[0121] The data acquired from the data collection unit (110) and selected by the first candidate gene selection unit (120) are divided into a certain ratio and used for learning and verifying the first prediction model. For example, 80% of the selected data are used for learning the first prediction model and the remaining 20% ​​are used for verifying the first prediction model, thereby selecting an optimal machine learning model. For example, among 10,068 samples collected from patients with various cancers, 80%, or 8,041 samples (20,502 genes), can be used for learning, and 20%, or 2,027 samples (2,538 genes), can be used for verification.

[0122] Machine learning algorithms used in the first prediction model generation unit (130) may include, but are not limited to, algorithms such as elasticnet regression, lasso regression, ridge regression, gradient boosting, support vector machine, and multi-layer perceptron.

[0123] The first prediction model generation unit (130) trains a machine learning model using the expression level data and transcript data of the first candidate gene as input and the HRD score as the correct answer, and evaluates the machine learning model using verification data to adjust the hyper parameters of the machine learning model to generate a first prediction model that predicts the association between homologous recombination deficiency and genes.

[0124] The first prediction model, which is a machine learning model generated by the first prediction model generation unit (130), calculates the weights of data used for learning, and based on the results, can identify important genes highly correlated with homologous recombination deficiency prediction.

[0125] The second candidate gene selection unit (140) is used to select genes that have a great influence on the construction of a homologous recombination deficiency prediction model by applying the genes selected by the first candidate gene selection unit (120) to the first prediction model, which is a machine learning model generated by the first prediction model generation unit (130).

[0126] The third candidate gene selection unit (150) selects the final gene, which is a biomarker, by performing a bootstrap analysis on the genes selected by the second candidate gene selection unit (140).

[0127] After fixing the parameters of the model generated by the first prediction model generation unit (130), the training samples used for model construction are randomly selected and bootstrap analysis is performed. For example, among the 4,436 genes selected by the first candidate gene selection unit (120), statistical analysis based on the gene importance obtained at each random step is performed, ultimately securing 356 genes. The third candidate gene selection unit (150) may use a bootstrap analysis method, but the analysis method is not limited thereto.

[0128] The second prediction model generation unit (160) generates a second prediction model based on the final biomarker gene data selected by the third candidate gene selection unit (140).

[0129] That is, a new machine learning model can be trained using the genetic data selected by the third candidate gene selection unit (150) to select an optimal learning model. Statistical verification can be used to select an optimal learning model. For each machine learning model, the R2 score and RMSE (Root Mean Square Error) can be calculated, and the model with the best performance, i.e., an R2 close to 1 and an RMSE close to 0, can be selected as the second prediction model.

[0130] For example, when Elasticnet, Lasso, Ridge, Gradient boosting (GBM), and Support vector machine (SVR) were used as machine learning models and Multi-layered perceptron (MLP) was used as a deep learning model, the R2 value and RMSE value of each model are as shown in Table 1 below.

[0131]

[0132] The prediction value generating unit (170) generates HRD prediction values ​​for individual patient genes using the second prediction model, which is the final selected machine learning model.

[0133] That is, the final biomarker genes selected by the third candidate gene selection unit (150) are selected as a group of genes for gene group analysis of individual samples. The selected genes have a degree of influence on homologous recombination deficiency, which is indicated by a positive or negative value, depending on their importance. The group of genes used for analysis is divided into positive and negative groups according to the value and calculated.

[0134] After correction for the calculated value, the final value is calculated as a ssGSEA (single sample Gene set Enrichment Analysis) score.

[0135] Figure 2 shows the correlation between the ssGSEA score, which is the result of the analysis, and the HRD score obtained using a single nucleotide sequence array.

[0136] The results obtained through Pearson correlation analysis demonstrate a high positive correlation and statistical significance. Specifically, the correlation coefficient (PCC) is 0.768, and the p-value for the correlation is 2.05E-12.

[0137] The prediction value generating unit (170) constructs a linear regression model based on the two values ​​derived in this way, and calculates the prediction value for the linear regression model as a homologous recombination deficiency prediction value (expHRD).

[0138]

[0139] Figure 3 is a schematic flowchart of a method for predicting homologous recombination deficiency according to another preferred embodiment of the present invention.

[0140] The method for predicting homologous recombination deficiency according to the present invention can be performed by a control unit including one or more processors and memory.

[0141] Data for generating a prediction model to predict homologous recombination deficiency is collected (S110).

[0142] Transcriptome (mRNA), gene expression, and homologous recombination deficiency (HRD) score data from various cancer patients are collected. The transcriptome data may be transcriptome sequence data. For example, transcriptome data from the Pan-Cancer database can be collected to train a machine learning model by selecting cancer types for which transcriptome data has been reported.

[0143] Data collection can be accomplished by receiving the required data from a database using wired or wireless communication technologies.

[0144] Additionally, single nucleotide variant array data can be received and HRD scores can be calculated based on this.

[0145] By normalizing the transcriptome and gene expression data collected during the data collection phase, transcriptome expression levels across various cancer types can be standardized. For example, data can be normalized by pooling all cancer transcriptome data into one and normalizing it to TPM (transcripts per million), or by taking the logarithm of each gene expression data and transforming it into a normal distribution. In addition, various normalization methods can be used, such as RPKM (reads per kilobase per millions mapped reads) and TMM (trimmed means of M values).

[0146] Among the data collected in this way, the first candidate gene related to homologous recombination deficiency is selected (S120).

[0147] For example, differentially expressed gene (DEG) analysis techniques can be used for genetic screening.

[0148] In the first prediction model generation step, a machine learning model is generated and trained using the selected first candidate genes (S130).

[0149] In the first candidate gene selection step, the gene expression data, transcriptome data, and HRD scores of the selected gene samples are used to generate and train the first prediction model, which is a homologous recombination deficiency prediction model. The variable used for training the homologous recombination deficiency prediction model is the normalized gene expression level, and the predicted value of the prediction model is the HRD score calculated for each patient. The selected data are divided into a certain ratio and used for training and validation of the first prediction model. For example, 80% is used for training and the remaining 20% ​​for validation, allowing the selection of the optimal machine learning model.

[0150] Machine learning algorithms used in the first prediction model creation step include, but are not limited to, elasticnet regression, lasso regression, ridge regression, gradient boosting, support vector machine, and multi-layer perceptron.

[0151] The first prediction model is trained using the expression level data and transcriptome data of the first candidate gene as input and the HRD score as the correct answer, and is evaluated using the validation data to adjust the hyper parameters of the machine learning model to predict the association between homologous recombination deficiency and the gene.

[0152] Once the first prediction model is generated, it is used to select a second candidate gene that has a great influence on the construction of a homologous recombination deficiency prediction model (S140).

[0153] Once the second candidate gene is selected, a third candidate gene to be used as the final biomarker is selected through bootstrap analysis (S150).

[0154] Figure 4 is a schematic flowchart of a bootstrap analysis method according to the present invention.

[0155] Among the genes selected as the second candidate genes, data is randomly selected (S151), and the machine learning result is generated by using the selected random data as input to the first prediction model (S152).

[0156] The importance of genes is evaluated based on the machine learning results (S153), and steps S151 to S153 are repeated to randomly select data and generate machine learning results to evaluate gene importance.

[0157] In this iterative process, the results of the gene importance evaluation are continuously saved (S154), and the results are subjected to statistical analysis (S155) to select the third candidate gene, which is the final candidate gene (S156).

[0158] Returning to Figure 4, a second prediction model is generated using the selected third candidate gene (S160).

[0159] The optimal learning model is selected by training a new machine learning model using the third candidate gene selected as the final biomarker gene. As previously discussed, the R2 score and RMSE values ​​are used to select the optimal learning model.

[0160] Finally, the selected final biomarker genes are used to generate a prediction value for homologous recombination deficiency of individual samples (S170).

[0161] The final biomarker genes selected are then used as a set of genes for the analysis of individual sample genes. The selected genes are assigned a positive or negative value, depending on their significance, indicating their influence on homologous recombination deficiency. The gene sets used in the analysis are then divided into positive and negative groups based on these values. After corrections are made to the calculated values, the final values ​​are calculated as ssGSEA (single sample Gene set Enrichment Analysis) scores. Using these ssGSEA values ​​and HRD scores, a linear regression model is constructed, and the predicted values ​​for individual sample genes using the linear regression model are calculated as the homologous recombination deficiency predicted value (expHRD).

[0162] The method for predicting homologous recombination deficiency according to the present invention can be implemented as computer-readable code on a computer-readable recording medium. The computer-readable recording medium can include any type of recording device that stores data that can be read by a computer system. Examples of the computer-readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical disk, HDD, SSD, USB memory, etc. In addition, the computer-readable recording medium can be distributed across network-connected computer systems or cloud environments, and can be written and executed as computer-readable code in a distributed manner.

[0163] FIG. 5 illustrates an embodiment of a method for predicting homologous recombination deficiency according to the present invention implemented as a computer code.

[0164] By inputting genetic data of individual samples and predicting the degree of homologous recombination deficiency using the homologous recombination deficiency prediction method according to the present invention, homologous recombination deficiency can be predicted in a user-friendly manner without complex genetic analysis.

[0165]

[0166] According to the above-described device and method for predicting homologous recombination deficiency using artificial intelligence of the present invention, genes related to the homologous recombination deficiency mechanism can be used as biomarkers to diagnose patients with abnormalities in the relevant function early, and homologous recombination deficiency traits of specific cells, such as stem cells and organoids, for which whole-genome or single-base sequencing is difficult, can be easily identified, thereby enabling personalized treatment utilizing PARP inhibitor sensitivity. In addition, by providing a computer-based, user-friendly homologous recombination deficiency prediction method, the time and cost required for genetic analysis in a laboratory environment can be significantly reduced.

[0167]

[0168] The scope of protection of the present invention is not limited to the description and expression of the embodiments explicitly described above. Furthermore, it should be noted that the scope of protection of the present invention may not be limited by obvious modifications or substitutions within the technical field to which the present invention pertains.

Claims

1. A homologous recombination deficiency prediction device comprising a control unit including one or more processors and a memory, wherein the control unit: A data collection unit that collects transcriptome data, gene expression level data, and homologous recombination deficiency (HRD) score data from various types of cancer patients; A first candidate gene selection unit for selecting a first candidate gene from the collected data using a differentially expressed gene (DEG) analysis technique; A first prediction model generation unit that uses the expression level data, transcriptome data, and HRD score data of the candidate gene among the collected data as learning data to train a machine learning model to generate a first prediction model for predicting key genes; A second candidate gene selection unit for selecting a second candidate gene through statistical analysis based on gene importance for the first candidate gene; A third candidate gene selection unit for selecting a third candidate gene, which is a final biomarker gene, from among the second candidate genes; A second prediction model generation unit for generating a second prediction model learned by the third candidate gene; and A homologous recombination deficiency prediction value calculation unit that calculates a prediction value for homologous recombination deficiency by applying the second prediction model to the third candidate gene selected by the third candidate gene selection unit among the individual sample gene groups; A homologous recombination deficiency prediction device, characterized by including a.

2. In paragraph 1, A homologous recombination deficiency prediction device, characterized in that the third candidate gene selection unit randomly selects the second candidate genes and performs bootstrap analysis to select the third candidate gene.

3. In paragraph 1, A homologous recombination deficiency prediction device, characterized in that the above-mentioned selected third candidate genes have influence values ​​of positive and negative values ​​according to their importance in homologous recombination deficiency.

4. In paragraph 1, A homologous recombination deficiency prediction device, characterized in that the transcriptome data of the various types of cancer patients is data normalized to TPM (transcripts per million).

5. In paragraph 1, A homologous recombination deficiency prediction device, characterized in that the above homologous recombination deficiency prediction value generating unit performs individual sample gene group analysis by the ssGSEA (single sample Gene set Enrichment Analysis) method.

6. A method for predicting homologous recombination deficiency performed by a control unit including one or more processors and memories: A step of collecting transcriptome data, gene expression level data and homologous recombination deficiency (HRD) score data of various types of cancer patients; A step of selecting a first candidate gene from the collected data using a differentially expressed gene (DEG) analysis technique; A step of generating a first prediction model for predicting key genes by training a machine learning model using the expression level data, transcriptome data, and HRD score data of the candidate gene among the collected data as learning data; A step of selecting a second candidate gene through statistical analysis based on gene importance for the first candidate gene; A step of selecting a third candidate gene, which is a final biomarker gene, among the second candidate genes; and A step of calculating a prediction value for homologous recombination deficiency by applying the second prediction model to the third candidate gene selected from the individual sample gene group; A method for predicting homologous recombination deficiency, characterized by including a.

7. In paragraph 6, A method for predicting homologous recombination deficiency, characterized in that the step of selecting the third candidate gene comprises randomly selecting the second candidate genes and performing bootstrap analysis to select the third candidate gene.

8. In paragraph 6, A method for predicting homologous recombination deficiency, characterized in that the selected third candidate genes have influence values ​​of positive and negative depending on their importance in homologous recombination deficiency.

9. In paragraph 6, A method for predicting homologous recombination deficiency, characterized in that the transcriptome data of the various types of cancer patients is standardized to TPM (transcripts per million).

10. In paragraph 6, A method for predicting homologous recombination deficiency, characterized in that the step of calculating a prediction value for homologous recombination deficiency comprises performing individual sample gene set analysis by the ssGSEA (single sample Gene set Enrichment Analysis) method.

Citation Information

Patent Citations

  • Ultrasonic cleaning device and ultrasonic cleaning method

    KR1020250032145A