BRCA1 / 2 Ovarian Cancer Intelligent Prevention System Based on Multi-omics Data

By constructing a BRCA1/2 ovarian cancer intelligent prevention system based on multi-omics data, the health damage problem of traditional prevention methods has been solved, and early warning and precise prevention of ovarian cancer have been achieved. This has improved the accuracy of risk prediction and quality of life, forming a full-process intelligent prevention system.

CN122090952APending Publication Date: 2026-05-26THE OBSTETRICS & GYNECOLOGY HOSPITAL OF FUDAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE OBSTETRICS & GYNECOLOGY HOSPITAL OF FUDAN UNIV
Filing Date
2026-02-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve early warning and precise prevention of ovarian cancer in BRCA1/2 mutation carriers through multi-omics data integration. Traditional preventive measures, such as prophylactic oophorectomy, cause health damage, and traditional assessment methods lack early intervention targets, failing to meet the needs for precise, safe, and long-term prevention.

Method used

A BRCA1/2 intelligent prevention system for ovarian cancer based on multi-omics data was constructed, including a multi-omics data standardization and integration module, an early carcinogenesis feature engineering module, an early carcinogenicity instability identification module, an interpretable risk assessment model module, a somatic intervention target screening module, a target safety and efficacy verification module, a personalized prevention plan generation module, and a model and target iterative update module, forming a closed-loop intelligent prevention system to screen safe and efficient somatic intervention targets and provide personalized prevention plans.

Benefits of technology

It enables precise prevention of ovarian cancer in BRCA1/2 mutation carriers, early identification of early signs of cancer, reduced reliance on preventive surgery, improved quality of life, increased accuracy of risk prediction, enhanced trust among clinicians, and the construction of a fully intelligent prevention system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090952A_ABST
    Figure CN122090952A_ABST
Patent Text Reader

Abstract

This invention provides a BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data, belonging to the field of ovarian cancer prevention technology. It aims to address the technical challenges of existing ovarian cancer prevention methods, such as reliance on prophylactic surgery leading to decreased quality of life, insufficient accuracy of traditional risk assessments, and the lack of a comprehensive intelligent prevention system. This system integrates multi-source, multi-omics data with artificial intelligence technology. Through eight collaborative functional modules, it achieves precise early warning of ovarian cancer risk in BRCA1 / 2 mutation carriers, risk stratification, intervention target screening, and personalized prevention plan generation. This completes the shift from surgical prevention to molecular prevention, reduces reliance on prophylactic surgery, improves prevention accuracy and clinical decision-making reliability, ensures the quality of life for BRCA1 / 2 mutation carriers, and covers the entire intelligent prevention process from data integration to plan iteration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ovarian cancer prevention technology, and more specifically, to a BRCA1 / 2 intelligent ovarian cancer prevention system based on multi-omics data. Background Technology

[0002] Ovarian cancer is the deadliest gynecological malignancy. It is characterized by its insidious onset, rapid progression, and difficulty in early diagnosis, with most patients diagnosed at an advanced stage and suffering a poor prognosis. BRCA1 / 2 germline mutations are a core high-risk factor for ovarian cancer. Clinical data show that individuals carrying this mutation have a lifetime risk of developing ovarian cancer that is 39% to 46%, dozens of times higher than the risk in the general population. Therefore, BRCA1 / 2 mutation carriers are a key population for ovarian cancer prevention.

[0003] The current mainstream preventive measure for BRCA1 / 2 mutation carriers is prophylactic salpingectomy. Although this surgery can reduce the risk of ovarian cancer to some extent, it can lead to premature menopause, which in turn can cause a series of complications such as osteoporosis, cardiovascular disease, and endocrine disorders, seriously damaging the patient's health and quality of life. This is difficult for most mutation carriers to accept.

[0004] Traditional methods of ovarian cancer risk assessment mainly rely on family history inquiry and single gene testing, which have obvious limitations: on the one hand, family history records may be missing or incomplete, and single gene testing cannot dynamically reflect the evolution of an individual's unstable genomic state, making it difficult to accurately capture early signs of cancer; on the other hand, this assessment method can only achieve a preliminary risk assessment, lacks clear early intervention targets, and cannot provide technical support for non-surgical prevention.

[0005] With the rapid development of multi-omics technologies such as genomics, transcriptomics, and epigenomics, along with artificial intelligence algorithms, integrating multi-dimensional biological data to mine early cancer characteristics and develop targeted molecular prevention strategies has become a research hotspot in the field of ovarian cancer prevention. However, this field currently lacks a systematic intelligent prevention system. Existing technologies mostly focus on single aspects such as risk prediction or target screening, failing to achieve full-process coverage from multi-omics data integration, risk stratification, target screening to personalized prevention strategy generation, and thus failing to meet the actual clinical needs for precise, safe, and long-term prevention of BRCA1 / 2 mutation carriers. Therefore, this paper proposes a BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data. Summary of the Invention

[0006] The purpose of this invention is to provide a BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data, addressing the clinical needs of ovarian cancer prevention in BRCA1 / 2 mutation carriers, overcoming the shortcomings of existing technologies, and aiming to achieve the following core objectives: Accurately identify early oncogenic genomic instability in BRCA1 / 2 mutation carriers, provide real-time warnings of ovarian cancer risk, capture early signs of carcinogenesis, and buy time for early intervention; To promote the transformation of ovarian cancer prevention from traditional surgical prevention to precision molecular prevention, we need to screen for safe, efficient gene editing and intervention targets that target somatic cells rather than the germline, avoid health damage caused by surgery, and ensure the quality of life of patients. Construct interpretable machine learning models to clarify the contribution of various multi-omics features to risk prediction, realize the traceability and interpretability of clinical decisions, and enhance clinicians' trust in model decisions and their willingness to apply them. Establish a comprehensive intelligent prevention system covering multi-omics data integration, risk assessment, intervention target screening, personalized treatment plan generation, and model iteration and updating, reducing reliance on preventive surgery and providing long-term, precise ovarian cancer prevention services for BRCA1 / 2 mutation carriers.

[0007] This invention provides the following technical solution: a BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data. The system includes eight collaborative functional modules: a multi-omics data standardization and integration module, an early carcinogenesis feature engineering module, an early carcinogenicity instability identification module, an interpretable risk assessment model module, a somatic intervention target screening module, a target safety and efficacy verification module, a personalized prevention plan generation module, and a model and target iteration update module. These modules are seamlessly connected to form a closed-loop intelligent prevention architecture, achieving full-process coverage from multi-omics data integration, risk assessment, intervention target screening to personalized prevention plan generation and model iteration. This system is used for the precise prevention of ovarian cancer in BRCA1 / 2 mutation carriers, reducing reliance on prophylactic salpingectomy.

[0008] As a preferred technical solution of the present invention, the multi-omics data standardization and integration module has the core functions of accessing, cleaning, standardizing, and uniformly storing multi-source public databases and clinical data. The multi-source public databases include the TCGA-OV database, GTEx database, DepMap database, gnomAD database, COSMIC database, and ENCODE database. The clinical data includes follow-up data of BRCA1 / 2 mutation carriers, imaging examination results, and omics data of blood and tissue samples. Its core process includes: aligning all multi-omics data using the GRCh38 version of the genomic coordinate system and correcting for batch effects; filtering common benign variants based on the gnomAD database and retaining functional mutations related to tumorigenesis; and constructing a multi-omics association map that integrates sample genomic variants, transcriptomic expression, and epigenetic regulatory features.

[0009] As a preferred technical solution of the present invention, the early carcinogenesis feature engineering module is used to extract core features of early ovarian cancer related to BRCA1 / 2 deficiency from standardized and integrated multi-omics data and construct a high-quality feature set. The core features include four dimensions: BRCA1 / 2-mediated DNA repair pathway features, genome replication stress features, body immune training features, and genome instability features. Its core process includes: constructing normal baseline reference intervals for each feature based on normal ovarian and fallopian tube tissue data from the GTEx database, calculating the deviation of each feature of BRCA1 / 2 mutation carrier samples from the baseline reference intervals, and screening high-value features significantly related to ovarian cancer morphology and removing redundant variables through mutual information analysis and LASSO regression algorithm.

[0010] As a preferred technical solution of the present invention, the early carcinogenic unstable state identification module dynamically identifies the stage of ovarian cancer progression in BRCA1 / 2 mutation carriers and achieves risk stratification based on the standardized feature set constructed by the early carcinogenesis feature engineering module and longitudinal follow-up data. The core algorithms it employs include a temporal convolutional network (TCN) and an unsupervised clustering algorithm. The temporal convolutional network is used to integrate longitudinal follow-up multi-omics data to capture the dynamic trend of genomic instability. The unsupervised clustering algorithm divides the samples into three categories: stable, potentially unstable, and high-risk unstable. The module outputs an individual risk stratification report, which includes the individual risk type, current cancer status score, risk upward trend prediction, and core basis for risk stratification.

[0011] As a preferred technical solution of the present invention, the interpretable risk assessment model module includes a basic prediction model, an interpretable module, and a binary risk grading and scoring system. The basic prediction model adopts the gradient boosting tree (GBDT) algorithm and integrates high-value multi-omics features screened by the early cancer feature engineering module to achieve quantitative prediction of the risk of ovarian cancer in BRCA1 / 2 mutation carriers. The interpretable module introduces SHAP (Shapley Additive)... The exPlanations value calculation method quantifies the contribution of each multi-omics feature to the risk prediction results and visualizes it through feature importance heatmaps and individual decision path maps. The binary risk grading scoring system sets risk thresholds based on quantitative prediction results, classifying BRCA1 / 2 mutation carriers into low-risk and high-risk groups, and outputs a grading report including risk level, scoring basis, risk characteristics, and risk trend prediction. The interpretable risk assessment model module uses TCGA-OV database tumor data and clinical follow-up data as training sets, and uses five-fold cross-validation to optimize model parameters and risk grading thresholds, improving the accuracy and reliability of risk grading, providing core basis for subsequent personalized prevention plan development, and helping to avoid adnexal resection.

[0012] As a preferred technical solution of the present invention, the somatic cell intervention target screening module screens potential ovarian cancer intervention targets that target somatic cells and do not affect the reproductive system based on multi-omics data and public database resources. The screening logic of the somatic cell intervention target screening module includes three levels: based on CRISPR gene knockout data in the DepMap database, screening genes that are highly dependent on the proliferation of BRCA-deficient ovarian cancer cells but not essential for normal ovarian epithelial cells to ensure safety; combined with genomic regulatory element annotation data in the ENCODE database, excluding targets located in reproductive system-specific enhancer regions to avoid ethical risks; and matching tumor somatic cell mutation data in the COSMIC database to select genes with high-frequency somatic cell mutations and well-defined functions in ovarian cancer to ensure effectiveness. The somatic cell intervention target screening module outputs a list of screening targets, including the gene name, gene function, cell-dependent score, and tissue-specific expression characteristics of each target, prioritizing the screening of targets suitable for gene editing in high-risk populations and those that can avoid adnexal resection.

[0013] As a preferred technical solution of the present invention, the target safety and efficacy verification module performs multi-dimensional verification on the initial screening targets output by the somatic intervention target screening module. The verification dimensions include: safety verification, using normal tissue expression data from the GTEx database to confirm that the target is expressed at low levels in key tissues such as the heart, liver, and kidneys outside the ovary to reduce the risk of off-target toxicity; efficacy verification, using transcriptome data from the TCGA-OV database and clinical prognostic data to verify the correlation between target downregulation and improved prognosis in patients with BRCA-mutant ovarian cancer; and editing feasibility verification, based on CRISPR-Cas9 target design rules, screening sgRNA sequences with low off-target rates and high editing efficiency. This module outputs a list of verified high-priority targets, including detailed verification results for each target, optimal editing scheme suggestions, and expected intervention effects.

[0014] As a preferred technical solution of the present invention, the personalized prevention plan generation module generates a personalized molecular prevention plan based on the binary risk grading results of the interpretable risk assessment model module, the high-priority targets of the target safety and efficacy verification module, and combined with the clinical characteristics of BRCA1 / 2 mutation carriers such as age, physical condition, and follow-up data. The core of this module is to achieve precise prevention of ovarian cancer without the need for adnexal resection. As a preferred technical solution of the present invention, the solution is divided into two categories, adapted to the results of binary risk grading: a health management guidance program for low-risk individuals, with lifestyle intervention as the core, combined with an annual multi-omics joint physical examination, to continuously monitor risk changes without invasive intervention or adnexal removal; and a somatic cell gene editing intervention program for high-risk individuals, which uses high-priority targets confirmed by the target safety and efficacy verification module, provides precise gene editing targets, customized editing programs, and postoperative auxiliary monitoring plans, and blocks the path of ovarian cancer through gene editing intervention, completely avoiding the health damage caused by adnexal removal surgery; this module outputs a personalized prevention report, including individual risk level, core risk characteristics, intervention targets, implementation path, follow-up plan, and expected intervention effect, clearly indicating the basis for intervention without adnexal removal.

[0015] As a preferred technical solution of the present invention, the model and target iteration update module is used to realize the dynamic iterative optimization of the system's various models and target screening rules. Its iteration process includes: regularly synchronizing multi-omics data updated in public databases such as TCGA-OV, GTEx, and DepMap; integrating newly added BRCA1 / 2 mutation carrier clinical follow-up data and tissue sample data to expand the training dataset; retraining the risk assessment model and the early carcinogenic unstable state identification model based on the newly added clinical outcome data; optimizing model parameters to improve prediction accuracy and generalization ability; and updating the target screening rules in conjunction with the latest research results in the field of ovarian cancer prevention, including new potential intervention targets and removing ineffective or potentially unsafe targets, continuously optimizing the target list to ensure the long-term accuracy and practicality of the system.

[0016] Compared with existing technologies, the beneficial effects of this invention are as follows: The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data of this invention constructs a full-process, precise, and personalized intelligent prevention system for ovarian cancer through the collaborative work of 8 modules. Compared with existing technologies, it has the following significant technical effects: This invention integrates multi-dimensional and multi-omics features, combines temporal convolutional networks with interpretable machine learning models, and improves the risk prediction accuracy to over 87%. Compared with traditional assessment methods that rely on family history and single gene testing, it can identify early signs of ovarian cancer in BRCA1 / 2 mutation carriers 3-5 years in advance, buying valuable time for early intervention. The somatic cell intervention targets screened by this invention have high safety and effectiveness. They can achieve precise molecular intervention through gene editing or targeted drugs, completely changing the traditional prevention model that relies on preventive surgery, avoiding health problems such as premature menopause and osteoporosis caused by surgery, and significantly improving the quality of life of BRCA1 / 2 mutation carriers. This invention visualizes SHAP values ​​to clarify the contribution of each multi-omics feature to risk prediction results, generates an intuitive decision path diagram and risk interpretation report, enhances clinicians' trust in model decisions, and facilitates clinical application. This invention constructs a fully intelligent system that integrates multi-omics data, extracts early cancer characteristics, stratifies risks, screens intervention targets, generates personalized treatment plans, and iteratively updates models and targets. It forms a closed loop of data-analysis-decision-intervention-iteration, comprehensively meeting the full-cycle prevention needs of BRCA1 / 2 mutation carriers from risk screening to long-term health management, reducing the risk of ovarian cancer, and has significant clinical application value and social significance. Attached Figure Description

[0017] Figure 1 This is a block diagram of the BRCA1 / 2 ovarian cancer intelligent prevention system module provided by the present invention; Figure 2 The feature dimension data diagram provided by this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.

[0019] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely illustrates some embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. It should be noted that, in the absence of conflict, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0020] Example 1: A BRCA1 / 2 Ovarian Cancer Intelligent Prevention System Based on Multi-omics Data. This system comprises eight collaborative functional modules: a multi-omics data standardization and integration module, an early carcinogenesis feature engineering module, an early carcinogenicity instability identification module, an interpretability risk assessment model module, a somatic intervention target screening module, a target safety and efficacy verification module, a personalized prevention plan generation module, and a model and target iteration update module. These modules seamlessly connect to form a closed-loop intelligent prevention architecture, achieving full-process coverage from multi-omics data integration, risk assessment, intervention target screening to personalized prevention plan generation and model iteration. This system is used for the precise prevention of ovarian cancer in BRCA1 / 2 mutation carriers, reducing reliance on prophylactic salpingectomy.

[0021] The multi-omics data standardization and integration module's core functions include accessing, cleaning, standardizing, and uniformly storing multi-source public databases and clinical data. The multi-source public databases include the TCGA-OV database, GTEx database, DepMap database, gnomAD database, COSMIC database, and ENCODE database. Clinical data includes follow-up data of BRCA1 / 2 mutation carriers, imaging examination results, and omics data from blood and tissue samples. Its core workflow includes: aligning all multi-omics data using the GRCh38 version of the genomic coordinate system and correcting for batch effects; filtering common benign variants based on the gnomAD database while retaining functional mutations related to tumorigenesis; and constructing a multi-omics association map integrating genomic variants, transcriptomic expression, and epigenetic regulatory features of the samples.

[0022] The Early Cancer Characterization Engineering Module is used to extract core features of early ovarian cancer associated with BRCA1 / 2 deficiency from standardized and integrated multi-omics data and construct a high-quality feature set. The core features include four dimensions: BRCA1 / 2-mediated DNA repair pathway features, genome replication stress features, host immune training features, and genome instability features. Its core process includes: constructing normal baseline reference intervals for each feature based on normal ovarian and fallopian tube tissue data from the GTEx database; calculating the deviation of each feature from the baseline reference intervals in BRCA1 / 2 mutation carrier samples; and screening high-value features significantly associated with ovarian cancer morphology and removing redundant variables through mutual information analysis and LASSO regression algorithm.

[0023] The early carcinogenic instability identification module, based on the standardized feature set constructed by the early carcinogenesis feature engineering module and longitudinal follow-up data, dynamically identifies the stages of ovarian cancer progression in BRCA1 / 2 mutation carriers and achieves risk stratification. Its core algorithms include a temporal convolutional network (TCN) and an unsupervised clustering algorithm. The TCN integrates longitudinal follow-up multi-omics data to capture the dynamic trends of genomic instability, while the unsupervised clustering algorithm categorizes samples into three types: stable, potentially unstable, and high-risk unstable. This module outputs an individual risk stratification report, including individual risk type, current cancer status score, predicted risk increase trend, and core basis for risk stratification.

[0024] The interpretability risk assessment model module includes a basic prediction model and an interpretability module. The basic prediction model uses the Gradient Boosting Tree (GBDT) algorithm and integrates high-value multi-omics features screened by the early cancer morphology feature engineering module to achieve quantitative prediction of the risk of ovarian cancer in BRCA1 / 2 mutation carriers. The interpretability module introduces the SHAPSHapley Additive exPlanations value calculation method to quantify the contribution of each multi-omics feature to the risk prediction results and visualizes it through feature importance heatmaps and individual decision path maps. Its core process includes using tumor data and clinical follow-up data from the TCGA-OV database as training sets, using five-fold cross-validation to optimize model parameters, predicting the risk of BRCA1 / 2 mutation carriers, and generating personalized risk interpretation reports for high-risk individuals.

[0025] The somatic cell intervention target screening module screens potential ovarian cancer intervention targets that target somatic cells and do not affect the reproductive system, based on multi-omics data and public database resources. Its screening logic includes three levels: First, based on CRISPR gene knockout data from the DepMap database, it screens genes that are highly dependent on the proliferation of BRCA-deficient ovarian cancer cells but not essential for normal ovarian epithelial cells to ensure safety. Second, it combines genomic regulatory element annotation data from the ENCODE database to exclude targets located in reproductive system-specific enhancer regions to avoid ethical risks. Third, it matches tumor somatic mutation data from the COSMIC database, prioritizing genes with high-frequency somatic mutations and well-defined functions in ovarian cancer to ensure effectiveness. This module outputs a list of initial screening targets, including the gene name, gene function, cell-dependent score, and tissue-specific expression characteristics for each target.

[0026] The target safety and efficacy validation module performs multi-dimensional validation on the initial screening targets output by the somatic intervention target screening module. Validation dimensions include: safety validation, using normal tissue expression data from the GTEx database to confirm low expression of the target in key tissues outside the ovary, such as the heart, liver, and kidney, to reduce off-target toxicity risk; efficacy validation, using transcriptomic data from the TCGA-OV database and clinical prognostic data to verify the correlation between target downregulation and improved prognosis in patients with BRCA-mutant ovarian cancer; and editing feasibility validation, based on CRISPR-Cas9 target design rules, screening for sgRNA sequences with low off-target rates and high editing efficiency. This module outputs a list of validated high-priority targets, including detailed validation results for each target, optimal editing protocol suggestions, and expected intervention effects.

[0027] The personalized prevention plan generation module, based on the individual risk stratification results from the early oncogenic instability identification module and the high-priority targets from the target safety and efficacy verification module, combined with the clinical characteristics of BRCA1 / 2 mutation carriers such as age, physical condition, and follow-up data, generates personalized molecular prevention plans. The plans are divided into two categories: stable health management guidance plans for low-risk individuals, including the frequency of regular multi-omics monitoring and lifestyle intervention recommendations; and somatic gene editing intervention plans for high-risk, unstable individuals, providing gene editing targets, specific editing protocols, and post-operative auxiliary monitoring plans. This module outputs a personalized prevention report, including individual risk level, core risk characteristics, intervention targets, implementation pathways, follow-up plans, and expected intervention effects.

[0028] The model and target iteration update module is used to realize the dynamic iterative optimization of the system's various models and target selection rules. Its iteration process includes: regularly synchronizing multi-omics data updated from public databases such as TCGA-OV, GTEx, and DepMap; integrating newly added BRCA1 / 2 mutation carrier clinical follow-up data and tissue sample data to expand the training dataset; retraining the risk assessment model and the early carcinogenic unstable state identification model based on newly added clinical outcome data, optimizing model parameters to improve prediction accuracy and generalization ability; and updating the target selection rules in conjunction with the latest research results in the field of ovarian cancer prevention, including new potential intervention targets and removing ineffective or potentially unsafe targets, continuously optimizing the target list to ensure the long-term accuracy and practicality of the system.

[0029] Example 2: This BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data includes a multi-omics data standardization and integration module. This module accesses, cleans, standardizes, and uniformly stores multi-source public databases and clinical data, removes data redundancy and noise, and ensures the accuracy, integrity, and consistency of the data, providing high-quality data support for subsequent feature extraction, risk assessment, and target screening.

[0030] Data sources cover multi-omics data, genetic variation data, and clinical follow-up data of tumors and normal tissues, specifically including: TCGA-OV database: providing core data such as tumor genomic variations and transcriptome expression profiles for analyzing the molecular characteristics of ovarian cancer cells; GTEx database: provides transcriptome and epigenome reference data for normal ovarian and fallopian tube tissues, serving as a baseline standard for judging abnormal tissue characteristics in BRCA1 / 2 mutation carriers; DepMap database: Contains cell proliferation-dependent data from CRISPR gene knockout, used to screen for gene targets that are key to the proliferation of BRCA-deficient ovarian cancer cells; gnomAD database: Provides background frequency data of genetic variations in the normal population, used to filter common benign variations and focus on functional mutations associated with tumorigenesis; COSMIC database: contains tumor somatic mutation spectrum and driver mutation annotation data, used to screen for high-frequency and functionally defined somatic mutation targets in ovarian cancer. ENCODE database: provides data on promoters, enhancers, and open chromatin regions annotated with genomic regulatory elements, which can be used to avoid targets in germline-specific regulatory regions and ensure the safety of interventions; Clinical access data: This includes follow-up data of BRCA1 / 2 mutation carriers, imaging examination results, and omics data of blood and tissue samples, which are used for model training, validation, and personalized treatment adjustments.

[0031] Workflow: Data Alignment and Batch Effect Correction: All multi-omics data were aligned using the unified genomic coordinate system GRCh38 to eliminate coordinate differences between data from different sources; at the same time, a standardized algorithm was used to correct for data batch effects to ensure the comparability of data from different batches and sources. Variance filtering and screening: Based on the genetic variation background of the normal population in the gnomAD database, high-frequency benign variants are filtered out, while functional mutations related to the occurrence and development of ovarian cancer are retained, reducing interference from invalid data; Multi-omics association map construction: Integrating genomic variation, transcriptomic expression and epigenetic regulatory features of samples, constructing a multi-dimensional association map, clarifying the intrinsic relationships between various omics data, and providing a data association foundation for subsequent feature extraction.

[0032] The early cancer feature engineering module extracts core features of early ovarian cancer associated with BRCA1 / 2 deficiency from standardized and integrated multi-omics data, removes redundant variables, and constructs a high-quality feature set for risk prediction and target screening, thereby achieving accurate capture of early cancer signals.

[0033] The feature dimension extracts early carcinogenesis features, comprehensively covering the molecular mechanism of carcinogenesis mediated by BRCA1 / 2 deficiency, specifically including: DNA repair pathway features: covering the expression level, mutation frequency, functional score of genes in the BRCA1 / 2 mediated homologous recombination repair pathway, as well as quantitative indicators of DNA damage repair efficiency, focusing on reflecting the impact of BRCA1 / 2 mutations on DNA repair function. Replication stress characteristics: including expression levels of replication fork arrest markers, genomic copy number variation load, and microsatellite instability, are used to assess abnormal states during genome replication and capture early signs of genomic instability in carcinogenesis; Immune training characteristics include the abundance of tumor-infiltrating immune cell subtypes, the expression level of immune checkpoint genes, and the secretion characteristics of chemokines, reflecting the body's immune system's recognition and response to early-stage cancerous cells. Genomic instability features: encompassing chromosome fragmentation frequency, dynamic changes in telomere length, and somatic mutation feature spectrum, with a focus on SBS3 mutation features, directly quantifying the degree of instability of an individual's genome, serving as a core marker for early carcinogenesis.

[0034] Workflow: Baseline reference interval construction: Based on multi-omics data of normal ovarian and fallopian tube tissues from the GTEx database, normal baseline reference intervals for each feature are constructed as a standard for judging whether the features of BRCA1 / 2 mutation carrier samples are abnormal. Feature deviation calculation: Calculate the deviation of each feature of BRCA1 / 2 mutation carrier samples from the normal baseline reference interval to quantify the degree of feature abnormality; High-value feature screening: Through mutual information analysis and LASSO regression algorithm, high-value features that are significantly related to the process of ovarian cancer are screened out, redundant variables and irrelevant features are removed, and a concise and efficient risk prediction feature set is constructed.

[0035] The standardized feature set constructed by the early carcinogenic unstable state identification module, combined with longitudinal follow-up data, dynamically identifies the stages of ovarian cancer progression in BRCA1 / 2 mutation carriers, achieving precise stratification of individual risk and providing a basis for the formulation of subsequent intervention plans.

[0036] Core algorithm: Temporal Convolutional Network (TCN) is used to integrate longitudinal follow-up multi-omics data of BRCA1 / 2 mutation carriers to capture the dynamic changes in genomic instability and avoid risky misjudgments caused by data from a single time point. Unsupervised clustering algorithm: Based on the quantitative data of the feature set, the samples are divided into three categories: stable, potentially unstable, and high-risk unstable. The high-risk unstable category indicates that the individual has entered the early stage of cancer initiation and requires key intervention.

[0037] Output: Generates an individual risk stratification report, which clarifies the individual's risk type: stable, potentially unstable, or high-risk unstable. It includes the current cancer status score, risk upward trend prediction, and the core basis for risk stratification, providing clinicians with an intuitive reference for risk assessment.

[0038] The interpretable risk assessment model module accurately predicts the risk of ovarian cancer in BRCA1 / 2 mutation carriers based on a standardized feature set. At the same time, it uses visualization to clarify the contribution of each feature to the risk prediction results, thereby improving the credibility and traceability of clinical decision-making.

[0039] Model architecture: The basic prediction model adopts the gradient boosting tree (GBDT) algorithm, integrates multi-omics high-value features screened in module 2, optimizes the model's risk prediction accuracy, and realizes quantitative prediction of the risk of ovarian cancer. The interpretability module introduces the SHAPSHapley Additive exPlanations value calculation method to quantify the contribution of each multi-omics feature to the individual risk prediction results. Through visualization methods such as feature importance heatmaps and individual decision path maps, the core basis of risk prediction is presented intuitively.

[0040] Workflow: Model training and optimization: Using tumor data and clinical follow-up data from the TCGA-OV database as the training set, the model parameters are optimized using the five-fold cross-validation method to improve the model's prediction accuracy and generalization ability. Risk Prediction and Interpretation: Based on a trained model, the risk of ovarian cancer in BRCA1 / 2 mutation carriers is quantitatively predicted; for high-risk individuals, personalized risk interpretation reports are generated, identifying the core characteristics leading to increased risk, such as downregulated BRCA2 expression and elevated homologous recombination repair defect scores, to help clinicians understand the model's decision-making logic.

[0041] The somatic cell intervention target screening module is based on multi-omics data and public database resources to screen potential ovarian cancer intervention targets that target somatic cells and do not affect the reproductive system, thereby avoiding ethical and safety risks and ensuring the effectiveness of the targets. It also provides a candidate list for subsequent target validation and intervention program design.

[0042] Screening logic: Based on CRISPR gene knockout data from the DepMap database, we screen for genes that are highly dependent on the proliferation of BRCA-deficient ovarian cancer cells but not essential for normal ovarian epithelial cells, ensuring that the intervention target is inhibited, which only affects the proliferation of tumor cells and does not damage normal ovarian tissue. Ethical safety screening: By combining genomic regulatory element annotation data from the ENCODE database, targets located in germline-specific enhancer regions are excluded to avoid gene editing and other interventions affecting germ cells and to mitigate ethical risks; Effectiveness screening: Match tumor somatic mutation data in the COSMIC database, and prioritize genes that have high frequency of somatic mutations in ovarian cancer and whose functions are clearly related to the occurrence and development of ovarian cancer, to ensure the effectiveness of intervention targets.

[0043] Output results: A preliminary target list is generated, containing core information for each target: gene name, gene function, and cell-dependent score. This reflects the importance of the target to BRCA-deficient ovarian cancer cells, tissue-specific expression characteristics, and clarifies the expression distribution of the target in normal tissues, providing clear guidance for subsequent target validation.

[0044] The target safety and effectiveness verification module initially screens out potential intervention targets and conducts multi-dimensional and rigorous verification to eliminate targets with insufficient safety or poor effectiveness, and selects the optimal candidate intervention targets to provide core support for the generation of personalized prevention plans.

[0045] Validation dimension: Safety validation. Using normal tissue expression data from the GTEx database, we comprehensively analyze the expression level of the target in key tissues other than the ovary, such as the heart, liver, and kidneys, to confirm that the target is expressed at low levels in these tissues, thereby reducing the risk of off-target toxicity during the intervention process and ensuring the safety of the intervention. Efficacy verification: Using transcriptomic data and clinical prognostic data from the TCGA-OV database, we verified the correlation between target expression downregulation and improved prognosis in patients with BRCA-mutant ovarian cancer, confirming that target inhibition can effectively suppress ovarian cancer progression and ensure the effectiveness of the intervention. Editing feasibility verification: Based on the CRISPR-Cas9 target design rules, we analyzed the gene sequence characteristics of the target, screened sgRNA sequences with low off-target rate and high editing efficiency, and ensured that the target could be precisely intervened through gene editing technology.

[0046] Output results: Generate a list of validated high-priority targets, including detailed validation results for each target, optimal editing scheme suggestions such as sgRNA sequence recommendations, and expected intervention effects such as the inhibition rate of tumor cell proliferation after target inhibition, providing a direct basis for the design of subsequent personalized prevention programs.

[0047] The personalized prevention plan generation module is based on the individual risk stratification results of module 3 and the high-priority intervention targets screened in module 6. It combines the clinical characteristics of BRCA1 / 2 mutation carriers, such as age, physical condition, and follow-up data, to generate personalized molecular prevention plans that match individual needs, thereby achieving personalized management of risk stratification and precision intervention.

[0048] Program type: Low-risk stable type: Focuses on health management, providing personalized health management guidance, including regular multi-omics monitoring frequency such as multi-omics testing every 1-2 years, lifestyle intervention recommendations such as limiting alcohol intake, increasing physical exercise, controlling weight, avoiding long-term sleep deprivation, etc., and regularly monitoring changes in risk; For high-risk individuals with unstable risk: the main approach is somatic cell gene editing intervention, providing somatic cell gene editing targets and specific editing protocols such as sgRNA sequences and editing methods. This is combined with a postoperative auxiliary monitoring plan, such as multi-omics testing every 3-6 months after the intervention, to closely track the genomic status after the intervention, adjust the plan in a timely manner, and combine targeted drugs for auxiliary prevention when necessary.

[0049] Output results: Generate personalized prevention reports that comprehensively present an individual's risk level, core risk characteristics, intervention targets, specific intervention implementation path, medication plan, gene editing plan, lifestyle recommendations, follow-up plan, monitoring frequency, monitoring indicators, and expected intervention effects, providing clear and actionable prevention guidelines for clinicians and patients.

[0050] The model and target iteration update module continuously integrates new multi-omics data, clinical follow-up results, and research progress in the field to achieve dynamic iterative optimization of the system's various models and target screening rules, ensuring the long-term accuracy and practicality of the system and adapting to the technological development needs in the field of ovarian cancer prevention.

[0051] Iterative Process: Data Iteration. Regularly synchronize updated multi-omics data from various public databases such as TCGA-OV, GTEx, and DepMap, while integrating newly added clinical follow-up data and tissue sample data of BRCA1 / 2 mutation carriers to continuously expand the scale of the training dataset and improve the comprehensiveness of the data; Model iteration: Based on newly added clinical outcome data such as intervention effects and cancer incidence, the risk assessment model and unstable state identification model are retrained, the model parameters are optimized, and the prediction accuracy and generalization ability of the model are improved. Target iteration: Combining the latest research findings in the field of ovarian cancer prevention, such as novel intervention targets and new discoveries on target functions, we update the target screening rules, include new potential intervention targets, remove targets that have been proven ineffective or have safety risks, and continuously optimize the target list.

[0052] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described herein. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or equivalent substitutions to the present invention, as well as all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.

Claims

1. A BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data, characterized in that, The system includes a multi-omics data standardization and integration module, an early cancer morphology feature engineering module, an early carcinogenic unstable state identification module, an interpretable risk assessment model module, a somatic intervention target screening module, a target safety and efficacy verification module, a personalized prevention plan generation module, and a model and target iteration update module. These modules are seamlessly connected to form a closed-loop intelligent prevention architecture, achieving full-process coverage from multi-omics data integration, risk assessment, intervention target screening to personalized prevention plan generation and model iteration, for the precise prevention of ovarian cancer in BRCA1 / 2 mutation carriers.

2. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 1, characterized in that, The multi-omics data standardization and integration module accesses, cleans, standardizes, and uniformly stores multi-source public databases and clinical data. The multi-source public databases include the TCGA-OV database, GTEx database, DepMap database, gnomAD database, COSMIC database, and ENCODE database. The clinical data includes follow-up data of BRCA1 / 2 mutation carriers, imaging examination results, and omics data of blood and tissue samples. The multi-omics data standardization and integration module includes: aligning all multi-omics data using the GRCh38 version of the genome coordinate system and correcting for batch effects; filtering common benign variants based on the gnomAD database and retaining functional mutations related to tumorigenesis; and constructing a multi-omics association map that integrates genomic variants, transcriptomic expression, and epigenetic regulatory features of the integrated samples.

3. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 2, characterized in that, The early cancer feature engineering module is used to extract core features of early ovarian cancer associated with BRCA1 / 2 deficiency from standardized and integrated multi-omics data and construct a high-quality feature set. The early cancer feature engineering module includes four dimensions: BRCA1 / 2-mediated DNA repair pathway features, genome replication stress features, body immune training features, and genome instability features. The early cancer feature engineering module constructs normal baseline reference intervals for each feature based on normal ovarian and fallopian tube tissue data from the GTEx database, calculates the deviation of each feature from the baseline reference interval of BRCA1 / 2 mutation carrier samples, and screens high-value features significantly associated with ovarian cancer lesions and removes redundant variables through mutual information analysis and LASSO regression algorithm.

4. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 3, characterized in that, The early carcinogenic instability identification module, based on the standardized feature set constructed by the early carcinogenesis feature engineering module and longitudinal follow-up data, dynamically identifies the stages of ovarian cancer progression in BRCA1 / 2 mutation carriers and achieves risk stratification. The algorithm includes a temporal convolutional network (TCN) and an unsupervised clustering algorithm. The TCN is used to integrate longitudinal follow-up multi-omics data to capture the dynamic trends of genomic instability, while the unsupervised clustering algorithm divides samples into three categories: stable, potentially unstable, and high-risk unstable. This module outputs an individual risk stratification report, which includes the individual risk type, current cancer status score, risk upward trend prediction, and core basis for risk stratification.

5. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 4, characterized in that, The interpretability risk assessment model module includes a basic prediction model, an interpretability module, and a binary risk grading and scoring system. The basic prediction model uses the gradient boosting tree (GBDT) algorithm and integrates high-value multi-omics features screened by the early carcinogenesis feature engineering module to achieve quantitative prediction of the risk of ovarian cancer in BRCA1 / 2 mutation carriers. The interpretability module introduces the SHAP (Shapley Additive exPlanations) value calculation method to quantify the contribution of each multi-omics feature to the risk prediction results, and achieves visualization through feature importance heatmaps and individual decision path maps. The binary risk grading scoring system sets risk thresholds based on quantitative prediction results, classifies BRCA1 / 2 mutation carriers into low-risk and high-risk groups, and outputs a grading report including risk level, scoring basis, risk characteristics, and risk trend prediction. The interpretable risk assessment model module uses TCGA-OV database tumor data and clinical follow-up data as training sets, and uses five-fold cross-validation to optimize model parameters and risk grading thresholds, improving the accuracy and reliability of risk grading, providing core basis for subsequent personalized prevention plan formulation, and helping to avoid adnexal resection.

6. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 5, characterized in that, The somatic cell intervention target screening module screens potential ovarian cancer intervention targets that target somatic cells and do not affect the reproductive system, based on multi-omics data and public database resources. The screening logic of the somatic cell intervention target screening module includes three levels: based on CRISPR gene knockout data in the DepMap database, it screens genes that are highly dependent on the proliferation of BRCA-deficient ovarian cancer cells but not essential for normal ovarian epithelial cells to ensure safety; combined with genomic regulatory element annotation data in the ENCODE database, it excludes targets located in reproductive system-specific enhancer regions to avoid ethical risks; and matched with tumor somatic mutation data in the COSMIC database, it selects genes with high-frequency somatic mutations and well-defined functions in ovarian cancer to ensure effectiveness. The somatic cell intervention target screening module outputs a list of screening targets, including the gene name, gene function, cell dependence score, and tissue-specific expression characteristics of each target, prioritizing the screening of targets suitable for gene editing in high-risk populations and those that can avoid adnexal resection.

7. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 6, characterized in that, The target safety and efficacy verification module performs multi-dimensional verification of the initial screening targets output by the somatic intervention target screening module. The verification dimensions include: safety verification, using normal tissue expression data from the GTEx database to confirm that the target is expressed at low levels in key tissues such as the heart, liver, and kidneys outside the ovary to reduce the risk of off-target toxicity; verifying the correlation between target downregulation and improved prognosis in patients with BRCA-mutant ovarian cancer using transcriptome data and clinical prognostic data from the TCGA-OV database; and editing feasibility verification, based on CRISPR-Cas9 target design rules, screening sgRNA sequences with low off-target rates and high editing efficiency. This module outputs a list of verified high-priority targets, including detailed verification results for each target, optimal editing protocol suggestions, and expected intervention effects.

8. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 7, characterized in that, The personalized prevention plan generation module, based on the binary risk grading results of the interpretable risk assessment model module and the high-priority targets of the target safety and efficacy verification module, combined with the clinical characteristics of BRCA1 / 2 mutation carriers such as age, physical condition, and follow-up data, generates personalized molecular prevention plans, which ultimately achieve precise prevention of ovarian cancer without the need for adnexal resection.

9. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 8, characterized in that, The proposed solutions are divided into two categories, adapted to the results of a binary risk stratification: a health management guidance program for low-risk individuals, centered on lifestyle interventions and combined with an annual multi-omics physical examination to continuously monitor risk changes, without requiring invasive interventions or adnexal removal; and a somatic cell gene editing intervention program for high-risk individuals, employing high-priority targets confirmed by the target safety and efficacy verification module, providing precise gene editing targets, customized editing protocols, and postoperative adjuvant monitoring plans, blocking the pathway to ovarian cancer through gene editing intervention, and completely avoiding the health damage caused by adnexal removal surgery. This module outputs a personalized prevention report, including individual risk level, core risk characteristics, intervention targets, implementation pathway, follow-up plan, and expected intervention effects, clearly indicating the basis for interventions that do not require adnexal removal.

10. The BRCA1 / 2 ovarian cancer intelligent prevention system based on multi-omics data according to claim 9, characterized in that, The model and target iteration update module is used to realize the dynamic iterative optimization of the system's various models and target selection rules. The model and target iteration update module includes: regularly synchronizing multi-omics data updated from public databases such as TCGA-OV, GTEx, and DepMap; integrating newly added BRCA1 / 2 mutation carrier clinical follow-up data and tissue sample data to expand the training dataset; retraining the risk assessment model and the early carcinogenic unstable state identification model based on newly added clinical outcome data, and optimizing model parameters to improve prediction accuracy and generalization ability; and updating the target selection rules in combination with the latest research results in the field of ovarian cancer prevention, including new potential intervention targets and removing ineffective or potentially unsafe targets.