Combinations of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer and their applications

CN122669080APending Publication Date: 2026-09-01NINGXIA MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610784949.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0011]为了解决现有技术存在的上述不足,本发明的目的是提供一种用于上皮性卵巢癌早期筛查和辅助诊断的cfRNA标志物组合及其应用,以解决现有上皮性卵巢癌早期诊断中血清学标志物灵敏度不足、特异度偏低、单一指标稳定性差、影像学检查对早期微小病灶识别能力有限以及组织活检侵入性强、不适于大规模早筛和动态监测等技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122669080A_ABST
    Figure CN122669080A_ABST
Patent Text Reader

Abstract

This invention discloses a combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer and their applications, belonging to the field of biomedical technology. The combination of cfRNA biomarkers includes: hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255, and DMXL1-DT. The biomarker combination employs combined analysis of 8 miRNAs and 2 lncRNAs, which can improve the sensitivity, specificity, and stability of early screening and auxiliary diagnosis of epithelial ovarian cancer, and has the advantages of being non-invasive, convenient, reproducible, and having high clinical translational value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedical technology, specifically relating to a combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer and their applications. Background Technology

[0002] Ovarian cancer (OC) is one of the most common malignant tumors of the female reproductive system and one of the deadliest gynecological malignancies, seriously threatening women's lives and health. In my country, the annual incidence rate of ovarian cancer ranks among the top of female reproductive system malignancies and is showing an increasing trend year by year; its mortality rate is also among the top of female reproductive tract malignancies. Due to the insidious onset and atypical early symptoms of ovarian cancer, most patients are diagnosed at an advanced stage, resulting in a poor overall prognosis and a high recurrence rate. Therefore, achieving early screening and early diagnosis of ovarian cancer is of great significance for improving patient survival and prognosis.

[0003] Malignant ovarian tumors include various pathological types, among which epithelial ovarian cancer (EOC) is the most common. Currently, there is a lack of reliable, stable, and large-scale population-based early screening methods for ovarian cancer. Most patients still rely on symptom onset for diagnosis, but early clinical manifestations of ovarian cancer are often nonspecific. Common symptoms include abdominal or pelvic pain, bloating, constipation, diarrhea, urinary frequency, abnormal vaginal bleeding, and fatigue. In advanced stages, ascites, abdominal masses, nausea, loss of appetite, indigestion, early satiety, and pleural effusion may also occur. Because these symptoms lack specificity, misdiagnosis and missed diagnoses are common, delaying optimal treatment.

[0004] Currently, the diagnosis of ovarian cancer mainly relies on imaging examinations, serological marker detection, and histopathological examination. Imaging examinations have limited ability to identify early, small lesions; the sensitivity and specificity of traditional serological markers in early diagnosis still fall short of clinical needs; while tissue biopsy can confirm the diagnosis, it is highly invasive and unsuitable for early screening and dynamic follow-up. Therefore, developing a novel molecular marker and its detection method that is highly sensitive, specific, non-invasive, repeatable, and suitable for early screening has become a crucial technical problem urgently needing to be solved in the field of ovarian cancer prevention and treatment.

[0005] In recent years, liquid biopsy technology has become a hot topic in early cancer screening and precision diagnosis research due to its advantages such as being non-invasive, repeatable, and able to dynamically reflect tumor changes. Liquid biopsy typically involves collecting bodily fluids such as peripheral blood, cerebrospinal fluid, urine, and saliva to detect changes in disease-related biomolecules, thereby achieving molecular diagnosis. Compared to traditional tissue biopsies, liquid biopsies are less invasive to patients, facilitate continuous monitoring, and can, to some extent, overcome sampling biases caused by tumor heterogeneity, showing promising clinical application prospects.

[0006] Currently, common targets for detection in liquid biopsies include circulating tumor cells (CTCs), circulating tumor DNA (ctDNA), circulating cell-free DNA (cfDNA), circulating cell-free RNA (cfRNA), exosomes, and proteins. Among these, cfDNA, as one of the earliest studied biopsy biomarkers, has shown certain application value in various tumors, such as detecting tumor-related mutations and methylation changes. However, the application of cfDNA in ovarian cancer, especially early-stage ovarian cancer, still faces several limitations: on the one hand, early-stage tumors have a low tumor burden, resulting in a limited amount of cfDNA released into the circulation system, leading to decreased detection sensitivity; on the other hand, cfDNA is susceptible to interference from factors such as DNA released during normal cell apoptosis, degradation during sample processing, and leukocyte genomic contamination, thus affecting detection accuracy and stability.

[0007] In contrast, circulating cell-free RNA (cfRNA), as free RNA molecules present in body fluids such as plasma, originates from active secretion by tissue cells or release during apoptosis and necrosis. It is characterized by its rich diversity, high information content, and ability to reflect gene expression activity. cfRNA includes various types such as mRNA, miRNA, lncRNA, circRNA, snRNA, snoRNA, tRNA, and other small RNAs. In recent years, the application value of cfRNA in pregnancy-related diseases, tumors, and other complex diseases has gradually attracted attention. Studies have shown that cfRNA exhibits specific abnormal expression in the plasma of various cancer patients, and can serve as a potential non-invasive diagnostic and disease monitoring biomarker.

[0008] In the field of ovarian cancer, current liquid biopsy research mainly focuses on ctDNA, circulating tumor cells (CTCs), and exosomes. Some studies suggest that ctDNA detection has high specificity in ovarian cancer diagnosis, but its detection rate for early-stage ovarian cancer remains limited. CTCs have shown high positive rates in some studies, but their clinical application is still limited due to their extremely low circulating numbers, complex detection techniques, and insufficient standardization. While exosomes have some diagnostic potential, current research suffers from insufficient sample size, a lack of standardized experimental methods, and unsatisfactory reproducibility and stability of results. Overall, the sensitivity, specificity, and clinical translatability of existing liquid biopsy biomarkers in early ovarian cancer screening still need improvement.

[0009] Of particular note is that, compared to cfDNA, cfRNA more directly reflects the dynamic changes in the expression of tumor-related genes and has a higher abundance in some body fluids, theoretically making it more suitable for detecting early stages of disease and identifying risks. However, systematic studies on peripheral blood cfRNA biomarkers for ovarian cancer are still limited, especially lacking cfRNA combination biomarkers and diagnostic models with high sensitivity and specificity for early screening. Therefore, screening for ovarian cancer-specific cfRNA biomarkers and constructing a non-invasive early screening and diagnosis model based on cfRNA has significant scientific and clinical application value.

[0010] Therefore, there is an urgent need to develop a new technical solution based on peripheral blood biopsy that can be used for early screening and auxiliary diagnosis of ovarian cancer, especially to screen out cfRNA molecular markers with high diagnostic efficacy and establish a joint diagnostic model by combining high-throughput sequencing and bioinformatics analysis, so as to improve the sensitivity, specificity and clinical applicability of early diagnosis of ovarian cancer. Summary of the Invention

[0011] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide a combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer and their application, thereby resolving technical problems in the early diagnosis of epithelial ovarian cancer, such as insufficient sensitivity and low specificity of existing serological biomarkers, poor stability of single indicators, limited ability of imaging examinations to identify early small lesions, and the high invasiveness of tissue biopsy, making it unsuitable for large-scale early screening and dynamic monitoring.

[0012] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer is provided. This combination of cfRNA biomarkers comprises 8 miRNAs and 2 lncRNAs, namely: hsa-miR-3184-3p (MIMAT0022731), hsa-miR-18a-3p (MIMAT0002891), hsa-miR-135b-5p (MIMAT0000758), hsa-miR-486-3p (MIMAT0004762), hsa-miR-320a-3p (MIMAT0 000510), hsa-miR-103a-1-5p (MIMAT0037306), hsa-miR-181a-5p (MIMAT0000256), hsa-miR-221-5p (MIMAT0004568), ENSG00000287255 and DMXL1-DT (ENSG00000249494).

[0013] The beneficial effects of this invention are as follows: This invention is the first to use the aforementioned 10 cfRNA molecules as combined biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer. Compared with traditional single serological indicators or single nucleic acid biomarkers, this invention adopts a multi-cfRNA combined detection approach, effectively integrating the main driving signal of lncRNA with the auxiliary fine-tuning signal of miRNA. This combination can significantly improve the early detection rate of epithelial ovarian cancer while ensuring extremely high specificity, and has excellent clinical application prospects. The aforementioned cfRNA biomarkers are derived from the blood samples of subjects, preferably plasma or serum samples. These 10 cfRNA biomarkers showed abnormal expression (specifically, upregulated expression) in the blood samples of patients with epithelial ovarian cancer compared to healthy controls.

[0014] The above-mentioned combination of cfRNA biomarkers is used in the preparation of products for epithelial ovarian cancer screening, auxiliary diagnosis, high-risk population identification, efficacy evaluation, follow-up monitoring, or recurrence early warning.

[0015] Furthermore, the products include testing reagents, kits, testing chips, data analysis systems, computer-readable storage media, or electronic devices.

[0016] The reagent for detecting the expression level of the above-mentioned cfRNA biomarker combination is used in the preparation of products for epithelial ovarian cancer screening, auxiliary diagnosis, high-risk population identification, efficacy evaluation, follow-up monitoring or recurrence early warning.

[0017] A diagnostic reagent or kit for screening or assisting in the diagnosis of epithelial ovarian cancer, the reagent or kit comprising reagents for detecting the expression levels of the above-mentioned combination of cfRNA biomarkers.

[0018] Furthermore, the detection reagents include specific primers, probes, capture probes, or combinations thereof for detecting the expression levels of the above-mentioned cfRNA biomarker combinations.

[0019] Furthermore, the kit includes reagents for detecting the expression levels of the above-mentioned cfRNA biomarker combination, and may also include at least one of the following components: internal control reagent, positive control, negative control, buffer, enzyme preparation, and standard.

[0020] The beneficial effects of this invention are as follows: the reagents or kits used to detect the expression levels of the above-mentioned cfRNA biomarker combinations are highly suitable for non-invasive detection based on body fluids (such as plasma) because the detection indicators are clear and the results are quantifiable. They can be widely applied to early screening of epithelial ovarian cancer, non-invasive auxiliary diagnosis of individuals with uncertain imaging results, and pre- and post-operative risk assessment and dynamic monitoring of patients, thus facilitating the standardization and widespread application of clinical testing.

[0021] A risk assessment model for early screening of epithelial ovarian cancer is an ensemble machine learning model trained based on a random forest algorithm, wherein the random forest model includes 500 decision trees with a maximum depth of 10. The model is configured to: acquire the expression levels of the 10 cfRNA biomarkers described in claim 1 in the sample to be tested as input variables; The results are evaluated sequentially by the 500 decision trees, and the outputs of each decision tree are integrated and weighted to output the epithelial ovarian cancer risk score of the sample to be tested. When the risk score is greater than the diagnostic cutoff value, it is considered a high risk of epithelial ovarian cancer; when the risk score is less than or equal to the diagnostic cutoff value, it is considered a low risk of epithelial ovarian cancer.

[0022] Furthermore, the diagnostic cutoff values ​​of the above risk assessment model are obtained through at least one of the following methods: (1) The threshold corresponding to the maximum point of the Youden exponent on the ROC curve; (2) The dividing point determined under preset sensitivity or specificity conditions; (3) The optimal classification threshold obtained through cross-validation or independent external validation queues.

[0023] Furthermore, the diagnostic cutoff value for the aforementioned risk assessment model is 0.5.

[0024] The method for constructing the risk assessment model for early screening of epithelial ovarian cancer described above includes the following steps: (1) Data collection: Obtain plasma cfRNA expression value matrix and corresponding clinical classification labels of samples from patients diagnosed with epithelial ovarian cancer and healthy controls; (2) Differential expression analysis: Differential analysis was performed on the expression value matrix to screen candidate cfRNAs that meet the conditions of P value < 0.05 and |log2FC| ≥ 0.05 between the two groups; (3) Feature screening: The candidate cfRNAs are quantified based on the SHAP algorithm, the average absolute SHAP value of each candidate cfRNA is calculated, and the candidate cfRNAs are sorted in descending order according to the average absolute SHAP value. The top 10 cfRNAs are selected as core cfRNA markers. The 10 core cfRNA markers are hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255 and DMXL1-DT. (4) Model construction: Using the expression data of the 10 core cfRNA markers as input features and the corresponding clinical classification labels as output targets, the random forest algorithm is used to train and generate an integrated model containing multiple decision trees; the model performs nonlinear mapping and probability calculation through the internally integrated multiple decision trees, takes the average probability of all decision trees predicting positive examples, and outputs a value between 0 and 1 as the final risk score to obtain a risk assessment model for early screening of epithelial ovarian cancer.

[0025] A system for screening or assisting in the diagnosis of epithelial ovarian cancer, the system comprising: The data acquisition module is used for sequencing or nucleic acid amplification of the combination of the 10 cfRNA biomarkers in the subject sample to be tested; The data preprocessing module is used to perform missing value processing, normalization processing and / or standardization processing on the sequencing or nucleic acid amplification data to generate quantitative expression profile features of the cfRNA biomarker combination; The model computation module is used to input quantitative expression profile features into the risk assessment model according to any one of claims 5-7, and calculate and output the risk score or predicted probability of epithelial ovarian cancer through the random forest algorithm; The threshold determination module is used to output a determination conclusion that the subject under test is at high or low risk of epithelial ovarian cancer based on the comparison result between the risk score or predicted probability and the preset diagnostic cutoff value. The results output module is used to display, store, print, and / or transmit the judgment conclusions and risk assessment reports.

[0026] Furthermore, the system is also equipped with a feature interpretation module, which is used to extract and display the contribution of each molecular feature in the 10 cfRNA biomarkers to the risk score of the current sample based on SHAP (SHapley Additive ex Planations) value or Gini Importance. Among them, biomarkers ENSG00000287255 and DMXL1-DT have higher feature contribution weights than other biomarkers.

[0027] The present invention has the following beneficial effects: (1) This invention is the first to use a combination of 10 cfRNA molecules, including hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255, and DMXL1-DT, for early screening and auxiliary diagnosis of epithelial ovarian cancer. This combination of 10 cfRNA biomarkers consists of 8 miRNAs and 2 lncRNAs. Using a combined detection method, it can simultaneously reflect multi-level transcriptional regulatory changes related to epithelial ovarian cancer. Compared with single indicators, it has higher diagnostic sensitivity, specificity, and accuracy, and has good clinical application prospects.

[0028] (2) The present invention constructs a detection reagent or kit, screening model and screening system based on the above 10 cfRNA biomarker combinations. The sample source is non-invasive, the results are objective and quantifiable, and it is very suitable for standardized interpretation. It has good clinical translational value.

[0029] (3) This invention breaks through the limitations of traditional linear models such as multifactor logistic regression and innovatively uses the Random Forest algorithm to construct a diagnostic model. This model can deeply capture the complex nonlinear relationships and variable interaction effects among 10 cfRNA biomarkers, effectively avoiding overfitting. Its generalization ability, area under the curve (AUC), sensitivity, and specificity in the independent validation cohort are significantly better than traditional linear models and some other machine learning algorithms, thus significantly improving the sensitivity and specificity of early screening for epithelial ovarian cancer, and has excellent prospects for widespread application. In addition, the computer-aided system established in this invention can not only automatically output risk scores, but also reveal the important driving role of key molecules such as ENSG00000287255 and DMXL1-DT through the feature interpretation module, which has both high diagnostic efficacy and excellent clinical interpretability. Attached Figure Description

[0030] Figure 1 This is a flowchart illustrating the process of using a combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer.

[0031] Figures 2-11 A volcano plot was used to analyze the differential expression of cfRNA in the plasma of ovarian cancer patients and healthy controls. The horizontal axis represents the fold change in difference, and the vertical axis represents statistical significance. This plot is used to show the cfRNA molecules that are significantly upregulated and downregulated in the plasma of ovarian cancer patients compared with healthy controls. Red indicates upregulation, blue indicates downregulation, and the gray area shows no significant difference.

[0032] Figures 12-20 After LASSO screening of different types of cfRNA, the AUC distribution of the ovarian cancer diagnostic model was constructed using logistic regression, random forest, and support vector machine algorithms after 100 iterations. In each sub-graph, the horizontal axis represents the training set, the testing set, and the validation set, and the vertical axis represents the AUC value obtained from 100 repeated modeling iterations.

[0033] Figure 21 The plot shows the receiver operating characteristic (ROC) curves of early ovarian cancer screening models constructed based on different types of cfRNA; where the horizontal axis represents 1-specificity and the vertical axis represents sensitivity.

[0034] Figure 22 This is a risk score distribution plot for candidate diagnostic panel lncRNA+miRNA (N=41) in the discovery and validation cohorts; plot A represents the discovery cohort with a sample size of 73 cases, and plot B represents the validation cohort with a sample size of 20 cases. The red box represents ovarian cancer patients, and the blue box represents healthy controls; the vertical axis represents the ovarian cancer risk score, and the horizontal axis represents the study subjects; the horizontal dashed line represents the pre-set diagnostic cutoff value.

[0035] Figure 23 This is a ranking of the importance of the top 10 cfRNA biomarkers among the candidate diagnostic panel lncRNA+miRNA (N=41) in an early ovarian cancer screening model using the SHAP (Shape Up) algorithm. The horizontal axis represents the mean absolute SHAP value, and the vertical axis represents each cfRNA biomarker.

[0036] Figure 24 The figure shows the ROC curves of an early ovarian cancer screening model based on 10 cfRNA biomarkers. The horizontal axis represents 1-specificity, the vertical axis represents sensitivity, and the gray diagonal line represents the random classification reference line. The blue curve, green curve, and orange dashed line represent the ROC curves of the models constructed by the Support Vector Machine (SVM), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost, xgb) algorithms, respectively.

[0037] Figure 25 The charts (bar charts and heatmaps) show the diagnostic efficacy of different machine learning models based on combinations of 10 cfRNA biomarkers. The evaluation metrics include accuracy, precision, sensitivity, specificity, F1 score, Matthews correlation coefficient (MCC), and area under the curve (AUC).

[0038] Figure 26 A radar chart comparing the diagnostic efficacy of different machine learning models based on combinations of 10 cfRNA biomarkers.

[0039] Figure 27 This is a bar chart comparing the diagnostic efficacy (in terms of AUC value) of the combined diagnostic model and the single biomarker; the red bars represent the combined diagnostic model built based on all 10 cfRNAs (including the random forest model and the linear logistic regression model of this invention), and the blue bars represent the 10 individual cfRNA biomarkers.

[0040] Figure 28 The box plot shows the risk score distribution of the random forest diagnostic model of this invention in the discovery and validation queues for the healthy control group and the ovarian cancer patient group. Detailed Implementation

[0041] The examples given below are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise specified, conditions in the examples are performed under standard conditions or as recommended by the manufacturer. Reagents or instruments whose manufacturers are not specified are all commercially available products.

[0042] Example 1: cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer I. Clinical Sample Collection and Processing 1. Sample collection The study included ovarian cancer patients and healthy female controls. All sample collection was ethically approved, and written informed consent was obtained from the participants.

[0043] Among them, ovarian cancer patients were those diagnosed with ovarian cancer through clinical imaging, serological markers, and / or postoperative histopathology. Healthy controls were women who underwent physical examinations during the same period and had no history of ovarian cancer or other malignant tumors.

[0044] In this embodiment, the study subjects were divided into a discovery cohort and a validation cohort. The discovery cohort contained 73 samples, and the validation cohort contained 20 samples. All samples included those from ovarian cancer patients and healthy controls.

[0045] Inclusion criteria for ovarian cancer patients include: (1) Postoperative pathology confirmed by two senior pathologists that the ovarian epithelial carcinoma included serous carcinoma, mucinous carcinoma and / or clear cell carcinoma. (2) The pathological diagnosis was epithelial ovarian cancer, and there were no other mixed pathological types; (3) The patient had not undergone surgery, radiotherapy, chemotherapy, targeted therapy, immunotherapy or other anti-tumor treatments before the blood sample was collected; (4) The patient's general information is complete; (5) The surgical pathological stage is before stage IIIC.

[0046] The inclusion criteria for the healthy control group included: individuals who underwent health checkups and for whom no evidence of ovarian cancer or other malignant tumors was found during the checkups.

[0047] The sample exclusion criteria include: (1) Those with a history of other cancers; (2) Those with concurrent acute or chronic inflammation; (3) Patients with serious illnesses other than ovarian cancer, especially those with major illnesses whose expected survival is less than 3 years; (4) Individuals with acute infection, severe inflammation, or active autoimmune disease prior to blood collection; (5) Individuals who have recently undergone major surgery, blood transfusion, or immunosuppressive therapy; (6) Those with a history of allogeneic organ transplantation or allogeneic hematopoietic stem cell transplantation; (7) Patients with severe liver or kidney dysfunction or hematological diseases; (8) Plasma samples showing obvious hemolysis, lipemia, or repeated freeze-thaw cycles; (9) Those with incomplete clinical data or unqualified sequencing quality control.

[0048] II. Peripheral blood collection and plasma separation Collect 2-10 mL of peripheral venous blood from the subject and place it in an EDTA anticoagulant blood collection tube. After blood collection, gently invert and mix the blood, avoiding vigorous shaking, and complete plasma separation within 4 hours after blood collection.

[0049] The plasma separation steps are as follows: (1) Centrifuge the anticoagulated blood at 1,600×g for 10 minutes at 4°C; (2) Carefully aspirate the upper plasma layer to avoid inhaling the white blood cell layer; (3) Transfer the obtained plasma to a new RNase-free centrifuge tube; (4) Centrifuge at 1,6000×g for 10 minutes at 4°C to remove residual cells, platelets and cell debris; (5) Collect the supernatant plasma and dispense it into tubes of 200 μL-1 mL each; (6) Store the dispensed plasma at -80°C to avoid repeated freeze-thaw cycles.

[0050] The above processing can effectively reduce cell-derived RNA contamination and improve the stability and reproducibility of plasma cfRNA detection results.

[0051] III. Extraction, library construction, and high-throughput sequencing of plasma cfRNA 1. Plasma cfRNA extraction: cfRNA was extracted from 200 μL of plasma using a cfRNA extraction kit.

[0052] 2. Library construction: The library is constructed using the SLiPiR-seq library construction method. The library fragment length is preferably distributed in the range of 100-300 bp. The number of library amplification cycles can be adjusted according to the amount of cfRNA starting, preferably 12-18 cycles.

[0053] 3. High-throughput sequencing: Sequencing was performed using the Illumina sequencing platform to obtain raw sequencing data.

[0054] IV. Quality control and expression quantification of sequencing data Quality control was performed on the sequencing data, and the quality-controlled clean reads were aligned to the human reference genome. Multiple RNA annotation databases were used to annotate cfRNA types, and the expression of various cfRNAs was then quantified to obtain a standardized expression matrix for subsequent differential analysis and model construction.

[0055] V. Differential cfRNA Screening Differential expression analysis was performed between ovarian cancer patients and healthy controls. The selection criteria for candidate differentially expressed cfRNAs were: P < 0.05; |log2FC| ≥ 0.05. Results are shown below. Figures 2-11 .Depend on Figures 2-11 It can be seen that there is a stable difference in expression between ovarian cancer patients and healthy controls.

[0056] VI. Comparison of modeling performance of different types of cfRNA LASSO feature screening was performed on nine different types of cfRNA, including mRNA, lncRNA, miRNA, piRNA, ysRNA, tsRNA, rsRNA, and snRNA. Figures 12-20 As shown. Figures 12-20 This study aims to demonstrate the differences in classification performance and stability of different types of cfRNA features against ovarian cancer and healthy controls under different machine learning algorithms.

[0057] Different types of cfRNAs were screened using LASSO features, and ovarian cancer diagnostic models were constructed using logistic regression (LR), random forest (RF), and support vector machine (SVM) algorithms, respectively. These models were used to evaluate the ability of different types of cfRNAs and different machine learning algorithms to distinguish between ovarian cancer and healthy controls. Results are shown in [Figure number missing]. Figure 21 .Depend on Figure 21 It can be seen that different types of cfRNAs all have certain discrimination capabilities. Among them, miRNA and lncRNA-related features are relatively stable in the model and are suitable as the focus of subsequent biomarker screening.

[0058] Example 2: Screening for the optimal cfRNA combination for epithelial ovarian cancer based on sensitivity and specificity comparison This embodiment systematically compares the classification efficacy of different cfRNAs used alone or in combination. Evaluation metrics primarily include sensitivity and specificity in the discovery and validation cohorts, and are further screened based on the overall performance under different machine learning algorithms.

[0059] I. Modeling and screening of different cfRNA combinations Based on the aforementioned plasma cfRNA sequencing and differential analysis results, to obtain a model with better overall performance, single or combined combinations of different types of cfRNA were constructed. The cfRNA types included nine types: lncRNA, miRNA, mRNA, piRNA, tsRNA, rsRNA, ysRNA, snRNA, and snoRNA.

[0060] By performing non-empty combinations on the above 9 cfRNA types, a total of 511 cfRNA combination methods were formed.

[0061] II. Calculation of Sensitivity and Specificity under Different Combinations Ovarian cancer screening models were constructed using logistic regression (LR), random forest (RF), support vector machine (SVM), and extreme gradient boosting (XGBoost, XGB) algorithms for each candidate cfRNA combination. The sensitivity and specificity of each model in the detection and validation cohorts were calculated.

[0062] III. Screening Results of Different cfRNA Combinations A comparison of all candidate combinations revealed significant differences in the performance of different cfRNA combinations in ovarian cancer screening.

[0063] 1. Among single cfRNA types, miRNA, tsRNA, and rsRNA exhibited high screening efficacy. Specifically, the miRNA combination, using logistic regression and support vector machine algorithms, achieved a sensitivity of 1.000 and a specificity of 1.000 in the detection cohort, and also maintained good classification ability in the validation cohort; lncRNA, used alone, also demonstrated good diagnostic performance.

[0064] 2. In some complex multi-type combinations, although the sensitivity and specificity were high in the cohort, the specificity of some models, such as snRNA and miRNA+mRNA, decreased significantly in the validation cohort, indicating the risk of overfitting, which is not conducive to the stable application of the model.

[0065] 3. The combined use of lncRNA and miRNA showed good overall stability under various algorithms, especially maintaining high sensitivity in the discovery cohort and good specificity in the validation cohort, demonstrating good generalization ability. Figure 22 As shown. Figure 22 Individuals with a diagnostic score higher than the diagnostic cutoff value were classified as high-risk individuals for ovarian cancer, while those with a score lower than the diagnostic cutoff value were classified as low-risk individuals for ovarian cancer. Figure 22 The results showed that the overall risk score of ovarian cancer patients was higher than that of healthy controls, indicating that the model constructed in this invention has good diagnostic and discriminative capabilities.

[0066] 4. Compared with complex combinations that include more cfRNA types, the lncRNA+miRNA combination maintains high classification performance while having a simpler feature composition, lower detection cost, and is more suitable for establishing a clinically translatable early ovarian cancer screening model.

[0067] After comprehensively comparing the sensitivity, specificity, and average performance of each combination in the discovery and validation queues under four machine learning algorithms, the lncRNA+miRNA combination was determined as the preferred cfRNA combination for subsequent feature screening and model construction.

[0068] IV. Explanation of Selection Criteria When screening for the optimal cfRNA combination, this invention does not rely solely on a single algorithm, a single cohort, or a single metric, but rather considers the following factors comprehensively: 1. Discover the sensitivity and specificity in the queue; 2. Verify the sensitivity and specificity in the queue; 3. Consistency of results among different machine learning algorithms; 4. Stability of the combination across different datasets; 5. Number of features and model complexity; 6. Feasibility of implementing subsequent clinical testing.

[0069] Based on the above comprehensive evaluation, the lncRNA+miRNA combination exhibits superior overall balance compared to single cfRNA types and more complex multi-type combinations, ensuring both high sensitivity and good specificity and validation stability. Therefore, the lncRNA+miRNA combination has been selected as the foundation for further screening of core biomarkers and constructing an early ovarian cancer screening model in this invention.

[0070] Example 3: Screening for early screening biomarkers for ovarian cancer based on SHAP feature contribution analysis This embodiment performs importance analysis and feature compression on candidate features in the lncRNA+miRNA combination to screen for core cfRNA biomarkers that can stably distinguish ovarian cancer patients from healthy controls.

[0071] I. Establishment of Candidate Feature Set Candidate features were extracted from the lncRNA+miRNA combination determined in Example 2 to construct an input feature matrix for ovarian cancer classification. The candidate features included lncRNA and miRNA molecules with significant differential expression, stable expression, and a certain detection rate.

[0072] Candidate features are first standardized to reduce the impact of differences in the expression levels of different molecules on model training and feature interpretation.

[0073] II. SHAP Feature Analysis Method To evaluate the contribution of each candidate cfRNA to the model's classification results, this embodiment employs the SHAP (Shapley Additive Explanations) method for model interpretation and analysis. Specifically, lncRNA+miRNA candidate features are input into an ensemble learning diagnostic model using Random Forest for training, and then the SHAP value corresponding to each feature is calculated based on the trained model. The SHAP value reflects the marginal contribution of a feature to the prediction result of a single sample, while the mean absolute SHAP value reflects the overall importance of that feature across all samples.

[0074] The candidate features were sorted in descending order based on their mean absolute SHAP values, and the core features that contributed significantly to the classification of ovarian cancer were gradually selected based on changes in model performance.

[0075] III. Screening of core cfRNA biomarkers Through SHAP feature contribution analysis, candidate features in the lncRNA+miRNA combination were sorted and compressed, and finally 10 cfRNA biomarkers were selected as the core input features of the early ovarian cancer screening model. The 10 cfRNA biomarkers include: hsa-miR-3184-3p, hsa-miR-18a-3p, ENSG00000287255, hsa-miR-135b-5p, DMXL1-DT, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, and hsa-miR-221-5p.

[0076] Among them, hsa-miR-3184-3p, hsa-miR-18a-3p, and ENSG00000287255 are features with high contribution; DMXL1-DT and hsa-miR-135b-5p also show high importance; although the remaining markers have relatively low contribution, they can improve the overall stability and classification ability of the model when combined with high-contribution features.

[0077] IV. SHAP Analysis Results like Figure 23 As shown, after sorting the above 10 cfRNA markers by their average absolute SHAP values, the following can be observed: 1. The hsa-miR-3184-3p has the highest mean absolute SHAP value, indicating that it contributes the most to the model output. 2. hsa-miR-18a-3p and ENSG00000287255 contributed the second most; 3. DMXL1-DT and hsa-miR-135b-5p are also of high importance; 4. hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p and hsa-miR-221-5p also make auxiliary contributions to the model.

[0078] The above results indicate that the 10 cfRNA biomarkers screened can reflect the expression differences between ovarian cancer patients and healthy controls at different levels, and have good combined discriminative value.

[0079] Example 4: Construction of an early screening model for epithelial ovarian cancer based on 10 cfRNA biomarkers I. Model Input Data Construction Expression values ​​of the 10 cfRNA biomarkers were extracted from each sample to construct the model input matrix. These expression values ​​can be normalized expression values ​​obtained from high-throughput sequencing.

[0080] Ovarian cancer patient samples were labeled as positive and healthy control samples as negative, and a classification model training dataset was constructed accordingly.

[0081] II. Model Building Methods Machine learning algorithms were used to jointly model the 10 cfRNA biomarkers. Applicable algorithms included: logistic regression (LR), support vector machine (SVM), extreme gradient boosting (XGBoost), and random forest (RF).

[0082] The expression values ​​of 10 cfRNA markers were used as input features, and the sample category was used as the output label. The model was trained in a discovery queue and tested in a validation queue.

[0083] III. Risk Score Output After the model is trained, an ovarian cancer risk score is generated for each subject. The higher the risk score, the higher the likelihood that the subject will have epithelial ovarian cancer.

[0084] Specifically, the expression values ​​of 10 cfRNA markers in the subjects (i.e., hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255, and DMXL1-DT) were used as feature vectors and input into a trained Random Forest model. The model performed nonlinear mapping and probability calculations through multiple decision trees integrated within it, and took the average probability of all decision trees predicting a positive example (i.e., cancer) to output a value between 0 and 1, which was used as the final risk score.

[0085] This embodiment establishes a random forest ovarian cancer early screening model based on 10 cfRNA biomarkers, overcoming the limitation of traditional linear summation formulas in failing to capture complex variable interactions, thus laying the foundation for subsequent model performance evaluation and sample classification validation. The model can be used to objectively predict the risk of ovarian cancer in subjects and can be further applied to early ovarian cancer screening and auxiliary diagnosis.

[0086] Example 5: ROC performance evaluation based on 10 cfRNA models I. Evaluation Methods An ovarian cancer classification model was constructed using the 10 cfRNA biomarkers selected in Example 3 as input features, employing Support Vector Machine (SVM), Logistic Regression (LR), and Extreme Gradient Boosting (XGBoost, XGB) algorithms respectively.

[0087] ROC curves were plotted on the prediction results of each model, and the area under the curve (AUC) was calculated to evaluate the model's ability to distinguish between ovarian cancer patients and healthy controls.

[0088] In the ROC curve, the horizontal axis represents 1 - specificity; the vertical axis represents sensitivity; the gray diagonal line represents the random classification reference line; the larger the AUC value, the better the model's classification performance.

[0089] II. ROC Analysis Results like Figure 24As shown, the early ovarian cancer screening model constructed based on 10 cfRNA biomarkers exhibited high diagnostic efficacy under different algorithms. Specifically, the AUC of the support vector machine model was 0.94; the AUC of the logistic regression model was 0.97; and the AUC of the extreme gradient boosting model was 0.96. Among them, the logistic regression model had the highest AUC, indicating that it performed best in classification on the combination of these 10 cfRNA biomarkers; the support vector machine model and the extreme gradient boosting model also showed high discriminative ability. These results indicate that the model constructed by combining the 10 cfRNA biomarkers selected in this invention has high sensitivity and specificity, and can effectively distinguish between ovarian cancer patients and healthy controls. This combination of 10 cfRNA biomarkers has good value for early ovarian cancer screening and has potential for auxiliary diagnosis.

[0090] Example 6: Comparison of Diagnostic Efficacy of Different Machine Learning Algorithms and Single / Joint Biomarker Models I. Establishment of Different Algorithm Models Using the 10 cfRNA biomarkers selected in Example 3 as input features and the disease status of the samples as output labels, four classification models were established: Support Vector Machine (SVM) model, Logistic Regression (LR) model, Extreme Gradient Boosting (XGBoost) model, and Random Forest (RF) model. Each model was trained in the discovery queue and independently validated in the validation queue.

[0091] II. Comparison of Diagnostic Efficiency of Different Machine Learning Algorithms Based on the four types of models constructed above, their diagnostic efficacy in the discovery and validation cohorts was evaluated. Evaluation metrics included accuracy, precision, sensitivity, specificity, F1 score, Matthews correlation coefficient (MCC), and area under the curve (AUC). The results are as follows: Figure 25 and 26 As shown.

[0092] like Figure 25 and 26 The bar charts, heatmaps, and radar charts in the data comprehensively show that among the four algorithms, the Random Forest (RF) model exhibits the best classification performance in both the discovery and validation queues. Particularly in the validation queue, the RF model maintains extremely high sensitivity, specificity, and AUC values ​​without significant performance degradation; in contrast, the LR, SVM, and XGBoost models showed fluctuations or declines in some metrics. This indicates that the diagnostic model built using the Random Forest algorithm has better generalization ability and classification stability.

[0093] III. Performance Comparison of Single Biomarker and Combined Diagnostic Models To further verify the necessity of multi-target joint detection and the superiority of the selected algorithm, this embodiment compares the AUC values ​​of 10 single cfRNA biomarkers with the multi-marker model. The results are as follows: Figure 27 As shown.

[0094] like Figure 27 As shown, the AUC values ​​of single cfRNA biomarkers range from 0.240 to 0.930 (the highest being ENSG00000287255 and DMXL1-DT, with an AUC of 0.930). When 10 biomarkers are combined using a traditional logistic regression model, the AUC value is 0.840; while when multiple targets are combined using the random forest (RF) algorithm preferred in this invention, the model's AUC value reaches 1.000.

[0095] IV. Results Explanation In summary, the diagnostic efficacy of the 10-cfRNA biomarker combination provided by this invention is significantly superior to that of a single nucleic acid biomarker. Furthermore, compared to traditional linear regression models and other machine learning algorithms, the joint model constructed based on the random forest (RF) algorithm can more effectively capture the complex characteristics among these 10 cfRNA molecules, significantly improving the accuracy and robustness of early screening for epithelial ovarian cancer, and possessing extremely high clinical application translational value.

[0096] Example 7: Ovarian Cancer Risk Score Assessment and Clinical Discrimination Validation Based on Random Forest Model I. Calculation and Distribution of Risk Scores Based on the optimal algorithm established in Example 6—the Random Forest (RF) joint diagnostic model—all subject samples (including healthy controls and ovarian cancer patients) in the Discovery and Validation cohorts are scored, and a Random Forest risk score (RF riskscore) is output for each sample. The score ranges from 0 to 1, and the results are visualized and analyzed using a box plot with overlaid scatter points, such as... Figure 28 As shown.

[0097] II. Results Analysis like Figure 28 As shown in the figure, the horizontal dashed line in the middle of the scatter box plot represents the optimal cutoff value that distinguishes between "high risk" and "low risk".

[0098] In the Discovery cohort: the median risk score of the healthy control group (blue box) was extremely low (close to 0.1), with the vast majority of samples falling below the cutoff value; while the median risk score of the ovarian cancer patient group (red box) was extremely high (close to 0.95), with all patient samples falling above the cutoff value. There was a highly statistically significant difference in risk scores between the two groups (p = 1.3e-12).

[0099] In the independent validation cohort: despite being an externally blinded sample, the model still demonstrated excellent discriminative ability. Risk scores in the healthy control group were predominantly below the cutoff value, while risk scores in the ovarian cancer patient group were predominantly above the cutoff value. The scores between the two groups also showed extremely high statistical significance (p = 0.00018).

[0100] III. Conclusion The above results fully demonstrate that the diagnostic model constructed based on 10 cfRNA biomarkers and combined with the random forest algorithm can output an intuitive and highly reliable "risk score." This score can significantly distinguish ovarian cancer patients from healthy individuals, and no serious overfitting was observed in either internal training or independent external validation. This scoring system can be directly used as a non-invasive and accurate quantitative indicator for early clinical screening and auxiliary diagnosis of ovarian cancer.

[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A combination of cfRNA biomarkers for early screening and auxiliary diagnosis of epithelial ovarian cancer, characterized in that, The cfRNA biomarker combination consists of 8 miRNAs and 2 lncRNAs, namely: hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255, and DMXL1-DT.

2. The use of the cfRNA biomarker combination of claim 1 in the preparation of products for epithelial ovarian cancer screening, auxiliary diagnosis, high-risk population identification, efficacy evaluation, follow-up monitoring or recurrence early warning.

3. A diagnostic reagent for screening or assisting in the diagnosis of epithelial ovarian cancer, characterized in that, Includes specific primers, probes, capture probes, or combinations thereof for detecting the expression level of the cfRNA biomarker combination of claim 1.

4. A kit for screening or assisting in the diagnosis of epithelial ovarian cancer, characterized in that, It includes the detection reagent as described in claim 3.

5. A risk assessment model for early screening of epithelial ovarian cancer, characterized in that, The model is an ensemble machine learning model trained based on the random forest algorithm. The random forest model includes 500 decision trees with a maximum depth of 10. The model is configured to: acquire the expression levels of the 10 cfRNA biomarkers described in claim 1 in the sample to be tested as input variables; The results are evaluated sequentially by the 500 decision trees, and the outputs of each decision tree are integrated and weighted to output the epithelial ovarian cancer risk score of the sample to be tested. When the risk score is greater than the diagnostic cutoff value, it is considered a high risk of epithelial ovarian cancer; when the risk score is less than or equal to the diagnostic cutoff value, it is considered a low risk of epithelial ovarian cancer.

6. The risk assessment model according to claim 5, characterized in that, Diagnostic cutoff values ​​are obtained through at least one of the following methods: (1) The threshold corresponding to the maximum point of the Youden exponent on the ROC curve; (2) The dividing point determined under preset sensitivity or specificity conditions; (3) The optimal classification threshold obtained through cross-validation or independent external validation queues.

7. The risk assessment model according to claim 5 or 6, characterized in that, The diagnostic cutoff value is 0.

5.

8. The method for constructing the risk assessment model according to any one of claims 5-7, characterized in that, Includes the following steps: (1) Data collection: Obtain plasma cfRNA expression value matrix and corresponding clinical classification labels of samples from patients diagnosed with epithelial ovarian cancer and healthy controls; (2) Differential expression analysis: Differential analysis was performed on the expression value matrix to screen candidate cfRNAs that meet the conditions of P value < 0.05 and |log2FC| ≥ 0.05 between the two groups; (3) Feature screening: The candidate cfRNAs are quantified based on the SHAP algorithm, the average absolute SHAP value of each candidate cfRNA is calculated, and the candidate cfRNAs are sorted in descending order according to the average absolute SHAP value. The top 10 cfRNAs are selected as core cfRNA markers. The 10 core cfRNA markers are hsa-miR-3184-3p, hsa-miR-18a-3p, hsa-miR-135b-5p, hsa-miR-486-3p, hsa-miR-320a-3p, hsa-miR-103a-1-5p, hsa-miR-181a-5p, hsa-miR-221-5p, ENSG00000287255 and DMXL1-DT. (4) Model construction: Using the expression data of the 10 core cfRNA markers as input features and the corresponding clinical classification labels as output targets, the random forest algorithm is used to train and generate an integrated model containing multiple decision trees; the model performs nonlinear mapping and probability calculation through the internally integrated multiple decision trees, takes the average probability of all decision trees predicting positive examples, and outputs a value between 0 and 1 as the final risk score to obtain a risk assessment model for early screening of epithelial ovarian cancer.

9. A system for screening or assisting in the diagnosis of epithelial ovarian cancer, characterized in that, include: The data acquisition module is used for sequencing or nucleic acid amplification of the combination of the 10 cfRNA biomarkers in the subject sample to be tested; The data preprocessing module is used to perform missing value processing, normalization processing and / or standardization processing on the sequencing or nucleic acid amplification data to generate quantitative expression profile features of the cfRNA biomarker combination; The model computation module is used to input quantitative expression profile features into the risk assessment model according to any one of claims 5-7, and calculate and output the risk score or predicted probability of epithelial ovarian cancer through the random forest algorithm; The threshold determination module is used to output a determination conclusion that the subject under test is at high or low risk of epithelial ovarian cancer based on the comparison result between the risk score or predicted probability and the preset diagnostic cutoff value. The results output module is used to display, store, print, and / or transmit the judgment conclusions and risk assessment reports.

10. The system according to claim 9, characterized in that, The system is also equipped with a feature interpretation module, which is used to extract and display the contribution of each molecular feature in the 10 cfRNA biomarkers to the risk score of the current sample based on SHAP value or Gini importance. Among them, biomarkers ENSG00000287255 and DMXL1-DT have higher feature contribution weights than other biomarkers.