Computer implemented method for non-invasive stepwise prediction of cancer risk, and system and storage medium

By combining a stepwise screening approach with biomarkers and NGS testing, the problems of high false positive rates and high costs in existing technologies have been solved, achieving highly specific and low-cost cancer screening, early detection and reducing the screening burden, and improving public acceptance of screening and overall health.

WO2025256121A1PCT designated stage Publication Date: 2025-12-18SEEKIN INC SHENZHEN CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/144586
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-14
Filing Date
2024-12-31
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

Existing cancer screening methods suffer from high false positive rates and high costs, leading to a waste of medical resources and a psychological and economic burden on individuals, making them difficult to apply in large-scale screening.

Method used

A stepwise approach to predicting cancer risk is adopted. First, the cancer signal score is determined by the level of biomarkers in blood samples. After preliminary screening, NGS testing is performed on positive subjects. By combining machine learning models and biomarker combinations, the specificity of screening is improved and the false positive rate is reduced.

Benefits of technology

It has achieved highly specific and low-cost cancer screening, enabling early detection of cancer, reducing unnecessary testing and treatment burdens, and improving public acceptance and overall health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024144586_18122025_PF_FP_ABST
    Figure CN2024144586_18122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a computer implemented method for non-invasive stepwise prediction of cancer risks, and a system and a storage medium. The aforementioned method comprises: (1) determining a cancer signal score on the basis of the level of a biomarker in a blood sample from a subject; (2) comparing the cancer signal score with a predetermined threshold value to determine a positive subject via preliminary screening; and (3) on the basis of the positive subject from the preliminary screening, further performing cancer risk determination and cancer type prediction by using an NGS method, wherein the biomarker includes at least one selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.
Need to check novelty before this filing date? Find Prior Art

Description

Computer-implemented method, system and storage medium for non-invasive, stepwise prediction of cancer risk

[0001] Priority information

[0002] No TECHNICAL FIELD

[0003] The present application relates to the field of bioinformatics, in particular, to a computer-implemented method, system and storage medium for non-invasive, stepwise prediction of cancer risk. BACKGROUND

[0004] Cancer is a global public health problem that seriously endangers human health. According to the latest global cancer data released by the International Agency for Research on Cancer (IARC) of the World Health Organization in 2020, there were 1929 million new cancer cases worldwide in 2020, of which 457 million new cancer cases occurred in China, accounting for 23.7% of the world. China is the world's largest population, and the number of new cancer cases far exceeds that of other countries in the world. With the progress of medical technology, although the treatment of various cancers has made some progress, more than two-thirds of patients have been diagnosed as advanced at the time of diagnosis due to the lack of obvious early symptoms, which greatly increases the difficulty of treatment, resulting in high mortality and poor prognosis. For cancer patients, their clinical stage is closely related to the prognosis of treatment. Clinical studies have shown that the 5-year survival rate of lung cancer patients in the IA stage is about 60%, while the 5-year survival rate of lung cancer patients in the II-IV stage is reduced to 40%. Early screening is also crucial for the successful treatment of colorectal cancer. The 5-year survival rate of stage I and II is more than 80%, and once distant metastasis occurs, it drops to about 10%.

[0005] To achieve universal cancer screening, two aspects need to be considered: one is the problem of false positives, and the other is the cost of screening. A high false positive rate will lead to a large waste of medical resources and unnecessary psychological and economic burden on individuals. Screening products must have a high enough specificity. Since the incidence of cancer in the population is about one percent, a low specificity will result in a large number of false positive results. This not only brings unnecessary psychological pressure and economic burden to individuals, but also causes a large waste of medical resources, because each false positive result requires further diagnostic procedures such as imaging and biopsy, which are both expensive and time-consuming. A screening method with high specificity can effectively reduce the number of false positives, thereby reducing these additional burdens and resource consumption. Therefore, specificity is a key factor in choosing and promoting cancer screening methods. Only in this way can we ensure that the screening project is not only economically efficient, but also widely applicable to the population, improving the overall effectiveness of screening and the health level of the public. Large-scale screening inevitably puts a huge financial pressure on the medical system, which is a key factor that prevents even high-income countries from conducting nationwide screening.

[0006] Therefore, in order to promote cancer screening nationwide, the screening method must be low in price, easy to operate, and acceptable and usable by everyone. Moreover, the product must have high enough specificity. SUMMARY

[0007] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a high-specificity, low-cost cancer screening method.

[0008] Specifically, the technical solution of the present application is as follows:

[0009] In the first aspect, the present application provides a computer-implemented method for step-by-step prediction of cancer risk. According to an embodiment of the present application, the method comprises: (1) determining a cancer signal score based on the level of a biomarker in a blood sample of a subject; (2) comparing the cancer signal score with a predetermined threshold to determine a preliminary screening positive subject; (3) further using NGS method to determine cancer risk and cancer type prediction based on the preliminary screening positive subject; wherein the biomarker comprises at least one selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

[0010] In some examples of the present application, the foregoing method effectively improves the specificity of screening, reduces the false positive rate, and reduces unnecessary follow-up examinations and waste of medical resources by step-by-step cancer risk detection of the subject. The preliminary screening uses low-cost and easy-to-operate blood detection, which is easy to popularize on a large scale, and further NGS detection is performed for the preliminary screening positive subjects, thereby optimizing the screening process and reducing the overall cost. Using multiple biomarkers to cover different types of cancer improves the universality and effectiveness of screening, which helps early detection and intervention, and improves patient survival rate and prognosis. In addition, the dynamic (time series) change of individualized risk assessment and cancer risk value based on NGS can timely and accurately evaluate the Treatment Response of cancer treatment, and improve the treatment effect and patient management level.

[0011] In the second aspect, the present application provides a cancer treatment monitoring method. According to an embodiment of the present application, the method comprises: predicting cancer risk based on the method of the first aspect; and monitoring the object of cancer treatment using the cancer risk.

[0012] In some examples of the present application, the foregoing method combines cancer risk prediction and treatment monitoring, providing a comprehensive solution from screening, diagnosis to treatment monitoring and cancer recurrence monitoring, with advantages of comprehensive monitoring, personalized treatment adjustment, early detection of recurrence, dynamic monitoring, reduction of over-treatment, etc. By regularly detecting the level of biomarkers, the treatment plan can be evaluated and adjusted in real time, improving the accuracy and effectiveness of treatment, avoiding unnecessary treatment burden and side effects, and improving the quality of life of patients.

[0013] In a third aspect, the present application provides a computer-implemented device for stepwise prediction of cancer risk. According to an embodiment of the present application, the device comprises: a cancer signal score determination unit configured to determine a cancer signal score based on the level of biomarkers in a blood sample of a subject; a preliminary screening positive subject determination unit configured to compare the cancer signal score with a predetermined threshold to determine a preliminary screening positive subject; and a cancer type determination unit configured to further determine cancer risk and predict cancer type based on the preliminary screening positive subject using a NGS method, wherein the biomarkers comprise at least one selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

[0014] In some examples of the present application, the foregoing device has the advantages of high efficiency, high specificity, sufficient sensitivity and low cost for cancer risk screening.

[0015] In a fourth aspect, the present application provides a computer program product. According to an embodiment of the present application, the computer program product comprises computer instructions, and when part or all of the computer instructions are executed on a computer, the computer-implemented method for stepwise prediction of cancer risk as described in the first aspect of the present application is executed.

[0016] In a fifth aspect, the present application provides a computing device. According to an embodiment of the present application, the computing device comprises a processor and a memory, wherein the memory is configured to store a computer program, and the processor is configured to execute the computer program to implement the computer-implemented method for stepwise prediction of cancer risk as described in the first aspect of the present application.

[0017] In a sixth aspect, the present application provides a computer-readable storage medium. According to an embodiment of the present application, the computer-readable storage medium stores computer instructions or programs, and when the computer instructions or programs are executed on a computer, the computer-implemented method for stepwise prediction of cancer risk as described in the first aspect of the present application is executed.

[0018] The foregoing system and computer-readable storage medium implement a computer-implemented method of predicting cancer risk through automatic execution of computer instructions, achieving high efficiency and automation, and improving efficiency and accuracy. Secondly, the characteristics of the instructions make the foregoing method highly consistent and reliable in complex experimental scenarios.

[0019] Additional aspects and advantages of the present application will be in part apparent and in part pointed out hereinafter. BRIEF DESCRIPTION OF DRAWINGS

[0020] The foregoing and / or additional aspects and advantages of the present application will become apparent and be readily appreciated when considered in connection with the following description of embodiments, taken in conjunction with the accompanying drawings.

[0021] FIG. 1 is a schematic diagram of a computer-implemented device for stepwise prediction of cancer risk according to embodiments of the present application;

[0022] FIG. 2 is a schematic diagram of an electronic device according to embodiments of the present application;

[0023] FIG. 3 is a schematic diagram of model integration according to an embodiment of the present application;

[0024] FIG. 4 is a schematic diagram of an outlier analysis method according to an embodiment of the present application;

[0025] FIG. 5 is a comparison of AUC before and after optimization according to an embodiment of the present application;

[0026] FIG. 6 is a schematic diagram of specificity results of secondary detection according to an embodiment of the present application;

[0027] FIG. 7 is a schematic diagram of terminal feature sequences according to an embodiment of the present application;

[0028] FIG. 8 is a schematic diagram of multi-dimensional cancer screening performance evaluation results according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] Embodiments of the present application are described in detail below with reference to the accompanying drawings, examples of which are shown in the drawings, wherein the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be understood as a limitation of the present application.

[0030] It should be noted that the terms "first", "second" are used only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. Further, in the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0031] In this paper, the term "biomarker" is used as tumor marker or protein marker unless otherwise specified.

[0032] In this paper, the term "screening positive subject" refers to the subject selected for secondary detection, not equivalent to cancer patients unless otherwise specified.

[0033] In this paper, the term "smoking status" not only refers to whether or not smoking, but also includes whether or not smoking, and the specific smoking frequency (such as the number of packs per year), the length of smoking (such as the number of years), and the length of time if quitting smoking.

[0034] Current clinical multi-cancer screening requires high specificity and low screening cost, while existing NGS-based detection costs are too high to be applied in large-scale universal screening. To solve this problem, the present application provides a computer-implemented method for stepwise prediction of cancer risk, a cancer treatment monitoring method, a computer-implemented device for stepwise prediction of cancer risk, a computer program product, a computing device and a computer readable storage medium. They are described as follows respectively:

[0035] Computer-implemented method for stepwise prediction of cancer risk

[0036] In one aspect of the present application, a computer-implemented method for stepwise prediction of cancer risk is provided, which comprises: (1) determining a cancer signal score based on the level of biomarkers in the blood sample of the subject; (2) comparing the cancer signal score with a predetermined threshold to determine a screening positive subject; (3) further determining the risk of cancer and predicting the cancer type by using NGS method based on the screening positive subject; wherein the biomarkers include at least one selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

[0037] The aforementioned method uses a cancer-specific biomarker protein assay in the preliminary screening stage, which has low detection costs and low environmental and instrument requirements, and does not require professional technical personnel. Therefore, based on this method, a larger and more extensive population, especially low-income populations in rural areas, can be greatly covered, medical costs are saved, and the public acceptance and coverage of the screening project are improved. Through preliminary screening, most normal subjects (preliminary screening negative) are excluded. For positive subjects (<20%), a high-specificity NGS method is used for further cancer detection, and cancer prediction is performed for subjects with positive detection. This step significantly reduces the false positive rate of cancer prediction, reduces unnecessary follow-up diagnosis-related detection and costs, reduces the physical burden and psychological stress of patients, saves medical costs, and improves the public acceptance and trust of the screening project. At the same time, cancer prediction can guide the subsequent diagnosis of subjects, and accelerate the diagnosis process. In summary, this step-by-step detection of cancer, based on certain medical resources, can greatly expand the coverage of the screening population, while the high specificity reduces the unnecessary physical burden and psychological stress of subjects, and accelerates the diagnosis of cancer, thereby achieving early detection and treatment of cancer and improving the overall public health level.

[0038] In some embodiments, the biomarkers include: AFP, CA125, CA15-3, CA19-9, CEA, and CYFRA 21-1. In some embodiments, the biomarkers include: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, and CYFRA 21-1. In some embodiments, the biomarker panel includes AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA.

[0039] In some cases, the subject is male, and the biomarker panel includes: AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA. In other cases, the subject is female, and the biomarker panel includes: AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA. In some cases, specific markers are selected for specific cancer screening, such as lung cancer screening, and the biomarker panel includes: CEA, CYFRA 21-1, ProGRP, and SCCA.

[0040] In some embodiments, the aforementioned step (1) is performed by: i) selecting a plurality of parameters to input into a first machine learning model, the aforementioned plurality of parameters comprising at least the quantitative level of the aforementioned biomarker; ii) selecting at least one machine learning algorithm to train the aforementioned first machine learning model; and iii) determining the aforementioned cancer signal score.

[0041] In some embodiments, the aforementioned plurality of parameters comprises at least one clinical parameter. The aforementioned clinical parameter comprises at least one of age, gender, and smoking status.

[0042] In some embodiments, the aforementioned plurality of parameters further comprises at least one of X-ray imaging, mammography, computed tomography, and magnetic resonance imaging.

[0043] In some embodiments, the quantitative level of the aforementioned biomarker is standardized by a modified Z-Score, the aforementioned modified Z-Score is obtained by dividing the difference between an observed value and a median by the median absolute deviation.

[0044] In some embodiments, the aforementioned machine learning algorithm comprises GLM, GBM, RF, and SVM.

[0045] In some embodiments, the aforementioned first machine learning model is trained using GLM.

[0046] In some embodiments, the aforementioned first machine learning model is obtained by training using at least two of GLM, GBM, RF, and SVM.

[0047] In some embodiments, the aforementioned method comprises applying GLM to an ensemble of at least two machine learning algorithms.

[0048] In some embodiments, a generalized linear model (GLM) is used. GLM is a flexible extension of ordinary linear regression. GLM allows linear models to be extended by relating the response variable through a link function and allowing the size of the variance for each measurement to be a function of its predicted value. Generalized linear models are constructed in a way that unifies various other statistical models, including linear regression, logistic regression, and Poisson regression. In some embodiments, the maximum likelihood estimation (MLE) of the model parameters is performed using iteratively reweighted least squares. MLE is widely applied and is the default method of many statistical computing software packages. Other methods have also been developed, including Bayesian regression methods and least squares fitting with variance-stabilizing responses.

[0049] In certain embodiments, Gradient Boosting Machines (GBMs) are used. Gradient Boosting Machines (GBMs) are a machine learning technique for regression and classification tasks that builds a series of weak prediction models (typically decision trees) to enhance the performance of the overall model. The core idea of GBMs is to train a set of weak learners iteratively, with each new learner trying to correct the errors of its predecessor, thereby gradually improving the accuracy of the model. GBMs train in stages and allow optimization of arbitrary variable loss functions, making them highly flexible and adaptable for cancer risk prediction.

[0050] In certain embodiments, the methods described herein involve gradient boosting. Gradient boosting is a machine learning technique for regression and classification tasks. It provides a prediction model in the form of an ensemble of weak prediction models, which are typically decision trees. In certain embodiments, gradient boosting tree models are built in stages like other boosting methods, but it generalizes the other methods by allowing optimization of arbitrary variable loss functions. In certain embodiments, the aforementioned methods herein involve random forests (RFs). Random forests are an ensemble learning method for classification, regression, and other tasks that operates by building a large number of decision trees at training time. For classification tasks, the output of a random forest is the class chosen by the majority of the trees. For regression tasks, the average prediction of the individual decision trees is returned. Random forests correct the problem of decision trees tending to overfit their training set.

[0051] In certain embodiments, the methods described herein involve support vector machines (SVMs, also known as support vector networks). SVMs are supervised learning models with a related learning algorithm that analyzes data for classification and regression analysis. Given a set of training examples (each marked as belonging to one of two categories), an SVM training algorithm builds a model that assigns new examples into one category or the other, making it a non-probabilistic binary linear classifier (although methods such as Platt scaling exist to use SVMs in a probabilistic classification setting). An SVM maps training examples into spaces in which they can be classified using a linear boundary. Then, new examples can be mapped into that same space and predicted to fall on the side of one category or the other, based on which side of the boundary they land on. In addition to performing a linear classification, SVMs can efficiently perform a non-linear classification using the so-called kernel trick, implicitly mapping their inputs into high-dimensional feature spaces.

[0052] In certain embodiments, the foregoing methods herein include one or more ensemble models, which improve overall predictive performance by combining models trained using different algorithms, parameters, and / or training data. Ensemble models do not rely on a single model, but rather leverage the diversity of multiple models to enhance the accuracy and robustness of predictions. Among other things, the foregoing accuracy is calculated by dividing the sum of true positive and true negative samples by the total number of samples.

[0053] In certain embodiments, the training dataset includes at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 blood samples from subjects. In certain embodiments, the true positive cancer patients comprise at least 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, or 80% of the number of samples. In certain embodiments, the training dataset does not exceed 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 blood samples from subjects.

[0054] In certain embodiments, the foregoing methods herein can achieve a sensitivity of at least 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, or 60%. In certain embodiments, the sensitivity is between about 30% to about 60%, about 30% to about 55%, about 30% to about 50%, about 30% to about 45%, about 30% to about 40%, about 30% to about 35%, about 35% to about 60%, about 35% to about 55%, about 35% to about 50%, about 35% to about 45%, about 35% to about 40%, about 40% to about 60%, about 40% to about 55%, about 40% to about 50%, about 40% to about 45%, about 45% to about 60%, about 45% to about 55%, about 45% to about 50%, about 50% to about 60%, about 50% to about 55%, or about 55% to about 60%.

[0055] In certain embodiments, the foregoing predetermined threshold is selected from determining upper and lower thresholds of cancer signal scores based on a predetermined specificity range.

[0056] In certain embodiments, the foregoing cancer signal score being within the foregoing predetermined range of cancer signal scores is an indication that the foregoing subject is positive for screening.

[0057] In certain embodiments, the foregoing methods herein can achieve a specificity of at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 95%. In certain embodiments, the specificity is between about 50% to about 95%, about 50% to about 90%, about 50% to about 85%, about 50% to about 80%, about 50% to about 75%, about 50% to about 70%, about 50% to about 65%, about 50% to about 60%, about 50% to about 55%, about 55% to about 95%, about 55% to about 90%, about 55% to about 85%, about 55% to about 80%, about 55% to about 75%, about 55% to about 70%, about 55% to about 65%, about 55% to about 60%, about 60% to about 95%, about 60% to about 90%, about 60% to about 85%, about 60% to about 80%, about 60% to about 75%, about 60% to about 70%, about 60% to about 65%, about 65% to about 95%, about 65% to about 90%, about 65% to about 85%, about 65% to about 80%, about 65% to about 75%, about 65% to about 70%, about 70% to about 95%, about 70% to about 90%, about 70% to about 85%, about 70% to about 80%, about 70% to about 75%, about 75% to about 95%, about 75% to about 90%, about 75% to about 85%, about 75% to about 80%, about 80% to about 95%, about 80% to about 90%, about 80% to about 85%, about 85% to 95%, about 85% to about 90%, or about 90% to about 95%.

[0058] In certain embodiments, the foregoing methods herein can achieve an accuracy of at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, or 90%. In certain embodiments, the accuracy is between about 50% to about 90%, about 50% to about 85%, about 50% to about 80%, about 50% to about 75%, about 50% to about 70%, about 50% to about 65%, about 50% to about 60%, about 50% to about 55%, about 55% to about 90%, about 55% to about 85%, about 55% to about 80%, about 55% to about 75%, about 55% to about 70%, about 55% to about 65%, about 55% to about 60%, about 60% to about 90%, about 60% to about 85%, about 60% to about 80%, about 60% to about 75%, about 60% to about 70%, about 60% to about 65%, about 65% to about 90%, about 65% to about 85%, about 65% to about 80%, about 65% to about 75%, about 65% to about 70%, about 70% to about 90%, about 70% to about 85%, about 70% to about 80%, about 70% to about 75%, about 75% to about 90%, about 75% to about 85%, about 75% to about 80%, about 80% to about 90%, about 80% to about 85%, or about 85% to about 90%.

[0059] In certain embodiments, the foregoing methods herein involve a cross-validation procedure. In certain embodiments, the cross-validation procedure is a 5-fold, 10-fold, 15-fold, 20-fold, 25-fold, 30-fold, 35-fold, or 40-fold cross-validation procedure. In certain embodiments, the cross-validation procedure is repeated at least 10 times, 20 times, 30 times, 40 times, 50 times, or 100 times.

[0060] In certain embodiments, the foregoing computer-implemented methods for early detection of abnormal signal quantification of a blood sample under test herein can reduce the false positive rate by at least 10%, 20%, 30%, 40%, or 50% compared to traditional methods (e.g., methods based on predetermined reference ranges for each protein tumor marker).

[0061] In certain embodiments, the foregoing methods herein can result in a false positive rate of less than 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 1.5%, 2%, 2.5%, 3%.

[0062] In certain embodiments, the foregoing methods herein can reduce the false positive rate to less than the original false positive rate by 1 / 5, 1 / 6 / , 1 / 7, 1 / 8 / , 1 / 9, 1 / 10, 1 / 15, 1 / 20, 1 / 25, 1 / 30, 1 / 35, 1 / 40, 1 / 45, or 1 / 50.

[0063] In certain embodiments, the foregoing methods herein can result in no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 false positives in 10,000 people.

[0064] In addition, in certain cases, certain protein biomarkers associated with multiple cancer types have a greater contribution or higher weight in the model. Conversely, certain protein biomarkers that are highly specific to a particular cancer type (e.g., AFP that is exclusively used for liver cancer detection) have a relatively lower contribution. This leads to a situation where certain protein biomarkers that are highly specific to a particular cancer type exhibit abnormally high levels (while other protein biomarkers remain normal), while the MCED model predicts a lower cancer signal score (POC).

[0065] To address this issue, the foregoing methods further comprise: biomarker outlier analysis. Outlier analysis focuses on identifying and analyzing cases where highly specific cancer protein biomarkers exhibit abnormal expression levels compared to normal cases. By incorporating this method into the MCED model, detection of cancer blood that can exhibit unique protein biomarker expression is improved. Here, based on over 6000 non-cancer samples, the applicants used the following three methods to determine the cutoff value for outlier analysis.

[0066] 1) Boxplot method: The boxplot method can identify outliers by plotting the boxplot of protein biomarkers of normal control samples. The boxplot shows the quartile range of the data, and the observation value exceeding the upper quartile value plus 1.5 times the interquartile range can be considered as the cutoff value of outliers;

[0067] 2) Modified Z-Score: Since some non-cancer diseases can also cause the expression level of protein biomarkers in the normal control group to rise, the protein expression level of the normal control group shows skewness. Therefore, by calculating the difference between the observation value and the median divided by the median absolute deviation (MAD), the modified Z-score is considered to be skewed data. The expression value of the protein biomarker with a modified Z-Score > 10 is defined as the cutoff value of outliers;

[0068] 3) Percentile: The percentile method is to compare the observation value with the percentile of the data, and the observation value exceeding the 99th percentile of the normal group can be regarded as the cutoff value of the outlier. According to the cutoff values obtained by the above three methods, the maximum value is selected as the final cutoff value of the high outlier. If the expression level of a specific protein biomarker in the test sample is higher than the corresponding abnormal cutoff value, it can be predicted that the sample has a higher cancer signal score. By developing this outlier analysis method, the applicant aims to enhance the identification of blood samples with significantly abnormal expression levels of a cancer-specific protein biomarker. By effectively predicting these special cases, the applicant can improve the sensitivity of detection for a specific type of cancer. Therefore, in certain embodiments, the foregoing methods herein further comprise outlier analysis. In certain embodiments, the outlier analysis described herein involves determining a cutoff value based on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, or 50000 non-cancer samples. In certain embodiments, the cutoff value is determined by selecting the maximum value obtained by the foregoing boxplot method, modified Z-Score, and / or percentile. In certain embodiments, a cutoff value can be determined for each protein biomarker described herein. In certain embodiments, the outlier analysis comprises comparing the quantitative level (e.g., expression level) of a selected protein biomarker (e.g., any protein biomarker described herein) with the corresponding cutoff value determined herein. For example, if the quantitative level of the selected protein biomarker is higher than the corresponding cutoff value, it indicates that the cancer signal score of the test sample is higher. In certain embodiments, the outlier analysis is performed by (a) determining the cutoff value of each protein biomarker (e.g., by the boxplot method, modified Z-Score, and / or percentile), and (b) comparing the quantitative level of each protein biomarker with its corresponding cutoff value. In certain embodiments, the higher the quantitative level of a protein biomarker relative to its corresponding cutoff value indicates that the cancer signal score of the test sample is higher.

[0069] In some embodiments, less than 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20% of the subjects will be further subjected to the second-step NGS detection. By reducing the number of subjects subjected to the second-step NGS detection, the resource utilization can be optimized, and the expensive and complicated NGS technology can be used only for high-risk individuals, thereby reducing the overall screening cost of cancer. In this way, the screening specificity can be improved, the false positive results can be reduced, and the unnecessary medical procedures and physical burden on the subjects can be reduced.

[0070] In some embodiments, the further cancer risk determination and cancer type prediction using NGS method comprises: constructing a cancer risk determination model and a cancer type prediction model using machine learning method based on the cancer-specific methylation site set, and determining the cancer risk and predicting the cancer type of the aforementioned preliminary screening positive subjects. In some embodiments, the aforementioned cancer-specific methylation site set can be a published site, such as a patent or a journal (Klein EA, Richards D, Cohn A, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Annals of Oncology. 2021; 32(9): 1167-1177. doi: 10.1016 / j.annonc.2021.05.806), etc. In some embodiments, it can also be obtained by self-test.

[0071] In some embodiments, the further cancer risk determination and cancer type prediction using NGS method comprises: constructing a multi-dimensional cancer risk determination model and a cancer type prediction model using machine learning method based on the multi-dimensional data of the blood sample of the aforementioned preliminary screening positive subjects, and determining the cancer risk and predicting the cancer type of the aforementioned preliminary screening positive subjects.

[0072] In some embodiments, the aforementioned multi-dimensional data comprises at least one of the following: a cancer signal score (POC) based on the aforementioned biomarker; and

[0073] The following information is obtained based on the cfDNA whole genome low-depth sequencing: end motif score (EMS) of the end motif of the cfDNA fragment; chromosome instability index (CNA); and fragment size (FS) measurement of the cfDNA fragment.

[0074] In some embodiments, the multi-dimension data further comprises one or more viral measures of viral abundance associated with cancer. In some embodiments, the one or more viruses associated with cancer comprises one or more of HBV, HPV, EBV, HHV, HIV, or JCPyV.

[0075] In some embodiments, the processing of the multi-dimension cancer risk determination model comprises EMS and model inputs selected from at least one of POC, CNA, FS, and viral measures.

[0076] In some embodiments, the viral measures are determined by: identifying, based on the sequencing data of the cfDNA, a first set of sequences that do not align to a human reference genome; identifying, from the first set of sequences, a second set of sequences that align to one of the viral reference sequences; and determining the viral measures based on a ratio of sequencing reads in the second set of sequences to total sequencing reads.

[0077] In some embodiments, the end motif score (EMS) of the end motif of the cfDNA fragment is obtained by: determining, based on each of the cfDNA sequencing data of the cfDNA, a length of the corresponding cfDNA sequence; extracting, based on each of the cfDNA sequencing data of the cfDNA, a corresponding end motif as (i) a subsequence comprising a first predetermined number of bases of a 5' end of the corresponding cfDNA sequence, or (ii) a subsequence comprising a second predetermined number of bases before or after a breakpoint of the corresponding cfDNA sequence; and determining, for a set of different cfDNA fragment sizes, a corresponding distribution of different types of end motifs in the test sample.

[0078] In some embodiments, the first predetermined number of bases is 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, 6 bp, 7 bp, or 8 bp.

[0079] In some embodiments, the second predetermined number of bases is 1 bp, 2 bp, 3 bp, 4 bp, 5 bp, or 6 bp.

[0080] In some embodiments, the determining of the end motif score further comprises: calculating the EMS of the test sample based on (i) the distribution of different types of end motifs of the cfDNA fragments in the test sample within a predetermined fragment size range, and (ii) the distribution of different types of end motifs in the control sample within the predetermined fragment size range.

[0081] In some embodiments, the predetermined fragment size range comprises (i) a first length range of 50 bp-250 bp; or (ii) a sub-range of the first length range within 50 bp-250 bp.

[0082] In certain embodiments, the predetermined fragment length range includes (i) a second length range of 250bp-450bp, or (ii) a sub-range of the second length range within 250bp-450bp.

[0083] In certain embodiments, the foregoing EMS calculation formula is:

[0084] where i denotes a type of terminal motif, pi denotes a proportion of terminal motif i among all terminal motifs in the test sample, is an average or median of the proportion of terminal motif i in the control sample, and m is a total number of terminal motifs.

[0085] In certain embodiments, the foregoing EMS calculation is a weighted sum based on (i) a first EMS determined based on a first predetermined fragment length range and (ii) a second EMS determined based on a second predetermined fragment length range.

[0086] In certain embodiments, the weight coefficients of the weighted sum are determined using a third machine learning model.

[0087] In certain embodiments, a fourth machine learning model is used to process a model input comprising (i) the EMS and (ii) a frequency of occurrence of a selected subset of terminal motifs in the test sample, to generate an output comprising a terminal motif metric.

[0088] In certain embodiments, the selected subset of terminal motifs is determined using one or more statistical tests that identify differences in frequencies of occurrence of terminal motifs between healthy subjects and cancer patients.

[0089] In certain embodiments, the FS is determined by determining one or more first distribution parameters from sequencing data of the cfDNA, characterizing proportions of fragments in a particular length range relative to other length ranges; determining one or more second distribution parameters from the sequencing data by binning, filtering, combining, and normalizing the sequencing data; processing a fifth model input comprising the first distribution parameters and the second distribution parameters using a fifth machine learning model, to generate an output specifying the FS.

[0090] In certain embodiments, prior to processing the first model input using the first machine learning model, a training process is performed to train the first machine learning model using a plurality of training examples, each training example comprising (i) a respective set of respective training inputs determined based on measurement data of a respective subject and (ii) a respective training output characterizing a diagnosis of the respective subject, the plurality of examples including (i) a set of training examples using measurement data of cancer patients and (ii) a set of training examples using measurement data of control subjects.

[0091] In some implementations, the aforementioned method is capable of early detection of at least one type of cancer. The aforementioned cancers include pancreatic cancer, ovarian cancer, liver cancer, lung cancer, stomach cancer, colorectal cancer, lymphoma, esophageal cancer, cervical cancer, prostate cancer, or breast cancer.

[0092] In some implementations, the aforementioned cancer may also be stage I, stage II, stage III, or stage IV cancer.

[0093] Markers in the measurement sample

[0094] As part of the aforementioned method, a set of markers from asymptomatic subjects can be measured. Many methods known in the art can be used in the aforementioned method to measure gene expression (e.g., mRNA) or the resulting gene product (e.g., peptide or protein).

[0095] In some embodiments, tumor antigen detection can be performed using an automated immunoassay analyzer. Representative analyzers include those from Roche Diagnostics. System or Abbott Diagnostics Analyzers. In some embodiments, the analyzers used herein include the Roche cobas e411 / e601 analyzer (Roche Diagnostics GmbH, Mannheim, Germany) and the Bioplex 200 platform. Using these standardized platforms allows results from one laboratory or hospital to be transferred to other laboratories worldwide. However, the methods presented herein are not limited to any single form of detection or any specific set of markers constituting a combination of markers.

[0096] The presence and quantification of one or more antigens or antibodies in a sample can be determined using one or more immunoassays known in the art. Immunoassays typically involve: (a) providing an antibody (or antigen) that specifically binds to a protein biomarker (i.e., an antigen or antibody); (b) contacting the sample with the antibody or antigen; and (c) detecting the presence of an antibody-antigen complex or an antigen-antibody complex in the sample.

[0097] Well-known immunoassays include, for example, enzyme-linked immunosorbent assay (ELISA), also known as the "sandwich method," enzyme immunoassay (EIA), radioimmunoassay (RIA), fluorescence immunoassay (HA), chemiluminescent immunoassay (CLIA), count immunoassay (CIA), filter paper media enzyme immunoassay (META), fluorescence-linked immunosorbent assay (FLISA), agglutination immunoassay, and multiplex fluorescence immunoassays (such as Luminex Lab MAP), as well as immunohistochemistry.

[0098] The immunoassay can be used to determine the amount of antigen detected in a sample from a subject. First, the amount of antigen detected in a sample can be detected using the immunoassay methods described above. If antigen is present in the sample, under the appropriate incubation conditions described herein before, it will form an antibody-antigen complex with the antibody that specifically binds to the antigen. The amount, activity, or concentration, etc. of the antibody-antigen complex can be determined by comparing the measured value to a standard or control.

[0099] The methods described herein before can be applied to any method that measures a marker or a set of markers using a sample from a human subject. In certain embodiments, the sample from a human subject is a tissue section, such as a section from a biopsy. In another embodiment, the sample from a human subject is a bodily fluid, such as blood, serum, plasma, or a portion or component thereof. In other embodiments, the sample is blood or serum, and the marker is a protein measured therefrom. Many other combinations of sample forms and marker forms from a human subject are contemplated in the methods described herein before.

[0100] Biomarker

[0101] However, prior to performing the measurements, a set of markers needs to be selected for the particular cancer being screened. For diseases including cancer, many markers are known, and a known set of markers can be selected, or as described in the examples. The set of markers can be selected based on measurements of individual markers in retrospective clinical samples, where a set of markers can be generated based on empirical data for the disease of interest, such as cancer, and in particular, pancreatic cancer, ovarian cancer, liver cancer, lung cancer, gastric cancer, colorectal cancer, lymphoma, cervical cancer, esophageal cancer, or breast cancer.

[0102] Examples of protein biomarkers that can be used include molecules that can be detected in a bodily fluid sample, such as antibodies, antigens, small molecules, proteins, hormones, enzymes, genes, etc. However, the use of tumor antigens has many advantages, as they have been used extensively for many years, and many tumor antigens have well-validated standardized detection kits available for use with the automated immunoassay platforms described above.

[0103] In one particular embodiment, a set of markers can be selected based on their association with a particular cancer type. For example, AFP is a specific protein biomarker for liver cancer. In addition, alpha-fetoprotein (AFP) can be used as a protein biomarker for hepatocellular carcinoma, CA125 for ovarian cancer, CA15-3 for breast cancer, CA19-9 for pancreatic cancer, CA72-4 for ovarian cancer, carcinoembryonic antigen (CEA) for digestive tract cancer, CYFRA 21-1 for breast cancer.

[0104] Among these biomarkers, CA125 is a repetitive peptide epitope of mucin MUC16, which can promote cancer cell proliferation and inhibit anti-cancer immune response. CA15-3 is from glycoprotein mucin-1 (MUC-1). CA19-9 is a tumor-associated antigen, which was originally defined by a monoclonal antibody produced by a hybridoma prepared from mouse spleen cells immunized with a human colorectal cancer cell line. CA19-9 exists in tissues as an epitope of sphingosine-1 -lewis A blood group antigen. CYFRA 21-1 is a fragment of cytokeratin 19 (KRT19).

[0105] ProGRP is related to gastrin-releasing peptide (GRP). GRP is an important regulatory molecule that is involved in many physiological and pathophysiological processes in humans. Its 148 amino acid proprotein is further processed to generate 27 amino acid GRP and 68 amino acid ProGRP after cleavage of the signal peptide. Since the half-life of GRP is very short, only 2 minutes, it is not possible to measure GRP in blood. Therefore, the detection method for measuring ProGRP is helpful for the study of GRP.

[0106] SCCA includes SCCA1 (also known as SERPINB3) and SCCA2 (also known as SERPINB4). These biomarkers are known in the art and reported in the literature, such as Locker, et al. "ASCO 2006 update of recommendations for the use of tumor markers in gastrointestinal cancer", Journal of clinical oncology 24.33 (2006): 5313-5327; Del Villano BC, et al. "Radioimmunometric assay for a monoclonal antibody-defined tumor marker, CA 19-9", Clin Chem 29:549, 1983 -552; Muraro, Raffaella, et al. "Generation and characterization of B72.3 second generation monoclonal antibodies reactive with the tumor-associated glycoprotein 72 antigen." Cancer research 48.16 (1988): 4588-4596; each of which is incorporated by reference herein in its entirety.

[0107] By selecting a combination of 7 markers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1) combined with clinical parameters, using an AI model for cancer screening, we have successfully validated in a large cohort of nearly 10,000 people, significantly higher than the clinically common single threshold-based method, see “Luan Y, Zhong G, Li S, et al. A panel of seven protein tumour markers for effective and affordable multi-cancer early detection by artificial intelligence: a large-scale and multicentre case-control study. eClinicalMedicine. 2023; 61: 102041. doi: 10.1016 / j.eclinm.2023.102041”.

[0108] In some embodiments, a panel of protein biomarkers can be selected from the following markers in combination with clinical parameters: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1. In other embodiments, a panel of protein biomarkers can be selected from the following markers: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCC, and PSA. In some embodiments, a panel of protein biomarkers can be selected from the following markers in combination with clinical parameters: AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCC, and PSA. In other embodiments, a panel of protein biomarkers can be selected from the following markers: AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCC, and PSA. For lung cancer screening, for example: CEA, CYFRA 21-1, ProGRP, and SCC can be added in some embodiments, with clinical parameters such as age and gender.

[0109] In some embodiments, wherein the subject is male and the set of biomarkers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA; or the subject is female and the set of biomarkers comprises or consists of AFP, CA125, CA15-3, CA19-9, CEA, CYFRA 21-1, ProGRP, and SCCA.

[0110] It is noted that the protein biomarker "SCC" or "SCCA" as used herein refers to squamous cell carcinoma antigen. The protein biomarker "PSA" or "TPSA" as used herein refers to total prostate specific antigen.

[0111] In certain embodiments, the marker combination can include protein biomarkers associated with one of the following cancers: pancreatic cancer, ovarian cancer, liver cancer, lung cancer, gastric cancer, colorectal cancer, lymphoma, esophageal cancer, cervical cancer, prostate cancer, or breast cancer. As a design choice, the marker panel can include any number of protein biomarkers, with the goal of maximizing the specificity or sensitivity of the detection. Thus, as a design choice, the detection of interest can require the presence of at least two or more protein biomarkers, three or more protein biomarkers, four or more protein biomarkers, five or more protein biomarkers, six or more protein biomarkers, seven or more protein biomarkers, eight or more protein biomarkers, nine or more protein biomarkers, ten or more protein biomarkers. Thus, in one embodiment, the protein biomarker panel can include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten or more different markers. In one embodiment, the protein biomarker panel includes about two to ten different markers. In another embodiment, the protein biomarker panel includes about four to eight different markers. In another embodiment, the protein biomarker panel includes about seven different markers. In another embodiment, the protein biomarker panel includes about ten different markers.

[0112] Typically, the sample is submitted for detection, and the result can be a series of numbers reflecting the presence and level (e.g., concentration, amount, activity, etc.) of each protein biomarker in the marker panel in the sample.

[0113] In certain embodiments, the methods described herein can be used to monitor the progression of a disease, determine the effectiveness of a treatment, and adjust treatment strategies. For example, cell free DNA can be collected from a subject to detect cancer, and this information can also be used to select an appropriate treatment for the subject. After the subject receives treatment, cell free DNA can be collected from the subject. Analysis of these cfDNA can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and / or adjust the treatment strategy. In certain embodiments, the test results are compared to earlier results. In certain embodiments, a sharp increase in circulating tumor DNA indicates that tumor cells are undergoing apoptosis, which can indicate that the treatment is effective.

[0114] Cancer treatment monitoring method

[0115] In another aspect of the present application, the present application provides a cancer treatment monitoring method. The method comprises: predicting cancer risk based on the method described in the first aspect; and monitoring the subject of cancer treatment using the risk of cancer. This method based on systematic data-driven decision support enhances the scientificity and reliability of the treatment plan, promotes patient compliance, improves long-term survival rate, and at the same time realizes higher economic benefits, which has significant clinical and economic value.

[0116] Computer-implemented device for stepwise prediction of cancer risk

[0117] In yet another aspect of the present application, the present application provides a computer-implemented device for stepwise prediction of cancer risk, referring to FIG. 1, the device comprises: a cancer signal score determination unit 100, a preliminary screening positive subject determination unit 200, and a cancer type determination unit 300. Wherein,

[0118] 100 unit, for determining a cancer signal score based on the level of biomarkers in the blood sample of the subject;

[0119] 200 unit, for comparing the cancer signal score with a predetermined threshold to determine a preliminary screening positive subject;

[0120] 300 unit, for further determining the risk of cancer and predicting the cancer type based on the preliminary screening positive subject using NGS method; wherein the biomarkers include at least one selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

[0121] In some examples of the present application, the device described above is used for cancer risk screening, which has the advantages of high efficiency, high specificity and low cost.

[0122] Computer program product, computing device and computer readable storage medium

[0123] In yet another aspect of the present application, a computer program product, a computing device and a computer readable storage medium are provided. Based on the aforementioned computer program product, computing device or computer readable storage medium, the computer-implemented method of predicting cancer risk in steps is executed.

[0124] The embodiments are described in detail with electronic device as an example.

[0125] The term "electronic device" is intended to refer to various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The computing device can also refer to various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown in the figures, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0126] Referring to FIG. 2, the electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded into a RAM (Random Access Memory) 503 from a storage unit 508. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0127] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0128] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the computer-implemented method of stepwise predicting cancer risk. For example, in some embodiments, the computer-implemented method of stepwise predicting cancer risk can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the aforementioned computer-implemented method of stepwise predicting cancer risk by other any appropriate means, such as by means of firmware.

[0129] In this application, logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of executable instructions for implementing the logic functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a product of the manufacturing and / or processing. The computer-readable medium can include, but is not limited to, the following: an electronic connection (electronic device) having one or more wires; a portable computer diskette (magnetic device); a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM or Flash memory); an optical fiber; and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium can even be paper or another suitable medium upon which the aforementioned program can be printed, because the aforementioned program can be electronically obtained, for example, by optically scanning the paper or other suitable medium, then

[0130] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, the hardware can include any or a combination of the following: discrete logic circuitry having logic gates for implementing logic functions upon data signals, application specific integrated circuits having logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and / or the like.

[0131] Those of skill in the art would understand that the various embodiments described above can be carried out by a computer program that can be stored in a computer readable storage medium. The computer readable storage medium can include a floppy disk, CD-ROM, hard disk, optical disk, DVD, RAM, ROM, PROM, EPROM, EEPROM, or other memory or computer readable medium suitable to the technical task. The computer program can be written in any of a number of suitable programming languages and can be executed by a computer of the type described above.

[0132] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing module, or each unit can exist physically separately, or two or more units can be integrated in one module. The above integrated module can be realized in the form of hardware, or in the form of a software functional module. When the foregoing integrated module is realized in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer readable storage medium.

[0133] It should be understood that various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which is not limited herein.

[0134] Cancer treatment method

[0135] In an aspect, the methods of the present disclosure comprise administering to a subject in need thereof a therapeutically effective amount of a therapeutic agent. In certain implementations, the subjects are judged to have cancer by the methods of the present disclosure. In certain embodiments, the subject has, for example, breast cancer (e.g., triple negative breast cancer), cervical cancer, endometrial cancer, liver cancer, lung cancer, small cell lung cancer, lymphoma, ovarian cancer, pancreatic cancer, prostate cancer, colorectal cancer, gastric cancer, or esophageal cancer. In certain embodiments, the cancer is unresectable metastatic non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), lymphoma, or metastatic hormone-refractory prostate cancer. In certain embodiments, the subject has a solid tumor. In certain embodiments, the cancer is esophageal cancer, gastric cancer, pancreatic cancer, or colorectal cancer. In certain embodiments, the subject has cervical cancer, ovarian cancer, breast cancer, or endometrial cancer. In certain embodiments, the subject has liver cancer, lung cancer, gastric cancer, colorectal cancer, prostate, or esophageal cancer.

[0136] As used herein, “effective amount” means an amount or dose sufficient to produce a beneficial or desired result, including an amount that stops, slows, delays, or inhibits progression of a disease (e.g., cancer). The effective amount will vary depending on factors such as the age, weight, severity of symptoms, and route of administration of the subject to whom the therapeutic agent is administered, and thus the amount administered can be determined on an individual basis. An effective amount can be achieved through one or more administrations. By way of example, an effective amount is an amount sufficient to improve, arrest, stabilize, reverse, inhibit, slow, and / or delay progression of cancer in a patient, or an amount sufficient to improve, arrest, stabilize, reverse, slow, and / or delay proliferation of a cell in vitro (e.g., a biopsy cell, any of the cancer cells or cell lines described herein (e.g., a cancer cell line)).

[0137] In certain embodiments, the methods described herein before can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and adjust the treatment strategy. For example, cell free DNA can be collected from a subject to detect cancer, and this information can also be used to select an appropriate treatment for the subject. After the subject receives a treatment, cell free DNA can be collected from the subject. Analysis of these cfDNA can be used to monitor the progression of the disease, determine the effectiveness of the treatment, and / or adjust the treatment strategy. In certain embodiments, the test results are compared to earlier results. In certain embodiments, a sharp increase in circulating tumor DNA indicates that tumor cells are undergoing apoptosis, which can indicate that the treatment is effective.

[0138] In certain embodiments, the therapeutic agent can include one or more inhibitors selected from the group consisting of inhibitors of B-Raf, inhibitors of EGFR, inhibitors of MEK, inhibitors of ERK, inhibitors of K-Ras, inhibitors of c-Met, inhibitors of anaplastic lymphoma kinase (ALK), inhibitors of phosphatidylinositol 3-kinase (PI3K), inhibitors of Akt, inhibitors of mTOR, dual PI3K / mTOR inhibitors, inhibitors of Bruton’s tyrosine kinase (BTK), and inhibitors of isocitrate dehydrogenase 1 (IDH1) and / or isocitrate dehydrogenase 2 (IDH2). In certain embodiments, the additional therapeutic agent is an inhibitor of indoleamine 2,3-dioxygenase-1 (IDO1) (e.g., epacadostat).

[0139] In certain embodiments, the therapeutic agent can include one or more inhibitors selected from the group consisting of inhibitors of HER3, inhibitors of LSD1, inhibitors of MDM2, inhibitors of BCL2, inhibitors of CHK1, inhibitors of activated hedgehog signaling pathway, and drugs that selectively degrade estrogen receptors.

[0140] In certain embodiments, the therapeutic agent can include one or more therapeutic agents selected from the group consisting of Trabectedin, nab-paclitaxel, Trebananib, Pazopanib, Cediranib, Palbociclib, Everolimus, fluoropyrimidine, IFL, regorafenib, Reolysin, Alimta, Zykadia, Sutent, temsirolimus, axitinib, sorafenib, Votrient, Pazopanib, IMA-901, AGS-003, cabozantinib, Vinflunine, Hsp90 inhibitors, Ad-GM-CSF, temazolomide, IL-2, IFNa, vincristine, thalidomide, dacarbazine, cyclophosphamide, lenalidomide, azacytidine, lenalidomide, bortezomid, amrubicine, carfilzomib, pralatrexate, and enzastaurin. In certain embodiments, the therapeutic agent can include one or more therapeutic agents selected from the group consisting of adjuvants, TLR agonists, tumor necrosis factor (TNF) alpha, IL-1, HMGB1, IL-10 antagonists, IL-4 antagonists, IL-13 antagonists, IL-17 antagonists, HVEM antagonists, ICOS agonists, therapeutic agents targeting CX3CL1, therapeutic agents targeting CXCL9, therapeutic agents targeting CXCL10, therapeutic agents targeting CCL5, LFA-1 agonists, ICAM1 agonists, and Selectin agonists, among others.

[0141] In certain embodiments, the subject is administered carboplatin, nab-paclitaxel, paclitaxel, cisplatin, pemetrexed, gemcitabine, FOLFOX, or FOLFIRI.

[0142] In certain embodiments, the therapeutic agent is an antibody or antigen-binding fragment thereof. In certain embodiments, the therapeutic agent is an antibody that specifically binds to PD-1, CTLA-4, BTLA, PD-L1, CD27, CD28, CD40, CD47, CD137, CD154, TIGIT, TIM-3, GITR, or OX40.

[0143] In certain embodiments, the therapeutic agent is an anti-PD-1 antibody, an anti-OX40 antibody, an anti-PD-L1 antibody, an anti-PD-L2 antibody, an anti-LAG-3 antibody, an anti-TIGIT antibody, an anti-BTLA antibody, an anti-CTLA-4 antibody, or an anti-GITR antibody.

[0144] In certain embodiments, the therapeutic agent is an anti-CTLA4 antibody (e.g., ipilimumab), an anti-CD20 antibody (e.g., rituximab), an anti-EGFR antibody (e.g., cetuximab), an anti-CD319 antibody (e.g., elotuzumab), or an anti-PD1 antibody (e.g., nivolumab).

[0145] Kit

[0146] The present disclosure also provides kits for collecting, shipping, and / or analyzing samples. Such kits can include materials and reagents needed to obtain appropriate samples from a subject or to measure levels of particular protein biomarkers. In certain embodiments, the kits include materials and reagents needed to obtain and store samples from a subject. The samples are then shipped to a service center for further processing (e.g., sequencing and / or data analysis).

[0147] In certain embodiments, the foregoing kits further include instructions for collecting the sample, performing the assay, and methods for interpreting and analyzing the assay results data.

[0148] Embodiments of the present application will be described in more detail by way of examples as shown in the accompanying drawings. The embodiments described below by reference to the drawings are illustrative, and are intended to explain the present application, and are not to be understood as limiting the present application.

[0149] Example 1: Cancer type detection cost comparison

[0150] In 2014, Exact Sciences published a comparison of the performance of their new product Cologuard and the traditional method FIT in the screening of colorectal cancer (Imperiale TF, Ransohoff DF, Itzkowitz SH, et al. Multitarget Stool DNA Testing for Colorectal-Cancer Screening. N Engl J Med. 2014; 370(14): 1287-1297. doi: 10.1056 / NEJMoa1311194), as shown in Table 1. Although Cologuard has slightly higher sensitivity in screening colorectal cancer than the traditional clinical method FIT, the cost of each test is much higher than FIT, 30 times that of FIT. Therefore, for each colorectal cancer patient screened, Cologuard costs ~100,000 USD, while FIT only costs ~4,000 USD. Moreover, in healthy populations, the specificity of FIT is 96.4%, slightly higher than that of Cologuard, effectively reducing the number of false-positive patients who need to undergo colonoscopy for diagnosis.

[0151] Table 1

[0152] Therefore, although early screening of colorectal cancer in the United States mainly uses Cologuard, in Europe and other developed countries, including the Netherlands, FIT is still the main method. The main reason is that the screening cost of FIT is only 20 dollars per test, and the price is relatively cheap, with high specificity. The benefits of using FIT for comprehensive screening are much greater than those of Cologuard. This is why FIT is still recommended as a method for screening colorectal cancer around the world, such as in the developed country of the Netherlands and the developing country of China.

[0153] The above method is based on fecal detection, and the compliance is not as high as blood; at the same time, it can only be used for colorectal cancer examination.

[0154] Example 2: Step-by-step method for predicting cancer risk

[0155] Here, the present application proposes a two-step screening process: first, use a cost-effective screening method (such as OncoSeek: an assay based on a panel of proteins) for preliminary screening to enrich cancer, and then use a more superior screening method (such as an NGS-based method: SeekInCare and Galleri) for secondary detection in the population that tested positive in the preliminary screening, greatly improving the specificity.

[0156] For the first step screening, based on the expression level of a set of selected protein biomarkers (PTMs) combined with clinical information (gender, age or smoking status, etc.) for cancer screening, the method in the article is named OncoSeek, and has been published in eClinicalMedicine, reference: Luan Y, Zhong G, Li S, et al. A panel of seven protein tumour markers for effective and affordable multi-cancer early detection by artificial intelligence: a large-scale and multicentre case-control study. eClinicalMedicine. 2023; 61: 102041. doi: 10.1016 / j.eclinm.2023.102041.

[0157] Specifically, the experimental process is as follows: blood is collected by routine venipuncture, and plasma or serum is obtained by separation within the specified time of the blood collection tube. The tumor protein markers are quantified using clinically common platforms and methods. For example: Roche cobas e411 / e601, Bio-Rad Bio-Plex 200 platform, and ABBOTT automatic immune analyzer ARCHITECT i2000SR, etc. According to the manufacturer's instructions, the expression levels of at least 6 PTMs in the 10 PTMs to be tested, including AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA, are determined. In this paper, the expression levels of 7 PTMs (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1) were determined using Roche cobas e411 and e601 platforms, and the clinical data and PTM quantitative data of 1005 cancer patients and 812 non-cancer individuals previously published by Cohen et al. were collected as an independent validation set 2 for analysis (Cohen JD, Li L, Wang Y, et al. Detection and localization of surgically resectable cancers with a multi-analyte blood test. Science. 2018; 359(6378): 926-930. doi: 10.1126 / science.aar3247). The published data from Johns Hopkins University only contains six biomarkers, and lacks the value of CA72-4. The inventors used the average value of CA72-4 in the non-cancer sample group in the training set to replace the missing value.

[0158] Using the AI method, a multi-cancer early screening model is constructed. Compared with the traditional method, the specificity of cancer screening can be greatly improved (as shown in the table below).

[0159] Table 2: Comparison of performance of conventional clinical methods and AI methods in multi-cancer early detection

[0160] *At least one individual with a tumor protein marker higher than the clinical threshold is considered positive (cancer); PPV, positive predictive value; NPV, negative predictive value.

[0161] Optimization 1: Model integration

[0162] In contrast to the article where GLM was chosen to build the AI screening model, here we innovatively introduce the idea of model ensemble: using different machine learning algorithms to build different early screening models, and finally ensembling these models (Figure 3).

[0163] Machine learning methods such as gradient boosting machine (GBM), generalized linear model (GLM), random forest (RF), and support vector machine (SVM) were used to build the MCED model. These machine learning methods were used to build different models that can distinguish cancer and non-cancer samples. To optimize the model performance, 10-fold cross-validation was performed for each machine learning method. The cross-validation process generated an average cancer prediction score from each machine learning method, ranging from 0 to 1. These average cancer prediction scores from all four machine learning methods were combined into a matrix as the input for a second layer GLM algorithm. The GLM algorithm integrated these scores from different machine learning methods and built the final ensemble model, generating a final cancer signal score for each test sample, called the POC index. A higher POC index indicates a higher cancer signal score for the test sample. The performance and robustness of the MCED model were verified using two independent cohorts.

[0164] This comprehensive approach aims to combine models trained using different methods. Model ensemble can capture different aspects of the data, reduce the risk of overfitting, and enhance the accuracy and reliability of the MCED model for determining whether a sample is a cancer sample. Compared to the article, an additional 18 cancer samples can be detected by the optimized MCED model. At the same specificity of 90.0%, the sensitivity is 61.3%, which is an increase of 3.1%.

[0165] Optimization II: Combining the optimized MCED model with outlier analysis

[0166] Certain protein biomarkers associated with multiple cancer types have a greater contribution or higher weight in the MCED model. Conversely, certain protein biomarkers that are highly specific for a specific cancer type (e.g., AFP specifically for liver cancer detection) have a relatively lower contribution. This leads to a situation where a highly specific protein biomarker for a certain specific cancer type shows an abnormally high level while other protein biomarkers remain normal, and the MCED model may predict a lower cancer signal score (POC).

[0167] To address this issue, an outlier analysis method was developed as a supplement to predict these types of cancer samples (as shown in Figure 4). The outlier analysis method focuses on identifying and analyzing cases that show abnormal expression levels of highly specific cancer protein biomarkers compared to normal cases. By incorporating this method into the MCED model, the detection ability for tumor samples that can exhibit abnormal protein biomarker expression is improved. Here, the inventors used the following three methods to determine the cutoff value for outlier analysis based on more than 6000 non-cancer samples.

[0168] 1) Boxplot method: The boxplot method identifies outliers by plotting the boxplot of protein biomarkers in normal control samples. The boxplot shows the quartile range of the data, and the observation value exceeding the upper quartile plus 1.5 times the interquartile range can be considered as the cutoff value of outliers;

[0169] 2) Modified Z-score: Since some non-cancer diseases can also cause the expression level of protein biomarkers in the normal control group to rise, the protein expression level in the normal control group shows skewness. Therefore, by calculating the difference between the observation value and the median divided by the median absolute deviation (MAD), the modified Z-score is considered to be skewed data. The expression of protein biomarkers with a modified Z-score > 10 is defined as the cutoff value of outliers;

[0170] 3) Percentile: The percentile method compares the observation value with the percentile of the data, and the observation value exceeding the 99th percentile of the normal control group can be considered as the cutoff value of outliers.

[0171] According to the cutoff values obtained from the above three methods, the maximum value is selected as the final high outlier cutoff value. If the expression level of a specific protein biomarker in a test sample is higher than the corresponding outlier cutoff value, then the sample can be predicted to carry a cancer signal. By developing this outlier analysis method, the inventors aim to enhance the identification of cancer samples with significantly abnormal expression levels of a cancer-specific protein biomarker. By effectively predicting these special samples, the inventors can improve the sensitivity of specific types of cancer detection.

[0172] For example, since alpha-fetoprotein (AFP) is a specific protein biomarker for liver cancer, it has a relatively low weight in the inventors’ multi-cancer early detection model. However, by incorporating the outlier analysis described above, the inventors can improve the sensitivity of liver cancer detection. With outlier analysis, the outlier cutoff value for AFP is 614.4 IU / ml. Based on the MCED model alone, 154 out of 244 liver cancer samples were successfully predicted as positive. With a specificity of 92.9% (95% confidence interval: 92.3% to 93.5%), the sensitivity was 63.1% (95% confidence interval: 56.7% to 69.2%). By combining outlier analysis with the optimized MCED model, an additional 23 positive samples were successfully predicted, resulting in a total of 177 positive samples. The sensitivity was increased to 72.5% (95% confidence interval: 66.5% to 78.0%), an increase of 9.4%, while the specificity also increased slightly to 93.4%. As shown in FIG. 5, the AUC increased from 0.865 to 0.907, and the Delong test showed that the AUC value was significantly improved (P < 0.001).

[0173] Example 3: Using 10 protein quantification results for cancer screening

[0174] PTMs quantification: Blood was collected using routine venipuncture, and plasma or serum was obtained by separation within the specified time of the collection tube. Standard tumor protein marker quantification methods were used for quantification. In this example, electrochemiluminescence immunoassay (Roche cobas e411 / e601) and commercially available kits compatible with this platform (e.g., when the inventors detected SCCA levels, Roche Elecsys SCC immunoassay kit was used) were used to determine the expression levels of 10 PTMs, including AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA, and PSA, according to the manufacturer’s instructions.

[0175] The inventors obtained the quantification of protein biomarkers of 706 samples, including 203 male and 503 female samples, according to the above method. Each male sample contains the expression level of 10 selected protein biomarkers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, Pro-GRP, SCCA, PSA), each female sample contains the expression level of 9 selected protein biomarkers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, Pro-GRP, SCCA), and clinical information such as age and gender. Because PSA is a prostate-specific antigen, it is not necessary for women to do this marker. Using the method in Example 1, a model is established using the GLM algorithm, and the cutoff value of the outlier of the protein biomarker is determined by combining the outlier analysis method in Example 2, which further improves the sensitivity of the model for early screening of multiple cancers.

[0176] The method of the 10-marker model combined with the abnormal cutoff obtains an AUC value of 0.797 for cancer detection in the above-mentioned samples, with a specificity of 89.9% (95% confidence interval: 84.1% to 94.1%) and a sensitivity of 55.1% (95% confidence interval: 50.8% to 59.3%). As a control, the above steps are repeated using only 7 protein biomarkers (AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1) on the same data set, with an AUC value of 0.786, a specificity of 89.9% (95% confidence interval: 84.1% to 94.1%), and a sensitivity of 52.9% (95% confidence interval: 48.6% to 57.2%), and there is a significant difference between the ROC curves of the two models (DeLong's test: p-value = 0.02). Since Pro-GRP, SCCA, and PSA have high specificity in small cell lung cancer (SCLC), squamous cell carcinoma, and prostate cancer, respectively, the performance of the 10-protein biomarker model is greatly improved in these cancer types. The detection rate in small cell lung cancer samples is 90.9% (20 / 22), which is 40.9% higher than the result of 7 protein biomarkers; the detection rate in cervical cancer samples is 50% (4 / 8), which is 37.5% higher than the result of 7 protein biomarkers; and the detection rate in prostate cancer samples is 85.7% (6 / 7), which is 42.9% higher than the result of 7 protein biomarkers.

[0177] Example 4: Specificity and sensitivity detection of cancer screening

[0178] According to the above method, the specificity of our new method in early screening of cancer is ~90%, but the incidence of cancer in the general population over the age of 50 is about ~1%. Therefore, the preliminary estimate in the real world is that the number of false positive patients is more than 10 times the number of true positive patients. Therefore, it is a big problem for subsequent diagnosis and the mental agent of the examinee. Therefore, it is necessary to further reduce the number of false positive patients. Here we collected 617 cancer samples of different cancer types and 580 healthy control human samples, and used the two-step method of the present application to perform cancer screening

[0179] First, use the method of Example 1 to obtain the POC value of each examinee. According to the article, the POC value of the training set at 90% specificity is used as the cutoff value. POC is greater than <0.5, which is predicted to be low risk, i.e. predicted negative; POC >= 0.5 and POC < 0.9 is predicted to be medium risk, and POC >= 0.9 is predicted to be high risk. Here, medium risk and high risk are both default initial screening positive patients. Therefore, the final multi-cancer screening performance is 91.0% specificity and 49.9% sensitivity. The performance is basically consistent with the performance reported in the article (“A panel of seven protein tumour markers for effective and affordable multi-cancer early detection by artificial intelligence: A large-scale and multicentre case-control study.”), which again proves the robustness of this method.

[0180] Secondly, the method of NGS is used for the second screening for the subjects with positive screening results (52 false positive patients and 308 true positive patients). The method of cfDNA-based NGS has relatively better performance and generally higher specificity. For example, the multi-dimensional multi-omics method developed by ourselves (named SeekInCare, see Example 6 below) has a sensitivity of 66.3% at a specificity of 97.9%, and the Galleri method of the American GRAIL company, which is published in (Klein EA, Richards D, Cohn A, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Annals of Oncology. 2021;32(9):1167-1177. doi:10.1016 / j.annonc.2021.05.806), has a sensitivity of 51.5% at a specificity of 99.5%.

[0181] In this embodiment, we use the multi-dimensional multi-omics method (SeekInCare) developed by ourselves for the second detection of the subjects with positive screening results, and finally the specificity of our method is increased to 99.0%. The number of false positive patients is greatly reduced (as shown in FIG. 6). The number of false positive patients is reduced from 52 to 6, about 8 times. The number of true positive patients is reduced from 308 to 260, only 48, accounting for 15.6%. Therefore, 42.1% sensitivity is obtained at a high specificity of 99.0%. Through this embodiment, it is proved that the step-by-step greatly reduces the screening cost, reduces the number of false positive patients in screening, and ensures sufficient screening sensitivity.

[0182] Alternatively, if we use the POC cutoff values of 0.5 and 0.9 corresponding to the specificities of 90% and 99% in the training set of the OncoSeek article. We define the subjects with POC between 0.5 and 0.9 as positive in the first screening, and perform the second detection. Here, the main assumption is that the subjects with POC greater than 0.9 (99% specificity) are basically cancer patients, and do not need to receive the second detection. Based on this assumption, although there are still a small number of false positives, most cancer patients do not need to be detected twice, which can speed up the diagnosis of cancer. The subjects with positive results in the first screening include 42 false positive patients and 113 true positive patients, and finally the number of false positive patients is reduced to 14 and the number of true positive patients is 277; 97.6% specificity and 44.9% sensitivity are finally obtained.

[0183] Optionally, for the primary screening positive patients, we can also choose to use the published methylation panel for secondary detection and cancer tracing. The specific process is described in the published article (Klein EA, Richards D, Cohn A, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Annals of Oncology. 2021;32(9): 1167-1177. doi: 10.1016 / j.annonc.2021.05.806)

[0184] Example 5: Real-world cancer screening application

[0185] The above examples are based on a retrospective case-control cohort study, and the performance of the two-step method based on the protein panel (OncoSeek) is improved. Further, we use theoretical calculations to estimate the performance and benefits of this method in the real world. To make the calculation more realistic and reliable, we refer to the real-world study published by the American GRAIL company (PATHFINDER, as shown in Table 3), and according to this real-world study, the incidence of cancer in people aged 50 and above is 1.9%. We also use the same cancer incidence rate for calculation. At the same time, we also found that the cancer screening performance of GRAIL's technology in the real world has a significant decrease in sensitivity compared to the retrospective cohort study (CCGA). Therefore, based on this result, we reduce the sensitivity of our method by the same proportion based on the sensitivity obtained in the retrospective cohort, as the final real-world sensitivity. We assume that we want to do a population of 5 million (require more than 50 years old), and finally we compare the OncoSeek method alone and the innovative use of the two-step method: although there is a slight decrease in sensitivity, the specificity has been significantly improved, and the number of false-positive patients has decreased from 441,450 to 49,050, a decrease of about 10 times. Greatly reduce the psychological problems and subsequent follow-up costs caused by false positives. Moreover, the 468,050 primary screening positive subjects account for about 9.4% of the total screening population, compared with the method of directly using NGS for cancer screening, which greatly reduces the screening cost (~6 times) and there is no significant difference in the final number of true-positive patients screened.

[0186] Currently, there is no method to promote early screening of multiple cancers in domestic and foreign clinical guidelines. Therefore, we compared the current technology of the GRAIL company in the United States, which is currently at the forefront and has published the largest sample size of multiple cancer early screening companies. Compared with this, in terms of cancer early screening performance, the sensitivity of our two-step method is slightly lower, and the specificity is basically the same. However, the cost of detecting 5 million people is about 6 times less than GRAIL. The cost of successfully detecting a cancer patient, we need $31,828, while GRAIL is 5 times more.

[0187] Table 3

[0188] 1. Klein EA, Richards D, Cohn A, et al. Clinical validation of a targeted methylation-based multi-cancer early detection test using an independent validation set. Annals of Oncology. 2021;32(9):1167-1177. doi: 10.1016 / j.annonc.2021.05.806

[0189] 2. Schrag, D., Beer, T. M., McDonnell, C. H., Nadauld, L., Dilaveri, C. A., Reid, R., Marinac, C. R., Chung, K. C., Lopatin, M., Fung, E. T., & Klein, E. A. (2023). Blood-based tests for multicancer early detection (PATHFINDER): A prospective cohort study. The Lancet, 402(10409), 1251-1260.

[0190] * The cancer incidence in the simulation of the real world study was according to the cancer incidence reported in the real world study (PATHFINDER) of GRAIL.

[0191] ** According to the proportion of sensitivity reduction from the retrospective cohort study (CCGA) to the real world study (PATHFINDER) of GRAIL, the sensitivity of our method in the real world was estimated in proportion.

[0192] Example 6: Multi-dimensional multi-omics method:

[0193] We developed a multi-dimensional multi-omics technology that includes proteomics and cancer cfDNA genomic features, and uses an AI model to ultimately predict the risk value of the subject suffering from cancer, wherein the proteomics POC calculation method is the same as in Example 1. And the genome includes cancer genomic copy number variation (CNA) features, fragmentomics (FS) features, differences in terminal sequence features, and cancer-related viral content. The experimental process of genomics: from blood (or similar cerebrospinal fluid and other liquids) separation, extraction of free cfDNA for library construction, sequencing. Here, the extraction, library construction, and sequencing methods are not specifically limited. You can also refer to the experimental methods in the patent number: CN112397143B, "Method for predicting tumor risk value based on plasma multi-omics multi-dimensional features and artificial intelligence", to obtain the sequencing results. You can also refer to our previous article on early screening of lymphoma (Chang Y, Li S, Li Z, et al. Non-invasive detection of lymphoma with circulating tumor DNA features and protein tumor markers. Frontiers in Oncology. 2024; 14. doi: 10.3389 / fonc.2024.1341997). Based on the methods in the article, we analyzed the 1197 samples (617 cancer samples and 580 healthy samples) collected in Example 3 above using the methods in the article to calculate the number of abnormal CNA bins (chromosome instability index) in the CNA dimension, and the long and short fragment ratio of the 5M region of the chromosome combined with the length fragment ratio to construct a training model to calculate the FS dimension value. For the calculation of the long and short fragment ratio of the 5M region of the chromosome, please refer to the method in the published article (Mathios D, Johansen JS, Cristiano S, et al. Detection and characterization of lung cancer using cell-free DNA fragmentomes. Nat Commun. 2021; 12(1): 5060. doi: 10.1038 / s41467-021-24994-w). Compared with the article, we added the short fragment ratio of the whole (all genomic regions) as a feature when calculating the FS value at the end. Finally, we calculated the sequencing read proportion of cancer-related viruses using the method in the article. The article is only for early screening of lymphoma, and lymphoma is only related to EBV, so the article only calculates the sequencing read proportion of EBV, while Example 3 is for early screening of multiple cancers, so we calculate the sequencing read proportion of cancer-related viruses EBV, HBV, and HPV using the same method.

[0194] Meanwhile, we also innovatively optimized the dimension of end motif, as follows:

[0195] Predicting cancer samples by the score of end motif (EMS)

[0196] According to the example of FIG. 7, several bases (1-8 bp) at the 5' end or a certain number of bases (1-6 bp) upstream and downstream of the breakpoint are extracted as the end motif. Specifically, as shown in FIG. 5, the information of 2 bp bases at the 5' end of reads is bases CG; the information of 3 bp bases at the 5' end of reads is bases CGA; the information of 4 bp bases at the 5' end of reads is bases CGAC. The information of 2 bp bases upstream and downstream of the breakpoint of reads is bases CGTC, which means that 2 bp upstream and downstream of the breakpoint are taken respectively to combine 4 bp base information; the information of 3 bp bases upstream and downstream of the breakpoint of reads is bases ACGTCG, which means that 3 bp upstream and downstream of the breakpoint are taken respectively to combine 6 bp base information.

[0197] In this embodiment, 4 bases at the 5' end are taken as an example of the end motif (the number of bases can be modified according to requirements). The length of each cfDNA fragment and the corresponding end motif are obtained. Finally, the number and proportion of the end motif of the cfDNA in the group under different fragment lengths are selected.

[0198] The score value (EMS) of the end distribution deviation of the plasma sample to be tested is obtained by using the newly developed EMS calculation method.

[0199] wherein i represents the type of end motif, Pi represents the proportion of end motif i in all end motifs of cfDNA fragments of the detection sample, is the average (or median) value of the proportion of end motif i in the control sample.

[0200] (1) According to the above method, the frequency of each end sequence of cfDNA fragments in the first nucleosome fragment size range (50-250 bp) and the score value EMS of the degree of deviation of the end sequence distribution are obtained. Wilcoxon rank sum test (or t-test) is used to determine the statistically significant difference of the frequency of each end motif between cancer patients and healthy individuals. The end motifs with corrected q-value less than 0.001 are selected to construct a classifier to distinguish cancer patients and healthy subjects. In order to minimize the problem of overfitting and reduce the number of end motifs, the feature importance of each end motif is evaluated. Subsequently, we arrange these features in descending order of importance and systematically merge them one by one until the classification performance tends to be stable.

[0201] (2) Based on the score value EMS of the degree of deviation of the end sequence distribution of cfDNA fragments in the first nucleosome fragment size range and the screened end sequence frequency features, a machine learning method such as logistic regression is used for modeling. In the modeling process, 10-fold cross-validation method is used and repeated 100 times to improve the detection accuracy, and the average value of 100 model cancer prediction values is defined as the end sequence prediction (EMP) value.

[0202] Finally, according to the values of protein dimension PTMs calculated above (POC values calculated from OncoSeek), the values of CNA and FS dimensions calculated in the reference, the proportion of sequencing reads reads of cancer-related viruses, and the optimized EMP value, as dimension feature values, a conventional linear fitting method or machine learning is used for multi-cancer screening. This embodiment uses a linear fitting model method to integrate 5 dimension features to calculate CRS value, i.e. CRS = a*CNA + b*FS + c*EndMotif + d*Virus + e*PTMs (a, b, c, d and e represent the weight coefficients of CNA, FS, EndMotif, Virus and PTMs, respectively). For 1197 samples (617 cancer samples and 580 healthy samples), the maximum AUC value is used as the evaluation standard, and the grid search method is used to determine the best weight coefficients of the 5 dimensions, and the CRS value is calculated by bringing the formula. The receiver operating characteristic curve (ROC) is used to evaluate the cancer screening performance of CRS, as shown in Figure 8. Finally we obtain 66.3% sensitivity under 97.9% specificity. Based on these predicted true positives, we select them to construct a cancer (TOO) prediction model using the above multi-dimensional features.

[0203] TOO cancer prediction

[0204] A prediction cancer tissue of origin model was constructed using the true positive cancer samples predicted from Example 6. With the idea of model ensemble, only using the samples of 8 major cancer types (the sample size of rare cancer types is too small), the RF method was used to establish a TOO prediction model for the three dimensions of CNA, fragmentation pattern and end sequence of the samples, respectively, while the protein dimension was the reference article for establishing a TOO prediction model (Luan Y, Zhong G, Li S, et al. A panel of seven protein tumour markers for effective and affordable multi-cancer early detection by artificial intelligence: a large-scale and multicentre case-control study. eClinicalMedicine. 2023; 61: 102041. doi: 10.1016 / j.eclinm.2023.102041). Since the sample size of each cancer type is unbalanced, a descending sampling method is used to balance the sample size of each cancer type. FS dimension: the ratio of short fragments to long fragments calculated for each 5Mb region based on Example 6 was used to construct the TOO model of the FS dimension, and the modeling was performed in a five-fold cross-validation manner and repeated 10 times. For the CNA dimension, the logR ratio of each 5Mb region was calculated according to the final CAN detection result, which was normalized by subtracting the mean value and dividing by the standard deviation, and then the normalized Z-score value of each 5Mb region was used as an input feature to construct the CNA dimension of the TOO model, and the modeling process was repeated 10 times. End motif dimension: the frequency of 256 end sequences was used to construct the TOO model of the end sequence dimension, and the modeling was performed in a five-fold cross-validation manner and repeated 10 times. The TOO prediction value from the protein dimension was integrated with the average prediction value of each cancer type of the CNA, FS and end sequence model by ensemble modeling method, which integrated the prediction results of the TOO model of each dimension and obtained the final prediction result. The highest score was selected as the most likely source site, and the second highest as the second possible source site (Top1 and Top2), and the standardized prediction score as the final TOO score. Based on the constructed TOO model, the 260 true positive cancer patients screened in Example 4 above were subjected to TOO prediction. Since the above method is a multi-cancer early screening method, some cancers are not included in the TOO model, and finally 207 positive true positive patients were selected for TOO prediction, and the cancer tracing accuracy of TOP1 was 69.6%, while the cancer tracing accuracy of TOP2 was 83.1%.

[0205] While the specific embodiments of the application have been described in detail, those skilled in the art will appreciate that various modifications and substitutions can be made thereto without departing from the scope of the application as set forth in the claims below. The scope of the application is defined by the appended claims and any equivalents thereof.

[0206] In the description of the specification, reference can be made to terms such as "one embodiment", "some embodiments", "certain embodiments", "an example", "a specific example", or "some examples" etc. It is understood that the specific features, structures, materials or characteristics that are described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. Descriptive terms such as "in one embodiment", "in some embodiments", "in certain embodiments", "in an example", "in a specific example", or "in some examples" etc. in the specification do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.

Claims

1. A computer-implemented method of stepwise prediction of cancer risk, characterized in that, The method comprises: (1) determining a cancer signal score based on the level of a biomarker in a blood sample of a subject; (2) comparing the cancer signal score with a predetermined threshold to determine a screening positive subject; (3) further determining a cancer risk and cancer type prediction for the screening positive subject using a NGS method; wherein the biomarker comprises at least one selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

2. The method of claim 1, wherein, The predetermined threshold is selected from a cancer signal score upper and lower threshold determined based on a predetermined specificity range; Optionally, the cancer signal score within the predetermined range of the cancer signal score is an indication that the subject is screening positive.

3. The method of claim 1, wherein, Step (1) is performed by: i) selecting a plurality of parameters including at least the quantitative level of the biomarker into a first machine learning model; ii) selecting at least one machine learning algorithm to train the first machine learning model; iii) determining the cancer signal score.

4. The method of claim 3, wherein, The plurality of parameters comprises at least one clinical parameter; Optionally, the clinical parameter comprises at least one selected from the group consisting of age, gender and smoking status; Optionally, the plurality of parameters further comprises at least one selected from the group consisting of X-ray imaging, mammography, computed tomography and magnetic resonance imaging.

5. The method of claim 3, wherein, The biomarker is selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CEA and CYFRA 21-1; Optionally, the biomarker is selected from the group consisting of AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

6. The method of claim 3, wherein, The quantitative level of the biomarker is normalized by a modified Z-Score, which is obtained by calculating the difference between the observed value and the median divided by the median absolute deviation.

7. The method according to any one of claims 3-6, characterized in that, The machine learning algorithm comprises GLM, GBM, RF and SVM; Optionally, the first machine learning model is trained using GLM; Optionally, the first machine learning model is obtained by training at least two selected from the group consisting of GLM, GBM, RF and SVM; Optionally, the method comprises applying GLM to an ensemble of at least two machine learning algorithms.

8. The method of claim 1, wherein, The method further comprises biomarker outlier analysis.

9. The method of claim 1, wherein, The further determining a cancer risk and cancer type prediction using a NGS method comprises: constructing a cancer risk model and a cancer type prediction model using a machine learning method based on a cancer-specific methylation site set, and determining a cancer risk and cancer type prediction for the screening positive subject.

10. The method of claim 1, wherein, The further determining a cancer risk and cancer type prediction using a NGS method comprises: Based on the multi-dimensional data of the blood sample of the initial screening positive subject, a multi-dimensional cancer risk determination model and a cancer type prediction model are constructed using a machine learning method, and the cancer risk determination and cancer type prediction of the initial screening positive subject are performed.

11. The method of claim 10, wherein, The multi-dimensional data includes at least one of the following: Based on the cancer signal score (POC) of the biomarker; and Based on the following information obtained by cfDNA whole gene low-depth sequencing: End motif score (EMS) of the end motif of the cfDNA fragment; Chromosome instability index (CNA); And fragment size (FS) measurement of the cfDNA fragment.

12. The method of claim 11, wherein, The multi-dimensional data further includes one or more viral measurements of cancer-related viral abundance.

13. The method of claim 12, wherein, The one or more cancer-related viruses include one or more of HBV, HPV, EBV, HHV, HIV or JCPyV.

14. The method according to any one of claims 10 to 13, characterized in that, The multi-dimensional cancer risk determination model is used to process the model input including EMS and at least one selected from POC, CNA, FS, and viral measurement.

15. The method of claim 1, wherein, The method can detect the presence of at least one type of cancer at an early stage.

16. A method of monitoring cancer therapy, characterized by, It comprises: Predicting cancer risk based on the method of any one of claims 1-15; And Monitoring the subject treated for cancer using cancer risk.

17. A computer-implemented apparatus for stepwise prediction of cancer risk, the apparatus comprising: It comprises: A cancer signal score determination unit for determining a cancer signal score based on the level of a biomarker in a blood sample of a subject; An initial screening positive subject determination unit for comparing the cancer signal score with a predetermined threshold to determine an initial screening positive subject; A cancer type determination unit for further determining cancer risk and predicting cancer type based on the initial screening positive subject using an NGS method; Wherein the biomarker includes at least one selected from AFP, CA125, CA15-3, CA19-9, CA72-4, CEA, CYFRA 21-1, ProGRP, SCCA and PSA.

18. A computer program product, characterised in that, The computer program product comprises computer instructions, when part or all of the computer instructions are run on a computer, so that the method of any one of claims 1-15 is performed.

19. A computing device, comprising: It comprises: A processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program to realize the method of any one of claims 1-15.

20. A computer-readable storage medium, characterized in that, The storage medium comprises computer instructions, when the instructions are executed by a computer, so that the computer realizes the method of any one of claims 1-15.

Citation Information

Patent Citations

  • Lung cancer and colorectal cancer gene detection method based on NGS method

    CN116042790A

  • Computer implementation method for detecting abnormal signal quantification of blood sample to be detected

    CN117831690A

  • Method for predicting cancer risk value based on multi-omics and multidimensional plasma features and artificial intelligence

    US20220136062A1

  • Methods and software systems to optimize and personalize the frequency of cancer screening blood tests

    US20230223145A1