Machine learning method for curating biomarker panel, for screening of cancer(s) and system therefor

A machine learning method using NLP and reinforcement learning curates a biomarker panel for accurate, non-invasive cancer screening, addressing the limitations of current diagnostic methods by providing a scalable and integrated solution for early cancer detection.

WO2026083439A1PCT designated stage Publication Date: 2026-04-23BIOMARKIQ SCIENTIFIC TECHNOLOGIES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIOMARKIQ SCIENTIFIC TECHNOLOGIES PTE LTD
Filing Date
2025-10-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current cancer diagnostic methods are invasive, costly, and lack a scalable, integrated solution for early detection across multiple cancer types, often leading to late-stage diagnoses and low survival rates.

Method used

A machine learning method for curating a biomarker panel using a computational system that integrates with existing lab infrastructure, employing NLP and reinforcement learning to identify relevant scientific literature, validate biomarkers, and integrate clinical data for accurate cancer screening.

Benefits of technology

Enables non-invasive, cost-effective, and efficient pan-cancer screening with high sensitivity and specificity, adaptable across various healthcare settings without specialized training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025051649_23042026_PF_FP_ABST
    Figure IN2025051649_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a machine learning method for curating biomarker panel, for screening of cancer(s) and system therefor, addressing the critical need for early detection through a combination of innovative biomarker research, advanced computational algorithms and integration with existing healthcare infrastructure. The biomarker panel is curated to facilitate pan-cancer screening of the patient with high sensitivity and specificity using the ML method deploying MoE (Mixture of Expert) models for accurate screening, thereby enhancing diagnostic accuracy and also empowering clinicians with actionable insights, thereby improving patient prognosis and facilitating personalized treatment strategies.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] BM131025

[0002] MACHINE LEARNING METHOD FOR CURATING BIOMARKER

[0003] PANEL, FOR SCREENING OF CANCER(S) AND SYSTEM THEREFOR

[0004] FIELD

[0005] The present disclosure in general relates to health sciences, and particularly to pancancer screening using biomarker panel and computer-implemented system and method.

[0006] BACKGROUND

[0007] Cancer remains one of the leading causes of mortality globally. With an increasing number of cases detected in advanced stages, there is a critical need for effective, non-invasive, and early diagnostic tools. Currently, diagnostic solutions vary by cancer type, but many rely on invasive procedures or advanced imaging techniques that come with several drawbacks, including late detection, specialized equipment, and high operational costs.

[0008] For example, over 90% of lung cancer cases in India are identified at advanced stages. Similar challenges exist for breast, prostate, and colorectal cancers, where late-stage detection often leads to lower survival rates. Early detection can drastically improve outcomes, but the lack of an integrated, scalable solution poses a significant barrier to healthcare systems worldwide.

[0009] Most current diagnostic methods focus on circulating tumor cells, cancer stem cells, or cell -free DNA and require specialized personnel and training. These methods do not offer comprehensive screening across multiple cancer types using a single platform.

[0010] The inventors of this disclosure propose a revolutionary solution — a universal screening platform applicable to prominent cancers. The system leverages advanced biomarker research and Computational models for early detection. The kit, system, and method / process are designed to integrate seamlessly with existing BM131025 lab infrastructure, allowing for rapid adoption in healthcare settings without specialized training.

[0011] OBJECTS

[0012] It is an object of the present disclosure to provide a machine learning method for curating biomarker panel for use in screening of cancer(s).

[0013] It is an object of the present disclosure to provide a biomarker panel that can be used for pan-cancer screening of the patient.

[0014] It is an object of the present disclosure to provide a method for screening of cancer(s) deploying a computer-implemented system.

[0015] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which accounts for individual risk factors and patient clinical parameters.

[0016] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which seamlessly integrates with standard or existing laboratory infrastructure.

[0017] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which can be handled without any specialized personnel training.

[0018] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which is scalable.

[0019] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which is adaptable across various healthcare settings.

[0020] It is an object of the present disclosure to provide a kit, system and method for the screening and diagnosis of prominent cancers which provides accessible, efficient, and effective diagnostic solutions to the healthcare community.

[0021] SUMMARY

[0022] The present invention, in a first aspect, provides a machine learning method (MLB) for curating biomarker panel. The biomarker panel being adapted for use in BM131025 screening of cancer(s) in a patient. The machine-learning method (MLB) is implemented by a computational system, the machine learning method (MLB) comprising the steps of: a) Retrieving and collecting scientific literature from one or more data sources to create a comprehensive repository of articles related to biomarkers associated with cancer using data acquisition techniques; b) Shortlisting and filtering out relevant scientific literature (or articles) based on title / abstract screening by applying pre-defined triage guidelines (or criteria) using Natural Language Processing (NLP) techniques via a machine learning (ML) model, wherein the ML model is trained to automatically predict triage decisions and classify the articles based on their pre-defined guidelines to arrive at a relevant dataset, and thereafter human expert reviews the relevant dataset to validate predictions made by the ML model, and wherein further the ML model is iteratively re-trained on reinforcement learning from human feedback (RLHF) to validate and enhance accuracy of predictions made therefrom; c) Reviewing the relevant dataset by screening full-text thereof by applying predefined secondary guidelines (or criteria) via the ML model trained thereon and performing quality assessment thereafter via manual review to ensure reliability and accuracy of information being processed, wherein the relevant dataset which passes through the quality assessment further undergoes a secondary analysis (manual review) for further validation thereby offering an additional layer of scrutiny; d) Performing in-depth analysis of the relevant dataset obtained at step c, against pre-defined detailed analysis guidelines (or criteria) by the ML model trained thereon to extract multiple data elements therefrom, and performing quality assessment thereafter via manual review followed by re-training of the ML model via RLHF to ensure reliability and accuracy of information being processed whilst ensuring adaptation to emerging terminologies; e) Assigning quantifiable score to each candidate biomarker identified at step d using a weighted scoring system adapted by the ML model, wherein the scores are based on comprehensive evaluation of their diagnostic performance in early-stage cancer detection, study robustness (evidence base), biological characteristics (age, BM131025 histology and stage, gender, race, ethinicity) and clinical implications thereby integrating the information obtained from the data elements extracted at step d while arriving at a final biomarker score to ensure a holistic assessment of each biomarker’s utility; f) Aggregating and synthesizing these final biomarker scores across the relevant dataset into a composite score that represents each candidate biomarker using the weighted scoring system by the ML model, wherein the composite score is computed by averaging and weighing the scores obtained at step e by the biomarker’s diagnostic performance and no. of evidences (no. of studies) in support thereof to arrive at the final composite score; g) Categorization of the candidate biomarkers into distinct performance profiles (or category) such as specific guardian (high specificity), sensitive detector (high sensitivity), balanced performer (balanced specificity and sensitivity) and early scout (in early-stage cancer detection) based on their diagnostic performance; h) Ranking of the candidate biomarkers within each of the performance category based on the biomarker’s composite scores using feature selection techniques, wherein the biomarkers with high composite scores within the said performance category along with a large evidence base gets a ‘high priority rank’ and vice-a- versa, and i) Validating the ranked biomarkers through cross-validation (or orthogonal validation) across standard clinical databases to confirm their diagnostic performance for early-stage cancer detection.

[0023] The steps (a) to (i) are repeated for each cancer type to arrive at the biomarker panel with high efficiency in screening of early-stage cancer(s) in the patient. The data sources include multiple journals, databases or repositories from where scientific articles are sourced. The data acquisition techniques are configured to access and integrate published scientific literature from multiple journals, databases, or repositories. The Natural Language Processing (NLP) techniques comprises of any one selected from Automated triage, Long Short-Term Memory (LSTM) Networks with Reinforcement Learning with Human in Loop Feedback mechanism and BM131025 combinations thereof. The data elements refer to multiple fields such as cohort characteristics including patient / study population, patient demographic data, and patient occupational history, disease histology (sub-type, stage), sample / specimen type and biomarker attributes such as name, type, diagnostic performance based on Area Under Curve (AUC), sensitivity, specificity, and statistical data (p-value), expression profile, and clinical implications (prognostic and therapeutic) thereof in early-stage cancer detection.

[0024] The present invention, in a second aspect provides, a biomarker panel for screening of cancer(s), wherein the panel comprises combination of:

[0025] CA19-9 (Carbohydrate Antigen 19-9), CA125 (Cancer Antigen 125), TAG-72 (Tumor-Associated Glycoprotein 72), CEA (Carcinoembryonic Antigen), HE4 (Human Epididymis Protein 4), SCC-Ag (Squamous Cell Carcinoma Antigen), PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist-II), NSE (Neuron-Specific Enolase), CYFRA-21 (Cytokeratin 19 Fragment (CYFRA 21-1)), AFP (Alpha-Fetoprotein), Total PSA (PSA) (Prostate-Specific Antigen (total)), CRP (C-Reactive Protein), ProGRP (Pro-Gastrin-Releasing Peptide), HSP90 AA1 ( Heat Shock Protein 90 Alpha Family Class A Member 1), AFP - L3 (Lens culinaris agglutinin-reactive fraction of Alpha-Fetoprotein), VEGFA ( Vascular Endothelial Growth Factor A), Neutrophil to Lymphocyte Ratio (NLR) and fPSA Free Prostate- Specific Antigen).

[0026] The cancer(s) that can be screened using said biomarker panel are: Breast cancer, Biliary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer. The present invention further provides cancerspecific individual biomarker panels for each of the afore-mentioned cancer screening thereby offering high accuracy. These cancer-specific panels are subset of the biomarker panel of 18 biomarkers shared above, from which these are derived.

[0027] The present invention, in a third aspect provides, a method (MLc) for screening of cancer(s). The method (MLc) being a computer-implemented method executed BM131025 on any computational system and adapted for screening of cancer(s), the method (MLc) comprising the steps of: a) Obtaining a biological sample from a patient; b) Measuring values (or levels) of one or more biomarkers in the biological sample using a kit, the biomarkers being selected from a biomarker panel for cancer(s); c) Obtaining clinical data from the patient, wherein the clinical data comprises of age, sex, family history, occupational exposure, co-morbidities and life style factors such as diet, exercise regimen, smoking history and intensity, alcohol consumption and Air Quality Index (AQI) for city and place of residence of the patient; d) Feeding the biomarker values and the clinical data into a machine-learning based system, and e) Analysing of the above data by the machine-learning based system using a computational machine learning model deploying a trained Mixture of Experts (MoE) framework to perform the steps of: i) Classifying the biological sample into any one category from cancer and noncancer (healthy) by giving a binary output (0-healthy, 1- specific cancer type) in form of a P_cancer score, and in event the cancer is detected, determining stage thereof based on the measured biomarker values by an expert model trained on patient dataset for stage-classification (Stage 1 to 4) and cancer subtype classification; ii) Computing a confidence score for TOO (Tissue of Origin), if presence of cancer is detected at step i); iii) Repeating the steps i) and ii) by each expert of a plurality of experts of the MoE framework to obtain a set of predictions with respect to P cancer and TOO confidence scores, wherein each expert functions as a submodel being trained on different datasets and specializes by organ, cancer-type and biomarker panel specific therefor; iv) Evaluating the set of predictions by a gating network (or Meta model / stacking model) based on the clinical data to compute a weighted sum thereof and arrive at a final confidence score indicative of probability of presence / absence of cancer, and informing the sub-type and stage, if present; and BM131025 v) Computing a final prediction score based on the final confidence score integrating with the clinical data, if outcome of the step iv) is positive indicating presence of cancer, and thereafter, vi) Generating a clinical report summarizing findings obtained at step v) and time stamped when the biological sample is screened.

[0028] The method (MLc), wherein the computational model operates in any of three modes- Mode A, Mode B and Mode C, wherein in Mode A, biomarker values for the biomarker panel (18 biomarkers) in entirety is fed into the machine -learning based system which performs the steps i) through vi) to output the final prediction score with clinical report summarizing the findings therefrom; in Mode B which operates as a tier-based approach, where, at level 1, measured values for all biomarkers selected from Tier 1 biomarker set (gender-specific) are fed into the machine-learning based system which follows through steps i) to vi) to predict probability of cancer, and thereafter, in an event output of the level 1 is indeterminate at step iv), then level 2 is initiated where measured values for one of more biomarkers drawn from Tier 2 biomarker set are fed into the machinelearning based system which follows through steps i) to vi) to predict probability of cancer, and further thereafter, in an event output of the level 2 is still indeterminate at step iv), then level 3 is initiated where measured values for system-recommended one or more biomarkers selected from the biomarker panel (of 18 biomarkers), and which were not measured earlier are fed into said system which follows through the steps i) to vi) to obtain the final prediction score and the clinical report, and in Mode C, operates as a restrictive cancer-specific approach, where measured values for biomarkers from cancer-specific biomarker panel of any of claims 10- 35, is fed into the machine -learning based system which activates the expert model specific to said cancer-type while restricting contribution from other expert models to perform steps of i) through vi) and output the final prediction score along with the clinical report, and in an event the prediction is indeterminate said system upgrades operation to Mode A or Mode B, thereby ensuring accurate screening of cancer. BM131025

[0029] The present invention, further provides a machine-learning based system for screening of cancer(s), the system being a computational system comprising of i) a processor, and ii) a memory, having a computer program product stored therewithin being executable on the processor, wherein, the processor executes the afore-mentioned computational machine learning model deploying a trained Mixture of Experts (MoE) framework consisting of a plurality of knowledgespecific expert model(s) with a gating network (meta model) for executing steps of the method (MLc).

[0030] Therefore, the present disclosure facilitates in accurate screening of any type of cancer in the patients with high sensitivity and specificity.

[0031] BRIEF DESCRIPTION OF DRAWINGS

[0032] The present disclosure is illustrated in the accompanying non-limiting drawings, throughout which reference letters indicate corresponding parts in the various figures.

[0033] Figure 1 is a flow-chart of the machine learning method for screening of cancer(s) in the patient using the biomarker panel, in accordance with the present invention.

[0034] DETAILED DESCRIPTION

[0035] The present disclosure provides a machine learning method for curating biomarker panel, for screening of cancer(s) and system therefor. Particularly, the present disclosure provides the biomarker panel, and machine learning method and system for pan-cancer screening and diagnosis of cancer implementing said biomarker panel.

[0036] The features, functionalities, raw material, components, dimensions, conditions of operation, steps of operation, end uses and the like of the kit, system(s) and the method(s) of the present disclosure includes, but are not limited to the disclosure provided herein below.

[0037] Definitions: BM131025

[0038] For convenience, the meaning of certain terms and phrases employed in the specification, examples, and appended claims are provided below.

[0039] The term ‘biological sample’, as used herein, refers to a non-invasive body fluid(s) sample obtained / derived from a patient. Such sample includes, but is not limited to, blood, serum, plasma, urine, saliva, semen, breast exudate, cerebrospinal fluid, tears, sputum, mucous, lymph, cytosols, ascites, pleural effusions, peritoneal fluid, amniotic fluid, bladder washes, and bronchioalveolar lavages, blood cells.

[0040] The term ‘prominent cancers’ or ‘cancer(s)’ includes, but is not limited to, the cancers of breast, respiratory system, brain, genitalia, eye, liver, oral, digestive system, and urinary tract, among others, and their distant metastases.

[0041] The term ‘Air Quality Index (AQI)’, is the system used to warn the public when air pollution is dangerous. It is an index for reporting air quality, and is used to provide information about how polluted the air currently is or how polluted it is forecasted to become. The AOI is location, city, and time- specific.

[0042] The term ‘patient’ or ‘subject’ as used herein, includes mammals (e.g. humans and animals).

[0043] The term “API” stands for Application Programming Interface. APIs are mechanisms that enable two software components to communicate with each other using a set of definitions and protocols (AWS).

[0044] The term “Asymptomatic” refers to having no signs or symptoms of disease.

[0045] AUC: The area under curve (AUC) of a classifier is equivalent to the probability that the classifier will rank a randomly chosen positive instance higher than a randomly chosen negative instance (Fawcett, 2006).

[0046] The term “Biomarker testing” is a laboratory method that uses a sample of tissue, blood, or other body fluid to check for certain genes, proteins, or other molecules that may be a sign of a disease or condition, such as cancer. Biomarker testing can also be used to check for certain changes in a gene or chromosome that may increase BM131025 a person’s risk of developing cancer or other diseases. Biomarker testing may be done with other procedures, such as biopsies, to help diagnose some types of cancer. It may also be used to help plan treatment, find out how well treatment is working, make a prognosis, or predict whether cancer will come back or spread to other parts of the body. Also called molecular profiling and molecular testing (NCI).

[0047] Biomarker: A biological molecule found in blood, other body fluids, or tissues that is a sign of a normal or abnormal process, or of a condition or disease. A biomarker may be used to see how well the body responds to a treatment for a disease or condition. Also called molecular marker and signature molecule (NCI).

[0048] CAP / AMP / ASCO guideline: In the AMP / ASCO / CAP Somatic Variants Guideline, the clinical significance of a variant, detected by either NGS or non-NGS assays, is defined in a tiered system in consideration of three categories of clinical and experimental evidence: diagnostic, prognostic, and therapeutic (PMID: 36503149).

[0049] Cell -free DNA (cfDNA): Circulating free DNA (cfDNA) are degraded DNA molecules that are released into the bloodstream by cells (Technology Network).

[0050] Circulating tumor cells (CTCs): A cancer cell that breaks away from the original (primary) tumor and enters the bloodstream. Circulating tumor cells can travel through the blood and form new tumors in other parts of the body. A sample of blood can be used to detect circulating tumor cells and learn more about the primary tumor. Circulating tumor cells are being used as a biomarker in some types of cancer to help plan treatment or make a likely prognosis. Also called CTC (NCI).

[0051] Data Mining: Data mining is a computer-assisted technique used in analytics to process and explore large data sets. With data mining tools and methods, organizations can discover hidden patterns and relationships in their data. Data mining transforms raw data into practical knowledge (AWS).

[0052] ELISA: A laboratory technique that uses antibodies linked to enzymes to detect and measure the amount of a substance in a solution, such as serum. The test is done BM131025 using a solid surface to which the antibodies and other molecules stick. In the final step, an enzyme reaction takes place that causes a color change that can be read using a special machine. There are many different ways that an ELISA can be done. ELISAs may be used to help diagnose certain diseases. Also called enzyme-linked immunosorbent assay (NCI).

[0053] GDPR: The term ‘personal data’ is the entryway to the application of the General Data Protection Regulation (GDPR). Only if a processing of data concerns personal data, the General Data Protection Regulation applies. The term is defined in Art. 4 (1). Personal data are any information which are related to an identified or identifiable natural person.

[0054] HIPAA: A 1996 U.S. law that allows workers and their families to keep their health insurance when they change or lose their jobs. The privacy rule of the HIPAA protects the privacy of a person’s health information and keeps it from being misused. It gives people the right to receive and review their health records and to choose with whom their health care providers and health insurance companies share their information (including friends, family members, and caregivers). The law also includes standards for setting up and maintaining secure electronic health records. Also called Health Insurance Portability and Accountability Act and Kassebaum Kennedy Act (NCI). The Health Insurance Portability and Accountability Act (HIPAA) sets the standard for sensitive patient data protection.

[0055] KM plotter: The Kaplan Meier plotter is capable of assessing the correlation between the expression of all genes (mRNA, miRNA, protein, & DNA) and survival in 35k+ samples from 21 tumor types (Kaplan Meier Plotter).

[0056] Liquid Biopsy: A laboratory test done on a sample of blood, urine, or other body fluid to look for cancer cells from a tumor or small pieces of DNA, RNA, or other molecules released by tumor cells into a person’s body fluids. Liquid biopsy allows multiple samples to be taken over time, which may help doctors understand what kind of genetic or molecular changes are taking place in a tumor. A liquid biopsy may be used to help find cancer at an early stage. It may also be used to help plan BM131025 treatment or to find out how well treatment is working or if cancer has come back (NCI).

[0057] Meta-analysis: A process that analyzes data from different studies done about the same subject. The results of a meta-analysis are usually stronger than the results of any study by itself (NCI).

[0058] Mixture of Experts: Mixture of experts (MoE) is a machine learning approach that divides an artificial intelligence (Al) model into separate sub-networks (or “experts”), each specializing in a subset of the input data, to jointly perform a task (IBM).

[0059] ML: Machine learning is a field of study that looks at using computational algorithms to turn empirical data into usable models.

[0060] NCCN: The National Comprehensive Cancer Network (NCCN) is a not-for-profit alliance of 33 leading cancer centers devoted to patient care, research, and education (NCCN).

[0061] NLP: Natural language processing, or NLP, combines computational linguistics — rule-based modeling of human language — with statistical and machine learning models to enable computers and digital devices to recognize, understand and generate text and speech (IBM).

[0062] Outcomes: A specific result or effect that can be measured. Examples of outcomes include decreased pain, reduced tumor size, and improvement of disease (NCI).

[0063] Pack year: A way to measure the amount a person has smoked over a long period of time. It is calculated by multiplying the number of packs of cigarettes smoked per day by the number of years the person has smoked. Lor example, 1 pack year is equal to smoking 1 pack per day for 1 year, or 2 packs per day for half a year, and so on (NCI). BM131025

[0064] Paucisymptomatic: Paucisymptomatic refers to a medical condition or disease where an individual experiences only a small number of symptoms or mild symptoms that may not be easily identifiable or noticeable.

[0065] Prospective Study: A research study that follows over time groups of individuals who are alike in many ways but differ by a certain characteristic (for example, female nurses who smoke and those who do not smoke) and compares them for a particular outcome (such as lung cancer) (NCI).

[0066] Retrospective Study: A study that compares two groups of people: those with the disease or condition under study (cases) and a very similar group of people who do not have the disease or condition (controls). Researchers study the medical and lifestyle histories of the people in each group to learn what factors may be associated with the disease or condition. For example, one group may have been exposed to a particular substance that the other was not. Also called case-control study (NCI).

[0067] Sensitivity: In medicine, sensitivity may describe how well a test can detect a specific disease or condition in people who actually have the disease or condition. No test has 100% sensitivity because some people who have the disease or condition will not be identified by the test (false-negative test result). Sensitivity may also refer to the way the body reacts to the environment or to drugs, chemicals, or other substances. For example, a person who is sensitive to the sun may have skin that bums easily or get a rash when exposed to the sun. A person who is sensitive to caffeine may need only small amounts of it to feel its effects (NCI).

[0068] Specificity: When referring to a medical test, specificity refers to the percentage of people who test negative for a specific disease among a group of people who do not have the disease. No test is 100% specific because some people who do not have the disease will test positive for it (false positive) (NCI). BM131025

[0069] Support Vector Machine: A support vector machine (SVM) is a supervised machine learning algorithm that classifies data by finding an optimal line or hyperplane that maximizes the distance between each class in an N-dimensional space (IBM).

[0070] Systematic review: A systematic review is a firmly structured literature review, undertaken according to a fixed plan, system or method. Systematic review methodology is explicit and precise because it aims to minimise bias, thereby enhancing the reliability of any conclusions. It is therefore considered an evidencebased approach. Systematic reviews are commonly used by health professionals, but also policy makers and researchers (Manchester Metropolitan University).

[0071] UAT: User Acceptance Testing (UAT), known as end-user or application testing, validates software in real-world conditions by its intended audience, an essential phase in software development cycle. UAT testing means the usage of the software by people from the target audience and recording and correcting of any defects.

[0072] I. First Aspect: Machine learning method for biomarker panel

[0073] In accordance with a first aspect of the present disclosure, a machine learning method, hereinafter referred to as “the method (MLB)” for curating a biomarker panel, is disclosed. The method (MUB) is implemented by a computational system. The biomarker panel curated using the method (MUB) finds application in pan-cancer screening and diagnosis of cancer(s) in patients. The method (MUB) is delineated into the phases / steps provided herein after, wherein each phase / step is crucial for the integrity and success of the process in identifying and validating biomarkers for early cancer detection. The steps of the method (MLB) are explained hereinbelow with reference to curating a biomarker panel specifically for lung cancer, as an example embodiment to explain in-detail the complete process. However, it is to be noted that these steps hereinbelow of the method (MLB) are to be repeated for curating a separate biomarker panel for each cancer-type (i.e., to arrive at a cancer-specific biomarker panel). Thereafter, the inventors by performing intensive experimentation have combined the cancerspecific biomarker panels to obtain a biomarker panel of the present disclosure, the BM131025 biomarker panel that can be independently used for pan-cancer-types screening of the patient and for diagnoses of any type / subtype of the cancer.

[0074] Steps for curating Biomarker Panel

[0075] Accordingly, the method (MLB) comprises the steps including, but not limited to:

[0076] 1. Article retrieval:

[0077] Objective: Retrieving and collecting scientific literature from one or more data sources to create a comprehensive repository of articles related to biomarkers associated with cancer(s) (in this example, lung cancer) using data acquisition techniques. To systematically gather a comprehensive collection of scientific literature essential for identifying and validating biomarkers for prominent cancers.

[0078] The data sources, for example, includes journals, databases, repositories etc. The data sources specifically includes all scientific literature / articles indexed on PubMed from 2018- 2023, in the following study.

[0079] The data acquisition techniques such as, for example, pulling data via available & ethical channels - like API, web-scrapping complying to the robot.txt guidelines) are configured to access and integrate published scientific literature from multiple journals, databases, or repositories.

[0080] Methodology:

[0081] Utilizing specifically designed keywords and queries, a bespoke web crawler is employed to systematically gather articles from publicly accessible databases. This ensures that a wide array of relevant research is captured. Steps as follows: a) Designing Comprehensive Query Sets: o Three-Tiered Query Structure (KI AND K2 AND K3):

[0082] ■ KI - Cancer Types: Includes a wide range of cancer subtypes such as, not limiting to, Adenocarcinoma, Squamous Cell Carcinoma, Small Cell Lung Cancer, Large Cell Neuroendocrine Carcinoma, among others. BM131025

[0083] ■ K2 - Diagnostic Terms: Incorporates terms related to biological samples and diagnostic processes like, not limiting to, Blood, Serum, Plasma, Sputum, Diagnostic, Detection, Diagnosis, Biomarker, and specific receptor and protein names.

[0084] ■ K3 - Biomarkers and Genetic Terms: Encompasses an extensive list of biomarkers, genes, proteins, and related molecular terms (e.g., C5AR1, CLEC4A, NLRP3, IL-6, TP53).

[0085] Example Query:

[0086] ((Adenocarcinoma AND LUNG) OR Adenosquamous OR Atypical Adenomatous Hyperplasia OR ... OR Mutation OR Substitution) AND (Blood OR Serum OR Plasma OR ... OR Expression OR Mutation OR Substitution) AND (Biomarker OR C5AR1 OR chemotactic receptor OR Mutation OR Substitution) b) Article sourcing: o API-Based Retrieval: The majority of articles are sourced using available APIs from designated databases. o Web Crawler Implementation: A Selenium-based web crawler was developed as a Proof of Concept (PoC) to identify and prevent duplication of articles in the database. The crawler cross-references retrieved papers to ensure they are not already present in the existing database. c) Standardized Nomenclature Handling: o Utilized NCBI and UniProt sources to categorize and streamline gene and protein aliases, ensuring consistency in search parameters. Alias names were also considered. d) Data Volume and Quality Assurance: BM131025 o Extensive Article Collation: A total of 312,658 articles were retrieved, providing a broad dataset for subsequent analysis. o Preliminary Filtering: Initial filters based on publication date, relevance scores, and citation counts were applied to prioritize high- impact and pertinent studies.

[0087] Outcome:

[0088] A comprehensive repository of over 312,658 scientific articles was established, forming the basis for subsequent biomarker evaluation and machine learning model development.

[0089] 2. Automated Triage via NLP

[0090] Objective: Shortlisting and filtering out relevant scientific literature (or articles) based on title / abstract screening by applying pre-defined triage guidelines (or criteria) using Natural Language Processing (NLP) techniques via a machine learning (ML) model. Thus, by leveraging Natural Language Processing (NLP), the present disclosure applies the pre-defined triage guidelines to sift through the fetched articles, identifying those pertinent to biomarkers in early prominent cancer(s).

[0091] The Natural Language Processing (NLP) techniques comprises of, not limiting to, Automated triage, Long Short-Term Memory (LSTM) Networks with Reinforcement Learning with Human in Loop Feedback mechanism.

[0092] Specifically, the ML model (NLP-based) is developed and trained as detailed out hereinbelow to automatically predict triage decisions and classify the articles based on their pre-defined guidelines to arrive at a relevant dataset, and thereafter human expert reviews the relevant dataset to validate predictions made by the ML model. The ML model is iteratively re-trained on reinforcement learning from human feedback (RLHF) to validate and enhance accuracy of predictions made therefrom. This step significantly enhances efficiency by prioritizing content for further review. The methodological steps of achieving the BM131025 above have been given out in section hereinbelow with reference to an example of lung cancer.

[0093] Referring again to our example of lung cancer, this step involves efficiently shortlisting the most relevant articles related to lung cancer biomarkers by applying predefined triage guidelines (or triage decision criteria) using Natural Language Processing (NLP) by following below steps. This ensures that only pertinent studies are selected for further analysis, optimizing the workflow and resource allocation. Methodology: a) Non-limiting embodiment of Natural Language Processing (NLP) Framework: o Libraries and Tools Used:

[0094] ■ spaCy: Utilized for advanced text parsing, tokenization, and entity recognition.

[0095] ■ Transformers (BERT-based models): Employed for contextual understanding and classification tasks.

[0096] ■ NLTK: Used for supplementary text processing and linguistic feature extraction.

[0097] ■ Scikit-learn: Applied for implementing machine learning algorithms and evaluation metrics.

[0098] ■ TensorFlow & PyTorch: For building and training transformer models. o Workflow Integration: The NLP pipeline integrates these libraries to process the titles and abstracts of retrieved articles, enabling automated decision-making based on the defined criteria. b) Text Parsing and Processing: o Preprocessing Steps:

[0099] ■ Tokenization: Breaking down text into tokens (words, phrases).

[0100] ■ Lemmatization: Reducing words to their base or root form. BM131025

[0101] ■ Stop-word Removal: Eliminating common words that do not contribute to the analysis o Entity Recognition:

[0102] ■ Identifying and extracting relevant entities such as diseases, biomarkers, sample types, and patient stages from the text. o Example Parsing:

[0103] ■ Title: "Early Detection of Lung Adenocarcinoma Using Protein Biomarkers in Blood Samples"

[0104] ■ Entities Extracted:

[0105] ■ Disease: Lung Adnocarcinoma

[0106] ■ Biomarkers: Protein Biomarkers

[0107] ■ Sample Type: Blood

[0108] ■ Stage: Early Detection c) Triage Decision Criteria: With reference to our current example of lung cancer, following are the example guidelines on which the ML model (NLP based) was trained to automate shortlisting / tagging of the articles based on title / abstract, and arrive at the relevant dataset. o Inclusion Criteria (select triage decision as accepted, if all of the following criteria are met):

[0109] ■ Language: English

[0110] ■ Diseases: Lung cancer, Pleural mesothelioma, Multiple cancers

[0111] ■ Biomarkers: Presence of broad categories such as proteins, cell-free RNAs, non-coding RNAs, biomarker panels, and exosome vesicles

[0112] ■ Article Type: All publications except Comment, Editorial, Perspective, Retraction

[0113] ■ Sample Collection Technique: Non-invasive methods (Blood, Serum, Plasma) BM131025

[0114] ■ Patients / Population span: Pre-diagnostic, screening, asymptomatic, symptomatic, early (0, I, II), late (III, IV, advanced, metastatic), all stages

[0115] ■ Additional Note: Detection of biomarkers using ELISA is within scope o Exclusion Criteria (Select triage decision as rejected, if any of the following criteria are met):

[0116] ■ Language: Non-English

[0117] ■ Biomarkers: Absence of biomarkers

[0118] ■ Sample Collection Technique: Invasive methods and biopsy

[0119] ■ Patient Number: Less than 10 (unless not mentioned)

[0120] ■ Malignancies: Other than specified lung-related cancers or not related to any cancer / malignancy

[0121] ■ Study Basis: Solely in vitro or in vivo experiments

[0122] ■ Article Type: Comment, Editorial, Perspective, Retraction o On-Hold Criteria (Select triage decision as on hold, if any of the following criteria are met):

[0123] ■ Management Index: Therapeutic, Prognostic, Prognostic / Therapeutic

[0124] ■ Disease: Pleural mesothelioma

[0125] ■ Biomarker Type: Metabolites, microbiome, cell-free DNA (cfDNA), circulating DNA (ctDNA), Circulating tumor cells (CTCs), Cytologically abnormal cells (CAGs), Cancer Stem Cells (CSCs), DNA

[0126] ■ Sample Type: Sputum, Urine, Breathe, Nasal swab, Lavage d) Implementation of Triage Guidelines: o Inclusion Evaluation:

[0127] ■ Articles are flagged as "Accepted" if they meet all inclusion criteria. BM131025

[0128] Example: An article titled "Detection of Protein Biomarkers in Serum for Early-Stage Lung Cancer" would be accepted. o Exclusion Evaluation:

[0129] ■ Articles are flagged as "Rejected" if they meet any exclusion criteria.

[0130] Example: An article published in a non-English journal or focusing solely on in vitro studies would be rejected. o On-Hold Evaluation:

[0131] ■ Articles meeting on-hold criteria are flagged for further review.

[0132] Example: Studies involving cell-free DNA biomarkers or those related to pleural mesothelioma are placed on hold for subsequent analysis. e) Technical Execution: o Model Training:

[0133] ■ Data Labelling: Articles are manually labelled based on the inclusion, exclusion, and on-hold criteria to create a training dataset.

[0134] ■ Model Selection: A BERT-based transformer model is finetuned for classification tasks to predict triage decisions. However, it is evident to a skilled person that any other NLPbased machine learning model can be trained to perform said analysis.

[0135] ■ Feature Engineering: Relevant features such as presence of specific biomarkers, sample types, and patient stages are extracted and used as input for the model. o Evaluation Metrics:

[0136] ■ Precision and Recall: Ensuring high precision to minimize false positives and high recall to capture all relevant articles.

[0137] ■ Fl-Score: Balancing precision and recall for optimal performance. BM131025 o Non-limiting examples of Frameworks Utilized:

[0138] ■ TensorFlow / PyTorch: For building and training transformer models.

[0139] ■ spaCy and NLTK: For preprocessing and entity extraction.

[0140] ■ Scikit-learn: For implementing classification algorithms and evaluating model performance. f) Human-in-the-Loop (HITL) Framework: o Integration of Human Oversight:

[0141] ■ Human experts review a subset of articles to validate and correct model predictions, ensuring the accuracy of triage decisions.

[0142] ■ Feedback from human reviewers is used to iteratively improve the model’s performance. o Reinforcement Learning from Human Feedback (RLHF):

[0143] ■ Process:

[0144] ■ Human feedback on model predictions is incorporated to refine the decision-making process.

[0145] ■ The model is trained using reinforcement learning techniques where human corrections serve as rewards or penalties to guide future predictions.

[0146] ■ Benefits:

[0147] ■ Enhances model accuracy by addressing nuanced cases that automated systems may misclassify.

[0148] ■ Continuously adapts to evolving terminology and criteria through ongoing human interaction. g) Integration and Deployment: o Pipeline Automation:

[0149] ■ The NLP-based triage system is integrated into the data processing pipeline, automatically processing newly retrieved articles and applying triage decisions. o Continuous Improvement: BM131025

[0150] ■ The ML model is periodically retrained with new labelled data and human feedback to enhance accuracy and adapt to evolving research terminology and biomarker discoveries.

[0151] Outcome:

[0152] The Automated Triage via NLP efficiently shortlists the most relevant articles (in case of current lung cancer example, 146,380 articles were shortlisted from out of 3L+ Articles from the complete set). In this example, these shortlisted articles refer to the relevant data which is completely focused on lung cancer biomarkers by accurately applying predefined inclusion, exclusion, and on-hold criteria listed above. Thus, incorporating a Human-in-the-Loop framework and Reinforcement Learning from Human Feedback (RLHF) ensures the continuous improvement and accuracy of the triage process, reducing the manual effort required for article review.

[0153] 3. Manual Review and Categorization

[0154] Objective: To meticulously assess articles deemed potentially relevant by the Automated Triage via NLP (step 2), ensuring accurate classification of the articles based on their significance and relevance to the inventor’s biomarker discovery objectives adhering to stringent selection or rejection criteria.

[0155] Therefore, this step involves reviewing of the relevant dataset (obtained at above step 2) by screening full-text thereof by applying pre-defined secondary guidelines (or criteria) via the ML model trained thereon. Thereafter, quality assessment (QA) review is performed via manual review to ensure reliability and accuracy of information being processed. The relevant dataset which passes through the quality assessment (QA) further undergoes a secondary analysis (again detail manual review and field population) for further validation thereby offering an additional layer of scrutiny, to land at only the most relevant dataset for further analysis.

[0156] The pre-defined secondary (analysis) guidelines are as follows:

[0157] Full texts of the relevant dataset were screened with the secondary bot and triage app. Articles were accepted only if all criteria listed is met: study compares controls BM131025

[0158] (healthy or healthy + benign) vs cancer patients; an AUC (Area Under Curve) is reported for a relevant, specific protein biomarker; and the sample is blood, serum, or plasma. Articles were rejected if any of the listed criteria is unmet / fail: if biomarker isn't protein, if no cancer patients, or if tissue / other fluids are studied.

[0159] Non-limiting example embodiment of methodology used and steps are as follows:

[0160] Methodology: a) In-House Application Development: o Custom Platform: An in-house application was developed to facilitate the manual review and categorization of articles. This platform streamlines the review process by providing a structured interface for reviewers to evaluate each article systematically. o Attached Resources:

[0161] ■ Screenshots: Illustrate the user interface and workflow of the manual review application.

[0162] ■ Review Sheet: A comprehensive spreadsheet detailing the fields populated during the review process. b) Review Process: o Article Assessment: Each article that passes the Automated Triage via NLP is subjected to a detailed manual review to verify its relevance and significance. o Field Population: Reviewers populate the following fields for each accepted article:

[0163] ■ PMID: PubMed Identifier for the article.

[0164] ■ Title: Title of the research article.

[0165] ■ Email: Contact email of the reviewer.

[0166] ■ Date: Date of the review.

[0167] ■ If Paid? Indicates whether the article is freely accessible or requires payment. BM131025

[0168] ■ Decision: Initial decision based on automated triage (e.g., Accepted, Rejected).

[0169] ■ QA - Email: Quality Assurance reviewer’s email.

[0170] ■ QA - Date: Date of Quality Assurance review.

[0171] ■ QA - Decision: Decision made during Quality Assurance (e.g., Approved, Needs Revision).

[0172] ■ QA - Comment: Comments or notes from the Quality Assurance review.

[0173] ■ SA - Email: Secondary Analysis reviewer’s email.

[0174] ■ SA - Date: Date of Secondary Analysis review.

[0175] ■ Secondary Analysis Decision: Decision from the Secondary Analysis (e.g., Accepted, On Hold, Rejected).

[0176] ■ SA Reason for Rejection / On Hold: Reason for rejecting or placing the article on hold during Secondary Analysis.

[0177] ■ Final Decision: Conclusive decision after all review stages.

[0178] ■ Need Repeat QA: Indicator if the article requires another round of Quality Assurance review. c) Categorization Criteria: o Significance and Relevance: Articles are classified based on how closely they align with the invention’s objectives, focusing on the presence and validity of biomarkers relevant to specific cancer, in this example, lung cancer detection. o Compliance with Guidelines: Adherence to predefined inclusion and exclusion criteria is verified to ensure consistency and reliability in article selection. d) Quality Assurance and Secondary Analysis: o QA Review: A separate Quality Assurance reviewer examines the initial decisions to ensure accuracy and consistency. o Secondary Analysis: Articles that pass QA undergo a secondary analysis (manual review once again) for further validation, providing an additional layer of scrutiny. BM131025 e) Documentation and Tracking: o Review Logs: All decisions and comments are meticulously documented within the in-house application and the attached review sheet, ensuring traceability and accountability. o Attached Sheet: A selected subset spreadsheet is provided, capturing all relevant fields for each reviewed article, facilitating easy reference and auditability.

[0179] Outcome:

[0180] The Manual Review and Categorization phase ensures that only the most relevant and high-quality articles are selected for further analysis. By utilizing a structured in-house application and comprehensive field population, the process maintains high standards of accuracy and consistency, supporting the integrity of the biomarker discovery process.

[0181] 4. In-Depth Article Analysis and Quality Assurance

[0182] Objective: To thoroughly analyse each selected article in alignment with detailed analysis guidelines (or framework) ensuring the reliability and accuracy of the extracted information for the biomarker discovery and validation processes of the present disclosure. Expert reviewers conduct quality assessments to ensure the reliability and accuracy of the information being processed.

[0183] Therefore, this step involves performing in-depth analysis of the relevant dataset (obtained at above step 3), against pre-defined detailed analysis guidelines (or criteria) by the ML model trained thereon to extract multiple data elements therefrom, and performing quality assessment thereafter via manual review followed by re-training of the ML model via RLHE to ensure reliability and accuracy of information being processed whilst ensuring adaptation to emerging terminologies.

[0184] The data elements, refers to multiple fields such as, but not limiting to, cohort characteristics including patient / study population, patient demographic data, and patient occupational history, disease histology (sub-type, stage), sample / specimen BM131025 type and biomarker attributes such as name, type, diagnostic performance based on Area Under Curve (AUC), sensitivity, specificity, and statistical data (p-value), expression profile, and clinical implications (for ex, prognostic and therapeutic implications of biomarker expression) thereof in early-stage cancer detection.

[0185] The pre-defined detailed analysis guidelines and the methodology with the current example of lung cancer, has been explained as follows.

[0186] Methodology: a) Data Analysis Guidelines: o Aim: To analyze and input data from shortlisted articles from the triage set based on comprehensive paper reading and interpretation. b) General Guidelines: o Order of Preference for Article Selection:

[0187] ■ Meta-analyses

[0188] ■ Systematic reviews

[0189] ■ Articles with high patient numbers (>100)

[0190] ■ Articles with high impact factors (>10) o Lower Preference: Publications with fewer patients (10-100) and low impact factors (<3) are given less priority. Reviews are considered after achieving sprint goals. o Data Integrity:

[0191] ■ No cells in the analysis fields should be left empty.

[0192] ■ Avoid special characters in manually filled cells. o Biomarker Data Handling:

[0193] ■ If biomarkers are analyzed in both training and validation cohorts, only data from the validation dataset is curated unless different biomarkers are analyzed in each cohort.

[0194] ■ Papers with only differential expression data without diagnostic parameters are placed on hold.

[0195] ■ Erratum papers require notification before analysis. BM131025 o Controlled Vocabulary Compliance: Do not modify or add new Permissible Values (PV). These PVs were created during data curation to remove subjectivity in annotations. c) Pre-Reading Checks: o Sections to Review:

[0196] ■ Materials & Methods

[0197] ■ Results o Information to Extract:

[0198] ■ Presence of lung cancer patients

[0199] ■ Sample type (blood, serum, plasma)

[0200] ■ AUC values

[0201] ■ Survival metrics (PFS, OS, DFS)

[0202] ■ Comparison groups (healthy controls vs. lung cancer) d) Manual Review Process: o In-House Application: An internally developed application facilitates the manual review and categorization of articles.

[0203] Fields Populated During Review:

[0204] • PMID: PubMed Identifier

[0205] • Analysis Decision: Final decision based on analysis (Accepted, Rejected, On Hold)

[0206] • On Hold Comments: Reasons for placing an article on hold

[0207] • Email: Reviewer’s Email

[0208] • Date: Date of Review

[0209] • Biomarker / s: Names of biomarkers analyzed

[0210] • Biomarker Type: Category of each biomarker (e.g., Proteins, noncoding RNAs)

[0211] • Stage of Cancer: Cancer stage analyzed (e.g., Early, Late)

[0212] • Disease: Specific disease studied (e.g., Lung cancer)

[0213] • Disease Subtype-Specific Population: Population details for each disease subtype BM131025

[0214] • Experimental Population: Number of patients in the experimental group

[0215] • Total Population: Sum of experimental and comparator arm populations

[0216] • Cohort Characteristics: Key inclusion / exclusion criteria succinctly described

[0217] • Gender: Gender distribution of the cohort

[0218] • Age: Age distribution or average age of the cohort

[0219] • Ethnicity: Ethnic background of the cohort

[0220] • If Biomarker Status is Dependent on Smoking History: Yes / No

[0221] • Family History: Presence of family history influencing biomarker status

[0222] • Occupational Hazard: Occupational exposures affecting biomarker status

[0223] • Carcinogen Exposure: Exposure to known carcinogens

[0224] • Area Under the Curve Value (AUC): AUC value for diagnostic performance

[0225] • AUC CI: Confidence interval for the AUC value

[0226] • Sensitivity: Sensitivity percentage

[0227] • Sensitivity CI: Confidence interval for sensitivity

[0228] • Specificity: Specificity percentage

[0229] • Specificity CI: Confidence interval for specificity

[0230] • Change in Expression Eevel: Upregulated, Downregulated, Not specified

[0231] • Statistical Significance of the Difference Between Control and Cancer Samples: p-value or Not specified

[0232] • Comparator Arm: Group against which the biomarker was compared (e.g., Healthy controls)

[0233] • Sample Type: Type of biological sample used (e.g., Blood)

[0234] • Analysis Bias Mentioned in the Study: Yes / No BM131025

[0235] • If Biomarker has Prognostic / Therapeutic Implication: Specific implication selected from dropdown

[0236] • Analysis Remarks: Additional comments or notes from the analysis

[0237] • QA - Email: Quality Assurance Reviewer’s Email

[0238] • QA - Date: Date of QA Review

[0239] • QA - Decision: QA Decision (e.g., Approved, Needs Revision)

[0240] • QA - Comment: QA Reviewer’s Comments

[0241] • Final Decision: Conclusive decision after all review stages e) Analysis Framework: o Stage of Cancer: Select the appropriate cancer stage from the dropdown menu. Multiple stages for the same biomarker require separate entries. o Disease Subtype: Categorize based on disease ontology, creating separate entries for different histologies and parent disease terms. o Biomarker Details:

[0242] ■ Name: Select from the dropdown or refer to the biomarker-controlled vocabulary.

[0243] ■ Type: Categorize as Proteins, non-coding RNAs, etc., based on the biomarker.

[0244] ■ Diagnostic Metrics: Enter AUC, sensitivity, specificity, and their confidence intervals as reported.

[0245] ■ Statistical Values: Record p-values related to biomarker expression differences. o Comparator Arm: Define the comparison group (e.g., Healthy controls, Benign controls). o Cohort Characteristics: Summarize key inclusion / exclusion criteria succinctly. o Sample Type: Select from available options or mark as not reported. o Analysis Bias: Indicate the presence of any study biases and provide brief descriptions if applicable. BM131025 o Clinical Implications: Record prognostic or therapeutic implications based on the study findings. f) Quality Assurance and Secondary Analysis: o QA Review: A separate Quality Assurance reviewer validates initial decisions to ensure consistency and accuracy. o Secondary Analysis: Further scrutiny by a secondary reviewer provides an additional layer of validation, especially for articles placed on hold. o Human-in-the-Loop Framework: Expert reviewers interact with the automated system to correct and refine triage decisions, enhancing overall accuracy. o Reinforcement Learning from Human Feedback (RLHF): Feedback from QA and secondary reviewers is used to iteratively improve the ML model based on NLP, ensuring the system / model adapts to nuanced cases and evolving terminology. g) Documentation and Tracking: o Review Logs: All decisions and comments are recorded within the in-house application and a review sheet, ensuring traceability. o Review sheet: Detailed spreadsheet capturing all populated fields for each reviewed article was created.

[0246] 5. Biomarker Evaluation

[0247] Objective: To assign comprehensive and quantifiable scores to identified biomarkers based on their diagnostic value for early-stage lung cancer detection, in current example. This evaluation framework integrates multiple data elements extracted from analysed articles, ensuring a holistic and robust assessment of each biomarker's utility.

[0248] Specifically, this step involves assigning quantifiable score to each candidate biomarker identified at initial step 4, using a weighted scoring system adapted by the ML model (NLP based), wherein the scores are based on comprehensive evaluation of diagnostic performance of the candidate / identified biomarkers in BM131025 early-stage cancer detection, study robustness (evidence base), biological characteristics (such as age, histology & stage, gender, race & ethnicity etc.) and clinical implications, thereby integrating the information obtained from the data elements extracted at step 4 while arriving at a final biomarker score to ensure a holistic assessment of each biomarker’s utility.

[0249] Methodology: a) Scoring Criteria: The evaluation incorporates the following key data elements extracted during the analysis phase, at initial stage: o Diagnostic Performance:

[0250] ■ Area Under the Curve (AUC): Measures the biomarker’s ability to distinguish between cancerous and non-cancerous cases.

[0251] ■ Sensitivity and Specificity: Assess the true positive and true negative rates, respectively.

[0252] ■ Statistical Significance: Indicates the reliability of the biomarker’s diagnostic performance (p-value). o Study Population:

[0253] ■ Number of Patients: Reflects the robustness of the study; larger cohorts provide more reliable results.

[0254] ■ Disease Subtype-Specific Population: Ensures the biomarker’s applicability across various lung cancer subtypes.

[0255] ■ Stage of Cancer: Evaluates effectiveness in early (0, I, II) versus late (III, IV) stages. o Biomarker Characteristics:

[0256] ■ Biomarker Type: Categorizes the biomarker (e.g., Proteins, non-coding RNAs).

[0257] ■ Change in Expression Level: Indicates whether the biomarker is upregulated or downregulated in cancer patients. o Clinical Implications: BM131025

[0258] ■ Prognostic / Therapeutic Implications: Assesses the biomarker’s role in prognosis and therapy, enhancing its clinical utility. o Sample and Cohort Characteristics:

[0259] ■ Sample Type: Ensures non-invasive collection methods (Blood, Serum, Plasma).

[0260] ■ Cohort Characteristics: Includes gender, age, ethnicity, smoking history, family history, occupational hazards, and carcinogen exposure. b) Scoring Framework: A weighted scoring system is employed to integrate the various criteria, ensuring each element contributes appropriately to the final biomarker score. The framework is designed to be adaptable and scalable, allowing for future enhancements as more data becomes available. o Non-limiting embodiment of Weight Allocation:

[0261] ■ AUC: 20%

[0262] ■ Sensitivity and Specificity: 20%

[0263] ■ Statistical Significance (p-value): 10%

[0264] ■ Number of Patients: 15%

[0265] ■ Biomarker Type: 10%

[0266] ■ Change in Expression Level: 10%

[0267] ■ Stage of Cancer: 10%

[0268] ■ Prognostic / Therapeutic Implications: 5%

[0269] ■ Sample and Cohort Characteristics: 10% o Scoring Function: The overall biomarker score is calculated using a comprehensive function that incorporates all relevant data elements:

[0270] Biomarker

[0271] Score=(AUC x 0 ,20)+((Sensitivity+Specificity)2 x 0 ,20)+(( 11 +e-p)x0.10)+(Number of PatientsMax Patients x 0.15 )+(Biomarker Type Score x0.10)+(Expression Level Scorex0.10)+( Stage of Cancer Score x0.10)+(Implications Scorex0.05)+(Sample and Cohort BM131025

[0272] Score><0.10)\text{Biomarker Score} = (AUC \times 0.20) + \lcft(\frac{(\tcxt {Sensitivity} + \text{Specificity})} {2} \times 0.20\right) + \left(\left(\frac{ 1 } { 1 + eA{-p} }\right) \times 0.10\right) + \left(\frac{\text{Number of Patients} } {\text{Max Patients}} \times 0. 15\right) + (\text{Biomarker Type Score} \times 0.10) + (\text{Expression Level Score} \times 0.10) + (\text{Stage of Cancer Score} \times 0.10) + (\text{ Implications Score} \times 0.05) + (\text{ Sample and Cohort Score} \times 0.10)Biomarker Score=(AUCx0.20)+(2(Sensitivity+Specificity)x0.20)+((l+e-pl )x0. 10)+(Max PatientsNumber of PatientsxO.15)+(Biomarker Type Score x0.10)+(Expression Level Scorex0.10)+( Stage of Cancer Score x0.10)+(Implications Scorex0.05)+(Sample and Cohort Score x 0.10)

[0273] Where:

[0274] ■ AUC: Direct value from the study.

[0275] ■ Sensitivity and Specificity: Averaged percentage values.

[0276] ■ p-value: Transformed using a logistic function to normalize significance: Transformed p- value=l l+e-p\text{ Transformed p-value} = \frac{ l} { l + eA{-p}}Transformed p-value=l+e-pl

[0277] ■ Number of Patients: Scaled relative to the maximum number of patients across all studies.

[0278] ■ Biomarker Type Score:

[0279] ■ 3 for Proteins

[0280] ■ 2 for non-coding RNAs

[0281] ■ 1 for other types

[0282] ■ Expression Level Score:

[0283] ■ 2 for Upregulated

[0284] ■ 1 for Downregulated

[0285] ■ 0 for Not Specified

[0286] ■ Stage of Cancer Score: BM131025

[0287] ■ 3 for Early Stage

[0288] ■ 2 for Mid Stage

[0289] ■ 1 for Late Stage

[0290] ■ Implications Score:

[0291] ■ 2 for Therapeutic

[0292] ■ 1 for Prognostic

[0293] ■ 0 for Not Applicable

[0294] ■ Sample and Cohort Score:

[0295] ■ 3 for comprehensive cohort characteristics (e.g., diverse gender, age, ethnicity)

[0296] ■ 2 for partial

[0297] ■ 1 for minimal or not specified c) Scoring Logic: o High Score (> 2.0): Biomarkers with high AUC, validated in large patient cohorts, applicable to multiple cancer types, effective in early-stage detection, and possessing prognostic or therapeutic implications. o Medium Score (1.0 - 1.99): Biomarkers with moderate diagnostic performance, validated in medium-sized cohorts, specific to lung cancer, effective in mid-stage detection, and limited clinical implications. o Low Score (< 1.0): Biomarkers with low diagnostic performance, validated in small patient cohorts, limited to specific cancer types, effective only in late-stage detection, and lacking clinical implications. d) Example Evaluation: o Biomarker Y:

[0298] ■ AUC: 0.90

[0299] ■ Sensitivity: 85%

[0300] ■ Specificity: 88%

[0301] ■ p-value: 0.002 BM131025

[0302] ■ Number of Patients: 150

[0303] ■ Biomarker Type: Proteins (Score = 3)

[0304] ■ Change in Expression Level: Upregulated (Score = 2)

[0305] ■ Stage of Cancer: Early Stage (Score = 3)

[0306] ■ Implications: Therapeutic (Score = 2)

[0307] ■ Sample and Cohort Score: Comprehensive (Score = 3) o Biomarker

[0308] Score=(0.90x0.20)+(85+882x0.20)+(l l+e-0.002x0.10)+(150300x 0.15)+(3x0.10)+(2x0.10)+(3x0.10)+(2x0.05)+(3x0.10)=0.18+17.6 5+0.10+0.075+0.30+0.20+0.30+0. 10+0.30=19.205\text{Biomarker Score} = (0.90 \times 0.20) + \left(\frac{85 + 88} {2} \times 0.20\right) + \left(\frac{ l} { l + eA{-0.002}} \times 0.10\right) + \left(\frac{ 150} {300} \times 0. 15\right) + (3 \times 0.10) + (2 \times 0.10) + (3 \times 0.10) + (2 \times 0.05) + (3 \times 0.10) = 0.18 + 17.65 + 0.10 + 0.075 + 0.30 + 0.20 + 0.30 + 0.10 + 0.30 = 19.205Biomarker Score=(0.90x0.20)+(285+88 x0.20)+(l+e-0.0021 x0.10)+(300150 x0.15)+(3x0.10)+(2x0.10)+(3x0.10)+(2x0.05)+(3x0.10)=0.18+17. 65+0.10+0.075+0.30+0.20+0.30+0.10+0.30=19.205

[0309] (Assuming the maximum number of patients across all studies is 300.) e) Framework Robustness: o Holistic Integration: Combines multiple dimensions of biomarker utility, ensuring a comprehensive evaluation that accounts for diagnostic performance, study reliability, biomarker characteristics, and clinical implications. o Advanced Techniques: Utilizes normalization, weighted scoring, and logistic transformation of p-values to balance different aspects of biomarker performance, reducing bias and ensuring fair assessment across diverse studies. f) Technical Execution: BM131025 o Automation:

[0310] Implemented through advanced data processing pipelines that automatically extract and normalize relevant data from reviewed articles, ensuring efficient and consistent scoring. o Machine Learning Integration: Potential integration with machine learning models executed on any computing systems to refine scoring based on emerging data patterns and expert feedback, enhancing the system’s predictive capabilities. o Data Validation: Ensures accuracy through cross-validation with manual reviews and adherence to predefined guidelines, maintaining the integrity of the scoring process.

[0311] Outcome:

[0312] Each biomarker across each article is assigned a quantifiable score based on a comprehensive evaluation of its diagnostic performance, study robustness, biological characteristics, and clinical relevance.

[0313] 6. Evidence Synthesis

[0314] Objective: To aggregate and synthesize biomarker scores from multiple studies, constructing a robust evidence base that prioritizes biomarkers based on their diagnostic value for early-stage lung cancer detection. This aids in the prioritization of biomarkers, focusing on those most extensively researched and quantitatively assessed. This process involves calculating composite scores for each biomarker, categorizing them into distinct performance profiles, and ranking them within each category based on the accumulated evidence.

[0315] Therefore, this step involves aggregating and synthesizing these final biomarker scores across the relevant dataset into a composite score that represents each candidate biomarker using the weighted scoring system by the ML model. The composite score is computed by averaging and weighing the scores obtained at initial step 5 by said biomarker’s diagnostic performance and no. of evidences (no. of studies) in support thereof to arrive at the final composite score. Pursuant to this step, each biomarker will have a single composite score assigned therewith. BM131025

[0316] Methodology: a) Data Aggregation: o Grouping by Biomarker:

[0317] ■ Biomarkers are grouped by their names and types to consolidate data from various studies (articles from the relevant set).

[0318] ■ Each biomarker's performance metrics (AUC, sensitivity, specificity, etc.) from different studies are aggregated to provide a comprehensive overview. o Disease Ontology Integration:

[0319] ■ A detailed disease ontology was developed to understand the hierarchical relationships among lung cancer subtypes.

[0320] ■ This ontology ensures accurate mapping and aggregation of biomarker data across related disease categories, facilitating nuanced analysis. b) Composite Score Calculation: o Composite Score Formula:

[0321] ■ The composite score for each biomarker is derived from the previously defined scoring criteria, integrating multiple performance metrics into a single, quantifiable value.

[0322] ■ Composite Score Calculation:

[0323] Composite Score B M=]yi= ln(B iomarker Scorei><Weighti)\text{ Composite Score }_{\text{BM}} = \sum_{i=l}A{n} \left(\text{Biomarker Score}_i Times \text{Weight}_i \right)Composite ScoreBM=i=l n (Biomarker ScoreixWeighti)

[0324] Where:

[0325] I. Each Biomarker Scorei\text{ Biomarker

[0326] Score}_iBiomarker Scorei is derived from individual studies. BM131025

[0327] II. Weights are based on the relative importance of each criterion as defined in the Biomarker Evaluation framework. o Evidence Counting:

[0328] ■ The number of supporting studies (PMIDs) (articles) for each biomarker's performance is counted.

[0329] ■ Biomarkers supported by a higher number of studies receive greater weight in the composite score, reflecting stronger evidence. c) Framework Initialization: o Performance Categories:

[0330] ■ Biomarkers (or the candidate biomarkers) are categorized into four distinct primary performance profiles (or categories) specifically curated by the inventors of the present disclosure based on the biomarker’s diagnostic performance, to facilitate targeted ranking and prioritization:

[0331] I. Specific Guardian:

[0332] • Definition: Biomarkers with high specificity, minimizing false positives.

[0333] • Criteria:

[0334] 1. High specificity (>90%)

[0335] 2. Moderate to high AU C (>0.80)

[0336] 3. Validated in multiple studies with large patient cohorts

[0337] 4. Effective across various lung cancer subtypes

[0338] II. Sensitive Detector:

[0339] • Definition: Biomarkers with high sensitivity, minimizing false negatives.

[0340] • Criteria:

[0341] 1. High sensitivity (>85%) BM131025

[0342] 2. Moderate to high AU C (>0.80)

[0343] 3. Validated in multiple studies with large patient cohorts

[0344] 4. Effective in early-stage lung cancer detection

[0345] III. Balanced Performer:

[0346] • Definition: Biomarkers with balanced sensitivity and specificity.

[0347] • Criteria:

[0348] 1. Sensitivity and specificity both between 75% and 90%

[0349] 2. High AUC (>0.85)

[0350] 3. Validated in multiple studies with medium to large patient cohorts

[0351] 4. Applicable to multiple lung cancer subtypes and stages

[0352] IV. Early Scout:

[0353] • Definition: Biomarkers particularly effective in early-stage cancer detection .

[0354] • Criteria:

[0355] 1. Effective in early-stage (0, I, II) lung cancer detection

[0356] 2. Moderate to high sensitivity and specificity

[0357] 3. High AUC (>0.80)

[0358] 4. Validated in studies focusing on early-stage cohorts d) Composite Score Calculation and Ranking: o Aggregation Process:

[0359] ■ For each biomarker, all individual scores from different studies are aggregated to compute a composite score. BM131025

[0360] ■ The aggregation method involves averaging the scores, weighted by the number of evidences (PMIDs) supporting each study. o Ranking within Categories:

[0361] ■ Candidate Biomarkers within each performance category are ranked based on their composite scores. Ranking is performed using feature selection techniques. The feature selection techniques are selected from a group consisting of, but not limiting thereto, recursive feature elimination with classifiers including random forest (RF), support vector machines (SVM), and logistic / linear models; tree-ensemble and gradient boosting methods including GBM, XGBoost, LightGBM, and CatBoost; embedded regularization methods including LASSO (LI), ridge (L2), and Elastic Net; filter methods including mutual information, minimum- redundancy-maximum-relevance (mRMR), Relief / ReliefF, %2tests, ANOVA / F-score, and Fisher score; permutation importance (if one biomarker is evaluated in combinations, i.e. BM1 + BM2, so what and which biomarker contributes how much to overall diagnostic performance of said panel); attribution-based metrics including SHAP (TreeSHAP / Kemal SHAP) and integrated gradients; wrapper methods including sequential forward selection (SFS), sequential backward selection (SBS), sequential floating forward / backward selection (SFFS / SFBS), genetic-algorithm search, and Bayesian optimization for subset selection; correlation / variance / VIF- based pruning; and functional equivalents and combinations thereof. BM131025

[0362] ■ High Priority: Candidate Biomarkers with the highest composite scores within their category, supported by numerous studies (large evidence base).

[0363] ■ Medium Priority: Candidate Biomarkers with moderate composite scores, supported by a fair number of studies.

[0364] ■ Low Priority: Candidate Biomarkers with lower composite scores, supported by fewer studies. e) Categorization and Prioritization: o Specific Guardian:

[0365] ■ Top-Ranked Biomarkers: Those with the highest specificity and strong diagnostic performance across multiple studies. o Sensitive Detector:

[0366] ■ Top-Ranked Biomarkers: Those with the highest sensitivity and significant effectiveness in early detection. o Balanced Performer:

[0367] ■ Top-Ranked Biomarkers: Those maintaining a balance between sensitivity and specificity, ensuring reliable diagnostic performance. o Early Scout:

[0368] ■ Top-Ranked Biomarkers: Those demonstrating exceptional utility in identifying early-stage lung cancer, crucial for improving patient outcomes. f) Database Creation: o Biomarker Database:

[0369] ■ A centralized database is established to store aggregated biomarker scores, associated PMIDs, AUC values, patient numbers, and other relevant data.

[0370] ■ This database facilitates easy access, querying, and further analysis, supporting scalability and future enhancements. o Disease Ontology Mapping: BM131025

[0371] ■ Biomarkers are accurately mapped to their respective disease subtypes using the developed disease ontology.

[0372] ■ This mapping ensures precise aggregation and analysis, accounting for hierarchical disease relationships. g) Technical Execution: o Data Processing:

[0373] ■ Pandas: Utilized for data manipulation, grouping, and aggregation.

[0374] ■ Python Functions: Implemented to calculate composite scores, aggregate evidence, and perform ranking based on predefined criteria. h) Ranking and Prioritization:

[0375] • Category-Based Ranking: o Within each performance category (Specific Guardian, Sensitive Detector, Balanced Performer, Early Scout), biomarkers are ranked based on their composite scores. o High-Ranked Biomarkers: Those with the highest composite scores within their category, supported by numerous studies and robust evidence. o Medium and Low-Ranked Biomarkers: Biomarkers are positioned accordingly based on their composite scores and the extent of supporting evidence.

[0376] • Prioritization Logic: o Specific Guardian:

[0377] ■ High specificity and strong diagnostic performance across multiple studies.

[0378] ■ Prioritized for minimizing false positives in diagnostic applications. o Sensitive Detector:

[0379] ■ High sensitivity and effectiveness in early detection. BM131025

[0380] ■ Prioritized for ensuring early-stage cancer identification with minimal false negatives. o Balanced Performer:

[0381] ■ Balanced sensitivity and specificity, ensuring reliable diagnostic performance.

[0382] ■ Prioritized for general diagnostic applications requiring both sensitivity and specificity. o Early Scout:

[0383] ■ Exceptional utility in identifying early-stage lung cancer.

[0384] ■ Prioritized for improving patient outcomes through early intervention.

[0385] Outcome:

[0386] • A comprehensive and prioritized list of biomarkers is established, categorized into Specific Guardian, Sensitive Detector, Balanced Performer, and Early Scout.

[0387] • Each biomarker is assigned a composite score reflecting its diagnostic performance, study robustness, biological relevance, and clinical implications.

[0388] • The evidence synthesis process ensures that biomarkers supported by extensive and high-quality studies are prioritized, facilitating informed decision-making for further validation and integration into the present invention’s diagnostic system.

[0389] 7, Orthogonal Validation

[0390] Objective: To corroborate biomarker findings through multiple independent sources, enhancing the robustness and credibility of the biomarker selection phase of the present method / process. This step ensures that selected biomarkers are consistently supported across diverse and authoritative databases, thereby validating their diagnostic utility for early-stage lung cancer detection.

[0391] Therefore, this stage involves validating the ranked biomarkers (obtained at previous step) through cross-validation (or orthogonal validation) techniques BM131025 across standard clinical databases to confirm their diagnostic performance for early-stage cancer detection.

[0392] The standard clinical databases are selected from the group consisting of, but not limiting thereto, biomarker repositories such as EDRN, OncoMX, Knapsack family biomarkers, and human protein atlas; clinical trial registries; FDA- approved datasets; oncology databases such as KM Plotter, TCGA, GEO databases and like and combinations thereof.

[0393] Methodology: a) Orthogonal Validation Framework: o Definition:

[0394] Orthogonal validation employs multiple independent approaches and databases to confirm or refute biomarker findings. This method ensures that selected biomarkers are reliable, accurate, and widely supported by the scientific community. o Validation Sources:

[0395] ■ Non-limiting embodiments of Primary Validation Databases used:

[0396] ■ KM Plotter: Provides comprehensive gene expression and survival analysis data.

[0397] ■ Web of Science & Scopus: Extensive citation databases for verifying study impact and biomarker relevance.

[0398] ■ FDA (Premarket Approval & Premarket Notification): Regulatory validation for clinical applicability.

[0399] ■ NCCN, ESMO, ASCO: Leading oncology organizations offering guidelines and validated biomarkers.

[0400] ■ Clinical Trial Registry, India: Clinical validation through ongoing or completed trials. BM131025

[0401] ■ Semantic Scholar & Medline: Academic and peer- reviewed sources ensuring research quality.

[0402] ■ Cochrane: Systematic reviews validating evidence strength.

[0403] ■ Preprint Databases (BioRXiv, MEDRXiv): Emerging research insights, requiring cautious validation.

[0404] ■ Examples of Supplementary Validation Sources:

[0405] ■ EDRN, Protein Atlas, OncoMX: Specialized databases providing additional context and validation for biomarkers not covered in primary sources. o Scoring Criteria for Validation Sources

[0406] ■ Dependability: Reliability and authority of the source.

[0407] ■ Accurateness: Precision and correctness of the data provided.

[0408] ■ Frequency of Updates: How often the database is updated with new information.

[0409] ■ Comprehensiveness: Extent to which the database covers relevant biomarkers and associated data. o Scoring Framework:

[0410] ■ Each validation source is assigned a score based on the above criteria.

[0411] ■ Scores are normalized to ensure consistency across different criteria.

[0412] ■ Composite Validation Score:

[0413] Validation Score= i=ln(Source

[0414] Scorei><Weighti)\text{Validation Score} = \sum_{i=l }A{n} (\text{Source Score}_i \times \text{Weight}_i)Validation Score=i=l n(Source ScoreixWeighti) Where each source's score is weighted based on its overall reliability and relevance. BM131025 b) Composite Score Calculation and Aggregation: o Composite Score Formula: The composite score for each biomarker is calculated by aggregating scores from all relevant validation sources:

[0415] Composite ScorcBM= yi= I m( Validation Scorei)\text{Composite Score }_{\text{BM}} = \sum_{i=l}A{m} (\text{Validation

[0416] Score }_i)Composite ScoreBM=i= 1 m( Validation Scorei) Where mmm is the number of validation sources used for each biomarker. o Evidence Counting:

[0417] ■ Number of Supporting Studies (PMIDs): The number of studies supporting each biomarker’s performance is counted. Biomarkers supported by a higher number of studies receive greater weight in the composite score, reflecting stronger evidence. o Disease Ontology Integration:

[0418] ■ Disease Hierarchy Mapping: A comprehensive disease ontology was developed to understand hierarchical relationships among lung cancer subtypes. This ensures accurate aggregation and analysis across related disease categories.

[0419] ■ Mapping Biomarkers to Disease Subtypes: Biomarkers are accurately mapped to specific lung cancer subtypes using the disease ontology, ensuring comprehensive coverage and precise evidence aggregation. c) Scoring Logic and Aggregation: o Framework Initialization:

[0420] Performance Categories:

[0421] ■ Specific Guardian: BM131025

[0422] ■ Definition: High specificity biomarkers minimizing false positives.

[0423] ■ Criteria: Specificity > 90%, AUC > 0.80, validated in multiple large cohorts, applicable across various subtypes.

[0424] ■ Sensitive Detector:

[0425] ■ Definition: High sensitivity biomarkers minimizing false negatives.

[0426] ■ Criteria: Sensitivity > 85%, AUC > 0.80, effective in early-stage detection, validated in multiple large cohorts.

[0427] ■ Balanced Performer:

[0428] ■ Definition: Biomarkers with balanced sensitivity and specificity.

[0429] ■ Criteria: Sensitivity and specificity between 75-90%, AUC > 0.85, validated in medium to large cohorts, applicable to multiple subtypes.

[0430] ■ Early Scout:

[0431] ■ Definition: Biomarkers particularly effective in early-stage cancer detection.

[0432] ■ Criteria: Effective in early stages (0, I, II), AUC > 0.80, validated in studies focusing on early-stage cohorts. o Composite Score Calculation:

[0433] ■ Normalization and Weighting: Each criterion is normalized and weighted to ensure uniform contribution to the final composite score.

[0434] ■ Advanced Scoring Formula:

[0435] Biomarker

[0436] Score=(AUC x 0 ,20)+(Sensitivity+Specificity2 x 0.20)+( 11 +e BM131025

[0437] -px0.10)+(Number of PatientsMax

[0438] Patients x 0.15 )+(Biomarker Type Score x0.10)+(Expression Level Score x0.10)+(Stage of Cancer

[0439] Score x0.10)+(Implications Scorex0.05)+(Sample and Cohort Scorex0.10)\text{Biomarker Score} = (AUC \times 0.20) + \left(\frac{\text{ Sensitivity} +

[0440] \text{ Specificity} } {2} \times 0.20\right) + \left(\frac{ 1 } { 1 + eA{-p}} Times O. lOTight) + \left(\frac{\text{Number of Patients}} {\text{Max Patients}} Times 0.15\right) + (Text {Biomarker Type Score} Times 0.10) +

[0441] (\text{Expression Level Score} Times 0.10) + (\text{Stage of Cancer Score} Times 0.10) + (Text {Implications Score} Times 0.05) + (\text{Sample and Cohort Score} Times 0.10)Biomarker

[0442] Score=(AUC x 0 ,20)+(2Sensitivity+Specificity x0.20)+(l+e-plx0.10)+(Max PatientsNumber of Patients x0.15)+(Biomarker Type Score x0.10)+(Expression Level Score x0.10)+( Stage of Cancer Score x0.10)+(Implications Scorex0.05)+(Sample and Cohort ScorexO.10)

[0443] ■ Biomarker Type Score:

[0444] ■ 3 for Proteins

[0445] ■ 2 for non-coding RNAs

[0446] ■ 1 for other types

[0447] ■ Expression Level Score:

[0448] ■ 2 for Upregulated

[0449] ■ 1 for Downregulated

[0450] ■ 0 for Not Specified

[0451] ■ Stage of Cancer Score:

[0452] ■ 3 for Early Stage

[0453] ■ 2 for Mid Stage BM131025

[0454] ■ 1 for Late Stage

[0455] ■ Implications Score:

[0456] ■ 2 for Therapeutic

[0457] ■ 1 for Prognostic

[0458] ■ 0 for Not Applicable

[0459] ■ Sample and Cohort Score:

[0460] ■ 3 for comprehensive cohort characteristics (e.g., diverse gender, age, ethnicity)

[0461] ■ 2 for partial

[0462] ■ 1 for minimal or not specified d) Ranking and Prioritization: o Category-Based Ranking:

[0463] ■ Biomarkers are ranked within each performance category (Specific Guardian, Sensitive Detector, Balanced Performer, Early Scout) based on their composite validation scores.

[0464] ■ Top-Ranked Biomarkers: Biomarkers with the highest composite scores within their respective categories are identified as top candidates. o Prioritization Criteria:

[0465] ■ High Priority: Biomarkers with high composite scores, supported by numerous studies, large patient cohorts, early- stage effectiveness, and significant clinical implications.

[0466] ■ Medium Priority: Biomarkers with moderate composite scores, validated in medium-sized cohorts, specific to lung cancer, and some clinical relevance.

[0467] ■ Low Priority: Biomarkers with low composite scores, limited validation, specific to late-stage cancer, and minimal clinical implications. e) Framework Robustness: BM131025 o Holistic Integration: Combines multiple dimensions of biomarker utility, ensuring a comprehensive evaluation that accounts for diagnostic performance, study reliability, biomarker characteristics, and clinical implications. o Advanced Techniques: Utilizes normalization, weighted scoring, and logistic transformation of p-values to balance different aspects of biomarker performance, reducing bias and ensuring fair assessment across diverse studies. o Flexibility and Scalability: Designed to accommodate additional criteria and adapt to evolving research findings, ensuring long-term applicability and scalability of the evaluation framework. o Human-in-the-Loop Enhancement: Expert reviewers provide continuous feedback, enabling iterative improvements to the scoring model and ensuring accuracy through Reinforcement Learning from Human Feedback (RLHF). f) Technical Execution: o Data Processing:

[0468] ■ Pandas: Utilized for data manipulation, grouping, and aggregation.

[0469] ■ Python Functions: Implemented to calculate composite scores, aggregate evidence, and perform ranking based on predefined criteria. g) Finalization and Shortlisting: o Top 29 Biomarkers

[0470] ■ Biomarkers with the highest composite validation scores across multiple reliable sources are shortlisted.

[0471] ■ Selection Criteria:

[0472] ■ Consistent high scores across multiple validation sources. BM131025

[0473] Demonstrated effectiveness in early-stage lung cancer detection.

[0474] Strong clinical and prognostic implications.

[0475] Outcome:

[0476] The Orthogonal Validation phase consolidates evidence from multiple independent sources, ensuring that selected biomarkers are reliable, accurate, and widely supported. By aggregating validation scores and leveraging a comprehensive disease ontology, the framework prioritizes biomarkers with the highest utility for early-stage lung cancer detection, leading to a robust and credible selection for further validation and integration into the present disclosure’s diagnostic system.

[0477] The method (MLB) ends at step 7 (Orthogonal validation), where a biomarker panel which is cancer-type specific is obtained. In accordance with the present disclosure, each of the steps 1 through 7 are repeated for each cancer-type to arrive at multiple biomarker panels which are highly specific with regard to said cancer-type and which has high efficiency in screening of early-stage cancer(s) in the patient. So, for each type of cancer, these steps 1 to 7 are repeated such that we obtain a biomarker panel separate for each cancer.

[0478] Thereafter, the method (MLB) further comprises an additional step of ‘technical annotation’ for quantification of biomarkers present in the biomarker panels obtained pursuant to step 7.

[0479] Technical Annotation for Quantification

[0480] Objective: To elucidate the methods used for biomarker quantification in laboratory settings, ensuring practical applicability and reproducibility. This step involves detailed documentation of the experimental procedures, reagents, and conditions under which biomarkers were measured, facilitating consistent replication and validation across studies. BM131025

[0481] Methodology: a) Guidelines for Technical Annotation: The technical annotation process was structured based on the following guidelines to ensure comprehensive and uniform data capture: o Non-limiting examples of Data Entry Fields:

[0482] 1. PMID: PubMed Identifier for the article.

[0483] 2. Biomarker (BM): Name of the biomarker, adhering to the nomenclature from the Biomarker Controlled Vocabulary. For biomarker combinations, the entire combination is recorded.

[0484] 3. Sample Type: Type of biological sample used (e.g., Blood, Serum, Plasma).

[0485] 4. Technique (e.g. ELISA / CLIA): Detection method employed, restricted to ELISA or CLIA techniques.

[0486] 5. Antibody Company and Catalogue No. / Kit Details: Name of the antibody manufacturer and corresponding catalogue number or kit details, prefixed with '#'.

[0487] 6. Dilution Used: Dilution factor of the antibody in the experiment.

[0488] 7. Stage of Cancer: Specific stage(s) of cancer analyzed. Multiple stages are recorded in separate rows if applicable.

[0489] 8. Change in Expression Level: Indicates whether the biomarker is upregulated or downregulated.

[0490] 9. p-Value: Statistical significance of the difference between control and cancer samples.

[0491] 10. Mean Concentrations (ng / ml): Concentration levels of the biomarker in healthy, benign, and malignant populations. If a category is not applicable, 'NA' is recorded.

[0492] 11. Comparison: Type of comparison made (e.g., Healthy vs. Malignant, Healthy vs. Benign). BM131025 o Data Population Process:

[0493] 1. PMID Assignment: Each annotated entry begins with the PMID of the relevant study.

[0494] 2. Biomarker Identification: The biomarker of interest is recorded using standardized nomenclature. For combinations, the full combination is noted.

[0495] 3. Sample and Technique Documentation: The sample type and detection technique (ELISA or CLIA) are specified. Only non-invasive sample types are considered.

[0496] 4. Reagent Details: The antibody's company name and catalogue number are documented to ensure reproducibility.

[0497] 5. Experimental Conditions: Dilution factors and cancer stages are recorded. Multiple stages within a single study are entered as separate records.

[0498] 6. Expression and Statistical Analysis: The direction of biomarker expression change (upregulated / downregulated) and corresponding p-values are noted.

[0499] 7. Concentration Measurements: Mean concentrations of the biomarker are captured for healthy, benign, and malignant populations. Missing categories are marked as 'NA'.

[0500] 8. Comparison Groups: The specific comparisons made in the study are detailed to contextualize the biomarker's diagnostic performance. b) In-House Application Utilization: An internally developed application was employed to facilitate the manual annotation and categorization of technical data. This application ensures structured data entry and consistency across all annotations. c) Biomarker Narrowing and Selection: Based on the composite scores derived from the technical annotation and scoring framework, biomarkers were ranked and prioritized. Initially, 29 biomarkers (obtained at step 7 in current example of lung cancer) were evaluated, and through this BM131025 rigorous scoring and validation process, the list was narrowed down to the top 12 biomarkers. These selected biomarkers demonstrated superior diagnostic performance, robustness across multiple studies, and significant clinical implications, making them prime candidates for further validation and integration into our diagnostic system.

[0501] Outcome:

[0502] The Technical Annotation for Quantification phase ensures that each biomarker is meticulously documented with detailed experimental conditions, reagents, and performance metrics. By applying a robust scoring framework, biomarkers are systematically evaluated and prioritized based on their diagnostic utility, reliability, and clinical relevance.

[0503] Therefore, pursuant this iterative exercise by following each of the above steps for each cancer-type, the inventors of the present disclosure via wide experimentation have arrived at the biomarker panel that can be independently used to diagnose pancancer types in the patients.

[0504] II. Second Aspect: Biomarker Panel and Cancer type(s), and Cancerspecific Biomarker panels

[0505] In accordance with a second aspect of the present disclosure, a biomarker panel, for pan-cancer screening of the patients, is disclosed.

[0506] It is a characterizing feature of the present disclosure that the biomarker panel can be used to diagnose several different types of prominent cancers and stages thereof in the patient just by using a biological sample obtained therefrom. This afore-mentioned synergistic combination of the known biomarkers within the biomarker panel that facilitates accurate prediction of early-stage cancers, rather than a single early-stage biomarker as used in prior art, is another characterizing feature of the present invention. BM131025

[0507] Accordingly, the biomarker panel of the present disclosure adapted for early- screening of cancer(s) includes combination of 18 biomarkers as follows:

[0508] CA19-9 (Carbohydrate Antigen 19-9), CA125 (Cancer Antigen 125), TAG-72 (Tumor-Associated Glycoprotein 72), CEA (Carcinoembryonic Antigen), HE4 (Human Epididymis Protein 4), SCC-Ag (Squamous Cell Carcinoma Antigen), PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist-II), NSE (Neuron-Specific Enolase), CYFRA-21 (Cytokeratin 19 Fragment (CYFRA 21-1)), AFP (Alpha-Fetoprotein), Total PSA (PSA) (Prostate-Specific Antigen (total)), CRP (C-Reactive Protein), ProGRP (Pro-Gastrin-Releasing Peptide), HSP90 AA1 ( Heat Shock Protein 90 Alpha Family Class A Member 1), AFP - L3 (Lens culinaris agglutinin-reactive fraction of Alpha-Fetoprotein), VEGFA ( Vascular Endothelial Growth Factor A), Neutrophil to Lymphocyte Ratio (NLR) and fPSA Free Prostate- Specific Antigen).

[0509] In a preferred embodiment, the afore-stated biomarker panel is particularly used for diagnosis / screening of plurality of cancer(s) in the patients. The cancer(s) are selected from a group comprising of, but not limited to, Breast cancer, Biliary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer.

[0510] In another embodiment, the biomarker panel may be customized for use in diagnosis and screening of cancer(s) of, but not limiting to, Larynx, Testis, Penis, Vagina, Fallopian tube, Other male genitalia, Placenta, Intestine, Anus, Interlobular region, Cholangioadenoma, Brain, and Skin.

[0511] It is another characterizing feature of the present invention that the inventors via wide intensive experimentation has obtained the following individual biomarker panels specific as per each cancer type. These biomarker panels, which are cancerspecific, often referred as, cancer-specific biomarker panel, are derived from the biomarker panel given above and are sub-sets thereof. BM131025

[0512] Enlisted hereinafter are 13 separate cancer-specific biomarker panels (for 13 cancer types) along with their thresholds indicating (malignant) cancer, non-cancer and benign status in patients, obtained via intensive experimentation by the inventors, which is yet another characterizing feature of the present disclosure. It is to be noted that these threshold ranges have been arrived, by detecting these biomarkers in the patient sample using detection techniques, preferably, ELISA and CLIA. However, it is pertinent to note that biomarkers can be detected using other techniques known in the art, with slight variations in these threshold ranges, dependent on the detection technique used, needs to be considered.

[0513] 1. Breast Cancer:

[0514] Biomarker panel for screening of Breast cancer (Breast-cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA19-9, CA125, CEA, HPS90AA1, VEGFA and HE4, with respective threshold ranges obtained are as follows:

[0515] • CA19-9, between 49-76 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 30-55 U / mL indicates benign cancer;

[0516] • CA125, between 12-65 U / mL indicates cancer, between 9-32 U / mL indicates healthy, and between 12-45 U / mL indicates benign cancer;

[0517] • CEA, between 1.5-30 ng / mL indicates cancer, between 0-8.3 ng / mL indicates healthy and between 0.4-6 ng / mL indicates benign cancer;

[0518] • HPS90AA1, between 310-700 pg / mL indicates cancer, between 180-320 pg / mL indicates healthy and between 220-360 pg / mL indicates benign cancer;

[0519] • VEGFA, between 57-219 pg / ml indicates cancer, between 11-47 pg / ml indicates healthy and between 19-83 pg / ml indicates benign cancer, and

[0520] • HE4, between 45-81 pmol / L indicates cancer, between 45-70 pmol / L indicates healthy and between 50-72 pmol / L indicates benign cancer.

[0521] 2. Biliary tract cancer BM131025

[0522] Biomarker panel for screening of Biliary tract cancer: (Biliary tract cancer biomarker panel) Biomarker panel comprises a curated combination of biomarkers such as CA19-9, CA125, TAG-72, CEA and PIVKA-II, with respective threshold ranges obtained are as follows:

[0523] • CAI 9-9, between 129-1200 U / mL indicates cancer, between 0-37 U / mL indicates healthy and between 37-1,000 U / mL indicates benign cancer;

[0524] • CA125, between 105-550 U / mL indicates cancer, between 0-35 U / mL indicates healthy and between 35-350 U / mL indicates benign cancer;

[0525] • TAG-72, between 3-50 U / mL indicates cancer, between 0-7 U / mL indicates healthy and between 3-39 U / mL indicates benign cancer;

[0526] • CEA, between 10-90 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 5-20 ng / mL indicates benign cancer;

[0527] • PIVKA-II, between 30-800 mAU / mL indicates cancer, between 0-40 mAU / mL indicates healthy and between 10-40 mAU / mL indicates benign cancer.

[0528] 3. Cervical-Endometrial cancer

[0529] Biomarker panel for screening of Cervical-endometrial cancer (Cervical-endometrial cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CAI 9-9, CAI 25, TAG-72, HE4, SCC-Ag, ProGRP, and VEGLA with respective threshold ranges obtained are as follows:

[0530] • CAI 9-9, between 20-100 U / mL indicates cancer, between 0-18 U / mL indicates healthy, and between 18-40 U / mL indicates benign;

[0531] • CA125, between 20-100 U / mL indicates cancer, between 0-15 U / mL indicates healthy, and between 15-40 U / mL indicates benign;

[0532] • TAG-72, between 7-40 U / mL indicates cancer, between 0-6.9 U / mL indicates healthy, and between 6.9-12 U / mL indicates benign; BM131025

[0533] • HE4, between 90-200 pmol / L indicates cancer, between 17-75 pmol / L indicates healthy, and between 60-90 pmol / L indicates benign;

[0534] • SCC-Ag, between 2.0-10 ng / mL indicates cancer, between 0-1.5 ng / mL indicates healthy, and between 1.5-2.0 ng / mL indicates benign;

[0535] • ProGRP, between 80-300 pg / mL indicates cancer, between 0-50 pg / mL indicates healthy and between 50-80 pg / mL indicates benign cancer;

[0536] • VEGLA, between 300-1,300 pg / mL indicates cancer, between 50-250 pg / mL indicates healthy, and between 200-350 pg / mL indicates benign.

[0537] 4. Colorectal cancer

[0538] Biomarker panel for screening of Colorectal cancer (Colorectal cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CAI 9-9, CAI 25, TAG-72, CEA, HE4, NSE, AFP, CRP, HSP90 AA1, and VEGFA with respective threshold ranges obtained are as follows:

[0539] • CA19-9, between 35-1,000 U / mL indicates cancer, between 0-50 U / mL indicates healthy, and between 37-180 U / mL indicates benign cancer;

[0540] • CAI 25, 35-250 U / mL indicates cancer, between 0-55 U / mL indicates healthy, and between 20-105 U / mL indicates benign cancer;

[0541] • TAG-72, 7-100 U / mL indicates cancer, between 0-20 U / mL indicates healthy, and between 2-45 U / mL indicates benign cancer;

[0542] • CEA, between 5-200 ng / mL indicates cancer, between 0.2-7 ng / mL indicates healthy, and between 3-120 ng / mL indicates benign cancer;

[0543] • HE4, between 70-200 pmol / L indicates cancer, between 0-85 pmol / L indicates healthy, and between 45-135 pmol / L indicates benign cancer;

[0544] • NSE, between 9-40 ng / mL indicates cancer, between 0-18 ng / mL indicates healthy, and between 12-26.5 ng / mL indicates benign cancer;

[0545] • AFP, between 10-400 ng / mL indicates cancer, between 0-18 ng / mL indicates healthy, and between 15-100 ng / mL indicates benign cancer; BM131025

[0546] • CRP, between 10-100 mg / L indicates cancer, between 0-10 mg / L indicates healthy, and between 3-30 mg / L indicates benign;

[0547] • HSP90 AA1, between 20-140 ng / mL indicates cancer, between 0-10 ng / mL indicates healthy, and between 5-45 ng / mL indicates benign cancer;

[0548] • VEGFA, between 200-1,000 pg / mL indicates cancer, between 0-50 pg / mL indicates healthy, and between 35-400 pg / mL indicates benign cancer.

[0549] 5. Esophageal cancer

[0550] Biomarker panel for screening of Esophageal cancer (Esophageal cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA19-9, CA125, TAG-72, CEA, SCC- Ag, and NSE with respective threshold ranges obtained are as follows:

[0551] • CAI 9-9, between 40-1,000 U / mL indicates cancer, between 0-105 U / mL indicates healthy, and between 35-330 U / mL indicates benign cancer;

[0552] • CA125, between 35-500 U / mL indicates cancer, between 0-65 U / mL indicates healthy, and between 20-135 U / mL indicates benign cancer;

[0553] • TAG-72, between 7-500 U / mL indicates cancer, between 0-30 U / mL indicates healthy, and between 6-190 U / mL indicates benign cancer;

[0554] • CEA, between 5.0-200 ng / mL indicates cancer, between 0-10 ng / mL indicates healthy, and between 3-45 ng / mL indicates benign cancer;

[0555] • SCC-Ag, between 2.0-50 ng / mL indicates cancer, between 0-3 ng / mL indicates healthy, and between 1.5-9 ng / mL indicates benign cancer;

[0556] • NSE, between 20-150 ng / mL indicates cancer, between 0-35 ng / mL indicates healthy, and between 10-60 ng / mL indicates benign cancer.

[0557] 6. Gastric cancer

[0558] Biomarker panel for screening of Gastric cancer (Gastric cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA19-9, CA125, TAG-72, CEA, PIVKA-II, AFP, and CRP with respective threshold ranges obtained are as follows: BM131025

[0559] • CAI 9-9, between 35-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 25-200 U / mL indicates benign cancer;

[0560] • CA125, between 35-200 U / mL indicates cancer, between 0-45 U / mL indicates healthy and between 35-95 U / mL indicates benign cancer;

[0561] • TAG-72, between 6-100 U / mL indicates cancer, between 0-9 U / mL indicates healthy and between 6-15 U / mL indicates benign cancer;

[0562] • CEA, between 5-30 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;

[0563] • PIVKA-II, between 40-1000 mAU / mL indicates cancer, between 0-40 mAU / mL indicates healthy and between 40-60 mAU / mL indicates benign cancer;

[0564] • AFP, between 20-500 ng / mL indicates cancer, between 0-30 ng / mL indicates healthy and between 10-60 ng / mL indicates benign cancer;

[0565] • CRP, between 10-200 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 5-35 mg / L indicates benign cancer;

[0566] 7. Liver cancer

[0567] Biomarker panel for screening of Liver cancer (Liver cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CAI 9-9, CEA, PIVKA-II, AFP, CRP, HSP90AA1, and AFP-L3 with respective threshold ranges obtained are as follows:

[0568] • CAI 9-9, between 35-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 25-200 U / mL indicates benign cancer;

[0569] • CEA, between 5-200 ng / mL indicates cancer, between 0-7 ng / mL indicates healthy and between 3-10 U / mL indicates benign cancer;

[0570] • PIVKA-II, between 40-1000 mAU / mL indicates cancer, between 0-40 mAU / mL indicates healthy and between 40-60 mAU / mL indicates benign cancer;

[0571] • AFP, between 20-500 ng / mL indicates cancer, between 0-35 ng / mL indicates healthy and between 10-20 ng / mL indicates benign cancer; BM131025

[0572] • CRP, between 10-200 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 3- 35 mg / L indicates benign cancer;

[0573] • HSP90AA1, between 60-200 ng / mL indicates cancer, between 0-100 ng / mL indicates healthy and between 30-120 ng / mL indicates benign cancer;

[0574] • AFP-L3, between 10-50% indicates cancer, between 0-10% indicates healthy and between 10-15% indicates benign cancer.

[0575] 8. Lung cancer

[0576] Biomarker panel for screening of Lung cancer (Lung cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA19-9, TAG-72, CEA, HE4, SCC-Ag, NSE, Cyfra21-1, and ProGRP. with respective threshold ranges obtained are as follows:

[0577] • CAI 9-9, between 37-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy, and between 35-200 U / mL indicates benign cancer;

[0578] • TAG-72, between 10-200 U / mL indicates cancer, between 0-10 U / mL indicates healthy, and between 6-20 U / mL indicates benign cancer;

[0579] • CEA, between 5-20 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;

[0580] • HE4, between 3-110 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy, and between 2.5-10 ng / mL indicates benign cancer;

[0581] • SCC-Ag, between 2.0-10 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy, and between 1.5-9 ng / mL indicates benign benign cancer;

[0582] • NSE, between 25-1500 ng / mL indicates cancer, between 0-60 ng / mL indicates healthy, and between 15-355 ng / mL indicates benign cancer;

[0583] • Cyfra21-1, between 3.3-20 ng / mL (and higher) indicates cancer, between 0-2.0 ng / mL indicates healthy, and between 2.0-5 ng / mL indicates benign cancer; and BM131025

[0584] • ProGRP, between 100-17,000 pg / mL indicates cancer, between 0-80 pg / mL indicates healthy, and between 50-100 pg / mL indicates benign cancer.

[0585] 9. Oral cancer

[0586] Biomarker panel for screening of Oral cancer (Oral cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as SCC-Ag and CRP with respective threshold ranges obtained are as follows:

[0587] • SCC-Ag, between 2.0-15.0 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 1.5-10.0 ng / mL indicates benign cancer;

[0588] • CRP, between 10-100 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 3-30 mg / L indicates benign cancer;

[0589] 10. Ovarian cancer

[0590] Biomarker panel for screening of Ovarian cancer (Ovarian cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA125, TAG-72, HE4, and CRP with respective threshold ranges obtained are as follows:

[0591] • CA125, between 65-1500 U / mL indicates cancer, between 0-80 U / mL indicates healthy and between 35-215 U / mL indicates benign cancer;

[0592] • TAG-72, between 7-100 U / mL indicates cancer, between 3-10 U / mL indicates healthy and between 3-20 U / mL indicates benign cancer;

[0593] • HE4, between 140-1500 pmol / L indicates cancer, between 0-140 pmol / L indicates healthy and between 70-250 pmol / L indicates benign cancer; and

[0594] • CRP, between 10-100 mg / L indicates cancer, between 0-5 mg / L indicates healthy and between 3-10 mg / L indicates benign cancer.

[0595] 11. Pancreatic cancer BM131025

[0596] Biomarker panel for screening of Pancreatic cancer (Pancreatic cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA19-9, CA125, TAG-72, CEA, NSE, CRP, and VEGFA with respective threshold ranges obtained are as follows:

[0597] • CAI 9-9, between 100-1000 U / mL indicates cancer, between 0-55 U / mL indicates healthy and between 35-20 U / mL indicates benign cancer;

[0598] • CAI 25, between 65-1500 U / mL indicates cancer, between 0-35 U / mL indicates healthy and between 35-65 U / mL indicates benign cancer;

[0599] • TAG-72, between 5-100 U / mL indicates cancer, between 0-10 U / mL indicates healthy and between 3-30 U / mL indicates benign cancer;

[0600] • CEA, between 10-200 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;

[0601] • NSE, between 20-200 ng / mL indicates cancer, between 0-32 ng / mL indicates healthy, and between 12-60 ng / mL indicates benign cancer;

[0602] • CRP, between 10-100 mg / L indicates cancer, between 0-15 mg / L indicates healthy, and between 3-30 mg / L indicates benign cancer;

[0603] • VEGFA, between 300-2000 pg / mL indicates cancer, between 10-385 pg / mL indicates healthy, and between 100-500 pg / mL indicates benign cancer;

[0604] 12. Prostate cancer

[0605] Biomarker panel for screening of Prostate cancer (Prostate cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as tPSA and fPSA with respective threshold ranges obtained are as follows:

[0606] • tPSA, between 10-100 ng / mL indicates cancer, between 4-10 ng / mL indicates healthy and between 0-4 ng / mL indicates benign cancer; and

[0607] • fPSA, between 0-10% indicates cancer, between 10-25% indicates healthy and between 10-25% indicates benign cancer. BM131025

[0608] 13. Vulvar cancer

[0609] Biomarker panel for screening of Vulvar cancer (Vulvar cancer biomarker panel): Biomarker panel comprises a curated combination of biomarkers such as CA125 and HE4 with respective threshold ranges obtained are as follows:

[0610] • CAI 25, between 35-250 U / mL indicates cancer, between 0-65 U / mL indicates healthy and between 20-95 U / mL indicates benign cancer; and

[0611] • HE4, between 90-200 pmol / L indicates cancer, between 0-140 pmol / L indicates healthy and between 70-200 pmol / L indicates benign cancer.

[0612] It is crucial to mention that the development of the biomarker panel was the most challenging part of the present disclosure. With over 1,500,000+ research articles publicly available, each exploring various biomarkers across different body fluids, with diverse clinical implications and regulatory statuses, the paramount challenge has been to curate a biomarker panel optimized for maximal sensitivity, specificity, and accuracy in detecting the earliest stages of prominent cancers. By leveraging extensive in-silico research and employing sophisticated Computational algorithms, the inventors of the present disclosure have successfully identified a set of biomarkers uniquely sensitive and specific to prominent cancers. It is further crucial to note that the diagnostic utility of these biomarkers extends beyond mere detection; the biomarkers also contribute to clinical decision-making by providing insights into the prognosis and potential therapeutic strategies. It is a characterizing feature that the biomarker panel of the present disclosure has been meticulously designed to not only predict the onset of prominent cancers but also to offer prognostic evaluations and guide the selection of the most effective treatment interventions.

[0613] The method (MLB) of the present disclosure facilitates the following: BM131025

[0614] • Biomarker Network Formation: A network to explore the prognostic and therapeutic significance of biomarkers is established, which analyzes their interactions and collective impact on cancer prognosis / therapy.

[0615] • Machine Learning (ML) Model Development: A machine learning model (computational model) is created to integrate and interpret the complex data, aiming to predict early prominent cancers with high sensitivity, specificity, and accuracy. Details of development of this model has been explained in sections below.

[0616] III. Third Aspect: Machine-learning method for cancer screening using the biomarker panel

[0617] In accordance with a third aspect of the present disclosure, a machine learning method for screening of cancer(s), hereinafter referred to as “the method (MLc), is disclosed. The method (MLc) is a computer-implemented method being executed on any computational system and is adapted for screening of cancer(s) in the patient. The method (MLQ is executed on the computational system (or referred to as a machine -learning based system) being trained on dataset consisting of, but not limiting to, the biomarker panels and their respective thresholds thereof with respect to each cancer type, as is indicated in the second aspect of the present disclosure illustrated in the section described hereinabove.

[0618] Accordingly, the method (MLc) comprises the steps as follows:

[0619] At step (210), the method (MLc) includes obtaining the biological sample from the patient.

[0620] At step (220), the method (MLc) includes measuring values (or levels) of one or more biomarkers in the biological sample using a kit. Specifically, the biomarkers present within the biological sample of the patient are quantified using the kit. In a preferred embodiment, the kit using a detection technique such as ELISA / CLIA is used. However, it is evident that any other detection technique known in the art BM131025 may be deployed for quantification of biomarkers. The kit is any one selected from a group consisting of an immunoassay, radio assay, or a mass spectrometry-based kit facilitating quantification of the biomarkers within the biological sample. One or more biomarkers which are quantified are selected from the biomarker panel of the second embodiment, for screening of cancer(s) comprising a combination of 18 biomarkers including CA19-9 (Carbohydrate Antigen 19-9), CA125 (Cancer Antigen 125), TAG-72 (Tumor-Associated Glycoprotein 72), CEA (Carcinoembryonic Antigen), HE4 (Human Epididymis Protein 4), SCC-Ag (Squamous Cell Carcinoma Antigen), PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist-II), NSE (Neuron-Specific Enolase), CYFRA-21 (Cytokeratin 19 Fragment (CYFRA 21-1)), AFP (Alpha-Fetoprotein), Total PSA (PSA) (Prostate-Specific Antigen (total)), CRP (C-Reactive Protein), ProGRP (Pro-Gastrin-Releasing Peptide), HSP90 AA1 ( Heat Shock Protein 90 Alpha Family Class A Member 1), AFP - L3 (Lens culinaris agglutinin-reactive fraction of Alpha-Fetoprotein), VEGFA ( Vascular Endothelial Growth Factor A), NLR and fPSA Free Prostate-Specific Antigen). This biomarker panel is used for screening any type of the cancer(s) selected a group comprising of Breast cancer, Bihary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer.

[0621] At step (230), the method (MLc) includes obtaining clinical data from the patient. In an embodiment, the clinical data comprises of, but not limiting thereto, age, sex, family history, occupational exposure, co-morbidities and life style factors such as diet, exercise regimen, smoking history and intensity, alcohol consumption; Co- morbid conditions and Air Quality Index (AQI) for city and place of residence of the patient, at given time when said cancer screening is performed. The co-morbid conditions such as, Smoking History, Packs consumed / day, Alcohol Consumption, Physical Activity, COVID-19 History, Genetic Cancer History, Any known genetic mutations, Previous Chronic Ailments, Specific Ailment, Occupational Hazard, Metabolic Disease History etc. However, it is evident that any other additional BM131025 patient variables may be obtained to facilitate comprehensive prediction / screening of cancer.

[0622] At step (240), the method (MLc) includes feeding the biomarker values and the clinical data obtained at above steps, into a machine-learning based system for further processing. The machine-learning based system is any computational system as is indicated in the first paragraph of the method (MLc).

[0623] At step (250), the method (MLc) includes analysing of the above data by the machine-learning based system, hereinafter referred to as “the system”, which uses a computational machine learning model deploying a trained Mixture of Experts (MoE) framework and a gating network for performing tasks illustrated in the following steps hereinafter. Specifically, this integrated data is analyzed using the computational machine learning model, specifically employing the Mixture of Experts (MoE) framework, which includes multiple specialized expert models to enhance predictive accuracy. The present disclosure makes use of machine learning algorithms integrated with biomarker expression data (or values) and patient variables to generate automated diagnostic reports indicating the probability or risk of early-stage cancer presence. The system is any computational / computing system (trained on datasets of cancers, biomarker panels and thresholds thereof), and comprising of a processor, a memory, having a computer program product stored therewithin being executable on the processor, wherein, the processor executes the computational machine learning model deploying the trained Mixture of Experts (MoE) framework consisting of a plurality of knowledge-specific expert model(s) with a gating network (meta model) to perform the steps (260- 310) as illustrated below and to obtain a final prediction of cancer in the patient.

[0624] Brief Illustration of the machine-learning MoE model deployed by system

[0625] The computational machine learning model (or referred to as a MoE model) is trained on patient dataset for each cancer-type, biomarkers and thresholds thereof differentiating cancer from non-cancer. A Mixture of Experts (MoE) model is a BM131025 meta-leaming structure where multiple specialized “expert” models contribute to a final prediction, and a separate “gating network” (or “combiner” model) determines how much weight to assign each expert’s output (adjustment in the MoE framework) for a given input sample. Expert Models

[0626] • Each expert is typically a different classification or regression architecture (e.g., logistic regression, random forest, neural network).

[0627] • In some MoE setups, each expert may also specialize in a particular data regime (e.g., early-stage vs. late-stage patients, or different patient demographics).

[0628] • Alternatively, experts can be all the same algorithm but trained on different subsets or aspects of the data, learning complementary patterns.

[0629] • Expert ensemble: We train a Mixture-of-Experts where each expert specializes by organ system, aligned to the biology and the marker coverage. Model class per expert: LightGBM (this is a ML model framework), here each expert emits: Cancer presence (binary), TOO (Tissue of Origin / Disease) distribution (within its scope + “none”), Stage (ordinal 0- 4).

[0630] • Below are the expert models: BM131025

[0631] Gating Network

[0632] • The gating network is fed the same input (patient clinical data / variables + biomarker values) and learns to predict the mixing weights for each expert’s output. • Intuitively, the gating network decides “which expert (or combination) is best for this particular patient’s profile.”

[0633] Final Prediction

[0634] • The experts’ outputs are combined (usually a weighted sum), weighted by the gating network’s “confidence” distribution. This yields the final classification or regression score.

[0635] The computational machine learning model (or the MoE model) is trained on dataset for each cancer-type, biomarkers and thresholds ranges thereof differentiating cancer from non-cancer type, and therefore the system using this MoE model has an approximate sensitivity =96.57, Specificity = 98.18, and AUC =0.88, in accurate screening of cancers in the patients. These values confirm that the MoE model is both highly sensitive (minimizing false negatives) and specific (minimizing false positives). BM131025

[0636] Illustrated below are the executed bv the of the machine- machine-learning MoE model for screening of cancers in

[0637] • At step (260), classifying the biological sample into any one category selected from cancer and non-cancer (healthy) by giving a binary output (0- healthy, 1 -cancer type) in form of a P_cancer score (presence of cancer score). Particularly, this is a binary classification, where the model classifies the patient either 0 - healthy or 1 - as Cancer (for example, lung cancer). Further, in case cancer is detected, at this step, the MoE model informs on the stage of cancer and its subtype based on the measured biomarker values given as input. The MoE model deploys the expert model trained on patient dataset for stage-classification (Stage 1 to 4) and cancer subtype classification.

[0638] • At step (270), computing a confidence score for TOO (Tissue of Origin), if presence of specific cancer is detected at earlier step. The confidence scoring is performed by using following equation:

[0639] TOO confidence = 1 - H(TOO) / log K (K = #organs incl. “none”).

[0640] • At step (280), repeating both of the earlier steps by each expert model of the plurality of experts of the MoE framework (i.e., the MoE model) to obtain a set of predictions with respect to: P_cancer and TOO_confidence scores, given by each expert about their own prediction of cancer and confidence score in support thereof, wherein each expert functions as a submodel being trained on different datasets and specializes by organ, cancer-type and biomarker panel specific therefor. Each expert model is trained on dataset for particular cancer-type and biomarker panel for said cancer-type. For example, expert model 1 is trained on lung cancer, expert 2 is trained on breast cancer etc. In an embodiment, the MoE model includes 13 experts trained to capture 13 cancer types and respective cancer-specific biomarker panels illustrated above. However, it is evident that the no. of expert models BM131025 can be increased to capture any other cancer known in the art. Each expert model, detects the cancer on which it is trained using the patient data and the cancer-specific biomarker panels with thresholds (as described in second aspect of present invention).

[0641] • At step (290), evaluating the set of predictions by the gating network (or Meta model / stacking model) based on the clinical data to compute a weighted sum thereof and arrive at a final confidence score indicative of probability of presence / absence of cancer, and informing the sub-type and stage, if present. The final confidence scoring is obtained by using the following equation:

[0642] Confidence = 1 - (H(TOO) / log K) x (1 - 2xPresence_margin).

[0643] • At step (300), computing a final prediction score based on the final confidence score integrating with the clinical data. This step is performed if outcome of previous step is positive indicating presence of cancer. The final prediction score is computed using the following equation:

[0644] > P = calibrated P cancer

[0645] > S = expected(stage) / 4 (0..1 )

[0646] > a = 0.65

[0647] > Prediction score raw = a- P + (l-a) (P S)

[0648] > Prediction score = round (100 Prediction score raw) wherein, the band / scale is as follows, the prediction score of 0-9 healthy, 10-29 high-risk, 30-49 stage-I, 50-69 stage-II, 70-89 stage-III, 90-100 stage-IV so, if prediction score within ±2 of a band boundary and Confidence < 0.65, label Indeterminate with reflex / next-step note, by the MoE model. BM131025

[0649] At step (310), generating a clinical report summarizing findings obtained at previous step, with time stamped when said biological sample is screened. The method (MLC) ends at this step.

[0650] Illustrated above are the generalized steps that the processor executes when the biomarker values for the entire biomarker panel of 18 biomarkers is given as input. Thus, it is a characterizing feature of the present disclosure that the system (the MoE model) operates in any of / combination of the following modes:

[0651] 1. Mode A: Full-panel: run all 18 biomarkers once, and execute the aforementioned steps. The biomarker values for the biomarker panel in entirety are fed into the machine-learning based system which performs the aforementioned steps (step 260 to step 310) to output the final prediction score with clinical report summarizing the findings therefrom.

[0652] 2. Mode B: Tiered Reflex (tier-based approach): start testing break the testing into stages or tiers, analyzing a subset of biomarkers first, then reflexively testing additional biomarkers only if needed. The idea is to mimic a decision tree or stepwise diagnostic process, which can be more efficient. Start testing with Tier-1; add Tier-2 / 3 only if needed (MoE gating decides). Details as follows:

[0653] • Level 1 : measured values for all biomarkers from ‘Tier 1 biomarker set (gender-specific)’ depending on the patient are fed / inputted into the machine-learning based system. Thereafter, the system runs through step 260-310, to output final prediction score. However, in an event output of step 300, which computes the final confidence score, is indeterminate, that is, the MoE model cannot predict particular cancer with confidence, then level 2 is initiated.

[0654] Tier 1 biomarker set (gender-specific) comprises of: BM131025 Male set: consisting of biomarkers which includes CA19-9, CA125, TAG-72, CEA, HE4, Total PSA, CRP, and VEGFA; and Female set: consisting of biomarkers which includes CAI 9-9, CA125, TAG-72, CEA, HE4, CRP, and VEGFA

[0655] These sets have the widest disease coverage across Breast, Biliary tract, Cervical, Colon, Esophageal, Gastric, Liver, Lung, Pancreatic (and PSA for Prostate in males). These are chosen based on broad coverage or personalized factors like the patient’s age / sex and common cancer risks forthat demographic.

[0656] This Tier 1 panel is designed to be highly sensitive - it casts a wide net to catch any hint of cancer signal. Most healthy individuals will test negative at this stage (true negatives ruled out confidently), and only a subset with any suspicious readings would “trigger” further analysis. Forthat subset, the results from Tier 1 are fed into the ML model which then indicate the possibility of cancer / s and even which direction to investigate (e.g., the model might say “there’s likely a cancer signal, leaning towards a gastrointestinal and colon cancer”). Based on that guidance, a Tier 2 test is then run, adding a few more biomarkers to either confirm the signal or narrow down the cancer type.

[0657] • Level 2: where measured values for one of more biomarkers drawn from ‘Tier 2 biomarker set’ are fed into the machinelearning based system which runs the steps 260 to 310, to output final prediction score. However, in an event output of step 300, which computes the final confidence score, is indeterminate, that is, the MoE model cannot predict particular cancer with desired confidence, then level 3 is initiated. BM131025

[0658] Tier 2 / tier 3 biomarker set comprises of: biomarkers including SCC-Ag, PIVKA-II, NSE, CYFRA-21, AFP, ProGRP, HSP90AA1, AFP-L3, and fPSA.

[0659] Tier2 / Tier-3 biomarkers are always drawn from this above set. It is to be noted that the Tier-1 and the Tier-2 biomarker sets are subsets of / derived from the biomarker panel of 18 biomarkers of second aspect of the present invention. The Tier 2 biomarkers are more cancer-specific markers related to the suspected tissue / s of origin. By the end of Tier 2, the model has both the Tier 1 and Tier 2 data and can provide a refined result: e.g. “Positive for cancer, tissue origin likely liver.” If Tier 2 comes back negative or still uncertain, one could even imagine a Tier 3, Tier 4, etc., adding tests until a confident conclusion is reached. In practice, we’d hope that by Tier 2 or Tier 3 the answer is clear, so that no patient undergoes more tests than necessary to get an accurate result.

[0660] • Level 3: where measured values for system-recommended one or more biomarkers selected from the biomarker panel (of 18 biomarkers) is given as input. The MoE model recommends biomarkers, which were not tested earlier, now requires testing, based on earlier predictions, and then once values for said biomarkers are given as input, the system runs the steps 260 to 310, to output the final prediction score and the clinical report, and

[0661] In a rare event that the MoE model is still indeterminate after reaching this last level / tier, then the system returns switching to either Mode C / Mode A.

[0662] 3. Mode C: Disease-specific: run a target panel (for example: physician- directed / suspected, if a doctor suspect breast cancer, so the patient only chooses to test breast cancer biomarker panel and then the MoE model BM131025 predicts if the patient has breast cancer or not); MoE restricts to relevant experts but still returns prediction score + TOO (Tissue of Origin). This mode is a restrictive cancer-specific approach, where measured values for biomarkers from cancer-specific biomarker panels (illustrated in second aspect of the present disclosure) is fed into the machinelearning based system which activates that particular expert model specific to said cancer-type while restricting contribution from other expert models to perform the steps of 260 to 310 and output the final prediction score along with the clinical report. For example, Biomarker values for all biomarkers from the lung cancer biomarker panel is obtained from the biological sample and given as input to the system. Thereafter, the system will switch operation to Mode C, calculate P cancer and TOO confidence score, and then based on these values, the MoE model will operate at Mode C, wherein only lung cancer expert model will contribute and run steps 260 to 310, to output the final prediction score, and no other expert model will give its predictions. So, Mode C, will not have a set of predictions, as only one expert model performs the computational prediction.

[0663] However, in a rare event, the prediction is still indeterminate, the system upgrades operation to Mode A or Mode B, thereby ensuring accurate screening of cancer with high confidence.

[0664] Therefore, the system deploying the computational machine learning model performs analysis at step (250) explained above, by deciding operation in any / combination of the Mode A / B / C to arrive at the final prediction score. So, the data each of the experts get is the Mode (A / B / C), the Biomarker values and the patient variable (the clinical data). The gating network decides and routes to interplay operation of the machine-learning based system across said three modes based on the biomarkers which are quantified / given as input to the system, at step (240), for further processing. BM131025

[0665] Below section explains operation of the MoE model in above 3 modes, with reference to an example.

[0666] If Mode A,

[0667] Mode A: run all 18 analytes; same MoE stack; skip reflex; apply final prediction score computation (step 300)

[0668] If Mode C,

[0669] Mode C: restrict active experts to the organ group (e.g., Prostate), but keep E HEALTHY always on; if confidence < 0.8, system recommends upgrade to Mode A or the relevant Tier-2 pack.

[0670] If Mode B,

[0671] Mode B: Based on the patient Gender (Male or Female) test Tier 1 Biomarkers and send the values to the MoE Model, MoE will score and will output - Presence of Cancer (Y es or No), TOO distribution (over organs + “none”), Stage (ordinal - 0,1, 2, 3, 4).

[0672] This mode is explained with reference to one non-limiting example embodiment as illustrated hereinafter.

[0673] Example output -

[0674] {TOO_Exp 1: cf_score = 0.8: Lung Stage 2,

[0675] TOO Exp 2: cf score = 0.6: Colon Stage 1,

[0676] TOO Exp 3: cf score = 0.6 : Healthy, }

[0677] TOO_confidance = 1- ((0.8, 0.6) / log 2 )

[0678] TOO_confidance = ((Log 2 * Log 0.8) * (Log 2 * Log 0.6) / 2) = Range (0 -1)

[0679] Then it Computes (for all the disease):

[0680] TOO confidence = 1 - H(TOO) / log K (K = #organs incl. “none”). BM131025

[0681] Presence_margin = |P_cancer - 0.51.

[0682] Confidence = 1 - (H(TOO) / log K) x (1 - 2xPresence_margin).

[0683] Final Output from Tier 1 - {|P_margin|= 0.50, Confidence = 60, TOO: lung Cancer stage 2.}

[0684] Then it Decides:

[0685] • Tier- 1 Negative STOP: if P_cancer < 0.08 and Confidence > 0.70 finalize (prediction score <10).

[0686] • Tier-1 Positive STOP (rare): if P cancer > 0.85 and TOO conf > 0.75 finalize.

[0687] • Else REFLEX to Tier-2.

[0688] Example output -

[0689] TOO Exp 1: cf score = 0.8: Lung Stage 2,

[0690] TOO Exp 2: cf score = 0.6: Colon Stage 1,

[0691] TOO Exp 3: cf score = 0.6: Healthy,

[0692] TOO Exp 4: cf score = 0.8: Oral Cancer}

[0693] Choose Tier-2 branch:

[0694] • Take Top-2 organs by TOO probability.

[0695] • Pick primary branch = organ with higher product (TOO_prob x ( 1 -entropy within branch)) .

[0696] In the example above- ((0.8 *0.8) / 2) * (1- Entropy Decimal (It ranges from 0 to 1) for the Model l(TOO_Exp 1) and Model 2(TOO_Exp 4))

[0697] • Order fallback = the other organ.

[0698] For the above example, below are the only paths that will be recommended-

[0699] • Breast - Check HSP90AA1 - Because it strengthens separation from colon and liver signals when CA125 / HE4 / VEGFA / CEA show overlap. BM131025

[0700] Biliary tract - Check PIVKA-II - Because it helps confirm biliary vs hepatic origin where CA19-9 / CEA can be nonspecific.

[0701] • Cervical - Check SCC-Ag and ProGRP - Because SCC-Ag adds cervix specificity, while ProGRP helps distinguish cervix from lung when CA125 / HE4 / VEGFA suggest gynecologic signal.

[0702] • Colon - Check NSE, AFP, HSP90AA1 - Because these refine colon vs gastric vs liver vs breast when CA19-9 / CEA / TAG-72 / CRP produce mixed patterns.

[0703] • Esophageal - Check SCC-Ag, NSE, CYFRA-21 - Because these clarify esophagus vs gastric vs lung when Tier-1 upper-GI markers overlap.

[0704] • Gastric - Check NSE and AFP - Because they improve discrimination between gastric vs colon vs pancreatic cancers when TAG-72 / CA19-9 / CEA alone are insufficient.

[0705] . Liver - Check AFP, AFP-L3, PIVKA-II, HSP90AA1 - Because this set definitively anchors hepatic signals and separates liver from biliary / pancreas / colon.

[0706] • Lung - Check CYFRA-21, ProGRP, NSE - Because this trio resolves lung vs esophageal vs oral and sharpens stage prediction beyond Tier-1 markers.

[0707] • Oral (Head & Neck) - Check SCC-Ag and CYFRA-21 - Because they refine oral vs esophageal vs lung when CRP / CEA are nonspecific.

[0708] • Ovary - Rely on CA125, HE4, VEGFA (Tier-1) ± SCC-Ag - Because these already carry ovarian signal; SCC-Ag is added when cervix overlap must be excluded.

[0709] • Pancreatic - Check NSE - Because it helps confirm pancreatic origin when CA19-9 / CEA / TAG-72 / CRP show GI ambiguity.

[0710] • Prostate (male) - Check fPSA (compute %fPSA) - Because %fPSA is a key reflex discriminator for prostate cancer when total PSA is elevated.

[0711] After Tier-2 results arrive, Rescoring with MoE happens. New Presence of cancer, TOO distribution, Stage are computed. So it Computes (for all the disease): BM131025

[0712] TOO Exp 1: cf score = 0.8: Lung Stage 2,

[0713] TOO Exp 2: cf score = 0.6: Oral Stage 1,

[0714] It will calculate all the below 3:

[0715] TOO confidence = 1 - H(TOO) / log K (K = #organs incl. “none”).

[0716] Presence_margin = |P_cancer - 0.5|.

[0717] Confidence = 1 - (H(TOO) / log K) x (1 - 2xPresence_margin).

[0718] Example Output: {|P_margin|= 0.50, Confidence = 60, TOO: lung Cancer stage 2.}

[0719] Then it Decides:

[0720] • Tier-2 Positive STOP: if P cancer > 0.65 and TOO conf > 0.60 finalize

[0721] Positive (organ + stage).

[0722] • Tier-2 Negative STOP: ifP_cancer< 0.10 and Confidence > 0.75 -^ finalize Negative.

[0723] • Conflict handler: if two organs within A < 0.07 TOO-prob and P cancer e (0.10, 0.65) go Tier-3 with disambiguator pack (see below).

[0724] • Low-confidence: if P cancer > 0.10 but TOO conf < 0.60 Tier-3.

[0725] Select Tier-3 disambiguator (fixed rules)

[0726] In the above example:

[0727] • If Lung still top / tied after Lung pack add SCC-Ag (if not already); if SCC-Ag already present and tie persists finalize “thoracic / upper- aerodigestive indeterminate” and recommend imaging.

[0728] • If Pancreas still suspected after its Tier-2 add HSP90AA1; if already present or still tied with Esophageal / Gastric finalize “upper-GI indeterminate” (endoscopy / imaging).

[0729] • Ensure ALP + HSP90AA1 are present; if already present and tie persists finalize “Gl / hepatic indeterminate” (CT / MRI ± colonoscopy). BM131025

[0730] • If both remain close after Liver (AFP, AFP-L3, PIVKA-II, HSP90AA1) and Biliary (PIVKA-II) signals no further analytes (space saturated) finalize “hepato-biliary indeterminate subtype.”

[0731] • Ensure SCC-Ag (cervix) and HE4 (ovary / vulva) are both measured; if both present and still tied finalize “gynecologic indeterminate,” recommend pelvic imaging + HPV / cytology where applicable

[0732] • Ensure %fPSA computed; if low confidence persists finalize “urologic risk — prostate favored.”

[0733] • If Lung pack lacked NSE (edge cases) add NSE; if NSE already present and tie persists with esophagus / oral follow Thoracic rule above.

[0734] • Confirm PIVKA-II and HSP90AA1 / AFP are present (from liver / pancreas branches); if still tied finalize “peri-hepato-pancreatic indeterminate.”

[0735] • Confirm HSP90AA1 (breast) and AFP / HSP90AA1 (Gl / liver) are present; if still tied finalize “breast vs Gl / hepatic indeterminate.”

[0736] Tier-3 is the last reporting stage, so the STOP scenarios are -

[0737] Example Output: {|P_margin|= 0.56, Confidence = 90, TOO: lung Cancer stage 2.}

[0738] • Positive stop: if P cancer > 0.55 Positive (report organ = argmax TOO, stage = ordinal argmax, with confidence).

[0739] • Negative stop: if P_cancer < 0.12 Negative.

[0740] • Otherwise: Indeterminate with the specific cluster label above and directed next steps.

[0741] Thereafter, compute the final prediction score - (formulae in Step 300), to obtain final prediction.

[0742] High-Level Architecture of the MoE model of the present disclosure

[0743] The MoE Model works in a two-stage pipeline, each employing a Mixture of Experts design: BM131025

[0744] 1. Detection Stage (binary classification: Cancer vs. Healthy)

[0745] 2. Refinement Stage (a multi-task system that jointly or sequentially predicts both Stage (1-4) and Subtype (For example, Adeno / Squamous / Large, in case of lung cancer, if the sample is deemed LC)

[0746] Each stage is itself a Mixture of Experts. This allows specialized submodels for different data regions (e.g., certain biomarkers, smoking status) and robust combination of multiple ML algorithms. Below is an overview:

[0747] 1. Input: o Patient Variables: Age, Gender, Ethnicity, Smoking_History, Packs_consumed_per_day, Alcohol Consumption o Biomarker Values: e g. NSE, ProGRP, HE4, etc.

[0748] 2. Detection MoE: o Experts: multiple classification submodels (e.g., logistic regression, random forest, neural net). o Gating Network: determines how to weight each expert’s output. o Output: Probability of Lung Cancer vs. Healthy.

[0749] 3. Refinement MoE (only triggered if Detection outcome suggests LC): o Experts: specialized for stage classification (1-4) and subtype classification (3 classes). o We combine these in a multi-class approach for stage and a separate multi-class approach for subtype. The gating network merges the relevant expert outputs. o Outputs: Predicted Stage G { 1,2, 3, 4} and Subtype G {Adeno, Squamous, Large}.

[0750] Hence, the pipeline is:

[0751] (Input) — > Detection MoE — > if LC — > Refinement MoE — > Stage + Subtype.

[0752] (Input) — > Detection MoE — > if HC — > Report. BM131025

[0753] To summarize, anon-limiting description on the Computational Model is presented below. The present invention uses a Mixture of Experts (MoE) model combined with stacking for model selection based on biomarker (BM) values. a. Step 1: Individual Expert Models

[0754] In the MoE framework, "experts" are specialized models that are trained on subsets of the data or are designed to capture different aspects of the data. In the case of the present disclosure, each expert model is tailored to interpret the data from one or more of the BMs. Multiple classification models (the experts) have been trained, where each model is specialized in handling the data or patterns associated with specific BMs or combinations thereof. b. Step 2: Meta Model (The Gating Network)

[0755] The meta model, often referred to as the "gating network" in a MoE setup, learns how to allocate or weigh the input data (BM values) among the expert models. It essentially decides which expert model is most appropriate for making a prediction based on the given input. The gating network is trained on the same dataset as the expert models but with the goal of optimizing the overall performance by learning the strengths and weaknesses of each expert model. c. Step 3: Stacking for Model Selection

[0756] Stacking is a model ensembling technique where the predictions of multiple models are used as input for a second-level model (the meta- leamer) to make the final prediction. In the context of MoE of the present disclosure, stacking can be seen as part of the gating mechanism or as a separate layer that refines the outputs of the expert models based on their performance. The stacking layer will evaluate the predictions from each expert model and learn the best way to combine them, potentially using additional BM values or other relevant features (Personal Risk factors including but not limited to age, gender, smoking, and the like) as inputs to make the final prediction more accurate. d. Step 4: Operation BM131025

[0757] • For a new patient sample, the BM values are input into the system.

[0758] • Each expert model processes the input based on its training, producing a set of predictions regarding the likelihood of cancer and its risk percent & stage.

[0759] • The gating network or stacking model then evaluates these predictions, also considering personal variable data, and decides how to best combine them into a final prediction. This involves weighing some expert models more heavily than others based on the current input data. e. Step 5: Output

[0760] The output is a refined prediction that leverages the specialized knowledge of each expert model as coordinated by the meta model.

[0761] Example embodiment: Explaining detection MoE in Detail with an nonlimiting example of: Lung cancer

[0762] Aim: Decide if the patient is likely to have Lung Cancer or is Healthy.

[0763] Experts

[0764] • Expert A: A logistic regression focusing heavily on NLR and ProGRP plus basic patient info (age, smoking). This captures the strong known correlation (e.g. NLR near 3-4 LC).

[0765] • Expert B: A gradient boosting model that handles outliers in NSE / ProGRP more gracefully and uses the entire set of biomarkers.

[0766] • Expert C: A random forest specialized in low-smoking or non-smoking subsets, possibly capturing Adenocarcinoma patterns.

[0767] Each expert is trained on the same data (all patients) but incorporates different hyperparameters or features more strongly. In practice, we intentionally emphasize certain features in each expert so they learn complementary strategies.

[0768] Gating Network BM131025

[0769] A small neural net or logistic gating function that sees the entire input vector (patient variables + biomarker values).

[0770] Outputs a 3D weight vector a=[aA,aB,aC]\alpha = [\alpha_A, \alpha_B, \alpha_C]a=[aA,aB,aC], with ai=l\sum \alpha_i = l ai=l .

[0771] Example of how this works - If the gating net “notices” a pattern typical of advanced disease, maybe Expert B or C gets a higher weight. If NLR is borderline, maybe Expert A is more relevant.

[0772] Final Output

[0773] • Weighted sum of each expert’s probability: P(LC) = aApA+aB pB+aC pC. P(\text{LC}) \;=\; \alpha_A \, p_A + \alpha_B \, p_B + \alpha_C \, p_C.P(LC)=aApA+aBpB+aCpC.

[0774] • If P(LC)>0.5 P(\text{LC}) > 0.5P(LC)>0.5, we label the sample as “Lung Cancer.” Otherwise, “Healthy.”

[0775] Training Approach

[0776] • Step 1: Train each expert model on the entire LC vs. Healthy dataset.

[0777] • Step 2: Train the gating network with a standard MoE objective (e.g., maximizing the log-likelihood of correct LC vs. HC predictions) by adjusting gating parameters. The experts’ weights can be fixed or fine-tuned jointly.

[0778] Refinement MoE for Stage + Subtype

[0779] Once a patient is predicted to have LC, we produce two multi-class classifications:

[0780] 1. Stage e { 1,2, 3, 4}

[0781] 2. Subtype G {Adenocarcinoma, Squamous, Large Cell} BM131025

[0782] These tasks are combined in a single MoE structure that outputs 4 + 3 = 7 probability scores (4 for stage, 3 for subtype). Alternatively, a two-branch approach can be done as:

[0783] • Branch 1: Stage classification (1-4)

[0784] • Branch 2: Subtype classification (3 classes)

[0785] Both branches share the same gating network but have separate sets of experts or separate heads. Here’s how its finalized:

[0786] Experts

[0787] • Expert SI: A random forest specialized in early stage (1-2). Also includes partial subtype logic.

[0788] • Expert S2: A gradient boosting approach specialized in advanced stage (3-4). Subtype might differ for advanced squamous.

[0789] • Expert S3: A neural net that tries to directly map (ProGRP, NLR, other BMs + patient data) (Stage distribution + Subtype distribution).

[0790] • Expert S4: A Classification model (SVM) focusing on never-smokers vs. heavy smokers, which can impact histology patterns.

[0791] Gating Network

[0792] • Another gating function sees the same input features (the patient is already predicted LC, so we might specifically weight high BM11 for advanced stage, etc.).

[0793] • It outputs P=[PS l,pS2,pS3,pS4]\beta = [\beta_{Sl}, \beta_{S2}, \beta_{S3}, \beta_{S4}]p=[pSl,pS2,pS3,pS4],

[0794] Output

[0795] • For stage: Weighted sum of each expert’s stage probabilities yields a final 4-class distribution: P(Stage=i)P(\text{ Stage} = i)P(Stage=i).

[0796] • For subtype: Similarly, a 3-class distribution.

[0797] Then, the most likely stage and subtype is picked. Alternatively, the probability distributions for each are kept. BM131025

[0798] Training

[0799] • Used the LC-only subset of data (since only stage / subtype is done for confirmed LC).

[0800] • Train each expert on all 20 (or more) LC examples, but the gating network learns to route certain biomarker / patient profiles to the best expert(s).

[0801] • We have a multi-class cross-entropy loss that sums across both stage and subtype tasks. The gating network + experts are optimized to minimize total error in stage + subtype classification simultaneously.

[0802] Data Handling

[0803] • Each expert is kept relatively simple (e.g., smaller tree depth, logistic regression).

[0804] • Incorporation of cross-validation and possibly regularization (L2 or early stopping in boosting) to ensure robust performance.

[0805] • Addition of new experts specialized for different sub-cohorts (e.g., specific ethnic backgrounds or new biomarker panels) is possible. The gating network is retrained to exploit newly minted experts while still using the older ones if they remain strong for certain data segments. Thus, scaling-up of the present MoE model possible. This flexible structure ensures the MoE evolves gracefully, maintaining performance and potentially improving as the dataset grow.

[0806] • At a threshold of 0.5, the MoE model achieves a balanced performance with 94% precision and 95% recall.

[0807] Therefore, integrating and weighing these biomarkers from the biomarker panels within the machine learning algorithm, rather than using just one biomarker uniquely tailored to early-stage disease detection, is a characterizing feature of the present disclosure. The present disclosure uses combined predictive power of multiple biomarkers rather than depending only on any single “stage-specific” biomarker. BM131025

[0808] The set of biomarkers and the computational model forms the characterizing feature of the present disclosure. The set of biomarkers and the computational model in combination gives a unique result with highest sensitivity, specificity and accuracy. It is important to note that any change in either of them may impact the precision of the output.

[0809] IV. Fourth Aspect: Kit for quantification of biomarkers

[0810] The kit, as described earlier, in one embodiment, is a diagnostic kit. The kit, is an immunoassay, radio assay, mass-spectrometry-based kit using any detection technique known in the art. It is evident to a skilled person that any other technique for detection / quantification of the biomarker levels can be employed by the kit. The kit includes but is not limited to essential reagents, antibodies, primers solutions to detect biomarkers in the biological sample. These are primarily the chemicals that are utilized by diagnostic labs to run the test for diagnosing prominent cancers. The kit, system and methods of the present disclosure address the critical challenges in early cancer detection through a novel clinical inference engine that integrates advanced biomarker analysis with computational algorithms. The kit, system and methods of the present disclosure is specifically designed to overcome the limitations of current diagnostic methodologies, which often lead to late-stage cancer detection, by offering a non-invasive, efficient, and accessible means of screening.

[0811] The following non-limiting points are the characterizing features that demonstrate the “technical effect” and “technical contribution” of the present disclosure, in addition to being the points that act as the technical differentiators and advancements of the present disclosure w.r.t the prior art:

[0812] Integration with Existing Infrastructure: The biomarker panel of the present disclosure is designed to work with regularly available infrastructure and machinery, which means hospitals and clinics can utilize their current laboratory equipment without the need for costly upgrades or specialized machines. This not only reduces the initial capital expenditure for healthcare providers but also BM131025 accelerates the adoption process, as there's no need to wait for the installation and setup of new equipment.

[0813] Reduced Operational Costs: By leveraging existing machinery, the kit, system and process of the present disclosure can significantly lower operational costs. Specialized equipment often requires unique maintenance, proprietary reagents, and trained personnel, all of which add to the operational expenses. In contrast, utilizing widely available infrastructure means maintenance and operational procedures remain standard, which can lead to lower running costs and simplified training for laboratory staff.

[0814] High Sensitivity and Specificity: Through rigorous in silico research and the application of multiple computational models, the inventors of the present disclosure have curated a biomarker panel that is highly sensitive and specific to prominent cancers. This ensures that the present system can accurately distinguish between healthy individuals and those with prominent cancers, reducing the likelihood of false positives and negatives that are common with less sophisticated screening methods.

[0815] Accessibility and Scalability: The compatibility with commonly available machines, the non-invasive, radiation-free nature of screening makes the kit, system and method / process of the present disclosure more accessible to a wider range of healthcare facilities, including those in resource-limited settings. This not only broadens the potential market but also supports scalability. Facilities can increase their screening and diagnostic capabilities without significant additional investments in hardware, making it easier to expand services in response to growing demand.

[0816] Faster Implementation and Training: Since the kit, system and method of the present disclosure does not require specialized machines, the implementation timeline can be significantly shorter. Facilities can quickly integrate the biomarker panel of the present disclosure into their existing workflows, minimizing downtime. Additionally, the learning curve for operating personnel is likely to be less steep, as BM131025 they will be working with familiar equipment. This can lead to quicker adoption and proficiency, further economizing on training costs.

[0817] Versatility and Adaptability: The kit, system and process of the present disclosure allows for easy updates and adaptations to the panel of biomarkers as scientific knowledge advances. This could be a cost-effective feature in the long run, as it enables healthcare providers to stay at the forefront of diagnostic technology without the need for new hardware purchases.

[0818] Enhanced Clinical Decision Making: Beyond mere detection, the biomarker panel of the present disclosure provides valuable prognostic information and insights into potential therapeutic interventions. This dual functionality supports clinicians in making more informed decisions regarding patient care, from predicting disease progression to identifying the most appropriate treatment options based on individual biomarker profiles.

[0819] To summarize, the kit, system and method of the present disclosure present a transformative approach to cancer screening and diagnosis, addressing the critical need for early detection through a combination of innovative biomarker research, advanced computational algorithms, and integration with existing healthcare infrastructure. This comprehensive solution not only enhances diagnostic accuracy but also empowers clinicians with actionable insights, thereby improving patient prognosis and facilitating personalized treatment strategies.

[0820] TECHNICAL ADVANTAGES AND ECONOMIC SIGNIFICANCE

[0821] The technical advantages and economic significance of the kit, system and method of screening of the present disclosure include but are not limited to:

[0822] • Early Detection

[0823] • High Sensitivity and Specificity

[0824] • Integration with Existing Infrastructure

[0825] • Reduced Operational Costs

[0826] • Faster Implementation and Training BM131025

[0827] • Versatility and Adaptability

[0828] • Enhanced Clinical Decision Making

[0829] • Accessibility and Scalability

[0830] • Non-invasive • Less false positives or negatives

[0831] • Provides insights into the prognosis and potential therapeutic strategies.

[0832] The embodiments described herein above are non-limiting. The foregoing descriptive matter is to be interpreted merely as an illustration of the concept of the present disclosure and it is in no way to be construed as a limitation. Description of terminologies, concepts and processes known to persons acquainted with technology has been avoided to preclude beclouding of the afore-stated embodiments.

Claims

BM131025WE CLAIM1. A machine learning method (MLB) for curating biomarker panel, the biomarker panel being adapted for use in screening of cancer(s) in a patient, the machine-learning method (MLB) is implemented by a computational system, the machine learning method (MLB) comprising the steps of: a) Retrieving and collecting scientific literature from one or more data sources to create a comprehensive repository of articles related to biomarkers associated with cancer using data acquisition techniques; b) Shortlisting and filtering out relevant scientific literature (or articles) based on title / abstract screening by applying pre-defined triage guidelines (or criteria) using Natural Language Processing (NLP) techniques via a machine learning (ML) model, wherein the ML model is trained to automatically predict triage decisions and classify the articles based on their pre-defined guidelines to arrive at a relevant dataset, and thereafter human expert reviews the relevant dataset to validate predictions made by the ML model, and wherein further the ML model is iteratively re-trained on reinforcement learning from human feedback (RLHF) to validate and enhance accuracy of predictions made therefrom; c) Reviewing the relevant dataset by screening full-text thereof by applying pre-defined secondary guidelines (or criteria) via the ML model trained thereon and performing quality assessment thereafter via manual review to ensure reliability and accuracy of information being processed, wherein the relevant dataset which passes through the quality assessment further undergoes a secondary analysis (manual review) for further validation thereby offering an additional layer of scrutiny; d) Performing in-depth analysis of the relevant dataset obtained at step c, against pre-defined detailed analysis guidelines (or criteria)BM131025 by the ML model trained thereon to extract multiple data elements therefrom, and performing quality assessment thereafter via manual review followed by re-training of the ML model via RLHF to ensure reliability and accuracy of information being processed whilst ensuring adaptation to emerging terminologies; e) Assigning quantifiable score to each candidate biomarker identified at step d using a weighted scoring system adapted by the ML model, wherein the scores are based on comprehensive evaluation of their diagnostic performance in early-stage cancer detection, study robustness (evidence base), biological characteristics (age, histology and stage, gender, race, ethinicity) and clinical implications thereby integrating the information obtained from the data elements extracted at step d while arriving at a final biomarker score to ensure a holistic assessment of each biomarker’s utility; f) Aggregating and synthesizing these final biomarker scores across the relevant dataset into a composite score that represents each candidate biomarker using the weighted scoring system by the ML model, wherein the composite score is computed by averaging and weighing the scores obtained at step e by the biomarker’s diagnostic performance and no. of evidences (no. of studies) in support thereof to arrive at the final composite score; g) Categorization of the candidate biomarkers into distinct performance profiles (or category) such as specific guardian (high specificity), sensitive detector (high sensitivity), balanced performer (balanced specificity and sensitivity) and early scout (in early-stage cancer detection) based on their diagnostic performance; h) Ranking of the candidate biomarkers within each of the performance category based on the biomarker’s composite scores using feature selection techniques, wherein the biomarkers with highBM131025 composite scores within the said performance category along with a large evidence base gets a ‘high priority rank’ and vice-a-versa, and i) Validating the ranked biomarkers through cross-validation (or orthogonal validation) across standard clinical databases to confirm their diagnostic performance for early-stage cancer detection, wherein, the steps (a) to (i) are repeated for each cancer type to arrive at the biomarker panel with high efficiency in screening of early-stage cancer(s) in the patient.

2. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the data sources include multiple journals, databases or repositories from where scientific articles are sourced.

3. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the data acquisition techniques are configured to access and integrate published scientific literature from multiple journals, databases, or repositories.

4. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the Natural Language Processing (NLP) techniques comprises of any one selected from Automated triage, Long Short-Term Memory (LSTM) Networks with Reinforcement Learning with Human in Loop Feedback mechanism and combinations thereof.

5. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the data elements refer to multiple fields such as cohort characteristics including patient / study population, patient demographic data, and patient occupational history, disease histology (subtype, stage), sample / specimen type and biomarker attributes such as name, type, diagnostic performance based on Area Under Curve (AUC), sensitivity, specificity, and statistical data (p-value), expression profile, and clinical implications (prognostic and therapeutic) thereof in early-stage cancer detection.BM1310256. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the feature selection techniques are selected from a group consisting of recursive feature elimination with classifiers including random forest (RF), support vector machines (SVM), and logistic / linear models; tree-ensemble and gradient boosting methods including GBM, XGBoost, LightGBM, and CatBoost; embedded regularization methods including LASSO (LI), ridge (L2), and Elastic Net; filter methods including mutual information, minimum-redundancymaximum-relevance (mRMR), Relief / ReliefF, %2tests, ANOVA / F-score, and Fisher score; permutation importance; attribution-based metrics including SHAP (TreeSHAP / Kemal SHAP) and integrated gradients; wrapper methods including sequential forward selection (SFS), sequential backward selection (SBS), sequential floating forward / backward selection (SFFS / SFBS), genetic-algorithm search, and Bayesian optimization for subset selection; correlation / variance / VIF-based pruning; and functional equivalents and combinations thereof.

7. The machine learning method (MLB) for curating biomarker panel as claimed in claim 1, wherein the standard clinical databases are selected from the group consisting of biomarker repositories such as EDRN, OncoMX, Knapsack family biomarkers, and human protein atlas; clinical trial registries; oncology databases such as KM Plotter, TCGA, GEO databases and combinations thereof.

8. A biomarker panel for screening of cancer(s), wherein the biomarker panel adapted for early-stage screening of cancer(s) includes one or more biomarkers selected from a group consisting of CA19-9 (Carbohydrate Antigen 19-9), CA125 (Cancer Antigen 125), TAG-72 (Tumor-Associated Glycoprotein 72), CEA (Carcinoembryonic Antigen), HE4 (Human Epididymis Protein 4), SCC-Ag (Squamous Cell Carcinoma Antigen), PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist- II), NSE (Neuron-Specific Enolase), CYFRA-21 (Cytokeratin 19 Fragment (CYFRA 21-1)), AFP (Alpha-Fetoprotein), Total PSA (PSA) (Prostate-BM131025Specific Antigen (total)), CRP (C-Reactive Protein), ProGRP (Pro-Gastrin- Releasing Peptide), HSP90 AA1 ( Heat Shock Protein 90 Alpha Family Class A Member 1), AFP - L3 (Lens culinaris agglutinin-reactive fraction of Alpha-Fetoprotein), VEGFA ( Vascular Endothelial Growth Factor A), Neutrophil to Lymphocyte Ratio (NLR) and fPSA Free Prostate-Specific Antigen).

9. The biomarker panel for screening of cancer(s) as claimed in claim 8, wherein the cancer(s) is selected from a group comprising of Breast cancer, Bihary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer.

10. A biomarker panel for screening of breast cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CA125, CEA, HPS90AA1, VEGFA and HE4.

11. The biomarker panel for screening of breast cancer as claimed in claim 10, wherein, threshold ranges of the biomarkers• CAI 9-9, between 49-76 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 30-55 U / mL indicates benign cancer;• CA125, between 12-65 U / mL indicates cancer, between 9-32 U / mL indicates healthy, and between 12-45 U / mL indicates benign cancer;• CEA, between 1.5-30 ng / mL indicates cancer, between 0-8.3 ng / mL indicates healthy and between 0.4-6 ng / mL indicates benign cancer;• HPS90AA1, between 310-700 pg / mL indicates cancer, between 180-320 pg / mL indicates healthy and between 220-360 pg / mL indicates benign cancer;• VEGFA, between 57-219 pg / ml indicates cancer, between 11-47 pg / ml indicates healthy and between 19-83 pg / ml indicates benign cancer, andBM131025• HE4, between 45-81 pmol / L indicates cancer, between 45-70 pmol / L indicates healthy and between 50-72 pmol / L indicates benign cancer.

12. A biomarker panel for screening of biliary tract cancer comprising a combination of biomarkers selected from a group consisting of CA19-9, CA125, TAG-72, CEA and PIVKA-II.

13. The biomarker panel for screening of biliary tract cancer as claimed in claim 12, wherein, threshold ranges of the biomarkers• CAI 9-9, between 129-1200 U / mL indicates cancer, between 0-37 U / mL indicates healthy and between 37-1,000 U / mL indicates benign cancer;• CAI 25, between 105-550 U / mL indicates cancer, between 0-35 U / mL indicates healthy and between 35-350 U / mL indicates benign cancer;• TAG-72, between 3-50 U / mL indicates cancer, between 0-7 U / mL indicates healthy and between 3-39 U / mL indicates benign cancer;• CEA, between 10-90 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 5-20 ng / mL indicates benign cancer;• PIVKA-II, between 30-800 mAU / mL indicates cancer, between 0- 40 mAU / mL indicates healthy and between 10-40 mAU / mL indicates benign cancer.

14. A biomarker panel for screening of Cervical-Endometrial cancer comprising a combination of biomarkers selected from a group consisting of CA19-9, CA125, TAG-72, HE4, SCC-Ag, ProGRP, VEGFA.

15. The biomarker panel for screening of Cervical-Endometrial cancer as claimed in claim 14, wherein, threshold ranges of the biomarkers• CAI 9-9, between 20-100 U / mL indicates cancer, between 0-18 U / mL indicates healthy, and between 18-40 U / mL indicates benign;• CAI 25, between 20-100 U / mL indicates cancer, between 0-15 U / mL indicates healthy, and between 15-40 U / mL indicates benign;BM131025• TAG-72, between 7-40 U / mL indicates cancer, between 0-6.9 U / mL indicates healthy, and between 6.9-12 U / mL indicates benign;• HE4, between 90-200 pmol / L indicates cancer, between 17-75 pmol / L indicates healthy, and between 60-90 pmol / L indicates benign;• SCC-Ag, between 2.0-10 ng / mL indicates cancer, between 0-1.5 ng / mL indicates healthy, and between 1.5-2.0 ng / mL indicates benign;• ProGRP, between 80-300 pg / mL indicates cancer, between 0-50 pg / mL indicates healthy and between 50-80 pg / mL indicates benign cancer; and• VEGFA, between 300-1,300 pg / mL indicates cancer, between 50- 250 pg / mL indicates healthy, and between 200-350 pg / mL indicates benign cancer.

16. A biomarker panel for screening of Colorectal cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CA125, TAG-72, CEA, HE4, NSE, AFP, CRP, HSP90 AA1, and VEGFA.

17. The biomarker panel for screening of Colorectal cancer as claimed in claim 16, wherein, threshold ranges of the biomarkers• CAI 9-9, between 35-1,000 U / mL indicates cancer, between 0-50 U / mL indicates healthy, and between 37-180 U / mL indicates benign cancer;• CA125, 35-250 U / mL indicates cancer, between 0-55 U / mL indicates healthy, and between 20-105 U / mL indicates benign cancer;• TAG-72, 7-100 U / mL indicates cancer, between 0-20 U / mL indicates healthy, and between 2-45 U / mL indicates benign cancer;• CEA, between 5-200 ng / mL indicates cancer, between 0.2-7 ng / mL indicates healthy, and between 3-120 ng / mL indicates benign cancer;BM131025• HE4, between 70-200 pmol / L indicates cancer, between 0-85 pmol / L indicates healthy, and between 45-135 pmol / L indicates benign cancer;• NSE, between 9-40 ng / mL indicates cancer, between 0-18 ng / mL indicates healthy, and between 12-26.5 ng / mL indicates benign cancer;• AFP, between 10-400 ng / mL indicates cancer, between 0-18 ng / mL indicates healthy, and between 15-100 ng / mL indicates benign cancer;• CRP, between 10-100 mg / L indicates cancer, between 0-10 mg / L indicates healthy, and between 3-30 mg / L indicates benign;• HSP90 AA1, between 20-140 ng / mL indicates cancer, between 0- 10 ng / mL indicates healthy, and between 5-45 ng / mL indicates benign cancer; and• VEGFA, between 200-1,000 pg / mL indicates cancer, between 0- 50 pg / mL indicates healthy, and between 35-400 pg / mL indicates benign cancer.

18. A biomarker panel for screening of Esophageal cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CA125, TAG-72, CEA, SCC-Ag, and NSE.

19. The biomarker panel for screening of Esophageal cancer as claimed in claim 18, wherein, threshold ranges of the biomarkers• CAI 9-9, between 40-1,000 U / mL indicates cancer, between 0-105 U / mL indicates healthy, and between 35-330 U / mL indicates benign cancer;• CA125, between 35-500 U / mL indicates cancer, between 0-65 U / mL indicates healthy, and between 20-135 U / mL indicates benign cancer;• TAG-72, between 7-500 U / mL indicates cancer, between 0-30 U / mL indicates healthy, and between 6-190 U / mL indicates benign cancer;BM131025• CEA, between 5.0-200 ng / mL indicates cancer, between 0-10 ng / mL indicates healthy, and between 3-45 ng / mL indicates benign cancer;• SCC-Ag, between 2.0-50 ng / mL indicates cancer, between 0-3 ng / mL indicates healthy, and between 1.5-9 ng / mL indicates benign cancer; and• NSE, between 20-150 ng / mL indicates cancer, between 0-35 ng / mL indicates healthy, and between 10-60 ng / mL indicates benign cancer.

20. A biomarker panel for screening of Gastric Cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CA125, TAG-72, CEA, PIVKA-II, ALP, and CRP.

21. The biomarker panel for screening of Gastric cancer as claimed in claim 20, wherein, threshold ranges of the biomarkers• CA19-9, between 35-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 25-200 U / mL indicates benign cancer;• CA125, between 35-200 U / mL indicates cancer, between 0-45 U / mL indicates healthy and between 35-95 U / mL indicates benign cancer;• TAG-72, between 6- 100 U / mL indicates cancer, between 0-9 U / mL indicates healthy and between 6-15 U / mL indicates benign cancer;• CEA, between 5-30 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;• PIVKA-II, between 40-1000 mAU / mL indicates cancer, between 0-40 mAU / mL indicates healthy and between 40-60 mAU / mL indicates benign cancer;• ALP, between 20-500 ng / mL indicates cancer, between 0-30 ng / mL indicates healthy and between 10-60 ng / mL indicates benign cancer; andBM131025• CRP, between 10-200 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 5-35 mg / L indicates benign cancer.

22. A biomarker panel for screening of Liver Cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CEA, PIVKA-II, AFP, CRP, HSP90AA1, and AFP-L3.

23. The biomarker panel for screening of Liver cancer as claimed in claim 22, wherein, threshold ranges of the biomarkers• CA19-9, between 35-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy and between 25-200 U / mL indicates benign cancer;• CEA, between 5-200 ng / mL indicates cancer, between 0-7 ng / mL indicates healthy and between 3-10 U / mL indicates benign cancer;• PIVKA-II, between 40-1000 mAU / mL indicates cancer, between 0-40 mAU / mL indicates healthy and between 40-60 mAU / mL indicates benign cancer;• AFP, between 20-500 ng / mL indicates cancer, between 0-35 ng / mL indicates healthy and between 10-20 ng / mL indicates benign cancer;• CRP, between 10-200 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 3- 35 mg / L indicates benign cancer;• HSP90AA1, between 60-200 ng / mL indicates cancer, between 0- 100 ng / mL indicates healthy and between 30-120 ng / mL indicates benign cancer; and• AFP-L3, between 10-50% indicates cancer, between 0-10% indicates healthy and between 10-15% indicates benign cancer.

24. A biomarker panel for screening of Lung Cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, TAG-72, CEA, HE4, SCC-Ag, NSE, Cyfra21-1, and ProGRP.

25. The biomarker panel for screening of Lung cancer as claimed in claim 24, wherein, threshold ranges of the biomarkersBM131025• CA19-9, between 37-1000 U / mL indicates cancer, between 0-60 U / mL indicates healthy, and between 35-200 U / mL indicates benign cancer;• TAG-72, between 10-200 U / mL indicates cancer, between 0-10 U / mL indicates healthy, and between 6-20 U / mL indicates benign cancer;• CEA, between 5-20 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;• HE4, between 3-110 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy, and between 2.5-10 ng / mL indicates benign cancer;• SCC-Ag, between 2.0-10 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy, and between 1.5-9 ng / mL indicates benign benign cancer;• NSE, between 25-1500 ng / mL indicates cancer, between 0-60 ng / mL indicates healthy, and between 15-355 ng / mL indicates benign cancer;• Cyfra21-1, between 3.3-20 ng / mL (and higher) indicates cancer, between 0-2.0 ng / mL indicates healthy, and between 2.0-5 ng / mL indicates benign cancer; and• ProGRP, between 100-17,000 pg / mL indicates cancer, between 0- 80 pg / mL indicates healthy, and between 50-100 pg / mL indicates benign cancer.

26. A biomarker panel for screening of Oral Cancer comprising a combination of biomarkers selected from a group consisting of SCC-Ag and CRP.

27. The biomarker panel for screening of Oral cancer as claimed in claim 26, wherein, threshold ranges of the biomarkers• SCC-Ag, between 2.0-15.0 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 1.5-10.0 ng / mL indicates benign cancer; andBM131025• CRP, between 10-100 mg / L indicates cancer, between 0-15 mg / L indicates healthy and between 3-30 mg / L indicates benign cancer.

28. A biomarker panel for screening of Ovarian Cancer comprising a combination of biomarkers selected from a group consisting of CAI 25, TAG-72, HE4, and CRP.

29. The biomarker panel for screening of Ovarian cancer as claimed in claim 28, wherein, threshold ranges of the biomarkers• CA125, between 65-1500 U / mL indicates cancer, between 0-80 U / mL indicates healthy and between 35-215 U / mL indicates benign cancer;• TAG-72, between 7-100 U / mL indicates cancer, between 3-10 U / mL indicates healthy and between 3-20 U / mL indicates benign cancer;• HE4, between 140-1500 pmol / L indicates cancer, between 0-140 pmol / L indicates healthy and between 70-250 pmol / L indicates benign cancer; and• CRP, between 10-100 mg / L indicates cancer, between 0-5 mg / L indicates healthy and between 3-10 mg / L indicates benign cancer.

30. A biomarker panel for screening of Pancreatic Cancer comprising a combination of biomarkers selected from a group consisting of CAI 9-9, CA125, TAG-72, CEA, NSE, CRP, and VEGFA.

31. The biomarker panel for screening of Pancreatic cancer as claimed in claim 30, wherein, threshold ranges of the biomarkers• CAI 9-9, between 100-1000 U / mL indicates cancer, between 0-55 U / mL indicates healthy and between 35-20 U / mL indicates benign cancer;• CA125, between 65-1500 U / mL indicates cancer, between 0-35 U / mL indicates healthy and between 35-65 U / mL indicates benign cancer;BM131025• TAG-72, between 5-100 U / mL indicates cancer, between 0-10U / mL indicates healthy and between 3-30 U / mL indicates benign cancer;• CEA, between 10-200 ng / mL indicates cancer, between 0-5 ng / mL indicates healthy and between 3-10 ng / mL indicates benign cancer;• NSE, between 20-200 ng / mL indicates cancer, between 0-32 ng / mL indicates healthy, and between 12-60 ng / mL indicates benign cancer;• CRP, between 10-100 mg / L indicates cancer, between 0-15 mg / L indicates healthy, and between 3-30 mg / L indicates benign cancer; and• VEGEA, between 300-2000 pg / mL indicates cancer, between 10- 385 pg / mL indicates healthy, and between 100-500 pg / mL indicates benign cancer.

32. A biomarker panel for screening of Prostate Cancer comprising a combination of biomarkers selected from a group consisting of tPSA and fPSA.

33. The biomarker panel for screening of Prostate cancer as claimed in claim 32, wherein, threshold ranges of the biomarkers• tPSA, between 10-100 ng / mL indicates cancer, between 4-10 ng / mL indicates healthy and between 0-4 ng / mL indicates benign cancer; and• fPSA, between 0-10% indicates cancer, between 10-25% indicates healthy and between 10-25% indicates benign cancer.

34. A biomarker panel for screening of Vulvar Cancer comprising a combination of biomarkers selected from a group consisting of CAI 25 and HE4.

35. The biomarker panel for screening of Vulvar cancer as claimed in claim 34, wherein, threshold ranges of the biomarkersBM131025• CA125, between 35-250 U / mL indicates cancer, between 0-65 U / mL indicates healthy and between 20-95 U / mL indicates benign cancer; and• HE4, between 90-200 pmol / L indicates cancer, between 0-140 pmol / L indicates healthy and between 70-200 pmol / L indicates benign cancer.

36. A method (MLc) for screening of cancer(s), the method (MLc) being a computer-implemented method executed on any computational system and adapted for screening of cancer(s), the method (MLc) comprising the steps of: a) Obtaining a biological sample from a patient; b) Measuring values (or levels) of one or more biomarkers in the biological sample using a kit, the biomarkers being selected from a biomarker panel for cancer(s); c) Obtaining clinical data from the patient, wherein the clinical data comprises of age, sex, family history, occupational exposure, co-morbidities and life style factors such as diet, exercise regimen, smoking history and intensity, alcohol consumption and Air Quality Index (AQI) for city and place of residence of the patient; d) Feeding the biomarker values and the clinical data into a machine-learning based system, and e) Analysing of the above data by the machine-learning based system using a computational machine learning model deploying a trained Mixture of Experts (MoE) framework to perform the steps of: i) Classifying the biological sample into any one category from cancer and non-cancer (healthy) by giving a binary output (0-healthy, 1- specific cancer type) in form of a P_cancer score, and in event the cancer is detected, determining stage thereof basedBM131025 on the measured biomarker values by an expert model trained on patient dataset for stageclassification (Stage 1 to 4) and cancer subtype classification; ii) Computing a confidence score for TOO (Tissue of Origin), if presence of cancer is detected at step i); iii) Repeating the steps i) and ii) by each expert of a plurality of experts of the MoE framework to obtain a set of predictions with respect to P cancer and TOO confidence scores, wherein each expert functions as a submodel being trained on different datasets and specializes by organ, cancer-type and biomarker panel specific therefor; iv) Evaluating the set of predictions by a gating network (or Meta model / stacking model) based on the clinical data to compute a weighted sum thereof and arrive at a final confidence score indicative of probability of presence / absence of cancer, and informing the sub-type and stage, if present; and v) Computing a final prediction score based on the final confidence score integrating with the clinical data, if outcome of the step iv) is positive indicating presence of cancer, and thereafter, vi) Generating a clinical report summarizing findings obtained at step v) and time stamped when the biological sample is screened, wherein, the biomarker panel for screening of cancer(s) comprises a combination of CAI 9-9 (Carbohydrate Antigen 19-9), CAI 25 (Cancer Antigen 125), TAG-72 (Tumor-Associated Glycoprotein 72), CEA (Carcinoembryonic Antigen), HE4 (Human Epididymis Protein 4), SCC-AgBM131025(Squamous Cell Carcinoma Antigen), PIVKA-II (Protein Induced by Vitamin K Absence or Antagonist-II), NSE (Neuron-Specific Enolase), CYFRA-21 (Cytokeratin 19 Fragment (CYFRA 21-1)), AFP (Alpha- Fetoprotein), Total PSA (PSA) (Prostate-Specific Antigen (total)), CRP (C- Reactive Protein), ProGRP (Pro-Gastrin-Releasing Peptide), HSP90 AA1 ( Heat Shock Protein 90 Alpha Family Class A Member 1), AFP - L3 (Lens culinaris agglutinin-reactive fraction of Alpha-Fetoprotein), VEGFA ( Vascular Endothelial Growth Factor A), NLR and fPSA Free Prostate- Specific Antigen), wherein, the cancer is selected from a group comprising of Breast cancer, Biliary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer, wherein, the computational machine learning model is trained on patient dataset for each cancer-type, biomarkers and thresholds thereof differentiating cancer from non-cancer, and wherein further, the computational machine learning model has high sensitivity and specificity thereby offering accurate screening of cancer in the patients.

37. The method (MLc) as claimed in claim 36, wherein the step e) is achieved by the machine-learning based system deploying the computational machine learning model operating in any of the three modes- Mode A, Mode B and Mode C, wherein, in Mode A, biomarker values for the biomarker panel (18 biomarkers) in entirety is fed into the machine -learning based system which performs the steps i) through vi) to output the final prediction score with clinical report summarizing the findings therefrom;BM131025 in Mode B, operates as a tier-based approach, where, at level 1, measured values for all biomarkers selected from Tier 1 biomarker set (gender-specific) are fed into the machine-learning based system which follows through steps i) to vi) to predict probability of cancer, and thereafter, in an event output of the level 1 is indeterminate at step iv), then level 2 is initiated where measured values for one of more biomarkers drawn from Tier 2 biomarker set are fed into the machine-learning based system which follows through steps i) to vi) to predict probability of cancer, and further thereafter, in an event output of the level 2 is still indeterminate at step iv), then level 3 is initiated where measured values for system- recommended one or more biomarkers selected from the biomarker panel (of 18 biomarkers), and which were not measured earlier are fed into said system which follows through the steps i) to vi) to obtain the final prediction score and the clinical report, and in Mode C, operates as a restrictive cancer-specific approach, where measured values for biomarkers from cancer-specific biomarker panel of any of claims 10- 35, is fed into the machinelearning based system which activates the expert model specific to said cancer-type while restricting contribution from other expert models to perform steps of i) through vi) and output the final prediction score along with the clinical report, and in an event the prediction is indeterminate said system upgrades operation to Mode A or Mode B, thereby ensuring accurate screening of cancer, wherein, the Tier 1 biomarker set (gender-specific) comprises of:Male set, consisting ofbiomarkers which includes CA19-9, CA125, TAG-72, CEA, HE4, Total PSA, CRP, and VEGFA, andBM131025 female set, consisting of biomarkers which includes CAI 9-9, CA125, TAG-72, CEA, HE4, CRP, and VEGFA thereby ensuring wide cancer coverage, and the Tier-2 biomarker set, comprises of biomarkers including SCC-Ag, PIVKA-II, NSE, CYFRA-21, AFP, ProGRP, HSP90AA1, AFP-L3, and fPSA, wherein, the Tier- 1 and the Tier-2 biomarker sets being subsets of the biomarker panel, and wherein, the gating network decides and routes to interplay operation of the machine -learning based system across said three modes based on the biomarkers which are quantified, and given as input at step d) for further processing.

38. The method (MLc) as claimed in claim 36, wherein the biological sample is a non-invasive body fluid(s) and is any one selected from the group comprising blood, serum, plasma, urine, saliva, semen, breast exudate, cerebrospinal fluid, tears, sputum, mucous, lymph, cytosols, ascites, pleural effusions, peritoneal fluid, amniotic fluid, bladder washes, bronchioalveolar lavages and blood cells.

39. The method (MLc) as claimed in claim 36, wherein the kit is any one selected from a group consisting of an immunoassay, radio assay, or a mass spectrometry-based kit facilitating quantification of the biomarkers within the biological sample.

40. A machine-learning based system for screening of cancer(s), the machine-learning based system being a computational system comprising of: i) a processor, and ii) a memory, having a computer program product stored therewithin being executable on the processor, wherein, the processor executes a computational machine learning model deploying a trained Mixture of Experts (MoE) framework consisting of aBM131025 plurality of knowledge-specific expert model(s) with a gating network (meta model) being configured to: a) receive measured biomarker values and clinical data from a patient; and b) analyse the above data by performing the steps of:• classifying the patient into cancer and non-cancer (healthy) by giving a binary output (0-healthy, 1- cancer-type) in form of a P_cancer score, and in case cancer is detected, determining stage thereof based on the measured biomarker values by an expert model trained on patient dataset for stageclassification (Stage 1 to 4) and cancer subtype classification;• computing a confidence score for TOO (Tissue of origin), if cancer is detected;• repeating above steps by each expert model to obtain a set of predictions, wherein each expert model specializes by organ, cancer-type and biomarker panel used for training thereof;• evaluating the set of predictions by the gating network (stacking or meta model) to compute a weighted sum thereof for arriving at a final confidence score indicative of probability of cancer, and sub-type / stage thereof if present;• computing a final prediction score based on the final confidence score and the clinical data, if cancer is detected at previous step; and• generating a clinical report summarizing findings with a time stamp when the biological sample was tested / screened,BM131025 wherein, cancer is selected from a group comprising of Breast cancer, Biliary tract cancer, Cervical cancer, Colon cancer, Esophageal cancer, Gastric cancer, Liver cancer, Lung cancer, Oral cancer, Ovary cancer, Pancreatic cancer, Prostate cancer and Vulva cancer, wherein further, the computational machine learning model is trained on dataset for each cancer-type, biomarkers and thresholds ranges thereof differentiating cancer from non-cancer type, and wherein further, said system has high sensitivity and specificity in screening of cancer in the patients.

Citation Information

Patent Citations

  • Methods of assessing breast cancer using machine learning systems

    WO2022029492A1

  • Breast cancer diagnostic and treatment

    WO2023170659A1