A multi-cancer early screening method based on fragmentomics and microbiomics features and application thereof
By combining low-depth whole-genome sequencing and machine learning models with plasma cfDNA and microbiome characteristics, the limited coverage and high cost of existing multi-cancer screening technologies have been solved, achieving high sensitivity and high specificity for early detection of multiple cancers, making it suitable for cost-effective screening in resource-limited areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XUTENG GENE TECHNOLOGY CO LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing cancer screening methods mainly target single cancer types, which have limited coverage, high invasiveness, high cost, and are difficult to popularize. In particular, there is a lack of effective early screening methods for multiple cancer types, which means that patients need to undergo multiple examinations and that these methods are difficult to implement in areas with limited resources.
By combining fragment omics features of plasma cfDNA with microbiome features, and employing low-depth whole-genome sequencing and machine learning models, a single test can screen for multiple cancers, including lung cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer, pancreatic cancer, breast cancer, and cervical cancer. A two-stage machine learning model is constructed using fragment omics features such as FSD, FSC, CNV, EDM, and BPM, and microbiome features such as microbial species abundance and community diversity, for cancer detection and classification.
It achieves high sensitivity and high specificity for early screening of multiple cancer types, reduces the false positive rate, is suitable for large-scale population screening, and provides a cost-effective and efficient non-invasive early cancer detection method, applicable to cancer prevention and control in asymptomatic populations.
Smart Images

Figure CN121393535B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical detection technology, specifically a method for early screening of multiple cancer types based on fragment omics and microbiome characteristics and its application. Background Technology
[0002] Current cancer screening methods primarily target single cancer types. For example, low-dose spiral CT (LDCT) is used for lung cancer screening, colonoscopy for colorectal cancer screening, mammography for breast cancer screening, and cervical cytology for cervical cancer screening. While these methods can improve the early diagnosis rate of specific cancer types, they have significant limitations: First, their coverage is limited, with each test only detecting one type of cancer, while patients may need to undergo multiple tests for a comprehensive evaluation. Second, many tests are invasive (such as colonoscopy) or involve radiation exposure (such as CT), leading to decreased patient compliance. Third, these screening methods are typically costly and require specialized medical equipment and personnel, making them difficult to implement in resource-constrained areas. Most importantly, effective early screening methods are still lacking for many high-mortality cancers (such as pancreatic cancer and esophageal cancer).
[0003] This situation highlights the urgent need for early detection technologies for multiple cancer types. An ideal solution should possess the following characteristics: the ability to screen for multiple cancers simultaneously with a single non-invasive test (such as a blood sample); sufficient sensitivity and specificity to reduce false positives and false negatives; reasonable cost-effectiveness for large-scale population screening; and applicability to early-stage, asymptomatic patients. Therefore, innovative methods are needed to address the shortcomings of existing single-cancer screening, becoming an important supplement to the clinical cancer screening system, ultimately improving early cancer diagnosis rates and patient prognosis.
[0004] Cell-free DNA (cfDNA) in liquid biopsies has been considered a promising non-invasive tumor biomarker. cfDNA released by cancer cells contains tumor-specific information, such as gene mutations, methylation alterations, and fragmentation characteristics. However, previous detection methods based on single genetic variant markers (such as a specific gene mutation or methylation) have limited sensitivity and are insufficient for early cancer screening. In recent years, researchers have discovered that fragmentomics features of cfDNA contain rich tumor signals, including copy number variation (CNV), fragment size distribution (FSD), fragment size coverage (FSC), end motif (EDM), and breakpoint motif (BPM). These fragmentomics features are closely related to nucleosome distribution, nuclease activity, and genomic instability in tumor tissue. Combining multiple fragment features with machine learning models can significantly improve the sensitivity and specificity of tumor detection.
[0005] In addition to human cfDNA signals, circulating microbial DNA (cmDNA) in plasma has recently been found to have cancer-indicating properties. Cancer patients may experience increased mucosal barrier permeability or immune alterations associated with tumors, allowing certain bacterial or viral DNA fragments to enter the bloodstream. Studies have shown that the plasma microbial composition varies among different cancer patients. Incorporating cmDNA signatures can further improve the accuracy of multi-cancer detection and provide additional clues about the location of cancer.
[0006] Therefore, combining the cfDNA fragment omics characteristics obtained from low-depth WGS sequencing with microbial metagenomic characteristics (cmDNA) holds promise for developing a highly sensitive and specific early screening method for multiple cancer types, serving as an auxiliary tool in clinical cancer screening systems. This method can screen for multiple common cancers and indicate possible tissue origins with a single blood sample, providing a cost-effective solution for cancer prevention and control in asymptomatic populations. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a multi-cancer early screening method based on fragmentomics and plasma microbiome profiling, to address the technical shortcomings of existing single-cancer screening methods, which are time-consuming, expensive, and have limited coverage. This invention innovatively integrates fragmentomics characteristics of plasma cfDNA with plasma circulating microbiome information, aiming to achieve non-invasive early detection and tissue tracing of major solid tumors in a single experimental reaction. This method improves detection sensitivity while maintaining high specificity, making it suitable for large-scale screening of multiple cancers.
[0008] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0009] This invention provides a multi-cancer early screening method based on fragment omics and plasma microbiome profiling, comprising the following steps:
[0010] Obtain a plasma sample to be tested, extract cfDNA from the plasma sample, and perform low-depth whole-genome sequencing on the cfDNA to obtain sequencing data;
[0011] Based on the sequencing data, fragment omics features from human cfDNA were extracted; based on sequences in the sequencing data that did not align to the human genome, microbiome features from circulating microbial DNA in plasma were extracted.
[0012] The fragment omics features and the microbiome features are fused to form a fused feature vector; the fused feature vector is input into a trained first-stage classification model to obtain a judgment result on whether the individual from which the plasma sample to be tested is cancer-positive; if the judgment result is positive, the fused feature vector or features re-extracted based on the sequencing data are input into a trained second-stage classification model to obtain the cancer type of the individual.
[0013] Preferably, the fragment omics features include at least one of fragment size distribution, fragment size coverage, copy number variation, fragment terminal sequence motif, and breakpoint sequence motif.
[0014] Preferably, the microbiome characteristics include the relative abundance of microbial species and / or the microbial community diversity index.
[0015] Preferably, the classification model in the first stage is an ensemble learning model.
[0016] More preferably, the ensemble learning model is configured to: directly adopt the judgment result when the confidence level of the output of a single base model is higher than 0.95 or lower than 0.05; and enable the stacked ensemble model for comprehensive judgment when the confidence level is between 0.05 and 0.95.
[0017] Preferably, the classification model in the second stage is a deep neural network model.
[0018] More preferably, the probability output by the second-stage classification model is calibrated by temperature scaling, wherein the temperature parameter τ ranges from 0.5 to 2.0.
[0019] Preferably, the decision threshold of the first-stage classification model is dynamically set based on Bayesian decision rules.
[0020] Preferably, the cancer types include at least two of the following: lung cancer, liver cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and cervical cancer.
[0021] The present invention also provides an application of the aforementioned multi-cancer early screening method in the preparation of a kit for early screening of multiple cancers.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] This invention provides a multi-cancer early screening method based on fragmentomics and microbiome features. It integrates five classes of cfDNA fragmentomics features (FSD, FSC, CNV, EDM, and BPM). These features originate from large-scale instability of the tumor genome and alterations in nucleosome localization, exhibiting signal strengths far exceeding those of point mutations, thus significantly amplifying the detection signal of early-stage tumors. Simultaneously, this invention utilizes unique microbial community differences in the plasma of patients with different cancer types as molecular tags, providing independent and complementary discriminative criteria for cancer classification. This invention is not a simple superposition of fragmentomics and microbiome features, but rather achieves deep integration and complementary advantages of these two types of information through a two-stage model design. Clinical validation shows that, through multimodal feature fusion and advanced machine learning algorithms, the model achieved a median specificity of 96.6% in a cohort of thousands of cases, effectively reducing the false positive rate. Attached Figure Description
[0024] Figure 1 A schematic diagram of the fragment omics and microbiome-based early tumor screening process based on low-depth WGS;
[0025] Figure 2 A schematic diagram illustrating typical trends in plasma cfDNA fragment length distribution between healthy controls (blue) and cancer patients (orange);
[0026] Figure 3 This is a schematic diagram of the fragment size coverage offset;
[0027] Figure 4 A graph showing genome-wide copy number variations observed in a plasma cfDNA sample from a cancer patient;
[0028] Figure 5 A schematic diagram comparing the frequencies of cfDNA terminal motifs;
[0029] Figure 6 This is a schematic diagram illustrating the results of a differential analysis of plasma microbial α-diversity between cancer patients and healthy controls in a research cohort.
[0030] Figure 7 This is a schematic diagram illustrating the results of a differential analysis of plasma microbial β-diversity between cancer patients and healthy controls in a research cohort.
[0031] Figure 8 This is a schematic diagram of a two-stage multi-cancer early screening model based on the fusion of fragment omics and microbiome features. Detailed Implementation
[0032] This invention provides a method for early screening of multiple cancer types based on fragment omics and microbiome features, the method preferably comprising the following steps:
[0033] Obtain a plasma sample to be tested, extract cfDNA from the plasma sample, and perform low-depth whole-genome sequencing on the cfDNA to obtain sequencing data;
[0034] Based on the sequencing data, fragment omics features from human cfDNA were extracted; based on sequences in the sequencing data that did not align to the human genome, microbiome features from circulating microbial DNA in plasma were extracted.
[0035] The fragment omics features and the microbiome features are fused to form a fused feature vector; the fused feature vector is input into a trained first-stage classification model to obtain a judgment result on whether the individual from which the plasma sample to be tested is cancer-positive; if the judgment result is positive, the fused feature vector or features re-extracted based on the sequencing data are input into a trained second-stage classification model to obtain the cancer type of the individual.
[0036] In this invention, those skilled in the art should understand that the fusion refers to combining features from different sources into a unified mathematical representation for machine learning models to learn and discriminate. In this invention, the fusion is preferably achieved through vector concatenation, as direct concatenation is efficient, stable, and does not require the introduction of additional complex parameters. Those skilled in the art will understand that the "fusion" is not limited to direct concatenation, but also includes, but is not limited to, weighted concatenation, such as assigning different weights to features of different categories based on their discriminative importance before concatenation; and model-based fusion, such as inputting two types of features into different sub-networks for nonlinear transformation, and then concatenating or combining the output higher-order representations.
[0037] In this invention, the fragment omics features preferably include at least one of fragment size distribution, fragment size coverage, copy number variation, fragment terminal sequence motif, and breakpoint sequence motif.
[0038] In this invention, the fragment size distribution refers to the statistical distribution of plasma cfDNA fragment lengths. This characteristic reflects changes in nucleosome localization and differences in nuclease activity within tumor tissue. The preferred method for extracting and calculating the fragment size is to statistically analyze the cfDNA fragments aligned to the human reference genome based on their inserted fragment lengths. As a preferred embodiment, the fragment length range is divided into several consecutive intervals (called "bins"), for example, one bin every 10 bp, totaling approximately 27 intervals. The percentage of fragments falling into each length interval is calculated relative to the total number of fragments. Further derived characteristics can be calculated from the above distribution, including but not limited to: short fragment ratio, long fragment ratio, short-to-long fragment ratio, and peak height. The short fragment ratio refers to the percentage of fragments <150 bp in length. The long fragment ratio refers to the percentage of fragments >220 bp in length. The short-to-long fragment ratio is the quotient of the short fragment ratio and the long fragment ratio. The peak height refers to the distribution peak around approximately 167 bp (core nucleosome protection length).
[0039] The fragment size coverage refers to the bias in coverage depth of fragments of different lengths within a specific region of the genome. This characteristic is closely related to the open state of chromatin. The preferred method for extracting and calculating the fragment size coverage is to divide the human genome into continuous, non-overlapping windows, preferably 50kb or 1Mb windows. Within each window, the number of sequencing reads for short fragments (e.g., 40-150bp) and long fragments (e.g., 151-300bp) is counted. The proportion of short fragments (number of short fragments / total number of fragments) or the coverage ratio of short to long fragments is calculated for each window. Finally, a vector of short fragment proportions across all windows in the whole genome is obtained for each sample. A Z-score or ratio can be calculated by comparing with a healthy control baseline to highlight tumor-specific shifts.
[0040] The copy number variation refers to the increase (amplification) or decrease (deletion) of the copy number of a large DNA fragment (usually >1Mb) in the genome represented by cfDNA relative to the normal diploid genome. The preferred method for extracting and calculating the copy number variation is as follows: the genome is divided into fixed-size windows, preferably 1Mb windows. The number of uniquely aligned sequencing reads within each window is calculated as the original coverage depth of that window. Since sequencing depth is significantly affected by GC content, methods such as LOESS local regression are used to correct the original coverage depth based on the GC content of each window. The corrected coverage depth is compared with a normal reference baseline (usually the median) constructed from samples from multiple healthy individuals, and the log2 ratio of each window is calculated. Further, algorithms such as cyclic binary segmentation can be used to smooth and segment the log2 ratio to identify significant amplification (log2 ratio > 0.2) and deletion (log2 ratio < -0.2) regions.
[0041] The terminology of the fragment terminal sequence motif refers to the frequency of occurrence of short sequence patterns at the ends (usually the 5' end) of cfDNA fragments. This characteristic reflects the sequence preference of nucleases during DNA cleavage. The preferred method for extracting and calculating the fragment terminal sequence is as follows: for each aligned cfDNA fragment, extract a number of bases from its 5' end as the terminal sequence, preferably 4-5 bp; count the number of occurrences of all possible short sequences at the ends of all cfDNA fragments; calculate the frequency of each motif: the number of occurrences of the specific motif / the total number of all terminal motifs; further, calculate the logarithmic ratio of this frequency to the frequency in healthy controls to correct for sequencing bias and amplify differences.
[0042] The breakpoint sequence motif refers to the nucleotide sequence preference upstream and downstream of the cfDNA breakpoint (i.e., the start and end positions of the fragment on the genome). The preferred method for extracting and calculating the breakpoint sequence motif is as follows: determine the precise start and end coordinates of each cfDNA fragment on the reference genome; for each breakpoint, extract several bases upstream (preferably 4 bp) and several bases downstream (preferably 4 bp) to form a longer sequence unit, such as an 8-mer; count the frequency of each 8-mer sequence at all breakpoints; similarly, breakpoint motifs that are significantly enriched or missing in tumor samples can be identified by comparing with a healthy control baseline.
[0043] In this invention, the microbiome features include the relative abundance of microbial species and / or the microbial community diversity index. In this invention, the microbiome features aim to characterize the community composition of circulating microbial DNA (cmDNA) in plasma. The relative abundance of microbial species refers to the proportion of a sequence of a specific microbial species or taxonomic level (genus, phylum, etc.) among all detected microbial sequences. This feature reflects the relative abundance of different members in the plasma microbial community. The preferred method for extracting and calculating the relative abundance of microbial species includes host sequence removal and quality control, species annotation, and abundance calculation and normalization. Specific steps are as follows: All sequencing reads are aligned with the human reference genome (e.g., GRCh38), and all unaligned reads are extracted. To strictly control false positives, these reads were further compared with environmental pollutant databases (such as lists of common microorganisms from sequencing reagents and the laboratory environment), and matching sequences were removed to obtain high-quality candidate microbial source reads. The filtered reads were then compared using efficient and accurate taxonomic annotation software (such as Kraken2 and MetaPhlAn3) with customized microbial genome databases (such as databases integrating bacterial, archaea, viral, and fungal genomes) to assign each read its most probable species classification label. The number of reads assigned to each species, or according to analytical needs, at the genus, phylum, etc., was counted. To avoid bias caused by differences in sequencing depth, the read count for each species was divided by the total number of reads for all microbial species to obtain the relative abundance of that species (usually expressed as a percentage). In this invention, the relative abundance of the specific microorganism can serve as a powerful biomarker for cancer differentiation.
[0044] In this invention, the first-stage classification model is preferably an ensemble learning model. This invention first trains multiple base classifiers of different types. These base classifiers include, but are not limited to: gradient boosting decision trees (such as XGBoost), support vector machines (SVM), random forests (RF), and multilayer perceptrons (MLP). Each base model uses the same training set to learn discriminative patterns from fragment omics and microbiome fusion features. A higher-order meta-model is constructed using a stacking method, and the out-of-bag prediction probabilities generated by each base model on the training set through k-fold cross-validation are used as meta-features to train a meta-learner. This invention preferably uses a generalized linear model (such as logistic regression) as the meta-learner, which learns how to optimally weight and combine the outputs of each base model to make the final decision.
[0045] In this invention, the ensemble learning model is preferably configured such that when the confidence level of the output of a single base model is higher than 0.95 or lower than 0.05, the judgment result is directly adopted; when the confidence level is between 0.05 and 0.95, the stacked ensemble model is enabled for comprehensive judgment.
[0046] In this invention, the second-stage classification model is a deep neural network model. In this second-stage classification model, for samples identified as positive in the first stage, cancer classification is primarily based on their microbiome characteristics, where the discriminative weight of microbiome characteristics is greater than that of fragment omics characteristics. In the second-stage cancer classification model, this invention preferably uses a deep neural network (DNN) as the core classifier and introduces temperature scaling technology to calibrate its output prediction probabilities, thereby improving its reliability and interpretability in real-world screening scenarios.
[0047] In this invention, the DNN is a feedforward neural network with a specially designed structure to fully exploit the complex, non-linear mapping relationship between microbiome features and cancer types. In a preferred embodiment, the DNN comprises an input layer, multiple hidden layers, and an output layer. For example, a typical structure could be: input layer → 256 neurons → 128 neurons → 64 neurons → output layer. The dimension of the input layer matches the dimension of the fused feature vector (e.g., 85 dimensions), and the number of neurons in the output layer is equal to the number of cancer types to be classified (e.g., 8). Hidden layers all employ the SiLU (or ReLU) activation function to introduce non-linearity. To enhance training stability, layer normalization can optionally be added after the hidden layers. To prevent overfitting, Dropout regularization is introduced after each hidden layer, with the dropout rate preferably set to 0.30 to 0.50. The training of the model preferably includes the following steps: using a labeled, smoothed multi-class cross-entropy loss function; employing the AdamW optimizer, which can adaptively adjust the learning rate and improve generalization ability through improved weight decay processing. Typical hyperparameters are: β1=0.9, β2=0.999, and weight decay rate=1e-4. A warm-up and cosine annealing strategy is employed; for example, the initial learning rate is 3e-4, and after 5 epochs of warm-up, it decays to 3e-6 using a cosine function. An early stopping mechanism (e.g., patience=15) is used to terminate training when performance on the validation set no longer improves, thus avoiding overfitting. Simultaneously, class weights are set for the loss function based on the frequency of each class in the training set to alleviate class imbalance.
[0048] In this invention, the probabilities output by the second-stage classification model are calibrated using temperature scaling, wherein the temperature parameter τ preferably ranges from 0.5 to 2.0. In this invention, uncalibrated DNN output probabilities often fail to accurately reflect their true confidence levels. Temperature scaling is a lightweight and effective post-processing technique used to calibrate the model's original output probabilities to a probability that better conforms to the true distribution, thereby improving its posterior interpretability.
[0049] In this invention, the second-stage classification model can also employ a more complex architecture, such as ensemble learning or graph neural networks, similar to the first-stage classification model. For example, sub-models targeting fragment omics features and microbial spectra can be trained, and their higher-order representations can be fused for cancer type prediction; or graph convolutional neural networks (GCNs) can be used to infer based on cancer-feature association maps. However, in this invention, the use of DNNs has achieved superior results, balancing model complexity and performance.
[0050] In this invention, the decision threshold of the first-stage classification model is dynamically set based on Bayesian decision rules. These Bayesian decision rules comprehensively consider the total prior probability π of the target cancer type and the relative cost ratio C between false positives and false negatives. FP / C FN In this invention, by adjusting π and the cost ratio, the same core model can be easily adapted to different screening scenarios. For example, a higher cutoff can be set for annual routine physical examinations (low π, high specificity requirements); a lower cutoff can be set for follow-up of high-risk groups (high π, high sensitivity requirements).
[0051] In this invention, the cancer species preferably include at least two of the following: lung cancer, liver cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and cervical cancer.
[0052] The present invention also provides an application of the aforementioned multi-cancer early screening method in the preparation of a kit for early screening of multiple cancers.
[0053] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0054] Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods. Unless otherwise specified, the experimental materials used in the following embodiments are commercially available products.
[0055] Example 1
[0056] This invention utilizes a single low-depth whole-genome sequencing to extract two types of information in parallel: fragment omics features of human cfDNA and microbiome features. These two types of features are then fused to construct a two-stage machine learning model, ultimately achieving cancer detection (determining whether a tumor is positive or negative) and tissue tracing (identifying the specific cancer type). The overall technical process is as follows: Figure 1 As shown.
[0057] (1) Sample collection and sequencing
[0058] This invention is mainly based on the Low Pass WGS technique. The wet laboratory procedure includes the following steps: ① Sample collection: 10 mL of peripheral blood should be collected using a Streck tube. The sample should be transported at 6℃-26℃ and plasma separation completed within 72 hours. Severely hemolyzed samples are unusable. ② Plasma separation uses a two-step centrifugation method: Step 1: Centrifuge at 4℃, 1600×g for 15 minutes to remove whole blood cells and collect the supernatant; Step 2: Centrifuge at 4℃, 16000×g for 15 minutes to remove platelets and cell debris. ③ Free nucleic acid extraction uses the QIAamp Circulating Nucleic Acid Kit (Kaigen, 55114). Vector RNA is added to improve the recovery rate of small fragments. The concentration of cfDNA is determined using Qubit. If the total amount is >300 ng, Agilent 2100 quality control is required. The experiment is terminated if small fragment cfDNA is lacking. ④ Library construction was performed using the xGen Prism DNA Library Prep Kit, with 50–200 ng of cfDNA added. The steps included end repair, double magnetic bead purification, ligation reaction (20℃ 15 min → 65℃ 15 min → 4℃ storage; 65℃ 30 min → 4℃ storage), library amplification (10⁻¹² cycles), and introduction of 8 bp paired-end UMI adapters. ⑤ Sequencing was performed on the MGISEQ-2000 (BGI Genomics) sequencing platform, using 150 bp reads and paired-end sequencing. Sterile water was added to each batch for contamination monitoring.
[0059] Plasma samples were subjected to cfDNA extraction and low-depth whole-genome sequencing. The sequencing platform converted the obtained optical signals into sequencing data in BCL format and split the sequencing data. Based on the sample index, the sequencing data of individual samples were separated and converted into FASTQ format for bioinformatics analysis.
[0060] (2) Fragment omics feature extraction
[0061] After sequencing data is aligned to the human genome GRCh38.p13, various fragment omics features can be extracted.
[0062] Fragment size distribution (FSD): such as Figure 2 As shown, this is a schematic diagram illustrating the typical trends in plasma cfDNA fragment length distribution between healthy controls (blue curve) and cancer patients (orange curve). It can be observed that cfDNA from cancer patients is relatively enriched in the short fragment range of 100-150 bp, while the fragment distribution in healthy individuals shows a significant nucleosome protective peak at approximately 167 bp, with a higher proportion of long fragments. This invention quantifies this continuous distribution into a usable feature for the model by dividing bin intervals and calculating the short / long fragment ratio.
[0063] Fragment Size Coverage (FSC): Figure 3 A schematic diagram illustrating FSC characteristics is shown. By statistically analyzing the coverage ratio of short fragments (40-150bp) and long fragments (151-300bp) across the entire genome, it can be found that tumor samples (red curve) exhibit significant enrichment of short fragments in specific genomic regions (such as open chromatin regions). This coverage shift is an important signal of tumor specificity.
[0064] Copy number variation (CNV): By dividing the genome into 1Mb windows, calculating the cfDNA alignment depth of each window, and correcting for GC content and sequencing batch, the copy number ratio distribution is compared with a normal reference control to extract copy number variation characteristics across the entire genome. Figure 4 The CNV diagram of plasma cfDNA from a cancer patient shows significant amplification or deletion in multiple chromosomal regions. Regions with significant copy number amplification (>3.5) include chr7 (potentially containing key driver genes such as EGFR or MET), chr8 (a common MYC amplification region), chr12 (common CCND2 or MDM2 amplification), and chromosomes chr10, chr14, chr3, chr16, chr17, and chrX. Conversely, regions with significant copy number deletion (<1.0) are widely distributed across chr1, chr2, chr3, chr6, chr7, chr9, chr10, chr11, chr14, chr15, chr17, chr18, chr19, chr22, and chrX, with the deletion at chr22 being the most significant, potentially affecting tumor suppressor genes such as SMARCB1 and EP300. These CNV features suggest that the patient's genome has extensive instability, and the regions of amplification and deletion cover multiple known cancer driver genes and tumor suppressor genes, which have potential clinical significance.
[0065] End-of-segment feature (EDM): Figure 5 This is a schematic diagram comparing the frequencies of cfDNA terminal motifs. The 5' end of a cfDNA fragment is extracted (5 bp), and the frequency percentage of each short sequence (motif) is calculated to form a "terminal motif" spectrum. The frequency of common terminal motifs (such as "CTCAG") in the cfDNA of cancer patients is significantly higher than that in healthy individuals. This difference can reflect the action pattern of tissue-specific nucleases and provide fragmentomics markers for tumor detection.
[0066] Breakpoint sequence characteristics (BPM): Analysis of upstream and downstream sequences (4 bp each, totaling 8-mer) at cfDNA breakpoints in the genome, and statistical analysis of their frequency. The preference of breakpoint motifs reflects the sequence fragility and enzymatic digestion characteristics of DNA cleavage; differences in BPM profiles between tumor and healthy individuals can enhance discriminative ability.
[0067] BPM and EDM together characterize the enzyme digestion pattern of cfDNA fragmentation and can serve as a tumor tissue-specific fragment signal.
[0068] (3) Extraction of microbiome features
[0069] We analyzed the characteristics of the microbial community in plasma using reads from WGS data that did not align to the human genome. Specific steps included:
[0070] Host sequence removal: All sequencing reads are first aligned to the human reference genome. Unmatched reads are extracted as candidate microbial source sequences. To strictly control false positives, a dual filtering strategy can be used, simultaneously aligning to both known human references and microbial reference lists present in the reagent / environment to exclude potentially contaminated sequences.
[0071] Species annotation and quantification: Filtered reads are aligned to microbial genome databases using efficient classification algorithms (such as Kraken2, MetaPhlAn, etc.) to identify the species originating from the sequences (bacteria, viruses, fungi, etc.). The sequence abundance of each microorganism is calculated (e.g., number of reads or proportion), yielding a relative abundance spectrum for each sample. To reduce sequencing bias, abundance can be expressed as a percentage of the total number of reads normalized to a normalized total.
[0072] α-diversity analysis: Figure 6 This diagram illustrates the analysis results of plasma microbial alpha diversity (represented by the Shannon index) in a research cohort of cancer patients and healthy controls. The results show that the microbial community diversity in cancer patients was generally lower than that in healthy controls, a trend that can serve as a valid discriminative feature.
[0073] β-diversity analysis: Figure 7 Principal coordinate analysis based on Bray-Curtis distance revealed differences in plasma microbial community structure among patients with different cancer types (e.g., colorectal cancer versus non-colorectal cancer), exhibiting a certain clustering trend in two-dimensional space. This provides important clues for subsequent cancer classification.
[0074] Microbial feature selection: Microbial features that are discriminative across different categories are selected from species abundance and diversity data. For example, a significant increase or decrease in the abundance of a specific genus can serve as a biomarker. Simultaneously, abundance patterns of multiple species, overall diversity indices, etc., are combined as input models for the microbial feature set.
[0075] (4) Model building and training
[0076] This invention involves obtaining cfDNA through a single blood test and sequencing it. In bioinformatics analysis, human fragment omics features and plasma microbiome features are extracted in parallel. The selected cfDNA fragment omics features and microbiome features are fused to construct a machine learning model for sample classification and prediction. First, a first-stage cancer detection model is used to determine if the sample is cancer-positive. If the result is positive, a second-stage cancer classification model is used to predict the specific cancer type, achieving simultaneous screening and tissue tracing for multiple cancers. A schematic diagram of the two-stage multi-cancer early screening model based on the fusion of fragment omics and microbiome features is shown below. Figure 8 .
[0077] First-stage classification model: Cancer detection model
[0078] Feature Selection: The first-stage classification model input includes five major categories of cfDNA fragmentomics features and alpha diversity indices from the plasma microbiome. Fragmentomics features include: fragment size distribution (FSD), fragment size coverage preference (FSC), copy number variation (CNV), fragment end sequence motif preference (EDM), and breakpoint sequence preference (BPM). These features reflect nucleosome localization, nuclease cleavage patterns, and genomic instability specific to tumor tissue. In addition to human cfDNA features, the model also incorporates the alpha diversity index of the plasma microbiome, as clinical studies have found that plasma microbial diversity is often reduced in cancer patients. Integrating these multimodal features as model input aims to maximize the differentiation signals between early-stage cancer patients and healthy individuals.
[0079] Model Algorithm: Due to the high dimensionality and complex patterns of the fused features, the first-stage classification model employs an ensemble learning strategy to improve its robustness and accuracy. Firstly, XGBoost (binary classification) is used as the base model for the first stage to handle high-dimensional features and capture non-linear relationships. Key hyperparameters are set as follows: learning_rate=0.05, max_depth=4, n_estimators=1200, subsample=0.80, colsample_bytree=0.70, min_child_weight=5, gamma=1.0, reg_lambda=3.0, reg_alpha=0.20, max_delta_step=1, tree_method="hist", random_state=2025. Training is completed within a 5-fold hierarchical cross-validation and early stopping (100 epochs) framework, and the validation set probabilities are calibrated using temperature scaling. The recommended parameter ranges for this embodiment are as follows: learning_rate 0.03-0.07; max_depth 3-6; subsample 0.7-0.9; colsample_bytree 0.6-0.9; min_child_weight 3-7; gamma 0.3-2.0; reg_lambda 1.0–5.0; reg_alpha 0-0.5. Multiple base models with different algorithms are trained simultaneously, such as support vector machines, random forests, and multilayer perceptrons. These models learn discrimination patterns for fragment omics or microbial features, respectively. The probability values output by each base model are used as input to a meta-learner, such as a generalized linear model (GLM), to train a stacked ensemble model for final discrimination. This two-layer structure effectively integrates the advantages of each model, improving the sensitivity and specificity of detection.
[0080] Mathematical Expression: Without loss of generality, the cancer detection task can be viewed as predicting a binary label (0 for healthy, 1 for cancer) given a feature vector. The output of a single model (XGBoost) can be represented as the following function:
[0081]
[0082] Where β0 is the bias term, β i Let be the coefficient of the i-th feature. Here, β0 is automatically obtained by the logistic regression model through maximum likelihood estimation within the training-validation framework; L2 regularization and early stopping are employed, and the features are standardized using Z-score.
[0083] The above five types of features, after being engineered, are mapped into d-dimensional vectors in this invention; in this embodiment, d=84. The common representation of EDM and BPM is obtained by constructing a product / bilinear term using selected basis pairs, illustrating the cooperative preference between endpoints and breakpoints. For β... i The GLM (Logistic Regression) is trained using the aforementioned 84-dimensional standardized features, and automatically learned from the training data through maximum likelihood estimation. The fused feature vector, after engineering, has a dimension of d=84, therefore β... i The index range is i=1…84. In the XGBoost gradient boosting tree model, the linear combination β0+i=1∑dβixi in the above formula is replaced by an additive model of a set of regression trees. The model gradually adds trees to improve prediction accuracy by optimizing the objective function (binary log loss). In most cancer detection scenarios, a single XGBoost model can already achieve high sensitivity and specificity.
[0084] To further enhance robustness, this invention introduces stacked ensemble. For the stacked ensemble model, based on XGBoost, the output probabilities of other algorithms such as Random Forest (RF), SVM, and Multilayer Perceptron (MLP) are added as "secondary features" and input into the meta-learner (GLM). Based on these outputs, a linear combination is constructed to obtain the final discriminant probability. The calculation method is as follows:
[0085]
[0086] Here, α0 and αi represent the biases of the meta-model and the weights learned by the meta-learner. Both α0 and αi come from the stacked meta-learner (GLM), whose input is the OOF probability (or its logit) of each base model (XGBoost, Random Forest, SVM, MLP, etc.), obtained through maximum likelihood estimation and cross-validation, and then subjected to L2 regularization. Through this ensemble strategy, the model integrates the judgments of multiple base classifiers, improving its ability to detect cancer signals.
[0087] Training and Validation: The training set was used to learn the parameters of the first-stage classification model. During training, cross-validation was used to adjust the hyperparameters of XGBoost and the ensemble model, and early stopping and regularization techniques were applied to prevent overfitting. In cases of class imbalance, strategies such as adjusting the loss function weights or downsampling healthy controls were employed to ensure sufficient sensitivity to positive samples. Finally, the model's detection performance, including sensitivity and specificity, was evaluated on an independent validation set. High sensitivity ensures the capture of as many cancer-positive individuals as possible, while high specificity avoids unnecessary panic and subsequent investigations caused by excessive false positives. This invention's model strives for an optimal balance among these metrics through multimodal feature fusion and complex algorithm design. In practice, a higher specificity is preferred.
[0088] If the confidence level of a single XGBoost output is extremely high or low (>0.95 or <0.05), the result of the single model is directly adopted (greater than 0.95 is judged as cancer, less than 0.05 is judged as normal) to simplify the judgment. If the output of a single XGBoost model is in the gray area (between 0.05 and 0.95), a stacked ensemble model is enabled to calibrate the results based on the consensus of multiple base models, reducing misjudgments.
[0089] This invention sets a data-driven decision threshold for the posterior probability output by the stacked ensemble model, and employs a Bayesian threshold based on the missed diagnosis / false alarm cost ratio and the prevalence of the disease.
[0090]
[0091] Cutoff = 1 / (1+C) FP / C FN *((1-π) / π))
[0092] Where π refers to the total incidence of all target cancer types; C FP The cost of false positives (false positives cause psychological burden on patients and incur economic costs such as retesting); C FN The cost of missed diagnoses (i.e. false negatives, which are usually much greater than CFP) is.
[0093] In this invention, C FP / C FN The value is set to 1 / 50, and the final cutoff is 0.61. That is, in the stacked ensemble model, a probability greater than 0.61 is considered cancer, and a probability less than or equal to 0.61 is considered normal.
[0094] This approach combines the efficiency of a single model with the robustness of an ensemble model. In simulation experiments, this "hierarchical integration" method improved the overall AUC by approximately 0.02-0.03 compared to XGBoost alone, while reducing the false positive rate by about 30%, which is particularly important for large-scale population screening.
[0095] Second-stage classification model: Cancer type classification model
[0096] For samples that are identified as positive by the first-stage classification model, the second-stage classification model is used to predict their specific cancer type.
[0097] Feature Selection: Considering that different tumor tissue origins may be accompanied by varying microbial community characteristics, this stage primarily employs highly discriminative microbiome features, including microbial abundance profiles and β-diversity indices specific to each class of positive samples. For example, the relative abundance of certain bacterial genera in the plasma of patients with different cancer types shows significant differences, which can serve as clues for tissue origin tracing; furthermore, microbial β-diversity analysis among tumor patients shows that gut-associated tumors cluster separately from other cancer types in terms of community structure. In addition, the second-stage classification model can also incorporate cancer-specific indicators from some fragment omics features to further improve classification accuracy. However, overall, microbial features carry greater weight in cancer type discrimination, consistent with their unique role in reflecting differences in the tumor host microenvironment.
[0098] Model Algorithm: The second-stage classification model addresses a multi-class classification problem. This invention employs a deep neural network (DNN) model as the classifier to fully explore the complex relationship between microbial features and cancer types. This DNN consists of multiple layers of fully connected neurons, using a non-linear activation function to model the high-dimensional interactions of features. The model output layer uses a softmax function to generate the probability distribution for each cancer type. Specifically, let C be the total number of categories. Given a sample x (x represents the feature vector of a single sample formed by a single test of a subject, which integrates fragment omics and microbial features), the final discrimination probability of category c∈{1,…,C} (where the value of C is 8, meaning C can be 1, 2,…, 8, representing 8 detectable cancer types, where 1 is lung cancer, 2 is breast cancer, 3 is colorectal cancer, 4 is stomach cancer, 5 is liver cancer, 6 is pancreatic cancer, 7 is esophageal cancer, and 8 is cervical cancer) is defined as:
[0099]
[0100] Where x is the input feature vector (normalized / standardized), which is obtained by splicing cfDNA fragment omics features (FSD, FSC, CNV, EDM, BPM, etc.) with plasma microbial features (α diversity, relative abundance of key species, etc.) or by nonlinear mapping.
[0101] S c (m) (X) is the standardized score of the m-th sub-model for category c, which is linearly additive.
[0102] M is the number of sub-models participating in the fusion. It can be 1 (single model) or greater than 1 (stacked / integrated), and this parameter makes the formula adaptive to "whether the first-stage integrated model is used".
[0103] β mc ,β 0c To integrate weights and class biases, β mcIt can measure the contribution strength of the m-th sub-model to category c; β 0c These are the baseline terms for category c. These parameters can all be obtained by minimizing the multi-class cross-entropy within the validation / training-validation framework.
[0104] π c These are the prior probabilities of various cancer types, obtained by reviewing incidence data of various cancer types in the Chinese population.
[0105] δ c For class correction bias (posterior calibration), since the cancer probability of the actual clinical screening population is different from that of the training set, setting this item can be achieved by a simple addition correction to realign the model output with the true distribution, thus realizing rapid recalibration of the posterior probability.
[0106] τ is a temperature parameter that controls the "sharpness" of the distribution. A value of 1 maintains the confidence level of the original model output. In screening scenarios with high uncertainty, a value greater than 1 makes the probability values of multiple cancer types closer, avoiding false positives. In high-precision re-examination scenarios, a value less than 1 makes the highest probability category stand out more. In this invention, 8 cancer types are screened, and the set value of τ is 1.5.
[0107] The model prediction selects the cancer type label that maximizes P(c|x) as the output, and the corresponding c value is the final reported cancer type. During training, a positive sample set containing known cancer type labels is used. The network parameters are learned by minimizing the multi-class cross-entropy loss, making the model output close to the true classification. To avoid overfitting, Dropout regularization and early stopping strategies are used during training, and the network structure and hyperparameters (such as the number of layers, the number of neurons per layer, and the learning rate) are adjusted using independent validation sets to ensure robust discrimination performance across different cancer types.
[0108] Feature Importance Analysis: After model training, this invention used Shapley value analysis and XGBoost's built-in feature importance metric to rank the importance of fused features. The results show that features from fragmentomics and microbiome both contribute to the cancer detection model, but each reflects different information dimensions and complements the others. Table 1 lists several key features and their relative contributions to the model. The values are normalized relative contribution rates; a higher value indicates a greater impact of the feature on the model's decision. The actual model contains hundreds of features; the table above only lists some representative features.
[0109] Table 1. Several key features and their importance examples in cancer detection models.
[0110]
[0111] It can be seen that the genome-wide CNV variation level is the most important single feature (with the highest contribution rate), reflecting large-scale copy number abnormalities in tumor DNA, which is crucial for distinguishing between cancer and health. Meanwhile, microbial Shannon diversity ranks highly, supporting the hypothesis of decreased plasma microbial diversity in cancer patients. The proportion of short fragments (FSD feature) and the frequency of terminal sequence motifs (EDM feature, such as the occurrence rate of CTCAG) also contribute significantly, indicating that tumor-induced shifts in cfDNA fragment length distribution and alterations in enzyme digestion patterns are important discriminative signals. Furthermore, changes in the abundance of certain specific microbial genera (such as the enrichment of *Fusobacterium* in colorectal cancer patients) are helpful in the model's identification of specific cancer types. In cancer classification models, characteristic microbial combinations of different cancer types play a major role; for example, gut-associated flora features help identify gastrointestinal tumors, while viral DNA (such as HPV sequences detected in the plasma of cervical cancer patients) suggests virus-associated tumors. The importance analysis results of these features confirm the design concept of this invention, which combines multi-source information to improve the accuracy of pan-cancer screening.
[0112] Example 2
[0113] This embodiment uses the method of Example 1 to analyze a sample from a colorectal cancer patient. First, bioinformatics analysis is used to extract the following features from the cfDNA sequencing data: fragment size distribution (FSD), fragment size coverage (FSC), copy number variation (CNV), end motif (EDM), breakpoint motif (BPM), and microbiome characteristics. These features together constitute the model's input vector (in this embodiment, the total dimension is 84). This embodiment uses a normal human sample (N) and a colorectal cancer patient sample (C) as examples to demonstrate the model's operation.
[0114] In the first-stage classification model, for a normal sample N, the XGBoost model outputs a cancer probability of only 0.02 (2%), far below 0.05. Due to the extremely low confidence level, the system directly outputs the judgment result as "normal," without needing to activate the ensemble model. In other words, the sample is identified as a healthy individual with high confidence, and no obvious tumor signal is detected.
[0115] For another patient sample C known to have colorectal cancer, the initial predicted probability of the single model was 0.60 (60%), falling within the gray area (0.05-0.95). Although the single model detected some tumor signals, its confidence level was insufficient. Therefore, the system employed stacked ensemble for composite judgment. Besides the XGBoost output of 0.60, other base models such as SVM, RF, and MLP output probabilities of 0.55, 0.70, and 0.67 for this sample, respectively (these models focus on different feature patterns, each giving a moderately high probability of cancer). The GLM meta-learner weighted and fused these multiple probabilities, ultimately outputting a combined probability of approximately 0.72 (72%), higher than the judgment threshold of 0.61. Therefore, the ensemble model classified this sample as "high-risk for cancer," and the first-stage classification model's detection result was positive. Through this hierarchical judgment strategy, the model used ensemble consensus to correct the uncertainty of the single model, reducing the risk of missed diagnoses.
[0116] After testing by the first-stage classification model, C samples that are deemed positive will proceed to the second-stage classification model, where a cancer classification model will predict the specific cancer type. The second stage involves an 8-category classification problem (this embodiment can detect a total of 8 cancer types). The model primarily relies on the differences in plasma microbiome characteristics among different cancers for discrimination, while also incorporating some cancer-specific fragment omics features to improve classification accuracy.
[0117] The second-stage classification model employs a deep neural network (DNN) as a multi-classifier. This DNN consists of multiple layers of fully connected neurons, utilizing non-linear activation functions to model the complex relationship between features and cancer types. The output layer uses softmax activation to generate a predicted probability distribution for each cancer category. In this invention, the model has eight classification outputs, corresponding to eight screenable cancers. For any sample entering the second stage, the DNN model outputs a set of probability values for these eight categories, representing the model's confidence that the sample belongs to each cancer type. Finally, the system selects the category with the highest probability as the classification result, meaning it considers the sample to be most likely to have that cancer type (while also outputting the probabilities of each category for reference).
[0118] A second-stage classification model was used to classify the feature vectors of sample C, yielding predicted probabilities for eight cancer types. The model output is shown in Table 2.
[0119] Table 2. Predicted probability results for 8 cancer types
[0120]
[0121] As can be seen from the table, the model assigned the highest probability of "colorectal cancer" at 78%. Therefore, the model ultimately determined that the sample belonged to colorectal cancer, which is consistent with the patient's actual situation.
[0122] In summary, this invention achieves initial screening of healthy individuals and suspected patients through a first-stage classification model for cancer detection, excluding healthy individuals with high specificity and capturing potential positives with high sensitivity. For samples determined to be positive, a second-stage classification model further subdivides the cancer type, enabling the tracing and determination of the tumor tissue origin. In this embodiment, healthy individuals were successfully excluded, while colorectal cancer patients were accurately identified and their cancer type was further determined, demonstrating the model's discrimination process and results in an early screening context. The probabilities and decision rules output at each stage improve the reliability of the judgment, reduce the risk of false positives, and contribute to providing reliable early screening results.
[0123] Example 3
[0124] To verify the effectiveness of the method of the present invention, this embodiment was tested on a retrospective cohort containing 1000 samples (700 patients / 300 healthy controls), and the steps are as follows:
[0125] The study included 1000 participants (700 cancer patients and 300 healthy controls), covering eight common cancer types. The sample size distribution was consistent with epidemiological characteristics (100 patients each for lung cancer, breast cancer, colorectal cancer, and gastric cancer; and 75 patients each for liver cancer, pancreatic cancer, esophageal cancer, and cervical cancer). To evaluate the model's performance in early-stage cancer patients, approximately 60% of the participants were in this stage. The healthy control group comprised 30%, effectively simulating a low-prevalence (<5%) population in real-world screening scenarios and enhancing the model's clinical applicability. Participants underwent Low Pass WGS testing with a sequencing depth of 5× to ensure analytical stability.
[0126] The bimodal model based on fragmentomics and microbiome demonstrated differentiated performance in multi-cancer detection: it performed best in lung cancer (sensitivity 85.0% / specificity 98.0%), colorectal cancer (sensitivity 88.0% / 96.5%), and cervical cancer (sensitivity 88.0% / specificity 97.1%), with a negative predictive value (NPV) > 98.1%, indicating that the model has a near-clinical gold standard ability to distinguish these cancers. Although it showed high sensitivity in breast cancer (sensitivity 90.0% / specificity 94.0%) and liver cancer (sensitivity 84.0% / specificity 96.0%), the positive predictive value (PPV) was relatively low (65.2% and 65.6%, respectively), suggesting the need for secondary validation with imaging. However, there is still room for improvement in the detection efficacy for pancreatic cancer (sensitivity 72.0%) and esophageal cancer (sensitivity 76.0%). It is worth noting that the specificity for all cancer types remained above 94% (median 96.6%), and the overall performance of the model was balanced and stable. This technological breakthrough, which enables high-precision screening of eight cancers with a single test, provides a brand-new solution for early detection of multiple cancer types.
[0127] Table 3. Model performance across 8 cancer types.
[0128]
[0129] As shown in Table 3, compared with baseline models using only fragment omics or only microbiome features, the dual-modal fusion model of this invention significantly improved the overall AUC from 0.92 and 0.85 to 0.96. This result fully demonstrates that combining fragment omics with microbiome features produces a synergistic effect.
[0130] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multi-cancer early screening method based on fragment omics and microbiome characteristics, characterized in that, Includes the following steps: Obtain a plasma sample to be tested, extract cfDNA from the plasma sample, and perform low-depth whole-genome sequencing on the cfDNA to obtain sequencing data; Based on the sequencing data, fragment omics features from human cfDNA were extracted; Based on sequences in the sequencing data that did not align to the human genome, microbiome features were extracted from circulating microbial DNA in the plasma. The fragment omics features and the microbiome features are fused to form a fused feature vector; the fused feature vector is input into a trained first-stage classification model to obtain a judgment result on whether the individual from which the plasma sample to be tested is cancer-positive; if the judgment result is positive, the fused feature vector is input into a trained second-stage classification model to obtain the cancer type of the individual; The fragment omics features include at least one of fragment size distribution, fragment size coverage, copy number variation, fragment terminal sequence motif, and breakpoint sequence motif; The microbiome characteristics include the relative abundance of microbial species and / or the microbial community diversity index; The first-stage classification model is an ensemble learning model; the ensemble learning model is configured to: directly adopt the judgment result when the confidence level of the output of a single base model is higher than 0.95 or lower than 0.05; and enable the stacked ensemble model for comprehensive judgment when the confidence level is between 0.05 and 0.
95. The second-stage classification model is a deep neural network model; the probability output by the second-stage classification model is calibrated by temperature scaling, where the temperature parameter τ ranges from 0.5 to 2.0; in the second-stage classification model, the discriminative weight of microbiome features is greater than that of fragment omics features; The cancer types include at least two of the following: lung cancer, liver cancer, colorectal cancer, stomach cancer, esophageal cancer, pancreatic cancer, breast cancer, and cervical cancer.
2. The method for early screening of multiple cancer types according to claim 1, characterized in that, The decision threshold of the first-stage classification model is dynamically set based on Bayesian decision rules.
3. The application of the multi-cancer early screening method according to claim 1 or 2 in the preparation of a kit for multi-cancer early screening.
Citation Information
Patent Citations
Marker for early intestinal cancer screening and adenoma diagnosis and applications thereof
CN109852714A
Early tumor screening method for multiple cancer species WGS
CN115831355A