Application of genetic markers in early screening of esophagus, stomach, intestine multiple cancers, early screening model construction method and detection device

By combining gene markers and using deep learning models for converters, the problem of low early diagnosis rate of gastrointestinal tumors in existing technologies has been solved, achieving high sensitivity and high specificity in early screening and source tracing analysis of multiple gastrointestinal cancers.

CN120690283BActive Publication Date: 2026-04-14GENESEEQ TECH INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GENESEEQ TECH INC
Filing Date
2025-08-21
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing gastrointestinal tumor screening technologies, such as endoscopy and imaging examinations, are highly invasive, expensive, or lack sufficient sensitivity. Laboratory tests also have unsatisfactory specificity and sensitivity, resulting in a low early diagnosis rate of gastrointestinal tumors.

Method used

By employing multi-omics integrated deep learning analysis technology based on gene markers, a highly sensitive and specific early screening system for multiple gastrointestinal cancers is constructed by detecting circulating cell-free DNA (cfDNA). The system utilizes features such as copy number variation, the proportion of short and long reads in DNA fragments, and the coverage of transcription start sites of preset genes, combined with a deep learning model of a converter, to classify and trace samples.

Benefits of technology

It achieves high sensitivity (81.8%) and high specificity (99.0%) for early screening of esophageal, colorectal, and gastric cancer, accurately distinguishing digestive tract tumors from healthy individuals, and maintaining high accuracy in testing (sensitivity 79.4%, specificity 99.2%).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690283B_ABST
    Figure CN120690283B_ABST
Patent Text Reader

Abstract

The application discloses a kind of gene markers in esophagus, stomach, intestine multiple cancer early screening application, early screening model construction method and detection device, belong to the early non-invasive detection technical field of digestive tract tumor.It establishes a new type of multiple cancer screening system by analyzing the whole genome characteristics of circulating free DNA in peripheral blood.Based on low-depth whole genome sequencing data, three dimensions of molecular markers are detected: genome copy number variation pattern, DNA fragment distribution characteristics of specific length and epigenetic signals of transcription initiation region.Advanced converter neural network architecture is used, and the model can efficiently capture the complex feature correlation in the whole genome range through its unique self-attention mechanism.The model design specially considers the particularity of genomic data, and introduces an adaptive position coding system to accurately reflect the spatial distribution relationship of DNA fragments on the chromosome.The system can still maintain excellent detection performance at very low sequencing depth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an early screening technology for multiple types of digestive tract cancers (esophageal, gastric, and colorectal cancer) based on gene markers, including its model construction method and detection system, belonging to the field of molecular diagnostic technology. Background Technology

[0002] Malignant tumors have become a major public health problem that seriously endangers health. The incidence and mortality rates of malignant tumors are showing a continuous upward trend. Among them, the disease burden of gastrointestinal tumors (esophageal cancer, colorectal cancer, and gastric cancer) is particularly heavy, accounting for about 30% of the incidence and 35% of the mortality of all malignant tumors.

[0003] In clinical diagnosis and treatment, early-stage gastrointestinal tumors often lack specific symptoms. Early esophageal cancer may only present as mild swallowing discomfort; early symptoms of colorectal cancer are easily confused with irritable bowel syndrome; and early symptoms of gastric cancer are often similar to those of chronic gastritis. This nonspecificity of symptoms often leads to delayed diagnosis. The insufficient early diagnosis rate of gastrointestinal tumors is an important reason for the differences in prognosis.

[0004] Currently, commonly used screening methods in clinical practice mainly include endoscopy, imaging examinations, and laboratory tests. Endoscopy, as the gold standard, allows for direct observation of mucosal lesions and biopsy, but its invasiveness and high cost limit its widespread use in population screening. Imaging examinations include CT and MRI, but their sensitivity for early lesions is limited. Laboratory tests include fecal occult blood tests and tumor marker detection, but their specificity and sensitivity are not ideal.

[0005] Novel screening technologies based on liquid biopsy are becoming a research hotspot due to their non-invasive and convenient advantages. By detecting biomarkers such as circulating tumor DNA (ctDNA) and exosomes in the blood, early screening and monitoring of gastrointestinal tumors can be achieved. In particular, multi-omics joint analysis strategies, combining genomics, epigenetics, and other multi-dimensional information, are expected to further improve the accuracy of early diagnosis. Summary of the Invention

[0006] This application proposes an innovative multi-omics genetic biomarker integrated deep learning analysis technique to construct a highly sensitive and specific early screening system for multiple gastrointestinal cancers by systematically detecting circulating cell-free DNA (cfDNA). This technical solution specifically targets esophageal cancer, colorectal cancer, and gastric cancer—three gastrointestinal malignancies with significant clinical burdens—achieving accurate detection and source tracing of early-stage tumors through non-invasive liquid biopsy.

[0007] To achieve the above objectives, the present invention proposes the following technical solutions:

[0008] The application of a combination of gene markers in the preparation of an early screening reagent for diagnosing multiple gastrointestinal cancers, wherein the early screening reagent is used to distinguish cancer patients with multiple gastrointestinal cancers from healthy individuals; or, the early screening reagent is used to trace the origin of cancer in cancer patients with multiple gastrointestinal cancers; wherein the multiple gastrointestinal cancers are esophageal cancer, colorectal cancer, and gastric cancer; wherein the combination of gene markers is derived from whole-genome sequencing data of the subject's cfDNA and consists of the following markers:

[0009] First marker: copy number variation;

[0010] The second biomarker: the proportion of short and long reads in a DNA fragment;

[0011] The third biomarker: coverage of the transcription start site of the preset gene.

[0012] The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window.

[0013] The proportions of short and long reads in DNA fragments are obtained through the following steps: The whole genome is divided into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window are counted and designated as short reads and long reads, respectively. The proportions of short reads and long reads within each window range are then obtained.

[0014] The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window.

[0015] The preset genes consist of the following genes: ABCB1, ABCC3, ABL1, ABL2, ABRAXAS1, ACKR3, ACTG1, ACVR1, ACVR1B, ACVR2A, ADARB2, ADGRA2, ADGRG4, ADHFE1, AFDN, AFF4, AGGF1, AGK, AGO1, AGO2, AIP, AJUBA, AKT1, AKT1S1, AKT2, AKT3, ALB, ALDH1L2, ALDH2, ALK, ALOX12B, ALOX15B, ALOX5, AMER1, ANKRD11, ANKRD26, APC, APEX1, APLNR, A R, ARAF, ARHGAP26, ARHGAP35, ARHGEF12, ARHGEF28, ARID1A, ARID1B, ARID2, ARID3A, ARID3B, ARID3C, ARID4A, ARID4B, ARID5A, ARID5B, ASCL1, ASXL1, A SXL2, ATF1, ATIC, ATM, ATMIN, ATP1A1, ATP6AP1, ATP6V1B2, ATR, ATRIP, ATRX, ATXN2, ATXN7, AURKA, AURKB, AXIN1, AXIN2, AXL, B2M, BAALC, BABAM1, BACH 2. BAP1, BARD1, BBC3, BCL10, BCL11B, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL3, BCL6, BCL7A, BCL9, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BR CA2, BRD3, BRD4, BRIP1, BRSK1, BTG1, BTG2, BTK, BUB1B, CACNA1D, CAD, CALR, CAMTA1, CANT1, CARD11, CARM1, CASP8, CASR, CBFA2T3, CBFB, CBL, CBLB, CCN 6. CCNB3, CCND1, CCND2, CCND3, CCNE1, CCNQ, CD19, CD22, CD274, CD276, CD28, CD58, CD70, CD74, CD79A, CD79B, CDC42, CDC73, CDH1, CDH11, CDH2, CDH4, C DK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2C, CEBPA, CENPA, CHD2, CHD4, CHEK1, CHEK2, CHTF8, CIC, CIITA, CLTCL1, CMTR2,CNTRL, COL1A1, COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, C RLF2, CRTC1, CSDE1, CSF1R, CSF3R, CTC1, CTCF, CTDNEP1, CTLA4, CTNNA1, CTN NB1、CTR9、CUL3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP 1、DCUN1D1、DDB2、DDIT3、DDR2、DDX3X、DDX4、DDX41、DDX5、DEK、DHX15、DICER 1、DIS3、DIS3L2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3 A, DNMT3B, DOT1L, DPYD, DROSHA, DTX1, DUSP22, DUSP4, E2F3, EBF1, ECSIT, EC T2L、EED、EGFL7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、 ELF3、ELF4、ELK4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EP CAM, EPHA3, EPHA5, EPHA7, EPHB1, EPHB4, EPOR, ERBB2, ERBB3, ERBB4, ERC1 ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, E TAA1, ETNK1, ETS1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT1, EZH1, EZH2, EZHIP 、FANCA、FANCB、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、F AS、FAT1、FBXO11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2 、FGF23、FGF3、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGF R4、FH、FHIT、FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOX L2、FOXN4、FOXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、GAB1、GAB2、GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11 、GNA12、GNA13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、G SK3B, GSTP1, GTF2I, H1-2, H1-3, H1-4, H1-5, H2AC11, H2AC16, H2AC17, H2AC6, H2BC11, H2BC12, H2BC17, H2BC4, H2BC5, H2BC8, H3-3A, H3-3B, H3-4, H3-5 H3C1、H3C10、H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H 3C7、H3C8、H3P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、 HLA-A, HLA-B, HLA-C, HMGA1, HMGA2, HNF1A, HNF1B, HOXA11, HOXB13, HRAS, HSD17B2, HSD3B1, HSP90AA1, HTATIP2, ICOSLG, ID1, ID3, IDH1, IDH2, IFNAR1 IFNGR1, IGF1, IGF1R, IGF2, IKBKE, IKZF1, IKZF3, IL10, IL3, IL6ST, IL7R, ING1, INHA, INHBA, INPP4A, INPP4B, INPPL1, INSR, INTS6, IQGAP1, IRF1, IRF 2, IRF4, IRF8, IRS1, IRS2, ITPKB, JAK1, JAK2, JAK3, JARID2, JAZF1, JUN, KAT6A, KAT7, KBTBD4, KDM5A, KDM5C, KDM5D, KDM6A, KDR, KEAP1, KEL, KIAA1549 KIT、KLF2、KLF3、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KN STRN、KRAS、KSR2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、 LRP1B、LRP5、LRP6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、M ALT1、MAML2、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7, MAP4K4, MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2 MDM4, MECOM, MED12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MI DEAS, MITF, MKI67, MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, M PL, MRE11, MS4A1, MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2 MTHFR, MTOR, MUTYH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYO D1, NAB2, NADK, NBN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NF ATC2, NFE2, NFE2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOT CH3, NOTCH4, NPM1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1 NTRK1, NTRK2, NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8 PAK1, PAK5, PALB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1 DCD1LG2, PDGFB, PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19.P HF6, PHLPP1, PHLPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK 3CB PIK3CD PIK3CG PIK3R1 PIK3R2 PIK3R3 PIM1 PLCG1 PLCG2 PLK2P MAIP1, PML, PMS1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3 F2, POU3F4, PPARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1 PRDM14, PREX2, PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8.PRSS1、PSMB2、PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2 、PTPRD、PTPRS、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、 RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1 、RB1、RBM10、RBM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA 、RICTOR、RIOK2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS 15、RPS6KA4、RPS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、R UNX1T1、RXRA、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDH C、SDHD、SERPINB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SE TD1B、SETD2、SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、S F3B2、SFRP1、SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT 3、SLX4、SMAD2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARC D1、SMARCE1、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX 17、SOX2、SOX9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRS F2、SS18、SSX1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、ST AT5B、STAT6、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、 TAL1、TAP1、TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT, TET1, TET2, TET3, TFE3, TGFBR1, TGFBR2, TIGAR, TLE1, TLE2, TLE3, TLE4, TLX1, TLX3, TMEM127, TMPRSS2, TNFAIP 3. TNFRSF14, TNFRSF17, TNFSF13, TONSL, TOP1, TOP2A, TP53, TP53BP1, TP63, TPM3, TPMT, TRA, TRAF2, TRAF3, TRAF5, TR AF7, TRB, TRD, TRG, TRIB3, TRIM27, TRIP13, TSC1, TSC2, TSHR, TYK2, TYMS, U2AF1, U2AF2, UBA1, UBE2A, UBR5, UBTF, UCH L1, UPF1, USP1, USP6, USP8, VAV1, VAV2, VEGFA, VHL, VTCN1, WEE1, WIF1, WRN, WT1, WWP1, WWTR1, XBP1, XIAP, XPA, XPC, PO1, XRCC1, 3. ZRSR2, CARS1, CDX2, CEP43, CREB3L1, DDX10, DDX6, FCGR2B, FOXO4, GAS7, GPHN, H4C9, HERPUD1, HLF, HOXA9, HOXC11, IGK, IGL, IL21R, ITK, KIF5B, LASP1, LPP, MLF1, MLLT6, MSN, MUC1, MYH9, NCOA2, NFKB2, NUMA1, PBX1, PER1, PLAG1, PRRX 1. PSIP1, RNF213, RPL22, RPN1, SRSF3, SSX4, TAF15, TAL2, TFG, TPM4, TRIM24, TRIP11, ZBTB16, ZMYM2, ZNF384, ZNF521. ,

[0016] A method for constructing a classification model to distinguish patients with esophageal cancer, colorectal cancer, or gastric cancer from healthy individuals includes the following steps:

[0017] a) Obtain plasma samples from the subject population and healthy individuals, extract cfDNA and perform whole-genome sequencing to obtain sequencing data;

[0018] b) Based on the sequencing data, a feature vector is extracted from each sample to obtain a combination of gene markers, which serves as an input variable. The combination of gene markers consists of the following markers:

[0019] First marker: copy number variation;

[0020] The second biomarker: the proportion of short and long reads in a DNA fragment;

[0021] The third biomarker: coverage of the transcription start site of the predefined gene;

[0022] c) Inputting the input variables into a first converter deep learning model for training, the first converter deep learning model comprising:

[0023] The embedding layer is used to map input variables to preset dimensions.

[0024] Multiple parallel converter encoders, each encoder receiving a feature vector output by the embedding layer, and processing the feature using a multi-head self-attention mechanism and a feedforward network;

[0025] The merging layer is used to combine the outputs of various converter encoders to form a unified feature representation.

[0026] The output layer represents the probability that the output sample belongs to cancer based on the unified features.

[0027] d) Obtain a classification model through training.

[0028] In the converter encoder, the number of attention heads for the multi-head self-attention mechanism used to process copy number variation and transcription start site coverage is set to 4-16, and the number of attention heads for the multi-head self-attention mechanism used to process the proportion of short and long reads of DNA fragments is set to 2-8. The training uses binary cross-entropy as the loss function.

[0029] The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window;

[0030] The proportions of short and long reads in DNA fragments were obtained through the following steps: the whole genome was divided into multiple 4-6 MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted and designated as short reads and long reads, respectively, to obtain the proportions of short reads and long reads in each window range.

[0031] The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window.

[0032] A classification device for distinguishing patients with esophageal cancer, colorectal cancer, or stomach cancer from healthy individuals, comprising:

[0033] The sequencing module is used to obtain plasma samples from the subject population and healthy individuals, extract cfDNA, and perform whole-genome sequencing to obtain sequencing data.

[0034] The gene biomarker combination acquisition module is used to extract feature vectors from each sample based on the sequencing data to obtain gene biomarker combinations as input variables. The gene biomarker combinations consist of the following biomarkers:

[0035] First marker: copy number variation;

[0036] The second biomarker: the proportion of short and long reads in a DNA fragment;

[0037] The third biomarker: coverage of the transcription start site of the predefined gene;

[0038] A first converter deep learning model module is used to input the feature vector into the first converter deep learning model for training. The first converter deep learning model includes:

[0039] The embedding layer is used to map input variables to preset dimensions.

[0040] Multiple parallel converter encoders, each encoder receiving a feature vector output by the embedding layer, and processing the feature using a multi-head self-attention mechanism and a feedforward network;

[0041] The merging layer is used to combine the outputs of various converter encoders to form a unified feature representation.

[0042] The output layer represents the probability that the output sample belongs to cancer based on the unified features.

[0043] The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window;

[0044] The proportions of short and long reads in DNA fragments were obtained through the following steps: the whole genome was divided into multiple 4-6 MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted and designated as short reads and long reads, respectively, to obtain the proportions of short reads and long reads in each window range.

[0045] The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window.

[0046] A method for constructing a source tracing model for esophageal cancer, colorectal cancer, and gastric cancer includes the following steps:

[0047] a) Obtain plasma samples from patients whose models are determined to be positive by the construction method described above, extract cfDNA and perform whole-genome sequencing to obtain sequencing data;

[0048] b) Based on the sequencing data, a feature vector is extracted from each sample to obtain a combination of gene markers, which serves as an input variable. The combination of gene markers consists of the following markers:

[0049] First marker: copy number variation;

[0050] The second biomarker: the proportion of short and long reads in a DNA fragment;

[0051] The third biomarker: coverage of the transcription start site of the predefined gene;

[0052] c) The feature vector is input into the second converter deep learning model for training. The output layer of the second converter deep learning model has three output nodes corresponding to esophageal cancer, colorectal cancer and gastric cancer respectively, and outputs the probability of each cancer type.

[0053] d) A source tracing model is obtained through training, which is used to distinguish cancer types from samples that are identified as positive by the classification model obtained by the construction method;

[0054] The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window;

[0055] The proportions of short and long reads in DNA fragments were obtained through the following steps: the whole genome was divided into multiple 4-6 MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted and designated as short reads and long reads, respectively, to obtain the proportions of short reads and long reads in each window range.

[0056] The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window.

[0057] A device for tracing the source of esophageal cancer, intestinal cancer, and gastric cancer, characterized in that it comprises:

[0058] The sequencing module is used to acquire plasma samples from patients identified as positive by the classification device, extract cfDNA and perform whole-genome sequencing to obtain sequencing data.

[0059] The gene biomarker combination acquisition module is used to extract feature vectors from each sample based on the sequencing data to obtain gene biomarker combinations as input variables. The gene biomarker combinations consist of the following biomarkers:

[0060] First marker: copy number variation;

[0061] The second biomarker: the proportion of short and long reads in a DNA fragment;

[0062] The third biomarker: coverage of the transcription start site of the predefined gene;

[0063] The second converter deep learning model module is used to input the input variables into the second converter deep learning model for training. The output layer of the second converter deep learning model has three output nodes corresponding to esophageal cancer, colorectal cancer and gastric cancer respectively, and outputs the probability of each cancer type.

[0064] The second converter deep learning model module is used to distinguish cancer types from samples that are determined to be positive by the classification device.

[0065] The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window;

[0066] The proportions of short and long reads in DNA fragments were obtained through the following steps: the whole genome was divided into multiple 4-6 MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted and designated as short reads and long reads, respectively, to obtain the proportions of short reads and long reads in each window range.

[0067] The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window.

[0068] The beneficial effects of this invention are:

[0069] Compared with existing technologies, the advantages of this invention include: the early screening model developed in this invention can not only perform early screening for multiple cancers of the digestive tract (including esophageal cancer, gastric cancer, and colorectal cancer), but also distinguish between digestive tract tumors and healthy individuals. This early detection model for multiple digestive tract cancers exhibits a sensitivity of 81.8% and a specificity of 99.0% in the training set; in the test set, the sensitivity reaches 79.4% and the specificity reaches 99.2%, with no significant inter-set differences, which helps improve the accuracy of early screening for multiple cancers. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the model building process.

[0071] Figure 2 This is a schematic diagram illustrating the process of constructing a multi-cancer early screening model for the digestive tract.

[0072] Figure 3 This represents the AUC performance of the early screening model for multiple digestive tract cancers on the training and validation sets.

[0073] Figure 4 This is the score distribution of the early screening model for multiple digestive tract cancers in the validation set.

[0074] Figure 5 This reflects the sensitivity of the early screening model for multiple cancers of the digestive tract in the training and validation sets for each type of cancer.

[0075] Figure 6 This is the distribution of prediction scores for the early screening model for multiple digestive tract cancers at different stages in the validation set.

[0076] Figure 7 It is a confusion matrix of the top 1 cancer detection results in a multi-cancer tissue model of the digestive tract. Detailed Implementation

[0077] Low-depth (approximately 5x) whole-genome sequencing (WGS) analysis was performed on 1478 patients with esophageal, gastric, or colorectal cancer and 1500 healthy individuals to analyze their circulating cell-free DNA (cfDNA). Three different features were incorporated: copy number variation (CNV) within a 1MB window, fragment size ratio (FSC), and transcription start site coverage (TSS). A gastrointestinal cancer detection model (GI DOC Model) was constructed using a transformer, and a gastrointestinal cancer tissue of origin tracing model (GI TOO Model) was trained using the transformer.

[0078] This invention utilizes low-depth cfDNA high-throughput whole-genome sequencing technology to develop a diagnostic model encompassing multi-molecular features and deep learning algorithms for the early detection of three major digestive tract cancers: esophageal, gastric, and intestinal cancers, and to trace their tissue origin. This model offers advantages such as non-invasiveness, high specificity, and high sensitivity. Its performance was validated using a dataset of 1478 cancer patients and 1500 healthy individuals, demonstrating continued high specificity and sensitivity on the validation set. The early screening model used samples from 1478 patients and 1500 healthy individuals, while the TOO model used samples from 837 cancer patients.

[0079] The implementation of this invention also includes key steps such as cfDNA extraction, library construction, and sequencing. The standard steps for extraction and library construction are not limited, and existing technologies can be appropriately adjusted as needed. During sequencing, existing sequencing technologies can be used to obtain the base information of the cfDNA. Furthermore, the sample data included in the implementation process shown below are from data collected by the applicant during its testing work. The training and validation set sample data used in the implementation process shown below are all from actual real clinical samples collected by multiple hospitals that independently collaborate with the applicant.

[0080] Table 1 Dataset Grouping Information

[0081]

[0082] Table 2 Cancer type information in the dataset

[0083]

[0084] Plasma cfDNA extraction and sequencing methods: 8 ml of whole blood was collected from the patient using an EDTA-anticoagulated tube (purple tube). Centrifugation was performed within 2 hours to separate the plasma, and the plasma was rapidly transported to the laboratory. In the laboratory, cfDNA was extracted using the QIAGEN Plasma DNA Extraction Kit according to the manufacturer's instructions. The extracted cfDNA underwent library construction and whole-genome sequencing (WGS), with a sequencing depth controlled at approximately 5X. After sequencing, the obtained data was mapped to a human reference genome to obtain base data for each read.

[0085] Data processing and molecular characterization:

[0086] 1. 1 Mb-Bin Copy Number Variation (CNV)

[0087] Copy number variation plays a crucial role in cancer diagnosis. Analyzing the copy number of key genes or genomic regions can help differentiate between different types of cancer. Furthermore, certain rare or under-studied genomic regions may also contain important copy number variation information. In data processing, the reference genome of chromosomes 1 to 22 in each sample's whole-genome sequencing (WGS) data was divided into 1 Mb non-overlapping windows. The read depth of each window was calculated using the bedtools coverage tool and corrected according to the GC content of the window (information from the UCSC BigWig file). Hidden Markov Models (HMMs) were used to compare the depth of each window with the population baseline, and the logarithmic ratio of copy number variation for each window (log2[corrected sample depth / population baseline depth]) was calculated to obtain copy number variation data with values ​​ranging from all real numbers.

[0088] 2. Fragmentation Size Coverage (FSC)

[0089] The proportion of DNA fragment sizes reflects the length distribution of circulating cell-free DNA (cfDNA) fragments, with eigenvalues ​​ranging from all real numbers. Analyzing the coverage depth proportions of DNA fragment sizes using machine learning techniques allows for the construction of a predictive model to identify different types of cancer. Within a 5Mb window, there are significant differences in the distribution of 100-150 base pair and 151-220 base pair cfDNA fragments on chromosomes; these differences can serve as important features for distinguishing different cancer types. The distribution characteristics of cfDNA fragment lengths were extracted and quantified using the following process: First, in the aligned BAM file, the quality information, fragment length, and alignment position of each read on the human reference genome (hg19) were extracted. Then, the entire genome was divided into non-overlapping windows of 5Mb each, resulting in a total of 541 independent windows. To eliminate biases in sequencing depth or fragment abundance between different windows, the number of fragments within each window was further Z-score normalized, i.e., normalized value = (original fragment count – average fragment count across all windows) / standard deviation. The final result is a 1082-dimensional standardized fragment length distribution feature vector, which is used for model training and discriminant analysis.

[0090] 3. Transcription Start Sites (TSS) Coverage

[0091] Transcription start regions (TSS) of 990 genes were selected from the RefSeq database (database accessed in September 2024). The selection was based on: 1) support from public cancer gene databases: candidate genes were selected from authoritative databases such as COSMIC, cBioPortal, and MSK-IMPACT, identifying known tumor-related genes; 2) prior literature support: reference to key cancer genes with driver mutations repeatedly reported in existing studies; and 3) internal data validation: preliminary mining based on our own cfDNA database revealed significant differences in TSS region coverage between healthy and cancer samples, indicating potential classification value. Therefore, 990 genes that were both relevant to tumor biology and exhibited good discriminative ability in their TSS regions were ultimately selected. These regions encompassed TSSs ranging from disease-related genes to genes with unknown functions. Given that some genes have multiple transcription start sites, we uniformly selected the longest transcript and used its corresponding TSS as the representative site for that gene. The included genes comprise 990 genes, with details as follows:

[0092] ABCB1、ABCC3、ABL1、ABL2、ABRAXAS1、ACKR3、ACTG1、ACVR1、ACVR1B、ACVR2A、ADARB2、ADGRA2、ADGRG4、ADHFE1、AFDN、AFF4、AGGF1、AGK、AGO1、AGO2、AIP 、AJUBA、AKT1、AKT1S1、AKT2、AKT3、ALB、ALDH1L2、ALDH2、ALK、ALOX12B、ALOX15B、ALOX5、AMER1、ANKRD11、ANKRD26、APC、APEX1、APLNR、AR、ARAF、ARHGAP 26、ARHGAP35、ARHGEF12、ARHGEF28、ARID1A、ARID1B、ARID2、ARID3A、ARID3B、ARID3C、ARID4A、ARID4B、ARID5A、ARID5B、ASCL1、ASXL1、ASXL2、ATF1、ATIC、ATM、ATMIN、ATP1A1、ATP6AP1、ATP6V1B2、ATR、ATRIP、ATRX、ATXN2、ATXN7、AURKA、AURKB、AXIN1、AXIN2、AXL、B2M、BAALC、BABAM1、BACH2、BAP1、BARD1、 BBC3、BCL10、BCL11B、BCL2、BCL2L1、BCL2L11、BCL2L2、BCL3、BCL6、BCL7A、BCL9、BCOR、BCORL1、BCR、BIRC3、BLM、BMPR1A、BRAF、BRCA1、BRCA2、BRD3、BRD4 、BRIP1、BRSK1、BTG1、BTG2、BTK、BUB1B、CACNA1D、CAD、CALR、CAMTA1、CANT1、CARD11、CARM1、CASP8、CASR、CBFA2T3、CBFB、CBL、CBLB、CCN6、CCNB3、CCND1 、CCND2、CCND3、CCNE1、CCNQ、CD19、CD22、CD274、CD276、CD28、CD58、CD70、CD74、CD79A、CD79B、CDC42、CDC73、CDH1、CDH11、CDH2、CDH4、CDK12、CDK4、CDK 6、CDK8、CDKN1A、CDKN1B、CDKN1C、CDKN2A、CDKN2B、CDKN2C、CEBPA、CENPA、CHD2、CHD4、CHEK1、CHEK2、CHTF8、CIC、CIITA、CLTCL1、CMTR2、CNTRL、COL1A1、COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, CRLF2, CRTC1, CS DE1、CSF1R、CSF3R、CTC1、CTCF、CTDNEP1、CTLA4、CTNNA1、CTNNB1、CTR9、CUL 3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP1、DCUN1D1、DD B2, DDIT3, DDR2, DDX3X, DDX4, DDX41, DDX5, DEK, DHX15, DICER1, DIS3, DIS3L 2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3A、DNMT3B、DOT 1L、DPYD、DROSHA、DTX1、DUSP22、DUSP4、E2F3、EBF1、ECSIT、ECT2L、EED、EGFL 7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、ELF3、ELF4、EL K4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EPCAM、EPHA3、EP HA5、EPHA7、EPHB1、EPHB4、EPOR、ERBB2、ERBB3、ERBB4、ERC1、ERCC1、ERCC2、 ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, ETAA1, ETNK1, ETS1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT1, EZH1, EZH2, EZHIP, FANCA, FANC B、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、FAS、FAT1、FBX O11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2、FGF23、FGF3 、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGFR4、FH、FHIT、 FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOXL2、FOXN4、F OXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、GAB1、GAB2、GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11、GNA12、GNA 13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、GSK3B、GSTP1、 GTF2I, H1-2, H1-3, H1-4, H1-5, H2AC11, H2AC16, H2AC17, H2AC6, H2BC11, H2BC12, H2BC17, H2BC4, H2BC5, H2BC8, H3-3A, H3-3B, H3-4, H3-5, H3C1, H3C10 H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H3C7、H3C8、H3 P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、HLA-A、HLA-B、 HLA-C, HMGA1, HMGA2, HNF1A, HNF1B, HOXA11, HOXB13, HRAS, HSD17B2, HSD3B1, HSP90AA1, HTATIP2, ICOSLG, ID1, ID3, IDH1, IDH2, IFNAR1, IFNGR1, IGF1 IGF1R、IGF2、IKBKE、IKZF1、IKZF3、IL10、IL3、IL6ST、IL7R、ING1、INHA、INH BA、INPP4A、INPP4B、INPPL1、INSR、INTS6、IQGAP1、IRF1、IRF2、IRF4、IRF8、I RS1, IRS2, ITPKB, JAK1, JAK2, JAK3, JARID2, JAZF1, JUN, KAT6A, KAT7, KBTBD4, KDM5A, KDM5C, KDM5D, KDM6A, KDR, KEAP1, KEL, KIAA1549, KIT, KLF2, KLF3 、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KNSTRN、KRAS、KSR 2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、LRP1B、LRP5、LR P6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、MALT1、MAML2、MAP 2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7、MAP4K4、MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2, MDM4, MECOM, ME D12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MIDEAS, MITF, MKI6 7. MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, MPL, MRE11, MS4A1 MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2, MTHFR, MTOR, MUT YH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYOD1, NAB2, NADK, N BN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NFATC2, NFE2, NFE 2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NP M1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1, NTRK1, NTRK2 NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8, PAK1, PAK5, PA LB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1, PDCD1LG2, PDGFB PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19, PHF6, PHLPP1, PH LPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK3CB, PIK3CD, PI K3CG, PIK3R1, PIK3R2, PIK3R3, PIM1, PLCG1, PLCG2, PLK2, PMAIP1, PML, PMS 1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3F2, POU3F4, PP ARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1, PRDM14, PREX2 PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8, PRSS1, PSMB2PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2、PTPRD、PTPRS 、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、RAD51、RAD51B 、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1、RB1、RBM10、R BM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA、RICTOR、RIOK 2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS15、RPS6KA4、R PS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、RUNX1T1、RXRA 、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDHC、SDHD、SERP INB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SETD1B、SETD2、 SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、SF3B2、SFRP1、 SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT3、SLX4、SMA D2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARCD1、SMARCE1 、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX17、SOX2、SO X9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRSF2、SS18、SS X1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、STAT5B、STAT6 、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、TAL1、TAP1、 TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT、TET1、TET2、TET3、TFE3、TGFBR1、TGFBR2、TIGAR、TLE1、TLE2、TLE3、TLE4、TLX1、TLX3、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TNFRSF17、TNFSF13、TONSL、TOP1、TOP2A、TP53、TP53BP1、TP63、TPM3、TPMT、TRA、TRAF2、TRAF3、TRAF5、TRAF7、TRB、TRD、TRG、TRIB3、TRIM27、TRIP13、TSC1、TSC2、TSHR、TYK2、TYMS、U2AF1、U2AF2、UBA1、UBE2A、UBR5、UBTF、UCHL1、UPF1、USP1、USP6、USP8、VAV1、VAV2、VEGFA、VHL、VTCN1、WEE1、WIF1、WRN、WT1、WWP1、WWTR1、XBP1、XIAP、XPA、XPC、XPO1、XRCC1、XRCC2、YAP1、YES1、YY1、ZBTB20、ZBTB7A、ZFHX3、ZFP36L1、ZFP36L2、ZMYM3、ZNF217、ZNF292、ZNF750、ZNRF3、ZRSR2、CARS1、CDX2、CEP43、CREB3L1、DDX10、DDX6、FCGR2B、FOXO4、GAS7、GPHN、H4C9、HERPUD1、HLF、HOXA9、HOXC11、IGK、IGL、IL21R、ITK、KIF5B、LASP1、LPP、MLF1、MLLT6、MSN、MUC1、MYH9、NCOA2、NFKB2、NUMA1、PBX1、PER1、PLAG1、PRRX1、PSIP1、RNF213、RPL22、RPN1、SRSF3、SSX4、TAF15、TAL2、TFG、TPM4、TRIM24、TRIP11、ZBTB16、ZMYM2、ZNF384、ZNF521。、

[0093] For each selected transcription start site, a range from -1kb to +1kb around it was defined as the analysis window. Within this window, all fragments that could be aligned to this region were collected and filtered, and low-quality fragments in the BAM file were filtered (MAPQ threshold 15). The BAM file was GC bias corrected using a GC bias parameter file generated by compute GC Bias, and then coverage analysis was performed on the corrected BAM file to obtain the coverage from the upper 1kb to the lower 1kb of the transcription start regions of 990 genes. This feature ranges from 0 to positive infinity. In the above coverage calculation, coverage refers to the depth to which a specific genomic region is covered by sequencing reads; it represents the number of times the region was sequenced (i.e., how many reads covered this location). Coverage = total number of reads covering this location (e.g., if there are 5,000 reads in a 1,000 bp window, the average coverage is 5×).

[0094] The three types of feature data obtained through the above steps form an initial data vector set. These data vectors are then input into a Transformer model to construct a multi-cancer early screening model for the digestive tract (GI DOCModel). This model is specifically designed to distinguish cancer patients (including esophageal, gastric, and colorectal cancer) from healthy individuals. Based on this, using the same three types of data, a second Transformer model is used for integration, constructing a more specialized multi-cancer tissue tracing model for the digestive tract (DOC TOO Model). This advanced model can further distinguish esophageal, gastric, and colorectal cancers within a confirmed patient population with digestive tract tumors. Both models constructed using the Transformer deep learning model are examples of this. During the modeling process, a Random Grid Search Parameters algorithm can also be used to optimize the model, based on existing parameter optimization methods.

[0095] The framework for the early screening model for multiple gastrointestinal cancers (GI DOC Model) is as follows:

[0096] The early screening model for multiple gastrointestinal cancers combines two transformer models. A meta-learning process is used to perform secondary learning on the base transformer model, aiming to integrate various features and optimize the combination of these features. The model's structural design is as follows:

[0097] Data input section: The model receives three sets of input feature vectors, which have 2475 dimensions, 1082 dimensions and 990 dimensions respectively.

[0098] Feature embedding processing: consists of an embedding layer.

[0099] Feature synthesis and classification section: Includes a converter encoder and a merging layer for higher-level integration and classification of features.

[0100] Output section: Outputs the probability of the sample having cancer.

[0101] Input details of the feature vector:

[0102] The first, second, and third features are directly fed into the input layer of the model as input vectors.

[0103] Data embedding processing flow:

[0104] Each set of features first passes through an embedding layer. This layer is a fully connected linear transformation (z = W·x + b, where x is the original feature vector, W is the weight matrix, b is the bias term, and the output z is the embedding vector), used to map the original features to a higher-dimensional space, unifying the input dimension and improving the model's processing efficiency. For example, for each sample, the first set of features (2475, 1) is mapped through an embedding layer to obtain a tensor of (2475, 512), the second set of features (1082 dimensions) is mapped to a tensor of (1092, 256), and the third set of features (990 dimensions) is mapped to a tensor of (990, 512).

[0105] Transformer encoder processing flow: Each set of embedded features is fed into an independent transformer encoder. Each encoder consists of multiple layers, and the processing flow is as follows:

[0106] Multi-head self-attention layer: The model learns the dependencies within features and enhances the capture of key information by weighing the importance of different features. Each attention module divides the input features into several subspaces (heads) according to their dimensions and executes the attention mechanism independently in each subspace, thereby achieving parallel capture of features of different dimensions. By optimizing different numbers of attention heads, it was found that the performance is best when the number of CNV and TSS heads is set to 8, and the performance is best when the number of FSC heads is set to 4. Each head receives a 64-dimensional sub-vector. During the model development process, the impact of different numbers of heads on model performance was tested, and the performance of the model under different numbers of attention heads is shown in detail in Table 3. In addition to the number of attention heads, the model also includes the following parameters: attention weight dropout attention_dropout=0.1, Q, V, K matrix bias use_bias_in_qkv=True, attention calculation method attention_type=cosine, and information masking causal_masking=False.

[0107] Table 3. Optimization of the number of attention heads parameter

[0108]

[0109] Feedforward Network Layer: Next is a feedforward network that further processes the self-attention output (each sample outputs a matrix of (L, d) dimensions, where L is the feature dimension and d is the embedding dimension), adding non-linear processing capabilities. This network consists of two fully connected layers with a ReLU activation function between them; the first fully connected layer acts as an expansion layer, mapping the input to a higher dimension (1024 for CNV and TSS, and 512 for FSC). The second fully connected layer acts as a compression layer, compressing the output back to the original dimension.

[0110] Residual connections and layer normalization: The output of each sublayer is added to the input via a residual connection, and then layer normalization is applied. This helps improve the training efficiency and model performance of deep networks. Normalization parameters include the normalization calculation method used (norm_type=LayerNorm), whether to perform normalization before the self-attention layer (pre_norm=True), and the handling of 0 values ​​(epsilon=1e-5).

[0111] Merging layer processing flow:

[0112] The outputs of each encoder are concatenated along the last dimension to form a unified feature representation. This merged feature incorporates information from all input features. The parameters of the merging layer include: output merging method `fusion_layer_type=concat`, number of merging layers `fusion_depth=1`, and whether to use cross-attention connections `cross_attention=False`.

[0113] Output layer processing:

[0114] The merged features are processed by a fully connected layer and then output as a final cancer probability prediction through a Sigmoid activation function, representing the likelihood of the sample having cancer.

[0115] The hyperparameters of the model during training also include: training epochs num_epochs=50, batch size batch_size=64, learning rate learning_rate=1e-4, optimizer optimizer=Adam, dropout probability dropout_rate=0.1, label smoothing label_smoothing=0.0, and loss function loss_type=binary_crossentropy.

[0116] The feature vectors of three gene biomarkers are each input into a Transformer deep learning model, constructing three basic transformer models. Each of these models outputs a probability prediction value for tumor detection. The output probabilities of the three sub-models are input into a fully connected layer, and a final combined prediction probability is generated through linear transformation and a sigmoid activation function to obtain the final judgment result.

[0117] Furthermore, the trained Transformer model was analyzed using the SHapley Additive exPlanations (SHAP) interpreter model. We evaluated the influence of different input features by assessing the importance of these three features and calculating the SHAP values ​​of the training set to obtain the feature values ​​that contribute the most to the model in each group. The specific feature performance and their importance ranking are shown in detail in Table 4.

[0118] Table 4 Key Characteristic Variables of Copy Number Variation (CNV)

[0119]

[0120] String meaning: Cnv.<chromosome number>.<start position>.<end position>.

[0121] Table 5 Key Characteristics of DNA Fragment Size Ratio (FSC)

[0122]

[0123] In Table 5, the numbers after the name indicate the sequential number of the window; "long" refers to the long read and "short" refers to the short read. The numbers 21q, 8q, etc., combined with p / q represent the long and short arms of the chromosome, and the numbers following them indicate the start and end positions of the window.

[0124] Table 6 Key characteristics of transcription start site coverage (TSS)

[0125]

[0126] In the string, the part between the two underscores is the transcript ID number, the part before and after the colon is the chromosome number and the start position, and the part after the period is the gene.

[0127] The TOO Model is designed specifically for samples initially identified as positive in early screening models for multiple gastrointestinal cancers, aiming to further refine the prediction of cancer type. This model uses all raw features from three genetic markers as input, trains a deep learning model using a second transformer, and its structure is similar to the first model, containing an embedding layer, an encoder-transformer, a merging layer, and an output layer. However, the output layer uses a fully connected layer with 3 output nodes (num_classes), employing a softmax function to represent the probability of the corresponding cancer type for each node's output. Hyperparameters of each algorithm are optimized using a grid search technique, training separate sub-models for different feature data.

[0128] The parameters of the multi-cancer tissue tracing model include: dropout for attention weights (attention_dropout=0.1), bias in the Q, V, and K matrices (use_bias_in_qkv=True), attention calculation method (attention_type=cosine), causal masking (causal_masking=False), normalization method (norm_type=LayerNorm), whether to perform normalization before the self-attention layer (pre_norm=True), and handling of 0 values ​​(epsilon=1e-5). Other parameters include: fusion_layer_type=concat, fusion_depth=1, whether to use cross-attention connections (cross_attention=False), training epochs=50, batch size=64, learning rate=1e-4, optimizer=Adam, and dropout probability (dropout_rate=0.1). Parameters different from the early screening model include: loss function (loss_function=Categorical Cross-Entropy), and label smoothing (label_smoothing=0.1).

[0129] The Gastrointestinal Multi-Cancer Tissue Origin Model (GI TOO Model) is a multi-classification prediction tool that outputs the probability of different cancer types, specifically predicting three major cancers: esophageal cancer, colorectal cancer, and gastric cancer. The final diagnosis is determined based on the highest probability (Top1) among the three cancer predictions. This invention utilizes a transformer model based on CNV, Frag, and TSS features to achieve highly sensitive and specific early screening for multiple gastrointestinal cancers. This model can not only effectively distinguish between cancer patients and healthy individuals but also trace the origin of cancer types, featuring non-invasiveness, low throughput, and high accuracy.

[0130] The early detection model for multiple types of gastrointestinal cancers can effectively distinguish between cancerous and healthy individuals. In the training set, the sensitivity was 81.8% and the specificity was 99.0%. In the test set, the model showed a sensitivity of 79.4% and a specificity of 99.2%, with no significant differences between sets. Specific results are shown in Tables 7-9.

[0131] Table 7. Sensitivity performance of the Gastrointestinal Multiple Cancer Early Screening Detection Model (GI DOC Model) for cancer detection.

[0132]

[0133] Table 8. Cancer detection specificity of the Gastrointestinal Multiple Cancer Early Screening Detection Model (GI DOC Model)

[0134]

[0135] Table 9. Sensitivity of the Gastrointestinal Multiple Cancer Early Screening Detection Model (GI DOC Model) for Staged Cancer Detection

[0136]

[0137] Table 10 Sensitivity Performance of Various Cancer Types in the Gastrointestinal Multiple Cancer Early Screening Detection Model (GI DOC Model)

[0138]

[0139] The Gastrointestinal Multi-Cancer Tissue Origin Model (GI TOO Model) was designed to track and identify cancer types in positive samples. After training, the model was used to predict different types of gastrointestinal cancers using validation samples, calculating the origin probability of each cancer type, and selecting the cancer with the highest probability to evaluate the model's accuracy. The model's performance is shown in Tables 10 and 11.

[0140] Table 11 Performance of the Gastrointestinal Multi-Cancer Tissue Source Tracing Model (GITOOModel)

[0141]

[0142] More specific test data are shown in Table 12:

[0143] Table 12 Performance data of source tracing models for different types of gastrointestinal cancers

[0144]

[0145] The method of this invention can classify various cancers and can be combined with relevant indicators in clinical practice to make more reasonable diagnostic choices.

[0146] Comparative experiment

[0147] In this invention, copy number variation (CNV), fragmentation size, and transcription start site coverage (TSS) were selected as the three features for joint input. This combination was the optimal result obtained after systematic feature evaluation and comparative analysis. All features were extracted from low-depth whole-genome sequencing data of the same group of gastrointestinal cancer and healthy subjects, with consistent sample data (887 cancer patients and 900 healthy individuals), differing only in feature processing methods. In the early stages of model development, we attempted to incorporate other potential biomarkers, including: 1) Nucleosome Positioning (NP), which uses fragments of 100-220 bp length obtained by taking a 5 kb range above and below the transcription factor site as a window to obtain the coverage pattern curve of each transcription factor; 2) Fragmentomics-based Methylation Analysis (FRAGMA), which reflects the changes in methylation levels across the entire genome based on the CGN / NCG base configuration ratio and their respective frequencies at the 5' end of cfDNA fragments; and 3) Mutation Context and Mutation Signature (MCMS), which reflects the specific mutagenesis mechanism in tumorigenesis by fitting the COSMIC Mutational Signature based on single base mutations and their contextual information. Using each feature as input, we constructed nine basic converter models. A consistent model structure was used in the modeling phase, and 10-fold cross-validation (10-fold CV) was used to compare and evaluate the model performance. After comparing the performance of different features in cross-validation, it was found that the AUC of the model using NP, FRAGMA, or MCMS was lower than that of CNV, Fragmentation Size, and TSS. The mean AUCs of 4, 8, and 16 attention heads for the three features were 0.917, 0.885, and 0.835, respectively. The performance of NP, FRAGMA, and MCMS under different numbers of attention heads is shown in detail in Table 1. In contrast, the basic transformer models of CNV, Fragmentation, and TSS performed well and had strong complementarity, enhancing the discriminative ability of the classification model. Therefore, they were selected as the input features of the final model.

[0148] Table 13 Classification performance of the comparison model

[0149]

[0150] It will be apparent to those skilled in the art that the present invention is not limited to the exemplary embodiments described above. The present invention may take many other forms without departing from its spirit and fundamental principles. Therefore, the above description should be considered exemplary and not restrictive.

Claims

1. A method for constructing a source tracing model for esophageal cancer, colorectal cancer, and gastric cancer, characterized in that, Includes the following steps: a) Obtain plasma samples from patients identified as positive by the model construction method used to distinguish between patients with esophageal cancer, colorectal cancer, or gastric cancer and healthy individuals, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; b) Based on the sequencing data, a feature vector is extracted from each sample to obtain a combination of gene markers, which serves as an input variable. The combination of gene markers consists of the following markers: First marker: copy number variation; The second biomarker: the proportion of short and long reads in a DNA fragment; The third biomarker: coverage of the transcription start site of the predefined gene; c) The feature vector is input into the second converter deep learning model for training. The second converter deep learning model includes an embedding layer, an encoder-transformer, a merging layer and an output layer. The output layer uses a fully connected layer. The output layer of the second converter deep learning model has three output nodes corresponding to esophageal cancer, colorectal cancer and gastric cancer respectively, and outputs the probability of each cancer type. d) A source tracing model is obtained through training, which is used to distinguish the cancer type in plasma samples from patients identified as positive in step a; The copy number variation value is obtained through the following steps: dividing chromosomes 1-22 of the genome into multiple non-overlapping windows of 0.8-1.2 MB, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportions of short and long reads in DNA fragments were obtained through the following steps: The whole genome was divided into multiple 4-6 MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted and designated as short reads and long reads, respectively, to obtain the proportions of short reads and long reads within each window range. The transcription start site coverage is obtained through the following steps: Select a preset gene, further select the longest transcript of each gene, and use the region from 1kb upstream to 1kb downstream of the corresponding transcription start site as the analysis window, calculate and correct the depth of coverage of the sequencing reads within each analysis window. The method for constructing a classification model to distinguish between patients with esophageal cancer, colorectal cancer, or gastric cancer and healthy individuals includes the following steps: a1) Obtain plasma samples from the subject population and healthy individuals, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; b1) Based on the sequencing data, a feature vector is extracted from each sample to obtain a combination of gene markers, which serves as an input variable. The combination of gene markers consists of the following markers: First marker: copy number variation; The second biomarker: the proportion of short and long reads in a DNA fragment; The third biomarker: coverage of the transcription start site of the predefined gene; c1) Inputting the input variables into a first converter deep learning model for training, the first converter deep learning model comprising: The embedding layer is used to map input variables to preset dimensions. Multiple parallel converter encoders, each encoder receiving a feature vector output by the embedding layer, and processing the feature using a multi-head self-attention mechanism and a feedforward network; The merging layer is used to combine the outputs of the encoders of various converters to form a unified feature representation. The output layer represents the probability that the output sample belongs to cancer based on the unified features. d1) Obtain a classification model through training; In the converter encoder, the number of attention heads for the multi-head self-attention mechanism used to process copy number variation values ​​and transcription start site coverage is set to 4-16. The number of attention heads for the multi-head self-attention mechanism used to process the proportion of short and long reads in DNA fragments is set to 2-8. Training uses binary cross-entropy as the loss function; The preset genes consist of the following genes: ABCB1, ABCC3, ABL1, ABL2, ABRAXAS1, ACKR3, ACTG1, ACVR1, ACVR1B, ACVR2A, ADARB2, ADGRA2, ADGRG4, ADHFE1, AFDN, AFF4, AGGF1, AGK, AGO1, AGO2, AIP, AJUBA, AKT1, AKT1S1, AKT2, AKT3, ALB, ALDH1L2, ALDH2, ALK, ALOX12B, ALOX15B, ALOX5, AMER1, ANKRD11, ANKRD26, APC, APEX1, APLNR, A R, ARAF, ARHGAP26, ARHGAP35, ARHGEF12, ARHGEF28, ARID1A, ARID1B, ARID2, ARID3A, ARID3B, ARID3C, ARID4A, ARID4B, ARID5A, ARID5B, ASCL1, ASXL1, A SXL2, ATF1, ATIC, ATM, ATMIN, ATP1A1, ATP6AP1, ATP6V1B2, ATR, ATRIP, ATRX, ATXN2, ATXN7, AURKA, AURKB, AXIN1, AXIN2, AXL, B2M, BAALC, BABAM1, BACH 2. BAP1, BARD1, BBC3, BCL10, BCL11B, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL3, BCL6, BCL7A, BCL9, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BR CA2, BRD3, BRD4, BRIP1, BRSK1, BTG1, BTG2, BTK, BUB1B, CACNA1D, CAD, CALR, CAMTA1, CANT1, CARD11, CARM1, CASP8, CASR, CBFA2T3, CBFB, CBL, CBLB, CCN 6. CCNB3, CCND1, CCND2, CCND3, CCNE1, CCNQ, CD19, CD22, CD274, CD276, CD28, CD58, CD70, CD74, CD79A, CD79B, CDC42, CDC73, CDH1, CDH11, CDH2, CDH4, C DK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2C, CEBPA, CENPA, CHD2, CHD4, CHEK1, CHEK2, CHTF8, CIC, CIITA, CLTCL1, CMTR2,CNTRL, COL1A1, COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, C RLF2, CRTC1, CSDE1, CSF1R, CSF3R, CTC1, CTCF, CTDNEP1, CTLA4, CTNNA1, CTN NB1、CTR9、CUL3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP 1、DCUN1D1、DDB2、DDIT3、DDR2、DDX3X、DDX4、DDX41、DDX5、DEK、DHX15、DICER 1、DIS3、DIS3L2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3 A, DNMT3B, DOT1L, DPYD, DROSHA, DTX1, DUSP22, DUSP4, E2F3, EBF1, ECSIT, EC T2L、EED、EGFL7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、 ELF3、ELF4、ELK4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EP CAM, EPHA3, EPHA5, EPHA7, EPHB1, EPHB4, EPOR, ERBB2, ERBB3, ERBB4, ERC1 ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, E TAA1、ETNK1、ETS1、ETV1、ETV4、ETV5、ETV6、EWSR1、EXT1、EZH1、EZH2、EZHIP 、FANCA、FANCB、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、F AS、FAT1、FBXO11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2 、FGF23、FGF3、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGF R4、FH、FHIT、FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOX L2、FOXN4、FOXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、GAB1、GAB2、GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11 、GNA12、GNA13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、G SK3B, GSTP1, GTF2I, H1-2, H1-3, H1-4, H1-5, H2AC11, H2AC16, H2AC17, H2AC6, H2BC11, H2BC12, H2BC17, H2BC4, H2BC5, H2BC8, H3-3A, H3-3B, H3-4, H3-5 H3C1、H3C10、H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H 3C7、H3C8、H3P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、 HLA-A, HLA-B, HLA-C, HMGA1, HMGA2, HNF1A, HNF1B, HOXA11, HOXB13, HRAS, HSD17B2, HSD3B1, HSP90AA1, HTATIP2, ICOSLG, ID1, ID3, IDH1, IDH2, IFNAR1 IFNGR1, IGF1, IGF1R, IGF2, IKBKE, IKZF1, IKZF3, IL10, IL3, IL6ST, IL7R, ING1, INHA, INHBA, INPP4A, INPP4B, INPPL1, INSR, INTS6, IQGAP1, IRF1, IRF 2, IRF4, IRF8, IRS1, IRS2, ITPKB, JAK1, JAK2, JAK3, JARID2, JAZF1, JUN, KAT6A, KAT7, KBTBD4, KDM5A, KDM5C, KDM5D, KDM6A, KDR, KEAP1, KEL, KIAA1549 KIT、KLF2、KLF3、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KN STRN、KRAS、KSR2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、 LRP1B、LRP5、LRP6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、M ALT1、MAML2、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7, MAP4K4, MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2 MDM4, MECOM, MED12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MI DEAS, MITF, MKI67, MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, M PL, MRE11, MS4A1, MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2 MTHFR, MTOR, MUTYH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYO D1, NAB2, NADK, NBN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NF ATC2, NFE2, NFE2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOT CH3, NOTCH4, NPM1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1 NTRK1, NTRK2, NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8 PAK1, PAK5, PALB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1 DCD1LG2, PDGFB, PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19.P HF6, PHLPP1, PHLPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK 3CB PIK3CD PIK3CG PIK3R1 PIK3R2 PIK3R3 PIM1 PLCG1 PLCG2 PLK2P MAIP1, PML, PMS1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3 F2, POU3F4, PPARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1 PRDM14, PREX2, PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8.PRSS1、PSMB2、PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2 、PTPRD、PTPRS、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、 RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1 、RB1、RBM10、RBM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA 、RICTOR、RIOK2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS 15、RPS6KA4、RPS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、R UNX1T1、RXRA、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDH C、SDHD、SERPINB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SE TD1B、SETD2、SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、S F3B2、SFRP1、SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT 3、SLX4、SMAD2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARC D1、SMARCE1、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX 17、SOX2、SOX9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRS F2、SS18、SSX1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、ST AT5B、STAT6、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、 TAL1、TAP1、TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT、TET1、TET2、TET3、TFE3、TGFBR1、TGFBR2、TIGAR、TLE1、TLE2、TLE3、TLE4、TLX1、TLX3、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TNFRSF17、TNFSF13、TONSL、TOP1、TOP2A、TP53、TP53BP1、TP63、TPM3、TPMT、TRA、TRAF2、TRAF3、TRAF5、TRAF7、TRB、TRD、TRG、TRIB3、TRIM27、TRIP13、TSC1、TSC2、TSHR、TYK2、TYMS、U2AF1、U2AF2、UBA1、UBE2A、UBR5、UBTF、UCHL1、UPF1、USP1、USP6、USP8、VAV1、VAV2、VEGFA、VHL、VTCN1、WEE1、WIF1、WRN、WT1、WWP1、WWTR1、XBP1、XIAP、XPA、XPC、XPO1、XRCC1、XRCC2、YAP1、YES1、YY1、ZBTB20、ZBTB7A、ZFHX3、ZFP36L1、ZFP36L2、ZMYM3、ZNF217、ZNF292、ZNF750、ZNRF3、ZRSR2、CARS1、CDX2、CEP43、CREB3L1、DDX10、DDX6、FCGR2B、FOXO4、GAS7、GPHN、H4C9、HERPUD1、HLF、HOXA9、HOXC11、IGK、IGL、IL21R、ITK、KIF5B、LASP1、LPP、MLF1、MLLT6、MSN、MUC1、MYH9、NCOA2、NFKB2、NUMA1、PBX1、PER1、PLAG1、PRRX1、PSIP1、RNF213、RPL22、RPN1、SRSF3、SSX4、TAF15、TAL2、TFG、TPM4、TRIM24、TRIP11、ZBTB16、ZMYM2、ZNF384、ZNF521。、

Citation Information

Patent Citations

  • Cancer-related biomarker based on cfDNA sequencing and data analysis as well as application of cancer-related biomarker based on cfDNA sequencing and data analysis in cfDNA sample classification

    CN111254194A

  • Application of gene marker in early screening of multiple cancer species, early screening model construction method and detection device

    CN118366547A

  • Application of gene marker in early screening of multiple cancer species in digestive tract, early screening model construction method and detection device

    CN119108017A