Application of gene marker in early screening of esophagus, stomach and intestine multiple cancer species, early screening model construction method and detection device

By integrating deep learning technology with multi-omics gene markers and utilizing cfDNA characteristics to construct an early screening model for multiple gastrointestinal cancers, the problem of low early diagnosis rate of gastrointestinal tumors in existing technologies is solved, and early detection and traceability analysis with high sensitivity and high specificity is achieved.

CN120690283AActive Publication Date: 2025-09-23GENESEEQ TECH INC +1

Patent Information

Application Number
CN202511174537.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-23
Estimated Expiration
2045-08-21

Smart Images

  • Figure CN120690283A_ABST
    Figure CN120690283A_ABST
Patent Text Reader

Abstract

The invention discloses application of a gene marker in early screening of esophagus, stomach and intestine multiple cancer species, an early screening model construction method and a detection device, and belongs to the technical field of early noninvasive detection of digestive tract tumors. By analyzing the whole genome characteristics of circulating free DNA in peripheral blood, a novel multi-cancer-species screening system is established. On the basis of low-depth whole genome sequencing data, molecular markers in three dimensions, namely a genome copy number variation mode, a DNA fragment distribution characteristic with a specific length and an epigenetics signal of a transcription initiation region, are emphatically detected. An advanced converter neural network architecture is adopted, and the model can efficiently capture complex feature association in a whole genome range through a specific self-attention mechanism. The model design particularly considers the particularity of genome data, introduces an adaptive position coding system, and accurately reflects the spatial distribution relationship of DNA fragments on chromosomes. Therefore, the system can still maintain excellent detection performance under extremely low sequencing depth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an early screening technology for multiple digestive tract cancers (esophageal, gastric, and intestinal cancers) based on gene markers, including a model construction method and a detection system thereof, and belongs to the field of molecular diagnostic technology. Background Art

[0002] Malignant tumors have become a major public health issue, posing a serious threat to health. The incidence and mortality rates of malignant tumors continue to rise, with digestive tract cancers (esophageal cancer, intestinal cancer, and gastric cancer) carrying a particularly heavy burden. Together, these three cancers account for approximately 30% of all malignant tumor incidence and 35% of all malignant tumor mortality.

[0003] In clinical diagnosis and treatment, early-stage digestive tract cancers often lack specific symptoms. Early esophageal cancer may present with only mild dysphagia; early symptoms of intestinal cancer can be easily confused with irritable bowel syndrome; and early gastric cancer symptoms often resemble chronic gastritis. This nonspecific nature of symptoms often leads to delayed diagnosis. The inadequate early diagnosis of digestive tract cancers is a major contributor to disparate prognoses.

[0004] Currently, commonly used screening methods in clinical practice include endoscopy, imaging, and laboratory tests. Endoscopy, the gold standard, allows for visual observation of mucosal lesions and biopsy, but its invasive nature and high cost limit widespread population screening. Imaging tests, including CT and MRI, have limitations in their sensitivity for detecting early lesions. Laboratory tests, including fecal occult blood tests and tumor marker detection, have suboptimal specificity and sensitivity.

[0005] New screening technologies based on liquid biopsies are becoming a research hotspot due to their non-invasive and convenient advantages. By detecting biomarkers such as circulating tumor DNA (ctDNA) and exosomes in the blood, early screening and monitoring of gastrointestinal tumors can be achieved. In particular, multi-omics analysis strategies, combining genomic, epigenetic, and other multi-dimensional information, are expected to further improve the accuracy of early diagnosis. Summary of the Invention

[0006] This application proposes an innovative multi-omics gene marker integrated with deep learning analysis technology. By systematically detecting circulating free DNA (cfDNA), it establishes a highly sensitive and specific early screening system for multiple gastrointestinal cancers. This technical solution specifically targets esophageal cancer, intestinal cancer, and gastric cancer, three digestive tract malignancies with significant clinical burden. Through non-invasive liquid biopsy, it enables accurate detection and traceability analysis of early tumors.

[0007] To achieve the above objectives, the present invention proposes the following technical solutions: A gene marker combination is used in the preparation of an early screening reagent for diagnosing multiple digestive tract cancers, wherein the early screening reagent is used to distinguish between patients with multiple digestive tract cancers and healthy people; or, the early screening reagent is used to trace the source of cancer in patients with multiple digestive tract cancers; the multiple digestive tract cancers are esophageal cancer, intestinal cancer, and gastric cancer; the gene marker combination is derived from whole-genome sequencing data of a subject's cfDNA and is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene.

[0008] The copy number variation value is obtained by dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window.

[0009] The proportion of short reads and long reads of DNA fragments was obtained by the following steps: the whole genome was divided into multiple 4-6MB non-overlapping windows, and the number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window were counted as short reads and long reads, respectively, to obtain the proportion of short reads and long reads in each window range.

[0010] The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

[0011] The preset genes are composed of the following genes: ABCB1, ABCC3, ABL1, ABL2, ABRAXAS1, ACKR3, ACTG1, ACVR1, ACVR1B, ACVR2A, ADARB2, ADGRA2, ADGRG4, ADHFE1, AFDN, AFF4, AGGF1, AGK, AGO1, AGO2, AIP, AJUBA, AKT1, AKT1S1, AKT2, AKT3, ALB, ALDH1L2, ALDH2, ALK, ALOX12B, ALOX15B, ALOX5, AMER1, ANKRD11, ANKRD26, APC, APEX1, APLNR, A R, ARAF, ARHGAP26, ARHGAP35, ARHGEF12, ARHGEF28, ARID1A, ARID1B, ARID2, ARID3A, ARID3B, ARID3C, ARID4A, ARID4B, ARID5A, ARID5B, ASCL1, ASXL1, A SXL2, ATF1, ATIC, ATM, ATMIN, ATP1A1, ATP6AP1, ATP6V1B2, ATR, ATRIP, ATRX, ATXN2, ATXN7, AURKA, AURKB, AXIN1, AXIN2, AXL, B2M, BAALC, BABAM1, BACH 2. BAP1, BARD1, BBC3, BCL10, BCL11B, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL3, BCL6, BCL7A, BCL9, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BR CA2, BRD3, BRD4, BRIP1, BRSK1, BTG1, BTG2, BTK, BUB1B, CACNA1D, CAD, CALR, CAMTA1, CANT1, CARD11, CARM1, CASP8, CASR, CBFA2T3, CBFB, CBL, CBLB, CCN 6. CCNB3, CCND1, CCND2, CCND3, CCNE1, CCNQ, CD19, CD22, CD274, CD276, CD28, CD58, CD70, CD74, CD79A, CD79B, CDC42, CDC73, CDH1, CDH11, CDH2, CDH4, C DK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2C, CEBPA, CENPA, CHD2, CHD4, CHEK1, CHEK2, CHTF8, CIC, CIITA, CLTCL1, CMTR2,CNTRL, COL1A1, COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, C RLF2, CRTC1, CSDE1, CSF1R, CSF3R, CTC1, CTCF, CTDNEP1, CTLA4, CTNNA1, CTN NB1、CTR9、CUL3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP 1、DCUN1D1、DDB2、DDIT3、DDR2、DDX3X、DDX4、DDX41、DDX5、DEK、DHX15、DICER 1、DIS3、DIS3L2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3 A, DNMT3B, DOT1L, DPYD, DROSHA, DTX1, DUSP22, DUSP4, E2F3, EBF1, ECSIT, EC T2L、EED、EGFL7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、 ELF3、ELF4、ELK4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EP CAM, EPHA3, EPHA5, EPHA7, EPHB1, EPHB4, EPOR, ERBB2, ERBB3, ERBB4, ERC1 ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, E TAA1、ETNK1、ETS1、ETV1、ETV4、ETV5、ETV6、EWSR1、EXT1、EZH1、EZH2、EZHIP 、FANCA、FANCB、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、F AS、FAT1、FBXO11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2 、FGF23、FGF3、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGF R4、FH、FHIT、FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOX L2、FOXN4、FOXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、<h2 style=";text-align:left;direction:ltr">GAB1、GAB2、GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11 、GNA12、GNA13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、G SK3B、GSTP1、GTF2I、H1-2、H1-3、H1-4、H1-5、H2AC11、H2AC16、H2AC17、H2AC6、H2BC11、H2BC12、H2BC17、H2BC4、H2BC5、H2BC8、H3-3A、H3-3B、H3-4、H3-5、 H3C1、H3C10、H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H 3C7、H3C8、H3P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、 HLA-A、HLA-B、HLA-C、HMGA1、HMGA2、HNF1A、HNF1B、HOXA11、HOXB13、HRAS、HSD17B2、HSD3B1、HSP90AA1、HTATIP2、ICOSLG、ID1、ID3、IDH1、IDH2、IFNAR1、 IFNGR1、IGF1、IGF1R、IGF2、IKBKE、IKZF1、IKZF3、IL10、IL3、IL6ST、IL7R、ING1、INHA、INHBA、INPP4A、INPP4B、INPPL1、INSR、INTS6、IQGAP1、IRF1、IRF 2、IRF4、IRF8、IRS1、IRS2、ITPKB、JAK1、JAK2、JAK3、JARID2、JAZF1、JUN、KAT6A、KAT7、KBTBD4、KDM5A、KDM5C、KDM5D、KDM6A、KDR、KEAP1、KEL、KIAA1549、 KIT、KLF2、KLF3、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KN STRN、KRAS、KSR2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、 LRP1B、LRP5、LRP6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、M ALT1、MAML2、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7, MAP4K4, MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2 MDM4, MECOM, MED12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MI DEAS, MITF, MKI67, MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, M PL, MRE11, MS4A1, MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2 MTHFR, MTOR, MUTYH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYO D1, NAB2, NADK, NBN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NF ATC2, NFE2, NFE2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOT CH3, NOTCH4, NPM1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1 NTRK1, NTRK2, NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8 PAK1, PAK5, PALB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1 DCD1LG2, PDGFB, PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19.P HF6, PHLPP1, PHLPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK 3CB PIK3CD PIK3CG PIK3R1 PIK3R2 PIK3R3 PIM1 PLCG1 PLCG2 PLK2P MAIP1, PML, PMS1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3 F2, POU3F4, PPARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1 PRDM14, PREX2, PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8.<h2 style=";text-align:left;direction:ltr">PRSS1、PSMB2、PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2 、PTPRD、PTPRS、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、 RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1 、RB1、RBM10、RBM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA 、RICTOR、RIOK2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS 15、RPS6KA4、RPS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、R UNX1T1、RXRA、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDH C、SDHD、SERPINB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SE TD1B、SETD2、SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、S F3B2、SFRP1、SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT 3、SLX4、SMAD2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARC D1、SMARCE1、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX 17、SOX2、SOX9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRS F2、SS18、SSX1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、ST AT5B、STAT6、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、 TAL1、TAP1、TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT, TET1, TET2, TET3, TFE3, TGFBR1, TGFBR2, TIGAR, TLE1, TLE2, TLE3, TLE4, TLX1, TLX3, TMEM127, TMPRSS2, TNFAIP 3. TNFRSF14, TNFRSF17, TNFSF13, TONSL, TOP1, TOP2A, TP53, TP53BP1, TP63, TPM3, TPMT, TRA, TRAF2, TRAF3, TRAF5, TR AF7, TRB, TRD, TRG, TRIB3, TRIM27, TRIP13, TSC1, TSC2, TSHR, TYK2, TYMS, U2AF1, U2AF2, UBA1, UBE2A, UBR5, UBTF, UCH L1, UPF1, USP1, USP6, USP8, VAV1, VAV2, VEGFA, VHL, VTCN1, WEE1, WIF1, WRN, WT1, WWP1, WWTR1, XBP1, XIAP, XPA, XPC, PO1, XRCC1, 3. ZRSR2, CARS1, CDX2, CEP43, CREB3L1, DDX10, DDX6, FCGR2B, FOXO4, GAS7, GPHN, H4C9, HERPUD1, HLF, HOXA9, HOXC11, IGK, IGL, IL21R, ITK, KIF5B, LASP1, LPP, MLF1, MLLT6, MSN, MUC1, MYH9, NCOA2, NFKB2, NUMA1, PBX1, PER1, PLAG1, PRRX 1. PSIP1, RNF213, RPL22, RPN1, SRSF3, SSX4, TAF15, TAL2, TFG, TPM4, TRIM24, TRIP11, ZBTB16, ZMYM2, ZNF384, ZNF521. ,

[0012] A method for constructing a classification model for distinguishing patients with esophageal cancer, intestinal cancer, or gastric cancer from healthy people comprises the following steps: a) Obtain plasma samples from the subject population and healthy subjects, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; b) extracting a feature vector for each sample based on the sequencing data to obtain a gene marker combination as an input variable, wherein the gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; c) inputting the input variables into a first converter deep learning model for training, where the first converter deep learning model includes: Embedding layer, used to map input variables to preset dimensions; Multiple parallel transformer encoders, each of which receives a feature vector output by the embedding layer and processes the feature using a multi-head self-attention mechanism and a feedforward network; The merging layer is used to concatenate the outputs of the encoders of each converter to form a unified feature representation; An output layer, which outputs the probability that the sample belongs to cancer based on the unified feature representation; d) Obtain a classification model through training.

[0013] In the transformer encoder, the number of attention heads of the multi-head self-attention mechanism for processing copy number variation values ​​and transcription start site coverage is set to 4-16, and the number of attention heads of the multi-head self-attention mechanism for processing short read and long read ratio features of DNA fragments is set to 2-8, and binary cross entropy is used as the loss function for training; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

[0014] A classification device for distinguishing patients with esophageal cancer, intestinal cancer, or gastric cancer from healthy people, comprising: A sequencing module is used to obtain plasma samples from the subject population and healthy people, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; The gene marker combination acquisition module is used to extract a feature vector for each sample based on the sequencing data and obtain a gene marker combination as an input variable. The gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; A first converter deep learning model module is configured to input the feature vector into a first converter deep learning model for training. The first converter deep learning model includes: Embedding layer, used to map input variables to preset dimensions; Multiple parallel transformer encoders, each of which receives a feature vector output by the embedding layer and processes the feature using a multi-head self-attention mechanism and a feedforward network; The merging layer is used to concatenate the outputs of the encoders of each converter to form a unified feature representation; An output layer, which outputs the probability that the sample belongs to cancer based on the unified feature representation; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

[0015] A method for constructing a traceability model for esophageal cancer, intestinal cancer, and gastric cancer comprises the following steps: a) obtaining a plasma sample from a patient determined to be positive by the model constructed using the method, extracting cfDNA and performing whole-genome sequencing to obtain sequencing data; b) extracting a feature vector for each sample based on the sequencing data to obtain a gene marker combination as an input variable, wherein the gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; c) inputting the feature vector into a second transformer deep learning model for training, wherein an output layer of the second transformer deep learning model has three output nodes corresponding to esophageal cancer, intestinal cancer, and gastric cancer, respectively, and outputs the probability of each cancer type; d) obtaining a traceability model through training, wherein the traceability model is used to distinguish cancer types by tracing the origin of samples determined to be positive by the classification model obtained by the construction method; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

[0016] A device for tracing the origin of esophageal cancer, intestinal cancer, and gastric cancer, comprising: A sequencing module is used to obtain plasma samples from patients determined to be positive by the classification device, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; The gene marker combination acquisition module is used to extract a feature vector for each sample based on the sequencing data and obtain a gene marker combination as an input variable. The gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; a second converter deep learning model module, configured to input input variables into the second converter deep learning model for training, wherein the output layer of the second converter deep learning model has three output nodes corresponding to esophageal cancer, intestinal cancer, and gastric cancer, respectively, and outputs the probability of each cancer type; The second converter deep learning model module is used to distinguish the cancer types of samples determined to be positive by the classification device; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

[0017] Beneficial effects of the present invention: Compared with existing technologies, the advantages of this invention include: the early screening model developed by this invention not only enables early screening for multiple digestive tract cancers (including esophageal cancer, gastric cancer, and intestinal cancer), but also distinguishes digestive tract tumors from healthy individuals. This multi-cancer early detection model demonstrated 81.8% sensitivity and 99.0% specificity in the training set; in the test set, the sensitivity reached 79.4% and the specificity reached 99.2%, with no significant inter-set differences. This helps improve the accuracy of early screening for multiple cancers. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the model building process.

[0019] Figure 2 This is a schematic diagram of the process of constructing a model for early screening of multiple digestive tract cancers.

[0020] Figure 3 It is the AUC performance of the gastrointestinal multiple cancer early screening detection model in the training set and validation set.

[0021] Figure 4 It is the score distribution of the gastrointestinal multi-cancer early screening detection model in the validation set.

[0022] Figure 5 It is the sensitivity performance of the gastrointestinal multi-cancer early screening detection model for each cancer type in the training set and validation set.

[0023] Figure 6It is the distribution of prediction scores of the gastrointestinal multi-cancer early screening detection model at different stages in the validation set.

[0024] Figure 7 This is the confusion matrix of the top 1 cancer detection results of the digestive tract multi-cancer tissue model. DETAILED DESCRIPTION

[0025] Low-depth (approximately 5x) whole-genome sequencing (WGS) was performed on circulating free DNA (cfDNA) from 1,478 patients with esophageal, gastric, or intestinal cancer and 1,500 healthy controls. Three features, including copy number variation (CNV) in a 1MB window, fragment size fraction (FSC), and transcription start site coverage (TSS), were incorporated into a Transformer to construct a gastrointestinal multi-cancer early screening model (GI DOC Model). The Transformer was also used to train a gastrointestinal cancer tissue of origin (GI TOO Model).

[0026] Based on low-depth, high-throughput whole-genome sequencing of cfDNA, this study developed a diagnostic model that incorporates multiple molecular signatures and deep learning algorithms for the early detection of three major digestive tract cancers: esophageal, gastric, and intestinal cancers, and for tracing their tissue origins. This model offers the advantages of non-invasiveness, high specificity, and high sensitivity. The model's performance was validated using a dataset from 1,478 cancer patients and 1,500 healthy individuals, with the validation set still achieving high specificity and sensitivity. The early screening model was constructed using samples from 1,478 patients and 1,500 healthy individuals, while the TOO model was constructed using samples from 837 cancer patients.

[0027] The implementation of this invention also includes key steps such as cfDNA extraction, library construction, and sequencing. The standard steps for extraction and library construction are not limited and can be appropriately adjusted based on existing technologies as needed. The sequencing process can utilize existing sequencing technologies to obtain cfDNA base information. Furthermore, the sample data included in the implementation process described below is derived from data collected by the applicant during testing. The sample data used in the training and validation sets for the implementation process described below are all derived from actual clinical samples collected by multiple hospitals that independently collaborate with the applicant.

[0028] Table 1 Dataset grouping information

[0029] Table 2. Cancer type information in the dataset

[0030] Plasma cfDNA sample extraction and sequencing method: 8 ml of whole blood is collected from the patient using an EDTA-anticoagulant tube (purple tube). The blood is centrifuged within 2 hours to separate the plasma and rapidly transported to the laboratory. In the laboratory, cfDNA is extracted using the QIAGEN Plasma DNA Extraction Kit according to the manufacturer's instructions. The extracted cfDNA is then used for library construction and whole-genome sequencing (WGS) at a sequencing depth of approximately 5X. Following sequencing, the sequenced data is mapped to the human reference genome to obtain base data for each read.

[0031] Data processing and molecular characterization analysis: 1. 1 Mb-Bin Copy Number Variation (CNV) Copy number variation plays a key role in cancer diagnosis. Analyzing the copy number of key genes or genomic regions can help distinguish different cancer types. Furthermore, certain rare or understudied genomic regions may also contain important copy number variation information. During data processing, the reference genome for chromosomes 1 to 22 in the whole-genome sequencing (WGS) data for each sample was divided into non-overlapping 1-Mb windows. The bedtools coverage tool was used to calculate the read depth for each window and corrected for the GC content of the window (information from the UCSC BigWig file). A hidden Markov model (HMM) was used to compare each window with the population baseline depth, and the logarithm ratio of copy number variation (log2[corrected sample depth / population baseline depth]) was calculated for each window, thereby obtaining copy number variation data with a range of all real values.

[0032] 2. Fragmentation Size Coverage (FSC) DNA fragment size fractions reflect the length distribution of circulating cell-free DNA (cfDNA) fragments, with characteristic values ​​ranging from all real numbers. Using machine learning techniques to analyze the coverage depth ratios of DNA fragment sizes, a predictive model can be constructed to distinguish different cancer types. Within 5Mb windows, the chromosomal distribution of cfDNA fragments 100-150 base pairs and 151-220 base pairs differ significantly, and these differences can serve as important features for distinguishing different cancer types. The cfDNA fragment length distribution features were extracted and quantified using the following pipeline: First, from the aligned BAM files, quality information, fragment length, and alignment position on the human reference genome (hg19) were extracted for each read. The whole genome was then divided into non-overlapping windows of 5Mb in length, resulting in a total of 541 independent windows. To eliminate variations in sequencing depth or fragment abundance between windows, the number of fragments within each window was further normalized using a Z-score: normalized value = (original number of fragments – average number of fragments across all windows) / standard deviation. Finally, a 1082-dimensional standardized fragment length distribution feature vector is formed for model training and discriminant analysis.

[0033] 3. Transcription Start Sites (TSS) Coverage The transcription start regions of 990 genes were selected from the RefSeq database (database access date: September 2024). These genes were selected based on: 1) Public cancer gene database support: candidate genes were identified from known tumor-related genes in authoritative databases such as COSMIC, cBioPortal, and MSK-IMPACT; 2) prior literature support: key cancer genes with driver mutations repeatedly reported in previous studies; and 3) internal data validation: preliminary mining based on our own cfDNA database revealed significant differences in TSS coverage between healthy and cancer samples, suggesting potential classification value. Therefore, 990 genes were selected for inclusion based on their tumor biology relevance and their TSS regions exhibiting good discriminative power. These regions encompass TSSs from disease-related genes to genes of unknown function. Given that some genes have multiple transcription start sites, we uniformly selected the longest transcript and used its corresponding TSS as the representative site for that gene. The 990 genes included are as follows: ABCB1、ABCC3、ABL1、ABL2、ABRAXAS1、ACKR3、ACTG1、ACVR1、ACVR1B、ACVR2A、ADARB2、ADGRA2、ADGRG4、ADHFE1、AFDN、AFF4、AGGF1、AGK、AGO1、AGO2、AIP 、AJUBA、AKT1、AKT1S1、AKT2、AKT3、ALB、ALDH1L2、ALDH2、ALK、ALOX12B、ALOX15B、ALOX5、AMER1、ANKRD11、ANKRD26、APC、APEX1、APLNR、AR、ARAF、ARHGAP 26、ARHGAP35、ARHGEF12、ARHGEF28、ARID1A、ARID1B、ARID2、ARID3A、ARID3B、ARID3C、ARID4A、ARID4B、ARID5A、ARID5B、ASCL1、ASXL1、ASXL2、ATF1、ATIC、ATM、ATMIN、ATP1A1、ATP6AP1、ATP6V1B2、ATR、ATRIP、ATRX、ATXN2、ATXN7、AURKA、AURKB、AXIN1、AXIN2、AXL、B2M、BAALC、BABAM1、BACH2、BAP1、BARD1、 BBC3、BCL10、BCL11B、BCL2、BCL2L1、BCL2L11、BCL2L2、BCL3、BCL6、BCL7A、BCL9、BCOR、BCORL1、BCR、BIRC3、BLM、BMPR1A、BRAF、BRCA1、BRCA2、BRD3、BRD4 、BRIP1、BRSK1、BTG1、BTG2、BTK、BUB1B、CACNA1D、CAD、CALR、CAMTA1、CANT1、CARD11、CARM1、CASP8、CASR、CBFA2T3、CBFB、CBL、CBLB、CCN6、CCNB3、CCND1 、CCND2、CCND3、CCNE1、CCNQ、CD19、CD22、CD274、CD276、CD28、CD58、CD70、CD74、CD79A、CD79B、CDC42、CDC73、CDH1、CDH11、CDH2、CDH4、CDK12、CDK4、CDK 6、CDK8、CDKN1A、CDKN1B、CDKN1C、CDKN2A、CDKN2B、CDKN2C、CEBPA、CENPA、CHD2、CHD4、CHEK1、CHEK2、CHTF8、CIC、CIITA、CLTCL1、CMTR2、CNTRL、COL1A1、COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, CRLF2, CRTC1, CS DE1、CSF1R、CSF3R、CTC1、CTCF、CTDNEP1、CTLA4、CTNNA1、CTNNB1、CTR9、CUL 3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP1、DCUN1D1、DD B2, DDIT3, DDR2, DDX3X, DDX4, DDX41, DDX5, DEK, DHX15, DICER1, DIS3, DIS3L 2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3A、DNMT3B、DOT 1L、DPYD、DROSHA、DTX1、DUSP22、DUSP4、E2F3、EBF1、ECSIT、ECT2L、EED、EGFL 7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、ELF3、ELF4、EL K4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EPCAM、EPHA3、EP HA5、EPHA7、EPHB1、EPHB4、EPOR、ERBB2、ERBB3、ERBB4、ERC1、ERCC1、ERCC2、 ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, ETAA1, ETNK1, ETS1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT1, EZH1, EZH2, EZHIP, FANCA, FANC B、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、FAS、FAT1、FBX O11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2、FGF23、FGF3 、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGFR4、FH、FHIT、 FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOXL2、FOXN4、F OXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、GAB1、GAB2、<h2 style=";text-align:left;direction:ltr">GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11、GNA12、GNA 13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、GSK3B、GSTP1、 GTF2I、H1-2、H1-3、H1-4、H1-5、H2AC11、H2AC16、H2AC17、H2AC6、H2BC11、H2BC12、H2BC17、H2BC4、H2BC5、H2BC8、H3-3A、H3-3B、H3-4、H3-5、H3C1、H3C10、 H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H3C7、H3C8、H3 P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、HLA-A、HLA-B、 HLA-C、HMGA1、HMGA2、HNF1A、HNF1B、HOXA11、HOXB13、HRAS、HSD17B2、HSD3B1、HSP90AA1、HTATIP2、ICOSLG、ID1、ID3、IDH1、IDH2、IFNAR1、IFNGR1、IGF1、 IGF1R、IGF2、IKBKE、IKZF1、IKZF3、IL10、IL3、IL6ST、IL7R、ING1、INHA、INH BA、INPP4A、INPP4B、INPPL1、INSR、INTS6、IQGAP1、IRF1、IRF2、IRF4、IRF8、I RS1、IRS2、ITPKB、JAK1、JAK2、JAK3、JARID2、JAZF1、JUN、KAT6A、KAT7、KBTBD4、KDM5A、KDM5C、KDM5D、KDM6A、KDR、KEAP1、KEL、KIAA1549、KIT、KLF2、KLF3 、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KNSTRN、KRAS、KSR 2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、LRP1B、LRP5、LR P6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、MALT1、MAML2、MAP 2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7、MAP4K4、MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2, MDM4, MECOM, ME D12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MIDEAS, MITF, MKI6 7. MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, MPL, MRE11, MS4A1 MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2, MTHFR, MTOR, MUT YH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYOD1, NAB2, NADK, N BN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NFATC2, NFE2, NFE 2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOTCH3, NOTCH4, NP M1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1, NTRK1, NTRK2 NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8, PAK1, PAK5, PA LB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1, PDCD1LG2, PDGFB PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19, PHF6, PHLPP1, PH LPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK3CB, PIK3CD, PI K3CG, PIK3R1, PIK3R2, PIK3R3, PIM1, PLCG1, PLCG2, PLK2, PMAIP1, PML, PMS 1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3F2, POU3F4, PP ARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1, PRDM14, PREX2 PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8, PRSS1, PSMB2<h2 style=";text-align:left;direction:ltr">PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2、PTPRD、PTPRS 、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、RAD51、RAD51B 、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1、RB1、RBM10、R BM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA、RICTOR、RIOK 2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS15、RPS6KA4、R PS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、RUNX1T1、RXRA 、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDHC、SDHD、SERP INB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SETD1B、SETD2、 SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、SF3B2、SFRP1、 SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT3、SLX4、SMA D2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARCD1、SMARCE1 、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX17、SOX2、SO X9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRSF2、SS18、SS X1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、STAT5B、STAT6 、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、TAL1、TAP1、 TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT、TET1、TET2、TET3、TFE3、TGFBR1、TGFBR2、TIGAR、TLE1、TLE2、TLE3、TLE4、TLX1、TLX3、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TNFRSF17、TNFSF13、TONSL、TOP1、TOP2A、TP53、TP53BP1、TP63、TPM3、TPMT、TRA、TRAF2、TRAF3、TRAF5、TRAF7、TRB、TRD、TRG、TRIB3、TRIM27、TRIP13、TSC1、TSC2、TSHR、TYK2、TYMS、U2AF1、U2AF2、UBA1、UBE2A、UBR5、UBTF、UCHL1、UPF1、USP1、USP6、USP8、VAV1、VAV2、VEGFA、VHL、VTCN1、WEE1、WIF1、WRN、WT1、WWP1、WWTR1、XBP1、XIAP、XPA、XPC、XPO1、XRCC1、XRCC2、YAP1、YES1、YY1、ZBTB20、ZBTB7A、ZFHX3、ZFP36L1、ZFP36L2、ZMYM3、ZNF217、ZNF292、ZNF750、ZNRF3、ZRSR2、CARS1、CDX2、CEP43、CREB3L1、DDX10、DDX6、FCGR2B、FOXO4、GAS7、GPHN、H4C9、HERPUD1、HLF、HOXA9、HOXC11、IGK、IGL、IL21R、ITK、KIF5B、LASP1、LPP、MLF1、MLLT6、MSN、MUC1、MYH9、NCOA2、NFKB2、NUMA1、PBX1、PER1、PLAG1、PRRX1、PSIP1、RNF213、RPL22、RPN1、SRSF3、SSX4、TAF15、TAL2、TFG、TPM4、TRIM24、TRIP11、ZBTB16、ZMYM2、ZNF384、ZNF521。、

[0034] For each selected transcription start site, an analysis window was defined from -1 kb to +1 kb around it. Within this defined window, all fragments aligning to this region were collected and filtered, and low-quality fragments in the BAM file were filtered (MAPQ threshold 15). The BAM file was corrected for GC bias using the GC bias parameter file generated by compute GC Bias. Coverage analysis was then performed on the corrected BAM file, obtaining coverage from 1 kb above to 1 kb below the transcription start region for 990 genes. This feature ranged from 0 to positive infinity. In the above coverage calculation, coverage refers to the depth of coverage of a specific genomic region by sequencing reads. It indicates the number of times the region was sequenced (i.e., the number of reads covering this position); coverage = the total number of reads covering that position (for example, if 5,000 reads are present in a 1,000 bp window, the average coverage is 5×).

[0035] The three types of feature data obtained through the above steps form an initial set of data vectors. These data vectors are then input into a Transformer model to construct a Digestive Tract Multi-Cancer Early Screening Model (GI DOCModel). This model is specifically designed to differentiate between cancer patients (including esophageal, gastric, and intestinal cancers) and healthy individuals. Building on this foundation, a second Transformer model was used to integrate the same three types of data to construct a more specialized Digestive Tract Multi-Cancer Tissue Origin Model (DOC TOO Model). This advanced model can further differentiate between esophageal, gastric, and intestinal cancers within a patient population confirmed to have a gastrointestinal tumor. Both models were constructed using the Transformer deep learning model. During the modeling process, the Random Grid Search Parameters algorithm was also employed, based on existing parameter optimization methods.

[0036] The framework of the Digestive Tract Multi-Cancer Early Screening Model (GI DOC Model) is as follows: The digestive tract multi-cancer early screening model combines two transformer models. Through a meta-learning process, the basic transformer model undergoes secondary learning, aiming to integrate various features and optimize their combination. The model's structural design is as follows: Data input part: The model receives three sets of input feature vectors with 2475 dimensions, 1082 dimensions, and 990 dimensions respectively.

[0037] Feature embedding processing: consists of an embedding layer.

[0038] Feature integration and classification part: Contains a transformer encoder and a merging layer for higher-level integration and classification of features.

[0039] Result output part: output the probability of the sample suffering from cancer.

[0040] Input details of feature vector: The first, second, and third features are directly fed into the input layer of the model as input vectors.

[0041] Data embedding process: Each set of features first passes through an embedding layer. This layer is a fully connected linear transformation (z = W·x+b, where x is the original feature vector, W is the weight matrix, b is the bias term, and the output z is the embedding vector). It is used to map the original features to a higher-dimensional space, unifying the input dimensions and improving model processing efficiency. For example, for each sample, the first set of features (2475, 1) is mapped through an embedding layer to a tensor of (2475, 512), the second set of features (1082 dimensions) is mapped to a tensor of (1092, 256), and the third set of features (990 dimensions) is mapped to a tensor of (990, 512).

[0042] Transformer Encoder Processing: Each set of embedded features is fed into a separate Transformer Encoder. Each encoder consists of multiple layers and the processing flow is as follows: Multi-head self-attention layer: The model learns dependencies within features and enhances the capture of key information by weighing the importance of different features. Each attention module divides the input features into several subspaces (heads) based on their dimensions and independently performs the attention mechanism in each subspace, thereby enabling parallel capture of features of different dimensions. By tuning different numbers of attention heads, optimal performance was found when 8 heads were set for CNV and TSS, and 4 heads were set for FSC, with each head receiving a 64-dimensional subvector. During model development, the impact of different numbers of heads on model performance was tested. The performance of the models with different numbers of attention heads is detailed in Table 3. In addition to the number of attention heads, the model also includes the following parameters: attention weight dropout (attention_dropout=0.1), Q, V, and K matrix biases (use_bias_in_qkv=True), attention calculation method (attention_type=cosine), and information masking (causal_masking=False).

[0043] Table 3 Attention head number parameter tuning

[0044] Feedforward Network Layer: This is followed by a feedforward network that further processes the self-attention output (each sample output is a matrix of (L, d), where L is the feature dimension and d is the embedding dimension), adding nonlinear processing capabilities. This network consists of two fully connected layers with a ReLU activation function between them. The first fully connected layer acts as an expansion layer, mapping the input to a higher dimension (1024 for CNV and TSS, and 512 for FSC). The second fully connected layer acts as a compression layer, compressing the output to the original dimension.

[0045] Residual connections and layer normalization: The output of each sublayer is appended to the input via a residual connection, followed by layer normalization, which helps improve the training efficiency and model performance of deep networks. Normalization parameters include the normalization calculation method used (norm_type=LayerNorm), whether to perform normalization before the self-attention layer (pre_norm=True), and the treatment of zero values ​​(epsilon=1e-5).

[0046] Merge layer processing flow: The outputs of each encoder are concatenated along the last dimension to form a unified feature representation. This merged feature contains information from all input features. The parameters of the merging layer include the output merging method fusion_layer_type=concat, the number of merging layers fusion_depth=1, and whether to use cross-attention connections cross_attention=False.

[0047] Output layer processing: The merged features are processed by a fully connected layer and then a Sigmoid activation function is used to output the final cancer probability prediction, indicating the possibility that the sample has cancer.

[0048] The model's hyperparameters during training also include num_epochs=50, batch size batch_size=64, learning rate learning_rate=1e-4, optimizer optimizer=Adam, Dropout probability dropout_rate=0.1, smoothing label label_smoothing=0.0, and loss function loss_type=binary_crossentropy.

[0049] The three gene marker feature vectors were each input into a transformer deep learning model, building three basic transformer models. These models each output a probability prediction value for tumor detection. The output probabilities of the three sub-models were input into a fully connected layer, where a linear transformation and a sigmoid activation function were used to generate the final combined prediction probability, resulting in the final judgment result.

[0050] We further analyzed the trained Transformer model using the SHapley Additive exPlanations (SHAP) interpreter model. We evaluated the influence of different input features by evaluating the importance of these three features and calculating the SHAP value for the training set, obtaining the feature value that contributed most to the model in each set. Table 4 details the feature performance and their importance ranking.

[0051] Table 4 Important characteristic variables of copy number variation (CNV)

[0052] String meaning: Cnv.<chromosome number>.<starting position>.<ending position>.

[0053] Table 5 Important characteristics of DNA fragment size fraction (FSC)

[0054] In Table 5, the numbers following the names indicate the sequential numbering of the windows. "Long" refers to long reads, and "Short" refers to short reads. Numbers like "21q" and "8q" plus "p" / "q" represent the long and short arms of the chromosome, and the numbers following them indicate the start and end positions of the window.

[0055] Table 6 Important characteristic variables of transcription start site coverage (TSS)

[0056] In the string, the two "_"s refer to the transcript ID number, the ":" before and after refer to the chromosome number and the starting position respectively, and the "." is followed by the gene.

[0057] The Multi-Cancer Tissue Origin Model (TOO Model) is designed specifically for samples initially identified as positive in the Digestive Tract Multi-Cancer Early Screening Model, with the goal of further accurately predicting cancer type. This model uses all the raw features of three genetic markers as input, trains a second transformer deep learning model, and utilizes this model. This model structure is similar to the first model, consisting of an embedding layer, an encoding transformer, a merging layer, and an output layer. However, the output layer uses a fully connected layer with 3 output nodes (num_classes). Using a softmax function, the output of each node represents the probability of the corresponding cancer type. Grid Search is used to optimize the hyperparameters of each algorithm, and each sub-model is trained for different feature data.

[0058] Parameters for the multi-cancer tissue tracing model include attention weight dropout (attention_dropout=0.1), Q, V, and K matrix bias (use_bias_in_qkv=True), attention calculation method (attention_type=cosine), information masking (causal_masking=False), normalization method (norm_type=LayerNorm), whether to perform normalization before the self-attention layer (pre_norm=True), and zero value handling (epsilon=1e-5). Fusion_layer_type=concat, fusion_depth=1, cross-attention connection (cross_attention=False), training epochs (num_epochs=50), batch size (batch_size=64), learning rate (learning_rate=1e-4), optimizer (optimizer=Adam), and dropout probability (dropout_rate=0.1). Parameters that differ from the early screening model include loss function (loss_function=Categorical Cross-Entropy) and label smoothing (label_smoothing=0.1).

[0059] The Digestive Tract Multi-Cancer Tissue Origin Model (GI TOO Model) is a multi-classification prediction tool that outputs probabilities for different cancer types, specifically predicting three major cancers: esophageal cancer, intestinal cancer, and gastric cancer. The final diagnosis is determined based on the highest probability (Top 1) of the three cancer predictions. This invention utilizes a transformer model based on CNV, Fragment, and TSS features to achieve high-sensitivity and high-specificity early screening for multiple GI cancers. This model not only effectively distinguishes cancer patients from healthy individuals but also allows for differentiation of cancer types by tracing their origins, demonstrating its non-invasive, low-throughput, and high-accuracy capabilities.

[0060] The early detection model for multiple digestive tract cancers can effectively distinguish between cancer and healthy individuals, with a sensitivity of 81.8% and a specificity of 99.0% in the training set. The model's performance in the test set was 79.4% sensitive and 99.2% specific, with no significant differences between sets. The specific results are shown in Tables 7-9. Table 7 Cancer detection sensitivity of the digestive tract multi-cancer early screening model (GI DOC Model)

[0061] Table 8 Cancer detection specificity of the digestive tract multi-cancer early screening model (GI DOC Model)

[0062] Table 9 Sensitivity of GI DOC Model for cancer detection by stage

[0063] Table 10 Detection sensitivity of various cancer types in the GI DOC Model

[0064] The Digestive Tract Cancer Tissue Origin Model (GI TOO Model) is designed to track and identify cancer types in positive samples. After training, the model was used to predict different GI cancer types using validation samples. The probability of tracing each cancer type was calculated, and the cancer with the highest probability was selected to evaluate the model's accuracy. Model performance is shown in Tables 10 and 11.

[0065] Table 11 Performance of the digestive tract multi-cancer tissue tracing model (GITOOModel)

[0066] More specific test data is shown in Table 12: Table 12 Performance data of the traceability model for different gastrointestinal cancer types

[0067] The method of the present invention can classify various cancers and can be combined with relevant indicators in clinical practice to make more reasonable diagnostic choices.

[0068] Comparative experiment In this study, three features—copy number variation (CNV), DNA fragment length distribution (Fragmentation Size), and transcription start site coverage (TSS)—were selected as the optimal combination after systematic feature evaluation and comparative analysis. All features were extracted from low-depth whole-genome sequencing data from the same cohort of gastrointestinal cancer and healthy subjects (887 cancer patients and 900 healthy subjects), with only the feature processing methods being different. In the early stages of model development, we attempted to incorporate other potential markers, including 1) the coverage feature of nucleosome positioning (NP), which is based on the coverage pattern curve of each transcription factor obtained by using 100-220 bp fragments within a 5 kb range above and below the transcription factor site; 2) the methylation analysis feature of fragment omics (FRAGmentomics-based Methylation Analysis (FRAGMA), which reflects the changes in methylation levels across the genome based on the CGN / NCG base configuration ratio and their respective frequencies at the 5' end of cfDNA fragments; and 3) the feature of the specific mutagenesis mechanism during tumorigenesis (Mutation Context and Mutation Signature (MCMS), which is based on single-base mutations and their contextual information by fitting the COSMIC Mutational Signature. Nine basic converter models were constructed using each feature as input. A consistent model structure was used during the modeling stage, and 10-fold cross-validation (10-fold CV) was used to compare and evaluate the model performance. Comparing the performance of different features in cross-validation revealed that the model's AUC using NP, FRAGMA, or MCMS was lower than that of CNV, Fragmentation Size, and TSS. The average AUCs for these three features with 4, 8, and 16 attention heads were 0.917, 0.885, and 0.835, respectively. The performance of NP, FRAGMA, and MCMS with different numbers of attention heads is detailed in Table 1. In contrast, the basic transformer model of CNV, Fragmentation, and TSS performed excellently and demonstrated strong complementarity, enhancing the discriminative ability of the classification model. Therefore, they were selected as the final model input features.

[0069] Table 13 Classification performance of control model

[0070] It will be apparent to those skilled in the art that the present invention is not limited to the exemplary embodiments described above. The present invention may also take other specific forms without departing from the spirit and underlying principles of the present invention. Therefore, the above description should be considered as illustrative rather than restrictive.

Claims

1. Use of a gene marker combination in the preparation of an early screening reagent for diagnosing multiple digestive tract cancers, characterized in that: The early screening reagent is used to distinguish patients with multiple digestive tract cancers from healthy people; or, the early screening reagent is used to trace the origin of cancer in patients with multiple digestive tract cancers; the multiple digestive tract cancers are esophageal cancer, intestinal cancer, and gastric cancer; the gene marker combination is derived from whole-genome sequencing data of the subject's cfDNA and is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

2. The use according to claim 1, characterized in that The preset genes are composed of the following genes: ABCB1, ABCC3, ABL1, ABL2, ABRAXAS1, ACKR3, ACTG1, ACVR1, ACVR1B, ACVR2A, ADARB2, ADGRA2, ADGRG4, ADHFE1, AFDN, AFF4, AGGF1, AGK, AGO1, AGO2, AIP, AJUBA, AKT1, AKT1S1, AKT2, AKT3, ALB, ALDH1L2, ALDH2, ALK, ALOX12B, ALOX15B, ALOX5, AMER1, ANKRD11, ANKRD26, APC, APEX1, APLNR, A R, ARAF, ARHGAP26, ARHGAP35, ARHGEF12, ARHGEF28, ARID1A, ARID1B, ARID2, ARID3A, ARID3B, ARID3C, ARID4A, ARID4B, ARID5A, ARID5B, ASCL1, ASXL1, A SXL2, ATF1, ATIC, ATM, ATMIN, ATP1A1, ATP6AP1, ATP6V1B2, ATR, ATRIP, ATRX, ATXN2, ATXN7, AURKA, AURKB, AXIN1, AXIN2, AXL, B2M, BAALC, BABAM1, BACH 2. BAP1, BARD1, BBC3, BCL10, BCL11B, BCL2, BCL2L1, BCL2L11, BCL2L2, BCL3, BCL6, BCL7A, BCL9, BCOR, BCORL1, BCR, BIRC3, BLM, BMPR1A, BRAF, BRCA1, BR CA2, BRD3, BRD4, BRIP1, BRSK1, BTG1, BTG2, BTK, BUB1B, CACNA1D, CAD, CALR, CAMTA1, CANT1, CARD11, CARM1, CASP8, CASR, CBFA2T3, CBFB, CBL, CBLB, CCN 6. CCNB3, CCND1, CCND2, CCND3, CCNE1, CCNQ, CD19, CD22, CD274, CD276, CD28, CD58, CD70, CD74, CD79A, CD79B, CDC42, CDC73, CDH1, CDH11, CDH2, CDH4, C DK12, CDK4, CDK6, CDK8, CDKN1A, CDKN1B, CDKN1C, CDKN2A, CDKN2B, CDKN2C, CEBPA, CENPA, CHD2, CHD4, CHEK1, CHEK2, CHTF8, CIC, CIITA, CLTCL1, CMTR2,CNTRL, COL1A1, COL2A1, COP1, CRBN, CREB1, CREB3L2, CREBBP, CREM, CRKL, C RLF2, CRTC1, CSDE1, CSF1R, CSF3R, CTC1, CTCF, CTDNEP1, CTLA4, CTNNA1, CTN NB1、CTR9、CUL3、CUL4A、CUX1、CXCR4、CYLD、CYP19A1、CYSLTR2、DAXX、DAZAP 1、DCUN1D1、DDB2、DDIT3、DDR2、DDX3X、DDX4、DDX41、DDX5、DEK、DHX15、DICER 1、DIS3、DIS3L2、DKK1、DKK2、DKK3、DKK4、DLL3、DNAJB1、DNM2、DNMT1、DNMT3 A, DNMT3B, DOT1L, DPYD, DROSHA, DTX1, DUSP22, DUSP4, E2F3, EBF1, ECSIT, EC T2L、EED、EGFL7、EGFR、EGR1、EGR2、EIF1AX、EIF2B1、EIF3E、EIF4A2、EIF4E、 ELF3、ELF4、ELK4、ELL、ELL2、ELN、ELOC、EML4、EMSY、EP300、EP400、EPAS1、EP CAM, EPHA3, EPHA5, EPHA7, EPHB1, EPHB4, EPOR, ERBB2, ERBB3, ERBB4, ERC1 ERCC1, ERCC2, ERCC3, ERCC4, ERCC5, ERCC6, ERF, ERG, ERRFI1, ESCO2, ESR1, E TAA1、ETNK1、ETS1、ETV1、ETV4、ETV5、ETV6、EWSR1、EXT1、EZH1、EZH2、EZHIP 、FANCA、FANCB、FANCC、FANCD2、FANCE、FANCF、FANCG、FANCI、FANCL、FANCM、F AS、FAT1、FBXO11、FBXW2、FBXW7、FES、FEV、FGF1、FGF10、FGF14、FGF19、FGF2 、FGF23、FGF3、FGF4、FGF5、FGF6、FGF7、FGF8、FGF9、FGFR1、FGFR2、FGFR3、FGF R4、FH、FHIT、FLCN、FLI1、FLT1、FLT3、FLT4、FOLH1、FOLR1、FOXA1、FOXF1、FOX L2、FOXN4、FOXO1、FOXO3、FOXP1、FRS2、FSTL1、FUBP1、FURIN、FUS、FYN、FZR1、<h2 style=";text-align:left;direction:ltr">GAB1、GAB2、GABRA6、GATA1、GATA2、GATA3、GATA4、GATA6、GEN1、GLI1、GNA11 、GNA12、GNA13、GNAQ、GNAS、GNB1、GPC3、GPS2、GRB7、GREM1、GRIN2A、GRM3、G SK3B、GSTP1、GTF2I、H1-2、H1-3、H1-4、H1-5、H2AC11、H2AC16、H2AC17、H2AC6、H2BC11、H2BC12、H2BC17、H2BC4、H2BC5、H2BC8、H3-3A、H3-3B、H3-4、H3-5、 H3C1、H3C10、H3C11、H3C12、H3C13、H3C14、H3C15、H3C2、H3C3、H3C4、H3C6、H 3C7、H3C8、H3P6、H4C6、HDAC1、HDAC2、HDAC4、HDAC7、HFE、HGF、HIF1A、HIRA、 HLA-A、HLA-B、HLA-C、HMGA1、HMGA2、HNF1A、HNF1B、HOXA11、HOXB13、HRAS、HSD17B2、HSD3B1、HSP90AA1、HTATIP2、ICOSLG、ID1、ID3、IDH1、IDH2、IFNAR1、 IFNGR1、IGF1、IGF1R、IGF2、IKBKE、IKZF1、IKZF3、IL10、IL3、IL6ST、IL7R、ING1、INHA、INHBA、INPP4A、INPP4B、INPPL1、INSR、INTS6、IQGAP1、IRF1、IRF 2、IRF4、IRF8、IRS1、IRS2、ITPKB、JAK1、JAK2、JAK3、JARID2、JAZF1、JUN、KAT6A、KAT7、KBTBD4、KDM5A、KDM5C、KDM5D、KDM6A、KDR、KEAP1、KEL、KIAA1549、 KIT、KLF2、KLF3、KLF4、KLF5、KLHL6、KMT2A、KMT2B、KMT2C、KMT2D、KMT5A、KN STRN、KRAS、KSR2、LARP4B、LATS1、LATS2、LCK、LEF1、LGR5、LMNA、LMO1、LMO2、 LRP1B、LRP5、LRP6、LTB、LTK、LYN、LZTR1、MAD2L2、MAF、MAFB、MAGI2、MAL2、M ALT1、MAML2、MAP2K1、MAP2K2、MAP2K4、MAP3K1、MAP3K13、MAP3K14、MAP3K21、MAP3K7, MAP4K4, MAPK1, MAPK3, MAPKAP1, MAX, MBD4, MBD6, MCL1, MDC1, MDM2 MDM4, MECOM, MED12, MEF2B, MEF2C, MEF2D, MEN1, MERTK, MET, MGA, MGAM, MI DEAS, MITF, MKI67, MLH1, MLH3, MLLT1, MLLT10, MLLT3, MN1, MOB3B, MPEG1, M PL, MRE11, MS4A1, MSH2, MSH3, MSH6, MSI1, MSI2, MST1, MST1R, MTAP, MTHFD2 MTHFR, MTOR, MUTYH, MYB, MYBL1, MYC, MYCL, MYCN, MYD88, MYH11, MYO5A, MYO D1, NAB2, NADK, NBN, NCOA3, NCOA4, NCOR1, NCOR2, NCSTN, NEGR1, NF1, NF2, NF ATC2, NFE2, NFE2L2, NFKBIA, NHERF1, NKX2-1, NKX3-1, NOTCH1, NOTCH2, NOT CH3, NOTCH4, NPM1, NQO1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1 NTRK1, NTRK2, NTRK3, NUF2, NUP214, NUP93, NUP98, NUTM1, ONECUT2, P2RY8 PAK1, PAK5, PALB2, PARP1, PAX3, PAX5, PAX7, PAX8, PBRM1, PCBP1, PDCD1 DCD1LG2, PDGFB, PDGFRA, PDGFRB, PDK1, PDPK1, PDS5B, PGBD5, PGR, PHF19.P HF6, PHLPP1, PHLPP2, PHOX2B, PIGA, PIK3C2B, PIK3C2G, PIK3C3, PIK3CA, PIK 3CB PIK3CD PIK3CG PIK3R1 PIK3R2 PIK3R3 PIM1 PLCG1 PLCG2 PLK2P MAIP1, PML, PMS1, PMS2, PNRC1, POLD1, POLE, POLG, POLH, POT1, POU2F2, POU3 F2, POU3F4, PPARG, PPM1D, PPP2R1A, PPP2R2A, PPP4R2, PPP6C, PRCC, PRDM1 PRDM14, PREX2, PRKACA, PRKAR1A, PRKCB, PRKCI, PRKD1, PRKDC, PRKN, PRPF8.<h2 style=";text-align:left;direction:ltr">PRSS1、PSMB2、PTCH1、PTEN、PTP4A1、PTPN1、PTPN11、PTPN13、PTPN14、PTPN2 、PTPRD、PTPRS、PTPRT、PUM1、QKI、RAB35、RAC1、RAC2、RAD17、RAD21、RAD50、 RAD51、RAD51B、RAD51C、RAD51D、RAD52、RAD54L、RAF1、RANBP2、RARA、RASA1 、RB1、RBM10、RBM15、RECQL、RECQL4、REL、RELN、REST、RET、REV3L、RHEB、RHOA 、RICTOR、RIOK2、RIT1、RNASEH2A、RNASEH2B、RNF43、ROBO1、ROS1、RPL5、RPS 15、RPS6KA4、RPS6KB1、RPS6KB2、RPTOR、RRAGC、RRAS、RRAS2、RTEL1、RUNX1、R UNX1T1、RXRA、RYBP、SAMD9、SAMD9L、SAMHD1、SCG5、SDHA、SDHAF2、SDHB、SDH C、SDHD、SERPINB3、SERPINB4、SESN1、SESN2、SESN3、SET、SETBP1、SETD1A、SE TD1B、SETD2、SETD3、SETD4、SETD5、SETD6、SETD7、SETDB1、SETDB2、SF3B1、S F3B2、SFRP1、SFRP2、SGK1、SH2B3、SH2D1A、SHOC2、SHQ1、SLFN11、SLIT2、SLIT 3、SLX4、SMAD2、SMAD3、SMAD4、SMARCA1、SMARCA2、SMARCA4、SMARCB1、SMARC D1、SMARCE1、SMC1A、SMC3、SMG1、SMO、SMYD3、SOCS1、SOCS3、SOS1、SOX10、SOX 17、SOX2、SOX9、SP140、SPEN、SPOP、SPRED1、SPRTN、SQSTM1、SRC、SRP72、SRS F2、SS18、SSX1、SSX2、STAG1、STAG2、STAT1、STAT2、STAT3、STAT4、STAT5A、ST AT5B、STAT6、STK11、STK19、STK40、SUFU、SUZ12、SYK、SZT2、TACSTD2、TAF1、 TAL1、TAP1、TAP2、TBL1XR1、TBX3、TCF3、TCF7L2、TCL1A、TCL1B、TEK、TENT5C、TERT、TET1、TET2、TET3、TFE3、TGFBR1、TGFBR2、TIGAR、TLE1、TLE2、TLE3、TLE4、TLX1、TLX3、TMEM127、TMPRSS2、TNFAIP3、TNFRSF14、TNFRSF17、TNFSF13、TONSL、TOP1、TOP2A、TP53、TP53BP1、TP63、TPM3、TPMT、TRA、TRAF2、TRAF3、TRAF5、TRAF7、TRB、TRD、TRG、TRIB3、TRIM27、TRIP13、TSC1、TSC2、TSHR、TYK2、TYMS、U2AF1、U2AF2、UBA1、UBE2A、UBR5、UBTF、UCHL1、UPF1、USP1、USP6、USP8、VAV1、VAV2、VEGFA、VHL、VTCN1、WEE1、WIF1、WRN、WT1、WWP1、WWTR1、XBP1、XIAP、XPA、XPC、XPO1、XRCC1、XRCC2、YAP1、YES1、YY1、ZBTB20、ZBTB7A、ZFHX3、ZFP36L1、ZFP36L2、ZMYM3、ZNF217、ZNF292、ZNF750、ZNRF3、ZRSR2、CARS1、CDX2、CEP43、CREB3L1、DDX10、DDX6、FCGR2B、FOXO4、GAS7、GPHN、H4C9、HERPUD1、HLF、HOXA9、HOXC11、IGK、IGL、IL21R、ITK、KIF5B、LASP1、LPP、MLF1、MLLT6、MSN、MUC1、MYH9、NCOA2、NFKB2、NUMA1、PBX1、PER1、PLAG1、PRRX1、PSIP1、RNF213、RPL22、RPN1、SRSF3、SSX4、TAF15、TAL2、TFG、TPM4、TRIM24、TRIP11、ZBTB16、ZMYM2、ZNF384、ZNF521。、 3. A method for constructing a classification model for distinguishing patients with esophageal cancer, intestinal cancer or gastric cancer from healthy people, characterized in that: The steps include: a) Obtain plasma samples from the subject population and healthy subjects, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; b) extracting a feature vector for each sample based on the sequencing data to obtain a gene marker combination as an input variable, wherein the gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; c) inputting the input variables into a first converter deep learning model for training, where the first converter deep learning model includes: Embedding layer, used to map input variables to preset dimensions; Multiple parallel transformer encoders, each of which receives a feature vector output by the embedding layer and processes the feature using a multi-head self-attention mechanism and a feedforward network; The merging layer is used to concatenate the outputs of the encoders of each converter to form a unified feature representation; An output layer, which outputs the probability that the sample belongs to cancer based on the unified feature representation; d) Obtaining a classification model through training; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

4. The construction method according to claim 3, characterized in that In the transformer encoder, the number of attention heads of the multi-head self-attention mechanism for processing copy number variation values ​​and transcription start site coverage is set to 4-16.

5. The construction method according to claim 3, characterized in that The number of attention heads of the multi-head self-attention mechanism used to process the short read and long read ratio features of DNA fragments is set to 2-8.

6. The construction method according to claim 3, characterized in that: Binary cross entropy is used as the loss function for training.

7. A classification device for distinguishing patients with esophageal cancer, intestinal cancer or gastric cancer from healthy people, characterized in that: include: A sequencing module is used to obtain plasma samples from the subject population and healthy people, extract cfDNA and perform whole-genome sequencing to obtain sequencing data; The gene marker combination acquisition module is used to extract a feature vector for each sample based on the sequencing data and obtain a gene marker combination as an input variable. The gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; A first converter deep learning model module is configured to input the feature vector into a first converter deep learning model for training. The first converter deep learning model includes: Embedding layer, used to map input variables to preset dimensions; Multiple parallel transformer encoders, each of which receives a feature vector output by the embedding layer and processes the feature using a multi-head self-attention mechanism and a feedforward network; The merging layer is used to concatenate the outputs of the encoders of each converter to form a unified feature representation; An output layer, which outputs the probability that the sample belongs to cancer based on the unified feature representation; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

8. A method for constructing a traceability model for esophageal cancer, intestinal cancer, and gastric cancer, characterized in that: The steps include: a) obtaining a plasma sample from a patient determined to be positive by the model constructed according to the method of claim 3, extracting cfDNA and performing whole genome sequencing to obtain sequencing data; b) extracting a feature vector for each sample based on the sequencing data to obtain a gene marker combination as an input variable, wherein the gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; c) inputting the feature vector into a second transformer deep learning model for training, wherein an output layer of the second transformer deep learning model has three output nodes corresponding to esophageal cancer, intestinal cancer, and gastric cancer, respectively, and outputs the probability of each cancer type; d) obtaining a traceability model through training, wherein the traceability model is used to distinguish cancer types by tracing the origin of samples determined to be positive by the classification model obtained by the construction method of claim 3; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

9. A device for tracing the origin of esophageal cancer, intestinal cancer and gastric cancer, characterized in that: include: a sequencing module for obtaining a plasma sample from a patient determined to be positive by the classification device according to claim 7, extracting cfDNA and performing whole-genome sequencing to obtain sequencing data; The gene marker combination acquisition module is used to extract a feature vector for each sample based on the sequencing data and obtain a gene marker combination as an input variable. The gene marker combination is composed of the following markers: First marker: copy number variation value; Second marker: the ratio of short reads to long reads of DNA fragments; The third marker: the coverage of the transcription start site of the preset gene; a second converter deep learning model module, configured to input input variables into the second converter deep learning model for training, wherein the output layer of the second converter deep learning model has three output nodes corresponding to esophageal cancer, intestinal cancer, and gastric cancer, respectively, and outputs the probability of each cancer type; The second converter deep learning model module is used to distinguish the cancer types of samples determined to be positive by the classification device according to claim 7; The copy number variation value is obtained by the following steps: dividing chromosomes 1-22 of the genome into multiple 0.8-1.2MB non-overlapping windows, calculating the read depth of each window, and comparing it with the population baseline depth to obtain the copy number variation value of each window; The proportion of short and long reads of DNA fragments was obtained by dividing the whole genome into multiple 4-6 MB non-overlapping windows. The number of cfDNA fragments with lengths of 100-150 base pairs and 151-220 base pairs in each window was counted as short reads and long reads, respectively. The proportion of short reads and long reads in each window was obtained. The transcription start site coverage is obtained by the following steps: selecting preset genes, further selecting the longest transcript of each gene, using the region from 1 kb upstream to 1 kb downstream of the corresponding transcription start site as the analysis window, and calculating and correcting the depth of coverage by sequencing reads in each analysis window.

Citation Information

Patent Citations

  • Cancer-related biomarker based on cfDNA sequencing and data analysis as well as application of cancer-related biomarker based on cfDNA sequencing and data analysis in cfDNA sample classification

    CN111254194A

  • Multi-cancer early screening model construction method and detection device

    CN114927213A

  • Application of gene marker in malignant pulmonary nodule screening, construction method of screening model and detection device

    CN115295074A

  • Cancer non-invasive early screening method based on cfDNA sequencing coverage depth characteristic near TSS

    CN117316281A

  • Method and device for constructing tumor classification model

    CN118135269A

Cited By

  • Multi-cancer-species early screening method based on fragment omics and microbiological omics characteristics and application of multi-cancer-species early screening method

    CN121393535A

  • Biomarker of complete grape tire

    CN121559090A