Predictive model for simultaneous detection of multiple tumors and histological tracing, training method and application thereof
By combining biomarkers and machine learning models, the sensitivity and specificity issues of early cancer screening have been addressed, enabling efficient and low-cost detection and accurate tissue tracing for various cancers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2026-03-17
AI Technical Summary
Existing cancer detection methods are mostly for detecting a single type of cancer, making it difficult to achieve early screening for multiple cancers. They also suffer from problems such as low sensitivity, poor specificity, high cost, and high invasiveness.
By combining biomarkers such as nucleosome distribution, fragment size distribution, terminal sequence distribution, genomic instability, gene expression prediction, and somatic copy number variation, and using low-depth whole-genome sequencing data, machine learning models are constructed to detect various tumors and trace their origins.
It achieves simultaneous detection of multiple cancers with high sensitivity, high specificity, and low cost, and can accurately trace the primary tissue of the tumor.
Smart Images

Figure CN116665771B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical technology, and more specifically, to a predictive model for simultaneously detecting multiple tumors and tracing their origins, as well as its training method and application. Background Technology
[0002] Liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, stomach cancer, ovarian cancer, and nasopharyngeal cancer are all cancers with high incidence and mortality rates.
[0003] For liver cancer, the commonly used clinical detection methods mainly include liver cancer marker detection, imaging examinations, and histopathological examination. Liver cancer marker detection: This mainly detects liver cancer markers in serum, such as AFP and DCP. This method is simple, convenient, and non-invasive, but its sensitivity and specificity are relatively low, which may lead to misdiagnosis or missed diagnosis. Imaging examinations: Such as ultrasound, CT, and MRI, can help doctors detect the presence, size, and location of tumors in the liver. This method can detect early lesions of liver cancer, but the sensitivity and specificity vary among different examination methods. Histopathological examination: This involves obtaining tumor tissue through biopsy and other methods for pathological examination and histological analysis. This method can provide the most accurate diagnostic results, but it requires resection or needle biopsy, which carries certain risks.
[0004] Colorectal cancer is a common malignant tumor with an increasing incidence rate. Statistics show that approximately 1.8 million people worldwide are diagnosed with colorectal cancer each year, with up to 900,000 deaths. Colonoscopy is one of the most commonly used methods for colorectal cancer screening. It can detect early-stage colorectal cancer or precancerous lesions, helping to improve cure rates. However, colonoscopy requires bowel preparation and can cause some discomfort for patients during the procedure. Fecal occult blood testing is a simple screening method that can detect the presence of occult blood in stool, but it may miss some early-stage colorectal cancer lesions.
[0005] Esophageal cancer remains a common cancer worldwide. Statistics show that approximately 570,000 new cases of esophageal cancer are diagnosed globally each year, with about 509,000 deaths. Esophageal cancer is generally classified into two types: squamous cell carcinoma and adenocarcinoma, with squamous cell carcinoma being the most common. Current clinical detection methods include esophagoscopy, one of the most commonly used methods for esophageal cancer detection. Esophagoscopy allows direct observation of lesions inside the esophagus and enables biopsy to determine the type and extent of cancer cells. However, this method requires patients to undergo endoscopic examination, which carries certain risks and discomfort. Furthermore, endoscopic examinations are expensive, and some regions lack the necessary medical facilities and technical expertise to offer this examination.
[0006] Pancreatic cancer is a highly malignant tumor originating from cells within the pancreas. Symptoms typically appear only in its advanced stages. While its incidence is relatively low, its mortality rate is extremely high, often referred to as a "silent killer." Currently, commonly used clinical detection methods include: Abdominal imaging examinations such as ultrasound, CT, and MRI. These examinations can determine the location, size, and presence of metastasis of the tumor, but are not sensitive for detecting early-stage pancreatic cancer. Blood biomarker tests such as CEA and CA19-9. These tests can aid in diagnosis but lack specificity, leading to a high rate of misdiagnosis. Endoscopic examination: An endoscope can be inserted through the mouth or nostril to directly observe pancreatic lesions. This method has a high detection rate for early-stage lesions, but requires a specialist and can be invasive and uncomfortable for the patient. In summary, current clinical methods for detecting pancreatic cancer have limitations, making early detection of the tumor difficult.
[0007] Lung cancer is one of the most common cancers worldwide and a leading cause of death. It is estimated that approximately 2 million people die from lung cancer globally each year, with a significant proportion of these deaths occurring in China. Currently, the main methods for detecting lung cancer include CT scans, chest X-rays, and bronchoscopy. CT scans are the most commonly used method and can detect lung cancer at an early stage. However, CT scans may also detect nodules that are not lung cancer, which could lead to unnecessary anxiety and further investigation. In addition, some biomarkers (such as CEA and SCC) can also be used for lung cancer screening and diagnosis. However, lung cancer biomarkers are not always very accurate and may result in false positives or false negatives.
[0008] Stomach cancer is a common malignant tumor of the digestive system, with high incidence and mortality rates worldwide. Early-stage stomach cancer often has no obvious symptoms, while late-stage symptoms may include indigestion, abdominal pain, loss of appetite, and weight loss. Stomach cancer can usually be screened through gastroscopy or blood tests. Gastroscopy is currently one of the most commonly used methods for stomach cancer screening. It can detect early-stage stomach cancer or precancerous lesions, helping to improve cure rates. However, gastroscopy requires fasting and can cause some discomfort for the patient during the procedure. Serological tests can detect certain tumor markers (such as CA72-4 and CA19-9), but their specificity and sensitivity are not very high, which may lead to some degree of misdiagnosis or missed diagnosis. CT scans can provide more detailed images, helping to assess the size and location of stomach cancer, but their effectiveness in screening for early-stage stomach cancer is not ideal.
[0009] Ovarian cancer is one of the most deadly cancers of the female reproductive system. Its incidence is relatively high, and it primarily affects women over 40. According to the World Health Organization, ovarian cancer is the seventh leading cause of cancer death among women worldwide. In the United States, it is the fifth leading cause of cancer death among women, with more than 20,000 women diagnosed with it each year. Currently, early screening and detection methods for ovarian cancer are not yet fully developed. Commonly used clinical methods include tumor marker testing: CA125 is one of the markers for ovarian cancer, but it is not highly specific and has a certain false positive and false negative rate. Ultrasound examination: Ultrasound examination is a very common method, but its accuracy is also relatively limited and cannot accurately diagnose all ovarian cancers. CT / MRI examination: CT / MRI examinations can more accurately determine the location, size, and extent of ovarian tumors, but they are expensive and not suitable for large-scale screening.
[0010] Nasopharyngeal carcinoma is a malignant tumor of the head and neck that seriously threatens people's lives and health. Nasopharyngoscopy is the primary method for early detection of nasopharyngeal carcinoma, offering advantages such as being non-invasive, simple, and accurate. However, it may miss cases with deep tumors or inconspicuous lesions. Imaging examinations can provide comprehensive tumor information but cannot determine the degree of malignancy. Serological testing can aid in early diagnosis, but relevant biomarkers are not detectable in the serum of all nasopharyngeal carcinoma patients.
[0011] Currently, most cancer or tumor detection products used in clinical practice are limited to a single specific type of cancer or tumor. The rise of multi-cancer early detection (MCED) technology is of epoch-making significance for the development of cancer or tumor screening and diagnosis. Compared with single-cancer or single-tumor detection, multi-cancer or multiple-tumor detection can achieve the detection of multiple cancers (including multiple cancers or tumors for which there are currently no recommended screening methods) in a single test, and has become an inevitable trend for the future development of the industry.
[0012] In view of this, the present invention is proposed. Summary of the Invention
[0013] The purpose of this invention is to provide a predictive model for simultaneously detecting multiple tumors and tracing their tissue origin, as well as its training method and application.
[0014] This invention is implemented as follows:
[0015] In a first aspect, embodiments of the present invention provide the application of reagent combinations for detecting biomarkers or combinations thereof in the preparation of products for detecting or assisting in the detection of tumors. The biomarkers or combinations thereof include any one or more combinations of the following six categories of biomarkers: nucleosome distribution biomarkers, fragment size distribution biomarkers, terminal sequence distribution biomarkers, genomic instability biomarkers, gene expression prediction biomarkers, and somatic copy number variation biomarkers. The nucleosome distribution biomarkers use the nucleosome distribution characteristics corresponding to the target gene or its transcript as biomarkers, including those listed in Table 2. The table shows any combination of one or more markers from items 1 to 247; the fragment size distribution markers use the fragment size distribution characteristics within a specified window as markers, including any combination of one or more markers from items 1 to 335 shown in Table 3, where chr1:6000000:8000000_frac_P160_180 represents the distribution of reads with a length between 160 and 180 bp within the chr1:6000000:8000000 window, and so on; the terminal sequence distribution markers use sequences with a length ≤ 170 bp. p And aligned to the position 4b upstream of the reference genome. p The distribution characteristics of the target terminal sequence reads are used as markers, including any combination of one or more markers from items 1 to 77 in Table 4; the target sequence refers to the 5' end of the read sequence being aligned 4 bp upstream of the reference genome position; the genomic instability markers use the genomic instability characteristics within the target window as markers, including any combination of one or more markers from items 1 to 78 in Table 5; the gene expression prediction markers use the gene expression prediction characteristics of the target gene or its transcript as markers, including any combination of one or more markers from items 1 to 258 in Table 6; the somatic copy number variation markers use the somatic copy number variation characteristics of the target copy number variation region as markers, and the target copy number variation region includes any one or more CNV regions from items 1 to 300 in Table 8; the above sequence information is in h g 19 is used as a reference genome.
[0016] Secondly, embodiments of the present invention provide a training method for a predictive model for detecting or assisting in the detection of tumors, comprising: acquiring the detection result of each biomarker or combination thereof in a training sample and the corresponding annotation result; wherein the biomarker or combination thereof is the biomarker or combination thereof described in the foregoing embodiments, and the annotation result is a label representing whether the sample has a tumor; inputting the detection result of the biomarker or combination thereof into a pre-constructed predictive model to obtain a prediction result; the predictive model is a machine learning model capable of predicting whether the sample has the tumor based on the detection result of the biomarker or combination thereof; and updating the parameters of the predictive model based on the annotation result and the prediction result.
[0017] Thirdly, embodiments of the present invention provide a predictive device for detecting or assisting in the detection of tumors, comprising: an acquisition module for acquiring the detection result of each biomarker or combination thereof in a sample to be tested; wherein the biomarker or combination thereof is the biomarker or combination thereof described in the foregoing embodiments; and a prediction module for inputting the acquired detection results of the biomarker or combination thereof into a prediction model trained by the training method described in the foregoing embodiments to obtain a prediction result.
[0018] Fourthly, embodiments of the present invention provide the application of reagent combinations for detecting biomarkers or combinations thereof in the preparation of products for tracing the origin of primary tumor tissue. The biomarkers or combinations thereof include any one or more combinations of the following five types of biomarkers: nucleosome distribution biomarkers, fragment size distribution biomarkers, terminal sequence distribution biomarkers, genomic instability biomarkers, and gene expression prediction biomarkers as described in the foregoing embodiments.
[0019] Fifthly, embodiments of the present invention provide a kit comprising: a reagent combination for detecting biomarkers or combinations thereof as described in the foregoing embodiments.
[0020] Sixthly, embodiments of the present invention provide a training method for a predictive model for tracing the origin of primary tumor tissue, comprising: acquiring the detection results and corresponding annotation results of each biomarker or combination thereof in a training sample; wherein the biomarker or combination thereof is any one or more of the five types of biomarkers described in the foregoing embodiments, and the annotation result is a label representing the tissue tracing result of the tumor in the sample; inputting the detection results of the biomarker or combination thereof into a pre-constructed predictive model to obtain a prediction result; the predictive model is a machine learning model capable of performing tissue tracing of the tumor in the sample based on the detection results of the biomarker or combination thereof; and updating the parameters of the predictive model based on the annotation result and the prediction result.
[0021] In a seventh aspect, embodiments of the present invention provide a predictive device for tracing the origin of primary tumor tissue, comprising: an acquisition module for acquiring the detection result of each biomarker or combination thereof in a sample to be tested; wherein the biomarker or combination thereof is any one or more of the five types of biomarkers described in the foregoing embodiments; and a prediction module for inputting the acquired detection results of the biomarker or combination thereof into a prediction model trained by the training method described in the foregoing embodiments to obtain a prediction result.
[0022] Eighthly, embodiments of the present invention provide an electronic device including a processor and a memory; the memory is used to store a program, which, when executed by the processor, causes the processor to implement any one of: a method for detecting or assisting in the detection of tumors, a method for tracing the primary tumor tissue, a training method for a predictive model for detecting or assisting in the detection of tumors as described in the foregoing embodiments, and a training method for a predictive model for tracing the primary tumor tissue as described in the foregoing embodiments.
[0023] The method for detecting or assisting in the detection of tumors includes: obtaining the detection result of each biomarker or combination thereof in the sample to be tested; wherein the biomarker or combination thereof is any one or more of the six types of biomarkers described in the foregoing embodiments; inputting the obtained detection results of the biomarker or combination thereof into the prediction model trained by the training method described in the foregoing embodiments to obtain the prediction result; the method for detecting or assisting in the detection of tumors is not for the direct purpose of disease diagnosis or treatment.
[0024] The method for tracing the primary tumor tissue includes: obtaining the detection result of each biomarker or combination thereof in the sample to be tested; wherein, the biomarker or combination thereof is the biomarker or combination thereof described in the foregoing embodiments (any one or more of the five types of biomarkers); inputting the obtained detection results of the biomarker or combination thereof into the prediction model trained by the training method described in the foregoing embodiments to obtain the prediction result; the method for tracing the primary tumor tissue does not have the direct purpose of diagnosing or treating the disease.
[0025] Ninthly, embodiments of the present invention provide a computer-readable medium storing a computer program, wherein the computer program, when executed by a processor, implements any one of the following: a method for detecting or assisting in the detection of a tumor, a method for tracing the primary tumor tissue, a method for training a predictive model for detecting or assisting in the detection of a tumor as described in the foregoing embodiments, and a method for training a predictive model for tracing the primary tumor tissue as described in the foregoing embodiments; wherein the method for detecting or assisting in the detection of a tumor is the method for detecting or assisting in the detection of a tumor as described in the foregoing embodiments; and the method for tracing the primary tumor tissue is the method for tracing the primary tumor tissue as described in the foregoing embodiments.
[0026] The present invention has the following beneficial effects:
[0027] This invention discovers nucleosome distribution (TSSNDR), fragment size distribution (Fragment), terminal sequence distribution (Motif), genome instability (Bincount), gene expression prediction (GeneEXP), and somatic copy number variation (…). s CNA (Corrective Nucleotide) can be used as a biomarker to identify the presence of any one or more of the following cancers in a sample: liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma. It can also be used for tissue tracing and has advantages such as high detection sensitivity, good specificity, high accuracy in tissue tracing, and low detection cost. Attached Figure Description
[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 Image showing high expression of the EIF3I gene in esophageal cancer patients;
[0030] Figure 2 Images showing low expression of the KCNK2 gene in patients with liver cancer;
[0031] Figure 3 A method for constructing predictive models for tumor detection;
[0032] Figure 4 A method for constructing predictive models for tracing the origin of primary tumor tissue;
[0033] Figure 5 The receiver operating characteristic (ROC) curve of the predictive model used for tumor detection on the training set;
[0034] Figure 6 The receiver operating characteristic (ROC) curve for a predictive model used to detect tumors on the test set;
[0035] Figure 7 To evaluate the confusion matrix of cancer sample tissue origination for the predictive model used for tracing the origin of primary tumor tissue;
[0036] Figure 8 The confusion matrix is used to evaluate the tissue origination of cancer samples in the test set for the predictive model used for tracing the origin of primary tumor tissue. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased commercially.
[0038] This novel liquid biopsy method detects circulating cell-free DNA (cfDNA) in body fluids, primarily derived from fragmented DNA during apoptosis, DNA fragments from necrotic cells, and exosomes. Among these, circulating tumor DNA (ctDNA) is particularly important, originating from the tumor genome and carrying tumor-related information. Existing tumor or cancer detection methods largely fail to effectively identify or detect multiple tumors or cancers simultaneously. To address this, the inventors of this application, through a series of inventive efforts, utilized low-pass whole-genome sequencing (low-pass WGS) data to perform bioinformatics analysis on the test samples, identifying biomarkers or combinations thereof capable of simultaneously identifying multiple tumors, including malignant tumors, specifically one or more of liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma. Furthermore, in addition to identifying the tumor state, it allows for further tracing of the primary tumor tissue, offering advantages such as high detection sensitivity, good specificity, high accuracy in tissue tracing, and low detection cost.
[0039] On one hand, embodiments of the present invention provide the application of reagent combinations for detecting biomarkers or combinations thereof in the preparation of products for detecting or assisting in the detection of tumors. The biomarkers or combinations thereof include any one or more combinations of the following six classes of biomarkers: nucleosome distribution biomarkers (TSSNDR), fragment size distribution biomarkers (Fragment), terminal sequence distribution biomarkers (Motif), and genomic instability biomarkers (Bi). nc o un t), gene expression prediction biomarkers (G) eneEXP) and somatic copy number variation markers ( s CNA).
[0040] The nucleosome distribution markers use the nucleosome distribution characteristics corresponding to the target gene or its transcript as markers. Specifically, these markers include any one or more of the items 1 to 247 shown in Table 2.
[0041] In this article, "characteristics" can be used as a symbol or sign of the features of things.
[0042] In this article, "window" refers to the window in bioinformatics analysis. n do ws orbi n ).
[0043] The fragment size distribution markers refer to the fragment size distribution characteristics within a specified window. Specifically, these markers include any combination of one or more markers from items 1 to 335 shown in Table 3. Here, Chr: X1 to X2 represents the specified window, frac_PX3_X4 refers to the distribution characteristics of reads with lengths within X3 to X4, and X1 to X4 are positive integers (specific values are shown in the table). The distribution characteristics of reads can be understood as the number of reads within a specific length range within the specified window or the percentage of all reads within that window. For example, chr1:6000000:8000000_frac_P160_180 refers to the distribution characteristics of reads with lengths between 160 and 180 bp within the chr1:6000000:8000000 window; other fragment size distribution markers follow the same logic.
[0044] The terminal sequence distribution markers use the distribution characteristics of reads with a sequence length ≤170bp and a target terminal sequence 4bp upstream of the alignment position to the reference genome as markers. These markers include any one or more combinations of items 1-77 shown in Table 4. Specifically, "4bp upstream of the alignment position to the reference genome" means: after aligning the read sequence to the reference genome, recording the genomic coordinates of the 5' end of the read (e.g., chr1:1000000), and using the reference sequence (e.g., chr1:999996-999999) on the human reference genome as the "4bp upstream of the alignment position to the reference genome" or the "terminal sequence". The same operation can be performed on each read, and the number of reads with any type of terminal sequence can be accumulated.
[0045] The genomic instability biomarkers use genomic instability features within the target window as biomarkers, including any combination of one or more biomarkers in windows 1 to 78 shown in Table 5.
[0046] The gene expression prediction biomarkers use the gene expression prediction characteristics of the target gene or its transcript as biomarkers, including any combination of one or more biomarkers from items 1 to 258 shown in Table 6.
[0047] The somatic copy number variation markers use the somatic copy number variation characteristics of the target copy number variation region as markers. The target copy number variation region includes any one or more combinations of items 1 to 300 shown in Table 7.
[0048] In some embodiments, the number of somatic cell copy number variation markers is 1.
[0049] The reference genome for the above sequence information is hg19.
[0050] It should be noted that in the analysis of nucleosome distribution (TSSNDR), fragment size distribution (Fragment), terminal sequence distribution (Motif), genome instability (Bincount), and gene expression prediction (G... ene EXP) and somatic copy number variation ( s When the detection gene or detection region corresponding to a CNA biomarker is already publicly available, the detection method, corresponding detection reagent, and calculation method can be obtained based on existing knowledge.
[0051] In some embodiments, the biomarkers or combinations thereof include any combination of three, four, five or six of the six categories of biomarkers.
[0052] In some implementations, the nucleosome distribution characteristics of any gene or transcript are calculated as follows:
[0053]
[0054] In the formula, Coverage NDR This refers to nucleosome loss and nucleosome deletion regions; NDR stands for nucleosome deletion region, specifically the corrected coverage of the region between 200 bp upstream and 100 bp downstream of the TSS; Mean (Coverage) TSS1 'Coverage TSS2Mean represents the average corrected coverage of the region between 2000 bp upstream and 200 bp upstream of the TSS, and the region between 100 bp downstream and 2000 bp downstream of the TSS; the TSSNDR Score is a calculated score used to quantify the nucleosome distribution characteristics of a gene or transcript. In some implementations, Mean can be replaced by Median (i.e., the median of the corrected coverage of the region between 2000 bp upstream and 200 bp upstream of the TSS, and the median of the corrected coverage of the region between 100 bp downstream and 2000 bp downstream of the TSS).
[0055] In some implementations, the method for calculating genomic instability features within any window can be selected from: a Z-value algorithm based on sequencing depth, a log2ratio algorithm based on control samples, an algorithm based on soft-clipped reads, and BinCount. i Any of the algorithms.
[0056] In some implementations, BinCount i The calculation formula is as follows:
[0057] Among them, BinCount i Genomic instability scoring; Fragment i is the number of reads within the i-th window; TotalMappedFragments is the total number of reads in the test sample aligned to the reference genome, preferably the total number of reads remaining after excluding reads aligned to the X, Y sex chromosomes, mitochondria, and other contigs; WindowLength i BinCount represents the length of the i-th window. Compared to other existing calculation methods, BinCount... i The inventors' optimized computational method can detect genomic instability in samples with low tumor concentrations.
[0058] In some implementations, the method for obtaining gene expression prediction features of any gene or transcript includes: for any target gene or transcript, predicting whether the gene is highly or poorly expressed by using the length and position information of reads around the gene transcription start site.
[0059] In some implementations, the acquisition of the gene expression prediction features includes: for any gene or transcript, extracting the positional information and length information of reads at various locations within the region between 1250 bp upstream and 1250 bp downstream of the TSS, constructing a three-dimensional vector, where the first dimension X-axis is the position index of each location relative to the TSS site within the region between 1250 bp upstream and 1250 bp downstream of the TSS, the second dimension Y-axis is the length information of reads from 0 bp to 400 bp, and the third dimension Z-axis is a normalized value that, after scaling with the second dimension using the Z-value, represents the severity of clustering of reads of arbitrary length at that location. Dimensionality reduction is performed on the normalized value of the severity of clustering of reads of arbitrary length at that location to obtain a three-dimensional matrix of the gene or transcript; the three-dimensional matrix of high-expression genes and low-expression genes with known gene expression levels (training set) is used as input values into a polynomial regression. (Regression) training yields a model capable of predicting the expression score or status of a gene or its transcript. A three-dimensional matrix of the target gene or its transcripts is input into the model to obtain the expression status of the target gene or its transcripts. Specifically, "high-expressing and low-expressing genes with known expression levels" can be obtained from known samples and / or public databases, such as downloading pan-cancer gene expression data (Batch effects normalized mRNA data*(n=11,060)) from the XENA database (Pan-Cancer AtlasHub). This allows obtaining the expression levels of any gene and all its transcripts, which are then used as learning labels for linear regression. In this training set, the number of high-expressing and low-expressing genes can be any one or any two of the following: 10, 20, 50, 100, 200, 300, and 400.
[0060] In some implementations, the dimensionality reduction method includes wavelet compression algorithms.
[0061] In some implementations, the somatic cell copy number variation features (comprehensive statistical analysis of the target copy number variation region; in this embodiment, this type of marker corresponds to a single result) are obtained as follows:
[0062]
[0063] In the formula, i represents the i-th CNV region (1≤i≤m); j represents the j-th gene in the i-th CNV region (1≤j≤n); n represents the total number of proto-oncogenes and tumor suppressor genes in the CNV region; and m represents the total number of all CNV regions in the subject. This represents the weight of the j-th tumor suppressor gene within the i-th CNV region in the tumor suppressor gene database; This represents the actual observed copy number of the j-th tumor suppressor gene within the i-th CNV region; This represents the weight of the j-th proto-oncogene within the i-th CNV region in the proto-oncogene database; This represents the actual observed copy number of the j-th proto-oncogene within the i-th CNV region. The weights of proto-oncogenes, tumor suppressor genes, and tumor suppressor genes in the tumor suppressor gene database, as well as the weights of proto-oncogenes in the proto-oncogene database, can be obtained from existing publicly available literature or databases.
[0064] In some implementations, the fragment size distribution feature (fragment size distribution feature within a certain window) is obtained as follows: the number of read segments with a length within a limited range within the limited window or the proportion of the number of read segments to the total number of read segments within the limited window is counted.
[0065] In some embodiments, the terminal sequence distribution markers (the distribution characteristics of reads whose sequences aligned to the reference genome within 4 bp of an arbitrary target terminal sequence) are obtained as follows: The number of reads with a sequence length ≤170 bp and aligned to the reference genome within 4 bp of a specific target terminal sequence, or the percentage of such reads relative to the total number of reads, is obtained. This total number of reads is the total number of all reads aligned to the reference genome in the sample, preferably the total number of reads remaining after excluding reads aligned to the X, Y sex chromosomes, mitochondria, and other contigs.
[0066] In some embodiments, the tumor includes malignant tumors, specifically including any one or more combinations of liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma. The specific selection of tumors in subsequent embodiments or implementations is the same as that described in this embodiment, and will not be repeated here.
[0067] In some implementations, the product is also used for tracing the origin of primary tumor tissue.
[0068] In some implementations, the product includes any one of reagents, reagent combinations, kits, and predictive models.
[0069] On the other hand, embodiments of the present invention provide a method for training a predictive model for detecting or assisting in the detection of tumors, comprising:
[0070] Obtain the detection results and corresponding annotation results of each biomarker or combination thereof in the training sample; wherein, the biomarker or combination thereof is the biomarker or combination thereof described in any of the foregoing embodiments (any one or more combinations of the six types of biomarkers), and the annotation result is a label representing whether the sample has a tumor;
[0071] The detection results of the biomarker or its combination are input into a pre-built prediction model to obtain the prediction result; the prediction model is a machine learning model that can predict whether a sample has the tumor based on the detection results of the biomarker or its combination.
[0072] The parameters of the prediction model are updated based on the annotation results and the prediction results.
[0073] In some implementations, the label can be a character or a string.
[0074] It is understood that the category and number of training samples are conventionally selectable by those skilled in the art. Optionally, the number of training samples may be greater than or equal to any value among 10, 50, 100, 200, 300, 400 and 500 or any range between two of them. Optionally, the sources of training samples include cancer patients and healthy individuals.
[0075] In some embodiments, the test sample or the training sample is independently selected from: plasma samples, serum samples, whole blood samples, negative standards, or positive standards. Optionally, the test sample or the training sample may also be selected from: environmental samples containing at least one of plasma samples or serum samples.
[0076] In some implementations, the machine learning model includes one or more classifiers. Both single-layer and multi-layer classifiers can detect or identify multiple types of tumors. A single-layer classifier can also detect multiple tumors simultaneously or trace their origin, but multi-layer classifiers generally offer better predictive sensitivity and accuracy.
[0077] In some embodiments, the multi-layer classifier includes: a first-layer classifier to an nth-layer classifier, where n is a positive integer greater than 1. n can specifically be any one or a range between 2, 3, 4, and 5. In some embodiments, for a predictive model for detecting or assisting in the detection of tumors, each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number) independently corresponds to at least one first-layer classifier. The input to the first-layer classifier is the detection result of each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number), and the output of the first-layer classifier is the classification result (first-layer score) of whether the sample has the tumor.
[0078] The n-layer classifier integrates the classification results of all categories of markers and the detection results of somatic cell copy number to output the final prediction result. Specifically, the input of the n-th layer classifier is: the detection result of the somatic cell copy number marker, and the output of the (n-1)-th layer classifier (n-1-th layer score) for all other categories of markers except somatic cell copy number. The output of the n-th layer classifier is: the final prediction result.
[0079] In some implementations, when n≥3, in addition to the first-layer classifier and the nth-layer classifier, an intermediate-layer classifier (n-1th-layer classifier) is also included. Each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number) independently corresponds to at least one intermediate-layer classifier. The input of the intermediate-layer classifier is the classification result of the previous-layer classifier on whether the sample has a tumor. The output of the intermediate-layer classifier for each biomarker is the classification result of that layer classifier on whether the sample has a tumor (intermediate-layer score).
[0080] In some embodiments, the multi-layer classifier includes a first-layer classifier, a second-layer classifier, a third-layer classifier, and a fourth-layer classifier. In application, the detection results of each biomarker or combination thereof, excluding somatic cell copy number, are input into the first-layer classifier to classify whether the sample has the tumor, obtaining a first-layer score for that biomarker. The first-layer scores of each biomarker are input into the second-layer classifier to classify whether the sample has the tumor, obtaining a second-layer score for that biomarker. The second-layer scores of each biomarker are input into the third-layer classifier to classify whether the sample has the tumor, obtaining a third-layer score for that biomarker. The third-layer scores of all biomarker categories and the detection results of the somatic cell copy number biomarker are input into the fourth-layer classifier for integration to obtain the final prediction result.
[0081] In some implementations, the first-layer classifier and / or the second-layer classifier includes any one or more combinations of Kneighbors, LightGBM, SVM, SGD, MLP, and Gaussian Process. The first-layer classifier and the second-layer classifier do not use the same algorithm model. Furthermore, there are no particular limitations on the order and combination of Kneighbors, LightGBM, SVM, SGD, MLP, and Gaussian Process for the first-layer classifier and / or the second-layer classifier; a preferred combination is KNN+LightGBM+SVM for the first layer and SGD+MLP+Gaussian Process for the second layer.
[0082] When the number of classifiers in any layer from the first layer to the (n-1)th layer is ≥2, the input and output of the classifier in that layer, as defined above, are the input and output of each classifier in that layer. For example, when the number of first-layer separators is ≥2, the step of inputting the first-layer classifier with the detection results of each marker except for somatic cell copy number includes: inputting the detection results of each marker into each first-layer classifier to obtain the first-layer score corresponding to each first-layer classifier. As another example, when the number of classifiers in both the first-layer classifier and the second-layer separator is ≥2, the step of inputting the first-layer score of each marker into the second-layer classifier includes: inputting the first-layer score corresponding to each first-layer classifier of each marker into each second-layer classifier to obtain the second-layer score corresponding to each second-layer classifier.
[0083] In some implementations, the first-layer classifier includes Kneighbors, LightGBM, and SVM.
[0084] In some implementations, the second-layer classifier includes SGD, MLP, and Gaussian Process.
[0085] In some implementations, the third-layer classifier includes a logistic regression model.
[0086] In some implementations, the fourth classifier includes a logistic regression model. Specifically, it can be a binary logistic regression model.
[0087] On the other hand, embodiments of the present invention provide a predictive device for detecting or assisting in the detection of tumors, comprising:
[0088] The acquisition module is used to acquire the detection result of each biomarker or combination thereof in the sample to be tested; wherein, the biomarker or combination thereof is any one or more of the six types of biomarkers described in any of the foregoing embodiments;
[0089] The prediction module is used to input the detection results of the obtained biomarkers or combinations thereof into the prediction model trained by the training method of the prediction model for detecting or assisting in detecting tumors as described in any of the foregoing embodiments, and obtain the prediction results.
[0090] In some embodiments, the modules described in this invention can be stored in memory or embedded in the operating system (OS) of the electronic device provided in this application, and can be executed by the processor in the electronic device. Meanwhile, the data, program code, etc., required to execute the above modules can be stored in memory.
[0091] On the other hand, embodiments of the present invention also provide the application of reagent combinations for detecting biomarkers or combinations thereof in the preparation of products for tracing the origin of primary tumor tissue. The biomarkers or combinations thereof include any one or more combinations of the following five types of biomarkers: nucleosome distribution biomarkers, fragment size distribution biomarkers, terminal sequence distribution biomarkers, genomic instability biomarkers, and gene expression prediction biomarkers as described in any of the foregoing embodiments.
[0092] On the other hand, embodiments of the present invention also provide a kit comprising: a combination of reagents for detecting biomarkers or combinations thereof as described in any of the foregoing embodiments.
[0093] On the other hand, embodiments of the present invention also provide a training method for a predictive model for tracing the origin of primary tumor tissue, comprising:
[0094] Obtain the detection results and corresponding annotation results of each biomarker or combination thereof in the training sample; wherein, the biomarker or combination thereof is the biomarker or combination thereof described in any of the foregoing embodiments (any one or more combinations of the five types of biomarkers), and the annotation result is a label representing the tissue origination result of the tumor in the sample;
[0095] The detection results of the biomarker or its combination are input into a pre-built prediction model to obtain the prediction results; the prediction model is a machine learning model that can trace the tissue origin of tumors in a sample based on the detection results of the biomarker or its combination.
[0096] The parameters of the prediction model are updated based on the annotation results and the prediction results.
[0097] It is understood that the type and number of training samples are conventionally selectable by those skilled in the art, and the number of training samples can be any value or range between any two of ≥10, 50, 100, 200, 300, 400 and 500; the source of training samples includes cancer patients; in optional embodiments, the source of training samples may also include healthy people.
[0098] In some embodiments, the test sample or the training sample is independently selected from: plasma samples, serum samples, whole blood samples, negative standards, or positive standards. Optionally, the test sample or the training sample may also be selected from: environmental samples containing at least one of plasma samples or serum samples.
[0099] In some implementations, the machine learning model includes one or more classifiers.
[0100] In some implementations, the prediction model and training method for tracing the origin of primary tumor tissue are largely the same as the prediction model and training method for detecting or assisting in the detection of tumors described in any of the foregoing embodiments. The main difference is that the prediction model for tracing the origin of primary tumor tissue does not use biomarkers such as somatic cell copy number as feature inputs, as such biomarkers are not specific for tracing the origin of tumor tissue.
[0101] In some embodiments, the multi-layer classifier includes: a first-layer classifier to an nth-layer classifier, where n is a positive integer greater than 1. n is the same as described in any of the foregoing embodiments and will not be repeated here.
[0102] In some implementations, for the predictive model for tracing the primary tumor tissue, each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number) independently corresponds to at least one first-layer classifier. The input to the first-layer classifier is the detection result of each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number), and the output of the first-layer classifier is the result of tracing the primary tumor tissue of the sample (first-layer score).
[0103] The input to the nth layer classifier is the output of the (n-1)th layer classifier (n-1th layer score) for all categories of markers except somatic cell copy number. The output of the nth layer classifier is the final prediction result.
[0104] In some implementations, when n≥3, in addition to the first-layer classifier and the nth-layer classifier, an intermediate-layer classifier (n-1th-layer classifier) is also included. Each biomarker (any one or more of the six biomarkers or combinations thereof, excluding somatic cell copy number) independently corresponds to at least one intermediate-layer classifier. The input of the intermediate-layer classifier is the tumor origin tracing result of the previous-layer classifier for the sample. The output of the intermediate-layer classifier for each biomarker is the tumor origin tracing result of that classifier for the sample (intermediate-layer score).
[0105] In the predictive model for tracing the primary tumor site, the specific results of each layer or a certain layer of classifier for tracing the primary tumor site of a sample can be: the probability or score of the sample having each type of tumor. The classifier in the predictive model for tracing the primary tumor site is mainly based on the traditional binary classifier, using a one-against-other (One vs. Rest / OvR) strategy for multi-classification, thereby realizing tissue tracing, i.e., locating the primary organ of the tumor signal.
[0106] In some implementations, for the predictive model of primary tumor tissue tracing, the multi-layer classifier includes: a first-layer classifier, a second-layer classifier, a third-layer classifier, and a fourth-layer classifier. In application, the detection results of each biomarker or combination thereof, excluding somatic cell copy number, are input into the first-layer classifier to perform tissue tracing of the tumor in the sample, obtaining a first-layer score for that biomarker. The first-layer scores of each biomarker are then input into the second-layer classifier to perform tissue tracing of the tumor in the sample, obtaining a second-layer score for that biomarker. The second-layer scores of each biomarker are then input into the third-layer classifier to perform tissue tracing of the tumor in the sample, obtaining a third-layer score for that biomarker. Finally, the third-layer scores of all biomarker categories are input into the fourth-layer classifier for integration to obtain the final prediction result.
[0107] In some implementations, the first-layer classifier and / or the second-layer classifier includes any one or more combinations of Kneighbors, LightGBM, SVM, SGD, MLP, and Gaussian Process. When the number of classifiers in the first to (n-1)th layers is ≥2, the input and output of the defined layer classifier are the input and output of each of the classifiers in that layer.
[0108] In some implementations, the first-layer classifier includes Kneighbors, LightGBM, and SVM.
[0109] In some implementations, the second-layer classifier includes SGD, MLP, and Gaussian Process.
[0110] In some implementations, the third-layer classifier includes a logistic regression model.
[0111] In some implementations, the fourth classifier includes a logistic regression model, specifically a multivariate logistic regression model.
[0112] On the other hand, embodiments of the present invention also provide a predictive device for tracing the origin of primary tumor tissue, comprising:
[0113] The acquisition module is used to acquire the detection result of each biomarker or combination thereof (any one or more of the six types of biomarkers or combinations thereof except somatic cell copy number) in the sample to be tested; wherein, the biomarker or combination thereof is the biomarker or combination thereof described in any of the foregoing embodiments;
[0114] The prediction module is used to input the detection results of the obtained biomarkers or combinations thereof into the prediction model trained by the training method of the prediction model for tracing the primary tumor tissue described in any of the foregoing embodiments, and obtain the prediction result.
[0115] On the other hand, embodiments of the present invention also provide an electronic device, which includes a processor and a memory; the memory is used to store a program, which, when executed by the processor, causes the processor to implement any one of: a method for detecting or assisting in the detection of tumors, a method for tracing the primary tumor tissue, a training method for a predictive model for detecting or assisting in the detection of tumors as described in any of the foregoing embodiments, and a training method for a predictive model for tracing the primary tumor tissue as described in any of the foregoing embodiments.
[0116] The methods for detecting or assisting in the detection of tumors include:
[0117] Obtain the detection result of each biomarker or combination thereof in the sample to be tested; wherein, the biomarker or combination thereof is any one or more of the six types of biomarkers described in any of the foregoing embodiments;
[0118] The detection results of the obtained biomarkers or combinations thereof are input into the prediction model trained by the training method of the prediction model for detecting or assisting in detecting tumors as described in any of the foregoing embodiments to obtain the prediction results.
[0119] The methods for detecting or assisting in the detection of tumors are not intended for the direct purpose of diagnosing or treating the disease.
[0120] The method for tracing the origin of the primary tumor tissue includes:
[0121] Obtain the detection result of each biomarker or combination thereof in the sample to be tested; wherein, the biomarker or combination thereof is any one or more of the six types of biomarkers or combinations thereof except somatic cell copy number as described in any of the foregoing embodiments;
[0122] The detection results of the obtained biomarkers or combinations thereof are input into the training method of the prediction model for tracing the primary tumor tissue described in any of the foregoing embodiments to obtain the prediction results.
[0123] The method for tracing the origin of the tumor is not for the direct purpose of diagnosing or treating the disease.
[0124] Electronic devices may include memory, processor, bus, and communication interface, which are electrically connected directly or indirectly to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more buses or signal lines.
[0125] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.
[0126] The processor can be an integrated circuit chip with signal processing capabilities. The processor 120 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0127] The electronic device can be a server, cloud platform, mobile phone, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), handheld computer, netbook, personal digital assistant (PDA), wearable electronic device, virtual reality device, etc. Therefore, the embodiments of this application do not limit the types of electronic devices.
[0128] Furthermore, embodiments of the present invention also provide a computer-readable medium storing a computer program, which, when executed by a processor, implements any one of the following: a method for detecting or assisting in the detection of tumors, a method for tracing the primary tumor tissue, a training method for a predictive model for detecting or assisting in the detection of tumors as described in any of the foregoing embodiments, and a training method for a predictive model for tracing the primary tumor tissue as described in any of the foregoing embodiments. Wherein, the method for detecting or assisting in the detection of tumors is the method for detecting or assisting in the detection of tumors as described in any of the foregoing embodiments; the method for tracing the primary tumor tissue is the method for tracing the primary tumor tissue as described in any of the foregoing embodiments.
[0129] Computer-readable media can be general-purpose storage media, such as removable disks and hard drives.
[0130] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0131] The clinical sample information used in the examples is shown in Table 1.
[0132] Table 1. Basic Information of Clinical Samples
[0133]
[0134]
[0135]
[0136] Example 1
[0137] 1. Low-pass whole-genome sequencing (WGS) library preparation and sequencing
[0138] (a) Extraction of cell-free DNA (cfDNA) from blood: Take 2-3 mL of peripheral blood (collected and stored in Streck cell-free DNA collection tubes), and centrifuge at 1600 g for 10 min at 4°C using an Eppendorf centrifuge (5810R and 5427R, German). Collect only the supernatant. Then centrifuge at 16000 g for 10 min and collect the supernatant to obtain the plasma sample. Extract cell-free DNA from the plasma using the MagMAX Cell-Free DNA Isolation Kit (Thermo) and a nucleic acid extractor (Thermo Kingfisher FLEX, USA). (b) DNA quality testing: DNA concentration was detected using a Qubit 3 nucleic acid / protein quantitative fluorometer (Thermo, USA), and DNA fragment distribution was detected using a Fragment Analyzer (Agilent, USA). (c) WGS Library Construction: Approximately 5 ng of cfDNA was used to construct a pre-library using a kit from Enzymatics (USA). This mainly involved two steps: end repair (5X ER / A-Tailing Enzyme Mix) and adapter ligation (WGS Ligase). The adapter sequences were suitable for the Illumina NovaSeq 6000 sequencing platform. After adapter ligation, purification was performed using XP magnetic beads (AgencourtAMPure XP beads, Beckman Coulter). The WGS library concentration was determined using qPCR (KAPA Library QuantKit, Roche), and the library size was determined using Fragment Analyzer (Agilent, USA). (d) Sequencing: Paired-end sequencing of 150 bp was performed on the Illumina NovaSeq 6000 sequencing platform. The average data volume per sample was 0.5 × 10⁻⁶ for the entire genome.
[0139] 2. Basic Analysis of Low-pass WGS Data
[0140] (a) Sequencing data was converted from BCL files to FASTQ files using bcl2fastq software; (b) CutAdapt software was used to remove low-quality reads, reads containing more than 5% N bases, and reads shorter than 50 bp from the sequencing data; (c) The above reads were aligned to the hg19 reference genome using bwa software to remove redundant PCR sequences; (d) SAMTools was used to remove DNA fragments with low alignment quality, those that were not aligned, and those with mismatched paired ends. The filtered DNA fragments were then sorted according to their alignment positions.
[0141] 3. Data analysis of six biomarkers used to identify the presence of any one of the following cancers in a test sample: liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma, and to trace the tissue origin of the cancer.
[0142] 3.1 Nucleosome Distribution: Nucleosomes are basic structural units of chromatin formed by DNA and histones, protecting DNA structure from damage by endogenous nucleases. Nucleosome distribution differs between regions of high transcriptional activity and low transcriptional activity. In regions with low transcriptional expression activity, nucleosomes are densely packed. However, in tissues with malignant tumors, chromosomal structures are usually loosely distributed, and the number of nucleosomes is sparse in regions of high gene transcriptional expression. Near transcriptionally active TSSs, nucleosomes exhibit certain localization characteristics: the promoter region of a gene contains a nucleosome deletion region (NDR), and the first nucleosome downstream of this region (+1 nucleosome) is stably located, with numerous histone modifications and histone variants. A series of regularly located nucleosomes exist downstream of this region. The barrier effect of the NDR causes downstream nucleosomes to locate based on this barrier, with the barrier effect weakening with distance. This combination of factors results in regular nucleosome localization. This invention uses whole-genome sequencing of cfDNA fragments to calculate the coverage of background regions and NDR regions at both ends of the TSS of each gene transcript, thereby inferring nucleosome localization and gene expression. Specifically, this is achieved through the following formula:
[0143]
[0144] In the formula, Coverage NDR This refers to nucleosome loss and deletion regions. NDR is a typical representative of most active eukaryotic promoters. Specifically, in this case, it refers to the corrected coverage of the region between 200 bp upstream and 100 bp downstream of the TSS. Mean (Coverage) TSS1 'Coverage TSS2 The TSSNDR Score refers to the average corrected coverage of the region between 2000bp upstream and 200bp upstream of the TSS, and the region between 100bp downstream and 2000bp downstream of the TSS; the TSSNDR Score refers to the score obtained from the above calculation process for quantifying the nucleosome distribution characteristics of genes or transcripts.
[0145] 3.2 Fragment Size Distribution: There are differences in cfDNA fragment omics characteristics between cancer patients and healthy individuals. Tumor-derived cfDNA fragments are shorter, and cancer patients exhibit higher diversity in the terminal base sequences of cfDNA. The genome-wide distribution and fragmentation of plasma cfDNA are not random. Plasma cfDNA fragments have a pattern size of 167 bp, do not degrade within nucleosomes, and nucleosome footprints from the source tissue can be captured at the beginning and end of the fragments. In cancer patients, these observations may aid in cancer detection, inferring the source tissue and gene expression. Furthermore, deviations from expected fragment size and location can be used to improve the signal-to-noise ratio of somatic genomic alterations in plasma cfDNA. Specifically, this is achieved in this embodiment as follows: the reads of the BAM file of the sample to be tested are restored to the original DNA template (fragment), the entire genome is divided into equal-length windows of 2M (two million) bases in length, and X, Y sex chromosomes, mitochondria, and other contigs are excluded. The number of read segments with lengths in the ranges of 20bp–150bp, 160bp–180bp, 100bp–150bp, 180bp–220bp, and 250bp–320bp within each window was counted separately. The ratios of the number of read segments with lengths in the range of 20bp–150bp to the number of read segments with lengths in the range of 160bp–180bp, the number of read segments with lengths in the range of 100bp–150bp to the number of read segments with lengths in the range of 163bp–169bp, and the number of read segments with lengths in the range of 20bp–150bp to the number of read segments with lengths in the range of 180bp–220bp were also counted. Different samples have varying library sizes, therefore, a homogenization operation needs to be performed within the samples. Specifically, this involves: dividing the number of reads within the window with a length range of 20bp to 150bp by the total number of fragments within that window to obtain P20_150; dividing the number of reads within the window with a length range of 160bp to 180bp by the total number of fragments within that window to obtain P160_180; and dividing the number of reads within the window with a length range of 100bp to 150bp by the total number of fragments within that window. The number of reads within a window with a length between 20bp and 150bp is divided by the total number of fragments within that window to obtain P20_150. The number of reads within a window with a length between 180bp and 220bp is divided by the total number of fragments within that window to obtain P180_220. The number of reads within a window with a length between 250bp and 320bp is divided by the total number of fragments within that window to obtain P250_320.Additionally, P20_150_vs_P160_180 is obtained by dividing the number of read segments with a length in the range of 20bp to 150bp by the number of read segments with a length in the range of 160bp to 180bp; P100_150_vs_P163_169 is obtained by dividing the number of read segments with a length in the range of 100bp to 150bp by the number of read segments with a length in the range of 163bp to 169bp; and P20_150_vs_P180_220 is obtained by dividing the number of read segments with a length in the range of 20bp to 150bp by the number of read segments with a length in the range of 180bp to 220bp.
[0146] 3.3 Terminal Sequence Distribution: The terminal base sequences (Motifs) of cfDNA reveal characteristics of a non-random fragmentation process, which may be related to tissue origin, disease state, nucleosome openness, and endonuclease activity. Characteristic size patterns of cfDNA from different tissues indicate that non-random DNA fragmentation occurs during genomic DNA entry into the bloodstream. Based on ultra-deep sequencing results of maternal plasma cfDNA, a subset of cfDNA fragments consistently appear at certain locations in the genome; these preferentially occurring cfDNA fragment ends are termed "Preferred Motifs." Preferred Motifs exhibit tissue specificity in terms of cfDNA origin; for example, in maternal cfDNA, cfDNA carrying preferred Motifs is shorter and positively correlated with the proportion of fetal origin. Furthermore, preferred Motifs appear in columns, closely resembling nucleosome patterns, and are widely distributed across various tissues, such as liver-specific preferred Motifs in the plasma of liver transplant patients and tumor-specific preferred Motifs in patients with primary liver cancer. Specifically, this is achieved as follows: The reads of the BAM file of the sample to be tested are restored to the original DNA template (fragment), and reads that align to the X, Y chromosome, mitochondria, and other contigs of the reference genome (in this embodiment, the reference genome version is hg19) are excluded. The 5' ends of all reads ≤170bp are counted to the position 4bp upstream of the reference genome position, including AAAA, AAAT, AAAC, AAAG, AATA, AATT, AATC, AATG, AACA, AACT, AACC, AACG, AAGA, AAGT, AAGC, AAGG, ATAA, ATAT, ATAC, ATAG, ATTA, ATTT, ATTC, ATTG, ATCA, ATCT, ATCC, ATCG, ATGA, ATGT, ATGC, ATGG, ACAA, ACAT, ACAC, ACAG, ACTA, ACTT, ACT. C. ACTG, ACCA, ACCT, ACCC, ACCG, ACGA, ACGT, ACGC, ACGG, AGAA, AGAT, AGAC, AGAG, AGTA, AGTT, AGTC, AGTG, AGCA, AGCT, AGCC, AGCG, AGGA, AGGT, A GGC, AGGG, TAAA, TAAT, TAAC, TAAG, TATA, TATT, TATC, TATG, TACA, TACT, TACC, TACG, TAGA, TAGT, TAGC, TAGG, TTAA, TTAT, TTAC, TTAG, TTTA, TTTT,TTTC, TTTG, TTCA, TTCT, TTCC, TTCG, TTGA, TTGT, TTGC, TTGG, TCAA, TCAT, TCAC, TCAG, TCTA, TCTT, TCTC, TCTG, TCCA, TCCT, TCCC, TCCG, TCGA, T CGT, TCGC, TCGG, TGAA, TGAT, TGAC, TGAG, TGTA, TGTT, TGTC, TGTG, TGCA, TGCT, TGCC, TGCG, TGGA, TGGT, TGGC, TGGG, CAAA, CAAT, CAAC, CAAG, CAT A. CATT, CATC, CATG, CACA, CACT, CACC, CACG, CAGA, CAGT, CAGC, CAGG, CTAA, CTAT, CTAC, CTAG, CTTA, CTTT, CTTC, CTTG, CTCA, CTCT, CTCC, CTCG, CTGA, CTGT, CTGC, CTGG, CCAA, CCAT, CCAC, CCAG, CCTA, CCTT, CCTC, CCTG, CCCA, CCCT, CCCC, CCCG, CCGA, CCGT, CCGC, CCGG, CGAA, CGAT, CGAC, CG AG, CGTA, CGTT, CGTC, CGTG, CGCA, CGCT, CGCC, CGCG, CGGA, CGGT, CGGC, CGGG, GAAA, GAAT, GAAC, GAAG, GATA, GATT, GATC, GATG, GACA, GACT, GACC , GACG, GAGA, GAGT, GAGC, GAGG, GTAA, GTAT, GTAC, GTAG, GTTA, GTTT, GTTC, GTTG, GTCA, GTCT, GTCC, GTCG, GTGA, GTGT, GTGC, GTGG, GCAA, GCAT, G Any one of the following motifs—CAC, GCAG, GCTA, GCTT, GCTC, GCTG, GCCA, GCCT, GCCC, GCCG, GCGA, GCGT, GCGC, GCGG, GGAA, GGAT, GGAC, GGAG, GGTA, GGTT, GGTC, GGTG, GGCA, GGCT, GGCC, GGCG, GGGA, GGGT, GGGC, and GGGG—is divided by the total number of reads remaining after excluding those aligned to the X, Y, mitochondria, and other contigs in the sample being tested. This yields the percentage of each of the 256 motifs in the sample being tested.
[0147] 3.4. Genomic Instability: Genomic instability and mutation are hallmarks of cancer and are also considered to promote other cancers. Genomic instability can be observed in various malignant tumors and precancerous lesions, and tumor drug resistance is also associated with genomic instability. Carcinogens induce genomic instability by chemically, physically, or biologically altering one or more proteins on the spindle apparatus or chromosomes, or by genotoxic agents specifically mutating mitotic genes. Genomic instability then disrupts the stability of the chromosome karyotype, leading to chromosome karyotype evolution and ultimately tumorigenesis. Specifically, in this embodiment, it is implemented as follows: the entire genome is divided into equal-length windows of 2M (two million) base pairs, excluding X and Y chromosomes, mitochondria, and other contigs. The featureCounts software is used to count the number of reads within any window. Due to the inconsistent library sizes of different samples, a homogenization operation is required within the samples. A "per million reads" homogenization method is introduced, specifically:
[0148]
[0149] Among them, Fragment i The number of reads within the i-th window is represented by ``, and `TotalMappedFragments` represents the total number of reads remaining after excluding reads aligned to the reference genome, excluding reads aligned to the X, Y, mitochondria, and other contigs. `WindowLength` represents the total number of reads remaining within the i-th window. i BinCount represents the length of the i-th window, specifically 2E6; i This refers to the score obtained from the above calculation process, which is used to quantify the characteristics of genomic instability.
[0150] 3.5 Gene Expression Prediction: Resting-state promoter regions (low expression) are protected from various endonucleases by their closed chromatin state and nucleosomes; while active promoter regions (high expression) have open chromatin and are more likely to be randomly cleaved. Therefore, cfDNA fragments in active promoter regions will exhibit more random fragmentation patterns, resulting in a wider variety of DNA fragment lengths. This embodiment uses the read length and position information around the cfDNA transcription start site, along with a public mRNA gene expression database as background, to infer the expression status of each gene in the sample. Specifically, this embodiment is implemented as follows: For any transcript, the positional information and length information of reads at various locations within the region 1250 bp upstream and 1250 bp downstream of the TSS are extracted to construct a three-dimensional vector. The first dimension (X-axis) is the positional index of each location relative to the TSS site within the region 1250 bp upstream and 1250 bp downstream of the TSS; the second dimension (Y-axis) is the length information of reads from 0 bp to 400 bp; and the third dimension (Z-axis) is a normalized value representing the severity of clustering of reads of arbitrary length at that location after Z-scaling with the second dimension as the index. The normalized value of the severity of clustering of reads of arbitrary length at that location is reduced in dimensionality using a wavelet compression algorithm. Therefore, expression prediction of any gene in the test sample can be performed on the 477 healthy controls in Table 1, based on 28,047 genes in the refseq database as a baseline. The three-dimensional image of highly expressed genes can be found in [reference needed]. Figure 1 As shown, a three-dimensional image of the low-expression gene can be referenced. Figure 2 As shown.
[0151] 3.6 Somatic Copy Number Variation: DNA copy number variation (CNV) is an important component of genetic variation, and its impact on the genome is greater than that of single nucleotide polymorphisms (SNPs). The role of germline (germ cell) CNV characteristics in various genetic diseases is well understood; however, research on somatic CNV is relatively limited. Somatic chromosome copy number alterations (SCNA), including the gain or loss of entire chromosomes or entire arms, affect large regions of the genome that may contain known oncogenes and tumor suppressor genes. Such genomic instability events promote tumorigenesis and development by leading to the rapid accumulation of mutations in other potential driver genes. Tumorigenesis is caused by a series of mutations or aberrations accumulating at the genomic level. Precise diagnosis of tumors often relies on gene-level testing, but increasing research shows that chromosomal-level testing is closely related to cancer prognosis. Tumor copy number variation (CNV) is closely related to tumorigenesis and development and is one of the most important mechanisms of somatic gene variation. Amplification of proto-oncogenes may lead to activation of the corresponding genes, such as MYC, MYCN, and ERBB2; while the deletion of certain tumor suppressor genes may lead to their inactivation, such as CDKN2A / B, TP53, and BRCA1 / 2. Specifically, this is achieved as follows: sequencing data is perfectly aligned to the hg19 reference genome using the k-mer method. The genome is divided into equal-length windows of 20K length, and the number of unique non-redundant reads within each window is counted. For each 20K-length window, regions with low coverage in the normal population and regions with extreme GC (GC% > 70% or GC% < 25%) are masked. GC preference correction is performed using the locally weighted regression LOESS method, specifically through the following formula:
[0152]
[0153] Where ri represents the number of reads actually observed in the i-th window; mGC represents the median number of reads actually observed in all windows with the same GC ratio as the i-th window; and m represents the median number of reads actually observed in all windows of the subject. This indicates the number of expected reads after GC correction predicted by the examinee in this window.
[0154] For any window, the actual observed copy number (CN) of the sample under test in that window is calculated as: (Number of actual observed reads of the sample under test in that window / (Total number of actual observed reads of the sample under test in all windows * Ratio of the average number of observed reads in that window to the total number of reads in all windows of the normal population). The CN, calculated using GC correction and the background pool of the normal population within any window of the sample under test, is used to search for candidate copy number variation regions across the entire genome using a circular binary cut algorithm. Specifically, this is achieved through the following formula:
[0155]
[0156] In the formula, i and j represent the i-th and j-th windows respectively; ni..j represents the total number of windows between the i-th and j-th windows (i.e., j-i+1); n-ni..j represents the total number of windows excluding the i-th and j-th windows. This represents the sum of the number of windows (CN) between the i-th window and the j-th window, divided by the total number of windows between the i-th window and the j-th window.
[0157] This represents the sum of the CN values of all windows except the i-th to j-th windows divided by the sum of the CN values of all windows except the i-th to j-th windows; specifically, in this embodiment, 1≤i≤j≤n, where n represents the number of the largest window.
[0158] For all combinations of i and j formed by 1≤i≤n and all combinations of 1≤j≤n, calculate T respectively. i.j Until max(T) is made i.j The copy number consistency contiguous region is established to obtain candidate copy number variant regions, and the average value of all window CNs within the copy number consistency contiguous region is used to represent the actual observed copy number of the copy number variant region. The oncogenicity of copy number variants is weighted and scored using multiple oncogenes and tumor suppressor genes, specifically through the following formula:
[0159]
[0160] In the formula, i represents the i-th CNV region (1≤i≤m); j represents the j-th gene in the i-th CNV region (1≤j≤n); n represents the total number of proto-oncogenes and tumor suppressor genes in the CNV region; and m represents the total number of all CNV regions in the subject. This represents the weight of the j-th tumor suppressor gene within the i-th CNV region in the tumor suppressor gene database; This represents the actual observed copy number of the j-th tumor suppressor gene within the i-th CNV region; This represents the weight of the j-th proto-oncogene within the i-th CNV region in the proto-oncogene database; This represents the actual observed copy number of the j-th proto-oncogene within the i-th CNV region. In this embodiment, the weights of tumor suppressor genes in the tumor suppressor gene database and the weights of proto-oncogenes in the proto-oncogene database are obtained from the appendix of the article Davoli T, Xu AW, Mengwasser KE, et al. Cumulative haploinsufficiency and triplosensitivity drive aneuploidy patterns and shape the cancer genome. Cell. 2013; 155(4): 948-962 doi: 101016 / j.cell201310011 (Table S3A Ranking of TSGs by TUSONExplorer and Lasso).
[0161] 4. Construction methods for feature vectorization, tumor status judgment classifier, and tumor signal primary organ localization classifier.
[0162] 41. Feature Vectorization: From step 3 above, for any test sample, five original feature vectors derived from various biomarkers can be obtained, for example: V TssNDR = <v1,v2,...,v n >, n = 28,047 represents the original feature vector of the nucleosome distribution (TSSNDR) biomarker, V Fraqment = <v1,v2,...,v n >, n = 11, 616 represents the original feature vector of the fragment size distribution (Fragment) biomarker, V Motif = <v1,v2,...,v n >, n = 256 represents the original feature vector of the terminal sequence distribution (Motif) biomarker, V Bincount = <v1,v2,...,v n >, n = 1,452 represents the original feature vector of the genomic instability (Bincount) biomarker, V GeneExP = <v1,v2,...,v n >, n = 28,047 represents the original feature vector of the Gene Expression Prediction (GeneEXP) biomarker.
[0163] 42. Selection or determination of feature vectors:
[0164] To construct a tumor status assessment model (a predictive model for tumor detection) and a tumor primary tissue tracing model, the training set in Table 1 included 106 cases of liver cancer (40 stage I, 53 stage II, 7 stage III, and 6 unstaged), 100 cases of colorectal cancer (97 stage I, 1 stage II, and 2 stage IV), 51 cases of esophageal cancer (10 stage I, 14 stage II, 8 stage III, 1 stage IV, 13 stage P0, and 5 unstaged), and 67 cases of pancreatic cancer (14 stage I, 2 stage II, and 2 stage III). A total of 571 patients were included in the cancer group: 5 cases of stage I cancer, 10 cases of stage II cancer, 2 cases of stage IV cancer, and 16 cases of unstaged cancer; 122 cases of lung cancer (96 cases of stage I cancer, 8 cases of stage II cancer, 8 cases of stage III cancer, 2 cases of stage IV cancer, and 8 cases of unstaged cancer); 50 cases of gastric cancer (44 cases of stage I cancer, 2 cases of stage II cancer, and 4 cases of stage IV cancer); 15 cases of ovarian cancer (9 cases of stage I cancer, 1 case of stage II cancer, 3 cases of stage IV cancer, and 2 cases of unstaged cancer); and 60 cases of nasopharyngeal carcinoma (52 cases of stage I cancer, 6 cases of stage II cancer, and 2 cases of stage IV cancer). 286 healthy controls served as the control group. To improve model accuracy, reduce the risk of overfitting, accelerate training speed, and increase model interpretability, the original feature vectors need to be dimensionality reduced to obtain feature vectors for training the model. Specifically, this is achieved as follows:
[0165] 421. The Mann-Whimey U test was used to determine whether there were significant differences in the ordered values of the various biomarkers between the cancer group (liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma) and the non-cancer group (P < 0.05 in this example). After statistical difference analysis, features without significant differences were filtered out, and features with significant differences were obtained. Then, three commonly used methods were used to evaluate the correlation between features and labels: First, chi-square filtering was used for correlation filtering of discrete labels. The chi-square test calculated the chi-square statistic between each non-negative feature and the label, and ranked the features according to the chi-square statistic from high to low. Then, the classes of the top K features (70% in this example) with the highest scores were selected, thereby excluding those most likely to be independent of the label; the F test, also known as ANOVA, is a filtering method used to capture the linear relationship between each feature and the label. It returns two statistics: F-value and p-value. Similar to chi-square filtering, features with p-values less than 0.05 are selected, indicating a significant linear correlation with the label. Features with p-values greater than 0.05 are considered to have no significant linear relationship with the label and are therefore removed. Mutual information is then used to capture any relationship (including linear and non-linear) between each feature and the label. Mutual information returns an estimate of the mutual information between each feature and the target, which takes values between [0, 1]. A value of 0 indicates that the two variables are independent, and a value of 1 indicates that the two variables are perfectly correlated. LASSO is then used to further reduce the dimensionality of the feature vectors selected by the Mann-Whitney U test, chi-square test, homogeneity of variance test, and mutual information method to obtain feature vectors suitable for tumor status detection. Specifically, LASSO constructs a penalty function to obtain a more refined model that compresses some regression coefficients, forcing the sum of the absolute values of the coefficients to be less than a fixed value (in this example, the absolute value is greater than or equal to 1e-5); while setting some regression coefficients to zero. Therefore, it retains the advantage of subset shrinkage and is a biased estimate for handling data with multicollinearity. The importance values shown in Tables 2 through 6 are obtained by the LASSO algorithm by attempting to minimize the sum of the loss function and the L1 regularization term. Each coefficient corresponds to a feature weight. A larger weight indicates a greater impact on the target variable, while a smaller weight may indicate a smaller or no impact on the target variable.
[0166] 422. By defining a tissue-specific index (τ), further filter the feature vectors obtained in step 421 for tumor status detection to obtain feature vectors (biomarker combinations) with tissue-based traceability characteristics. Specifically, this is the union of the biomarker sets for liver cancer (τ>0.85), colorectal cancer (τ>0.85), esophageal cancer (τ>0.85), pancreatic cancer (τ>0.85), lung cancer (τ>0.85), gastric cancer (τ>0.85), ovarian cancer (τ>0.85), and nasopharyngeal cancer (τ>0.85). This union (biomarker combination) can represent the genomic variation characteristics driving the occurrence of different cancers. τ is defined as:
[0167]
[0168]
[0169] In the formula, n represents the number of samples for any cancer type; x i This represents the methylation level of any sample; This represents the maximum methylation level in the same type of sample for any cancer type.
[0170] The final nucleosome distribution markers (feature vectors) and their feature importance scores can be obtained by the above method, as shown in Table 2.
[0171] Table 2. Nucleosome distribution feature vectors between cancer and non-cancer groups (nucleosome distribution biomarkers)
[0172]
[0173]
[0174]
[0175]
[0176] The final fragment size distribution class markers (feature vectors) and their feature importance scores can be obtained by the above method, as shown in Table 3.
[0177] Table 3. Fragment size distribution characteristics between cancer and non-cancer groups (fragment size distribution biomarkers)
[0178]
[0179]
[0180]
[0181]
[0182]
[0183] The final terminal sequence distribution class markers (feature vectors) and their feature importance scores can be obtained by the above method, as shown in Table 4.
[0184] Table 4. Distribution characteristics of terminal sequences between cancer and non-cancer groups (terminal sequence distribution biomarkers)
[0185]
[0186]
[0187] The final genomic instability markers (feature vectors) and their feature importance scores can be obtained using the above methods, as shown in Table 5.
[0188] Table 5. Genomic instability characteristics between cancer and non-cancer groups (genomic instability biomarkers)
[0189] Serial Number Features (Target Window) importance Serial Number Features (Target Window) importance 1 chr1_8000000_10000000 -3.965037632 40 chr7_52000000_54000000 -1.244657853 2 chr1_170000000_172000000 0.04202938 41 chr7_62000000_64000000 0.232323425 3 chr1_172000000_174000000 0.02723584 42 chr7_106000000_108000000 0.448714133 4 chr1_186000000_188000000 -0.229061207 43 chr7_116000000_118000000 0.016961028 5 chr1_214000000_216000000 1.304066163 44 chr7_152000000_154000000 0.328807254 6 chr2_12000000_14000000 -0.605427065 45 chr7_154000000_156000000 1.976290121 7 chr2_14000000_16000000 0.600823443 46 chr8_30000000_32000000 1.006142657 8 chr2_20000000_22000000 1.291222572 47 chr8_54000000_56000000 1.084412524 9 chr2_32000000_34000000 -0.356017908 48 chr8_72000000_74000000 0.001670313 10 chr2_110000000_112000000 1.368657082 49 chr9_66000000_68000000 0.003820202 11 chr2_118000000_120000000 0.241326705 50 chr9_102000000_104000000 0.462558067 12 chr2_124000000_126000000 -1.418205309 51 chr10_36000000_38000000 -1.810495947 13 chr2_130000000_132000000 1.035206591 52 chr10_74000000_76000000 -0.127804952 14 chr2_136000000_138000000 0.048399325 53 chr11_50000000_52000000 -0.04798088 15 chr2_168000000_170000000 3.853167961 54 chr11_90000000_92000000 -1.365860086 16 chr2_226000000_228000000 0.304863426 55 chr11_100000000_102000000 1.45030226 17 chr3_6000000_8000000 -1.559624405 56 chr12_34000000_36000000 1.154478575 18 chr3_66000000_68000000 0.217756503 57 chr12_44000000_46000000 2.392680889 19 chr3_130000000_132000000 0.590419622 58 chr13_102000000_104000000 0.064890804 20 chr3_158000000_160000000 1.661247842 59 chr13_110000000_112000000 1.337299099 21 chr3_178000000_180000000 0.432420893 60 chr14_76000000_78000000 4.948737204 22 chr4_18000000_20000000 0.535840497 61 chr15_40000000_42000000 -1.071434962 23 chr4_38000000_40000000 -2.076529629 62 chr15_64000000_66000000 -0.497998164 24 chr4_52000000_54000000 0.24693357 63 chr16_68000000_70000000 -4.352109758 25 chr5_2000000_4000000 1.512340795 64 chr16_82000000_84000000 0.166086294 26 chr5_22000000_24000000 -0.214184418 65 chr17_1_2000000 -1.939416575 27 chr5_34000000_36000000 0.249650808 66 chr17_28000000_30000000 -3.013673759 28 chr5_46000000_48000000 2.768834839 67 chr18_34000000_36000000 0.715397639 29 chr5_110000000_112000000 1.990224749 68 chr18_48000000_50000000 0.284235548 30 chr5_154000000_156000000 1.24212029 69 chr20_14000000_16000000 -0.926839362 31 chr6_6000000_8000000 0.758365475 70 chr20_16000000_18000000 0.756788606 32 chr6_54000000_56000000 -0.41577285 71 chr20_18000000_20000000 0.492997831 33 chr6_86000000_88000000 1.444997148 72 chr20_24000000_26000000 0.337704997 34 chr6_90000000_92000000 3.480630191 73 chr20_26000000_28000000 0.912406596 35 chr6_114000000_116000000 -4.344309527 74 chr20_52000000_54000000 -0.237805833 36 chr6_116000000_118000000 2.390876704 75 chr19_24000000_26000000 0.117570586 37 chr7_2000000_4000000 0.27144896 76 chr19_26000000_28000000 0.683131825 38 chr7_24000000_26000000 0.980288164 77 chr19_28000000_30000000 1.187146599 39 chr7_44000000_46000000 1.46750581 78 chr22_18000000_20000000 1.008988808
[0190] The final gene expression prediction markers (feature vectors) and their feature importance scores can be obtained using the above methods, as shown in Table 6.
[0191] Table 6. Predictive characteristics of gene expression between cancer and non-cancer groups (gene expression predictive biomarkers)
[0192]
[0193]
[0194]
[0195] The above methods show that there are differences in the mean of somatic copy number variation (sCNA) between the cancer group (liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma) and the non-cancer group, as shown in Table 7. The final somatic copy number variation class markers (feature vectors) are shown in Table 8.
[0196] Table 7. Differences in somatic copy number variation (sCNA) between cancer and non-cancer groups
[0197]
[0198]
[0199] Table 8. Regions with high copy number variation in cancer samples and their variation forms (somatic copy number biomarkers)
[0200]
[0201]
[0202]
[0203]
[0204] 43. Methods for constructing or training predictive models for tumor detection.
[0205] The six types of markers (feature vectors) obtained in step 42 are used for model construction, specifically including:
[0206] For the training samples, three algorithms were used as the first-layer classifier to generate classification results (classifying whether a sample has a tumor). K-Nearest Neighbors (KNN, also known as the K-nearest neighbor algorithm, has advantages such as fast model training time, O(n) training time complexity, algorithm laziness, no assumptions about the data, high accuracy, and insensitivity to outliers), LightGBM (Light Gradient Boosting Machine, a framework for implementing the GBDT algorithm, supporting efficient parallel training, and offering advantages such as faster training speed, lower memory consumption, better accuracy, and support for distributed processing of massive amounts of data), and SVM (Support Vector Machine, which maps vectors to a higher-dimensional space with a maximum margin hyperplane; a strong classifier with high accuracy and the ability to handle non-linear feature interactions) were used to generate classification results. For any feature vector, KNN, LightGBM, and SVM probabilities were generated, and SGD (Stochastic Gradient Boosting Machine) was applied to each. Stochastic gradient descent (SGD), compared to non-stochastic algorithms, utilizes information more effectively and performs exceptionally well in early iterations. Other algorithms used in the second layer include: Stochastic Gradient Descent (SGD), Multiprocessor Flow Perception (MLP), and Gaussian Process (Gaussian Processes). Gaussian Processes are a type of stochastic process in probability theory and mathematical statistics, an extension of the multivariate Gaussian distribution, applied in machine learning and signal processing. These algorithms generate probabilities for KNN-SGD, KNN-MLP, KNN-GP, LightGBM-SGD, LightGBM-MLP, LightGBM-GP, SVM-SGD, SVM-MLP, and SVM-GP. The output of the second-layer classifier is then used by logistic regression to generate TSSNDR, Fragment, Motif, Bincount, and GeneEXP probabilities, forming the third-layer classifier and its classification results. Furthermore, a binary logistic regression model is used to perform stacking learning on the TSSNDR probability, Fragment probability, Motif probability, Bincount probability, and GeneEXP probability to form a fourth-layer classifier and its classification results.
[0207] The construction or training method of this tumor detection prediction model (tumor state judgment classifier) can be found in [reference needed]. Figure 3The optimal parameters for the first, second, third, and fourth classifiers are found through iterative training. The Youden index method is used to determine the optimal threshold for the fourth classifier as the positive judgment value. If the sample is greater than or equal to the positive judgment value, it is judged as "positive (tumor or cancer)"; otherwise, it is judged as "healthy". The optimal parameters and threshold of the determined model are validated on an independent test set. Receiver Operating Characteristic (ROC) curves for the training and test sets are plotted, and the AUC value under the curve is calculated. The prediction model for tumor detection is finally predicted using the following formula (HIFI Score):
[0208]
[0209] In the formula, e is the natural logarithm (e≈2718);
[0210] z = coef i *P TssNDR +coef j *P Fragment +coef k *P Motif +coef l *P Bincount +coef m *P GeneEXP +coef n *CNvscore+intercept, formula 7:
[0211] In the formula, P TSSNDR The third-level classification model score representing nucleosome distribution markers; P Fragment The third-level classification model score represents the fragment size distribution class marker; P Motif The third-level classification model score representing nucleosome distribution markers; P Bincount The third-level classification model score represents the biomarker of genomic instability; P GeneEXP The third-level classification model score represents the gene expression prediction biomarker score; CNVscore represents the score of the somatic copy number variation biomarker score; coef i ,coef j ,coef k ,coef l ,coef m ,coef n represents the weights of nucleosome distribution markers, fragment size distribution markers, terminal sequence distribution markers, genomic instability markers, gene expression prediction markers, and somatic copy number variation markers in logistic regression, respectively; intercept represents the intercept term.
[0212] In this embodiment, coef i ,coef j ,coef k ,coef l ,coef m ,coef n The values are 3.49, 3.01, 0.46, 1.88, 5.57, and 0.18, respectively, with an intercept of -2.71. In other embodiments, these parameters may vary based on the specific composition of the training set, but will not significantly affect the final results. Based on this, the sensitivity of nucleosome distribution markers (Table 2), fragment size distribution markers (Table 3), terminal sequence distribution markers (Table 4), genomic instability markers (Table 5), gene expression prediction markers (Table 6), somatic copy number variation markers (Table 8), and the HIFI Score for detecting (detecting but not distinguishing) cancer samples in the training set, as well as their specificity for detecting healthy samples, can be evaluated. The prediction results are shown in Table 9, and the receiver operating characteristic (ROC) curve is shown in Table 9. Figure 5 As shown.
[0213] Table 9. Prediction Results for Training Set
[0214] Classifier Sensitivity [Detection Positive / True Positive] Specificity [negative test / true negative] Nucleosome distribution markers 73.73%[421 / 571] 95.10%[272 / 286] Fragment size distribution markers 64.62%[369 / 571] 95.10%[272 / 286] Terminal sequence distribution markers 60.25%[344 / 571] 95.45%[273 / 286] biomarkers of genomic instability 66.20%[378 / 571] 95.45%[273 / 286] Gene expression prediction biomarkers 75.13%[429 / 571] 95.45%[273 / 286] Somatic cell copy number variation markers 27.50%[157 / 571] 97.20%[278 / 286] HIFI Score 93.70%[535 / 571] 95.10%[272 / 286]
[0215] Furthermore, the test set in Table 1 included 99 cases of liver cancer (29 stage I, 51 stage II, 12 stage III, and 7 unstaged), 133 cases of colorectal cancer (123 stage I, 7 stage II, and 3 stage IV), 56 cases of esophageal cancer (11 stage I, 11 stage II, 11 stage III, 3 stage IV, 10 stage P0, and 10 unstaged), 157 cases of pancreatic cancer (29 stage I, 54 stage II, 17 stage III, 14 stage IV, and 43 unstaged), and 67 cases of lung cancer (51 stage I, 5 stage II, 2 stage III, and 43 unstaged). A total of 655 patients were included in the cancer group: 5 cases of stage I cancer, 4 cases of unstaged cancer, 55 cases of gastric cancer (54 cases of stage I, 0 cases of stage II, 1 case of stage IV), 23 cases of ovarian cancer (13 cases of stage I, 5 cases of stage II, 1 case of stage IV, 4 cases of unstaged cancer), and 65 cases of nasopharyngeal carcinoma (61 cases of stage I, 2 cases of stage II, 2 cases of stage IV). 191 healthy controls served as the control group. The sensitivity (detection but not differentiation) of nucleosome distribution markers, fragment size distribution markers, terminal sequence distribution markers, genomic instability markers, gene expression prediction markers, somatic copy number variation markers, and HIFI Score for detecting (but not differentiating) cancer samples in the test set, as well as the specificity for detecting healthy samples, were evaluated. The results are shown in Table 10. The receiver operating characteristic (ROC) curve is shown in Table 10. Figure 6 As shown.
[0216] Table 10. Prediction Results for the Test Set
[0217] Classifier Sensitivity [Detection Positive / True Positive] Specificity [negative test / true negative] Nucleosome distribution markers 65.19%[427 / 655] 95.29%[182 / 191] Fragment size distribution markers 65.04%[426 / 655] 95.81%[183 / 191] Terminal sequence distribution markers 41.22%[270 / 655] 95.29%[182 / 191] biomarkers of genomic instability 59.39%[389 / 655] 95.29%[182 / 191] Gene expression prediction biomarkers 76.49%[501 / 655] 95.29%[182 / 191] Somatic cell copy number variation markers 29.92%[196 / 655] 98.43%[188 / 191] HIFI Score 92.82%[608 / 655] 95.81%[183 / 191]
[0218] 4.4. Construction or training methods for predictive models used for tracing the origin of primary tumor tissue (tumor signal primary organ localization classifier):
[0219] The five biomarkers (excluding somatic cell copy number) obtained in step 4.2 were used for model construction. The training set data consisted of 106 cases of liver cancer (40 cases of stage I, 53 cases of stage II, 7 cases of stage III, and 6 cases of unstaged cancer), 100 cases of colorectal cancer (97 cases of stage I, 1 case of stage II, and 2 cases of stage IV), 51 cases of esophageal cancer (10 cases of stage I, 14 cases of stage II, 8 cases of stage III, 1 case of stage IV, 13 cases of stage P0, and 5 cases of unstaged cancer), and 67 cases of pancreatic cancer (Stage I, stage II, stage III, stage IV, stage P0, and staged cancer). The study included 571 cases: 14 cases of stage I, 25 cases of stage II, 10 cases of stage III, 2 cases of stage IV, and 16 cases unstaged; 122 cases of lung cancer (96 cases of stage I, 8 cases of stage II, 8 cases of stage III, 2 cases of stage IV, and 8 cases unstaged); 50 cases of gastric cancer (44 cases of stage I, 2 cases of stage II, and 4 cases of stage IV); 15 cases of ovarian cancer (9 cases of stage I, 1 case of stage II, 3 cases of stage IV, and 2 cases unstaged); and 60 cases of nasopharyngeal carcinoma (52 cases of stage I, 6 cases of stage II, and 2 cases of stage IV). The model construction process can be found in [reference needed]. Figure 4 The construction method of the tumor primary tissue tracing model is roughly the same as step 4.3, the main difference being that 4.3 is used to predict the tumor status, while 4.4 is used to trace the primary tumor tissue. The construction (training) method of the tumor primary tissue tracing prediction model includes the following:
[0220] For the training samples, K-Nearest Neighbors, LightGBM, and SVM algorithms are used as the first-layer classifiers to generate classification results (outputting the probability of a sample having each type of tumor, including liver cancer, colorectal cancer, esophageal cancer, pancreatic cancer, lung cancer, gastric cancer, ovarian cancer, and nasopharyngeal carcinoma). For any feature vector, KNN, LightGBM, and SVM probabilities are generated, and SGD, MLP, and Gaussian Process are used as the second-layer classifiers to generate KNN-SGD, KNN-MLP, KNN-GP, LightGBM-SGD, LightGBM-MLP, LightGBM-GP, SVM-SGD, SVM-MLP, and SVM-GP probabilities, respectively. The output of the second-layer classifiers is then used via logistic regression to generate TSSNDR, Fragment, Motif, Bincount, and GeneEXP probabilities, thus forming the third-layer classifier and its classification results. Furthermore, a multivariate logistic regression model is used to perform stacking learning on the TSSNDR probability, Fragment probability, Motif probability, Bincount probability, and GeneEXP probability to form a fourth-layer classifier and its classification results.
[0221] The 608 cancer samples (129 colorectal cancer, 52 esophageal cancer, 99 liver cancer, 53 lung cancer, 59 nasopharyngeal cancer, 17 ovarian cancer, 149 pancreatic cancer, and 50 gastric cancer) that were identified as "positive" by the HIFI Score in the test set of Table 1 were compared for each cancer type, as shown in Formula 10:
[0222] TOO Score = max(P 肝癌 P 结直肠癌 P 食管癌 P 胰腺癌 P 肺癌 P 胃癌 P 卵巢癌 P 鼻咽癌 ), formula 10.
[0223] Predictive models for tracing the origin of primary tumor tissue using TOO sc o re The primary organ localization of tumor signals is performed to assess tissue attribution accuracy. Attribution accuracy is defined as the percentage of samples predicted as having a particular cancer that actually contains that cancer. The training set tissue attribution assessment is presented as a multi-class confusion matrix, as shown below. Figure 7As shown, specifically, the accuracy of source tracing for colorectal cancer (COAD in the figure) is 84.88%, for esophageal cancer (ESCA in the figure) it is 81.67%, for liver cancer (LIHC in the figure) it is 100.00%, for lung cancer (LUAD in the figure) it is 98.36%, for nasopharyngeal carcinoma (NPC in the figure) it is 100.00%, for ovarian cancer (OV in the figure) it is 81.25%, for pancreatic cancer (PAAD in the figure) it is 88.89%, and for gastric cancer (STAD in the figure) it is 70.59%.
[0224] The test set organization and source tracing evaluation are presented using a multi-class confusion matrix, as shown in the following example. Figure 8 As shown, specifically, the accuracy rates for tracing the origins of colorectal cancer (COAD in the figure) are 81.89%, esophageal cancer (ESCA in the figure) are 84.21%, liver cancer (LIHC in the figure) are 100.00%, lung cancer (LUAD in the figure) are 89.66%, nasopharyngeal carcinoma (NPC in the figure) are 98.33%, ovarian cancer (OV in the figure) are 78.57%, pancreatic cancer (PAAD in the figure) are 97.93%, and gastric cancer (STAD in the figure) are 66.00%.
[0225] Example 2
[0226] Verify the impact of six types of biomarkers on detection results.
[0227] This embodiment sets up multiple experimental groups and constructs a predictive model for detecting tumors. The model construction process for multiple experimental groups is the same as in Embodiment 1. The difference between the experimental groups lies in the different feature vectors (biomarker combinations) of the models. The six biomarkers (nucleosome distribution, fragment size distribution, terminal sequence distribution, genomic instability, gene expression prediction, and somatic copy number variation) in Embodiment 1 are used as the basis. On this basis, any one or more biomarkers are deleted. The specific feature vectors of the models of multiple experimental groups can be referred to Table 11.
[0228] The sample data in this embodiment is based on the training set data in Embodiment 1. The training set samples in Table 1 are split into new training set samples and test set samples. The new training set samples include 74 cases of liver cancer, 70 cases of colorectal cancer, 35 cases of esophageal cancer, 46 cases of pancreatic cancer, 85 cases of lung cancer, 35 cases of gastric cancer, 10 cases of ovarian cancer, 42 cases of nasopharyngeal carcinoma, and 200 healthy controls. The new test set samples include 32 cases of liver cancer, 30 cases of colorectal cancer, 16 cases of esophageal cancer, 21 cases of pancreatic cancer, 37 cases of lung cancer, 15 cases of gastric cancer, 5 cases of ovarian cancer, 18 cases of nasopharyngeal carcinoma, and 86 healthy controls. The training set samples are used for model training, and the test set samples are used to verify the model's performance. The detection results of multiple experimental groups on the new test set samples are shown in Table 11.
[0229] Table 11 Results
[0230]
[0231]
[0232] As shown in Table 11, HIFI is the combination mode with the highest overall AUC, sensitivity, and specificity.
[0233] Example 3
[0234] The impact of the method for calculating the validation markers on the model's prediction results.
[0235] Based on the tumor detection prediction model of Example 1, a control group was set up. The model construction process of the control group was the same as that of Example 1, except that the calculation method of some feature vectors of the model was different, as follows:
[0236] The calculation method for somatic copy number variation characteristics adopts the CNV Z-value calculation method currently used in clinical practice, referring to patent: CN103080336B; the fragment size distribution marker adopts the conventional fragment size calculation method (calculating the ratio of small fragments (<160bp) to long fragments (>220bp) in different chromosome arms), referring to patent CN104254618B; the calculation method for terminal sequence distribution marker adopts the conventional motif terminal motif calculation method (calculating the proportion of 5' and 3' terminal 4mer sequences), referring to patent: CN108026572B.
[0237] The training set samples and test set samples in this embodiment are the same as the new training set samples and test set samples in embodiment 2. After the model is trained using the training set samples, the detection data of the test set is as follows.
[0238] Table 12 Results
[0239] Predictive Model AUC CI_Lower CI_Upper Sensitivity Specificity CNV_Z + fragment size + Motif 0.912 0.899 0.926 0.678 0.952 This invention (HIFI) 0.986 0.981 0.99 0.931 0.952
[0240] Example 4
[0241] Based on the predictive model for tumor detection in Example 1, multiple sets of experiments were set up. The model construction process for the multiple sets of experiments was the same as in Example 1, except that the number of a certain type of biomarker was different, as shown in the table below.
[0242] The training set samples and test set samples in this embodiment are the same as the new training set samples and test set samples in embodiment 2. After the model is trained using the training set samples, the detection data of multiple experimental groups on the test set are as follows.
[0243] Table 13 Results
[0244]
[0245] Example 5
[0246] Based on the tumor detection prediction model of Example 1, multiple sets of experiments were set up. The model construction process of the multiple experimental groups was roughly the same as that of Example 1. The main difference was the number of layers in the model classifier. Please refer to the table below for details.
[0247] The training set samples and test set samples in this embodiment are the same as the new training set samples and test set samples in embodiment 2. After the model is trained using the training set samples, the detection data of multiple experimental groups on the test set are as follows.
[0248] Table 14 Results
[0249]
[0250] Note: SVM stands for Support Vector Machine, SGD stands for Stochastic Gradient Descent, LR stands for Logistic Regression, MLP stands for Multilayer Perceptron, and RF stands for Random Forest.
[0251] Example 6
[0252] To verify the impact of five biomarkers (excluding somatic cell copy number) on tissue traceability.
[0253] This embodiment sets up multiple experimental groups, and constructs a prediction model (tumor signal primary organ localization classifier) for tracing the origin of tumor primary tissue. The model construction process of multiple experimental groups is the same as in embodiment 1. The difference between the experimental groups lies in the different feature vectors (biomarker combinations) of the models. Based on the five biomarkers in embodiment 1 (nucleosome distribution, fragment size distribution, terminal sequence distribution, genomic instability and gene expression prediction), any one or more biomarkers can be deleted on this basis. For details, please refer to Table 15.
[0254] To evaluate the impact of different biomarker combinations on predictive models for tracing the primary tumor origin, the training set included 106 cases of hepatocellular carcinoma (40 stage I, 53 stage II, 7 stage III, and 6 unstaged), 100 cases of colorectal cancer (97 stage I, 1 stage II, and 2 stage IV), 51 cases of esophageal cancer (10 stage I, 14 stage II, 8 stage III, 1 stage IV, 1 stage IV, 13 stage P0, and 5 unstaged), and 67 cases of pancreatic cancer (14 stage I, 1 stage II, 1 stage III, 1 stage IV, 1 stage P0, and 5 unstaged). The training set consisted of 571 cases: 25 cases of stage I, 10 cases of stage III, 2 cases of stage IV, and 16 cases of unstaged cancer; 122 cases of lung cancer (96 cases of stage I, 8 cases of stage II, 8 cases of stage III, 2 cases of stage IV, and 8 cases of unstaged cancer); 50 cases of gastric cancer (44 cases of stage I, 2 cases of stage II, and 4 cases of stage IV); 15 cases of ovarian cancer (9 cases of stage I, 1 case of stage II, 3 cases of stage IV, and 2 cases of unstaged cancer); and 60 cases of nasopharyngeal carcinoma (52 cases of stage I, 6 cases of stage II, and 2 cases of stage IV). The test set in Table 1 includes 99 cases of liver cancer (29 cases of stage I, 51 cases of stage II, 12 cases of stage III, and 7 cases of unstaged cancer), 133 cases of colorectal cancer (123 cases of stage I, 7 cases of stage II, and 3 cases of stage IV), 56 cases of esophageal cancer (11 cases of stage I, 11 cases of stage II, 11 cases of stage III, 3 cases of stage IV, 10 cases of stage P0, and 10 cases of unstaged cancer), and 157 cases of pancreatic cancer (29 cases of stage I, 54 cases of stage II, and 17 cases of stage III). A total of 655 cases were included in the test set: 14 cases of stage I and IV cancer, 43 cases of unstaged cancer, 67 cases of lung cancer (51 cases of stage I, 5 cases of stage II, 2 cases of stage III, 5 cases of stage IV, and 4 cases of unstaged cancer), 55 cases of gastric cancer (54 cases of stage I, 0 cases of stage II, and 1 case of stage IV), 23 cases of ovarian cancer (13 cases of stage I, 5 cases of stage II, 1 case of stage IV, and 4 cases of unstaged cancer), and 65 cases of nasopharyngeal carcinoma (61 cases of stage I, 2 cases of stage II, and 2 cases of stage IV). Training samples were used to train the model, and test samples were used to validate the model's performance. The detection results of multiple experimental groups on the new sequencing set samples are shown in Table 15.
[0255] Table 15 Results
[0256]
[0257]
[0258] As shown in Table 15, HIFI is the combination mode with the highest precision, recall and F1 score for tracing the primary tumor tissue.
[0259] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. Use of a reagent combination for detecting a biomarker or a combination thereof in the manufacture of a product for detecting or aiding in the detection of a tumor, characterized in that, The biomarkers or combinations thereof are combinations of six types of markers: nucleosome distribution type markers, fragment size distribution type markers, end sequence distribution type markers, genomic instability type markers, gene expression prediction type markers, and somatic copy number variation type markers. The nucleosome distribution type markers take nucleosome distribution characteristics of a target gene or its transcript as markers, and are combinations of markers 1-247 shown in the following table. The fragment size distribution type markers take fragment size distribution characteristics in a specified window as markers, and are combinations of markers 1-335 shown in the following table, wherein Chr: X1-X2 is the specified window, and frac_PX3-X4 refers to the distribution characteristics of reads with lengths within X3-X4. The end sequence distribution type markers take the distribution characteristics of reads with target end sequences as markers, and are combinations of markers 1-77 shown in the following table, wherein the target end sequences are sequences with lengths ≤ 170 bp and aligned to positions 4 bp upstream of the reference genome. The genomic instability type markers take genomic instability characteristics in a target window as markers, and are combinations of markers 1-78 shown in the following table. The gene expression prediction type markers take gene expression prediction characteristics of a target gene or its transcript as markers, and are combinations of markers 1-258 shown in the following table. The somatic copy number variation type markers take somatic copy number variation characteristics of a target copy number variation region as markers, and the target copy number variation region is a combination of CNV regions 1-300 shown in the following table. The sequence information takes hg19 as the reference genome. The tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer, and nasopharyngeal cancer.
2. Use according to claim 1, characterized in that, The calculation method of the nucleosome distribution characteristics is as follows: ; wherein, Coverage NDR NDR refers to the region between 200bp upstream of the TSS and 100bp downstream of the TSS after correction; Mean (NDR) refers to the mean of the coverage after correction of the region between 2000bp upstream of the TSS and 200bp upstream of the TSS and the region between 100bp downstream of the TSS and 2000bp downstream of the TSS; Coverage TSS1’ Coverage TSS2 NDR refers to the region between 200bp upstream of the TSS and 100bp downstream of the TSS after correction; Mean (NDR) refers to the mean of the coverage after correction of the region between 2000bp upstream of the TSS and 200bp upstream of the TSS and the region between 100bp downstream of the TSS and 2000bp downstream of the TSS; TSSNDR Score is a score calculated to quantify the nucleosome distribution profile of a gene or transcript.
3. Use according to claim 1, characterized in that, The calculation method of the genomic instability characteristics is as follows: ; wherein, Fragment i represents the number of reads within the i-th window, i TotalMappedFragments represents the total number of reads of the sample to be tested which are aligned to the reference genome, preferably the total number of reads of the sample to be tested which are aligned to the reference genome after excluding reads aligned to X, Y sex chromosomes, mitochondria and other contigs; WindowLength i represents the length of the i-th window; BinCount i is the genomic instability score. 4. Use according to claim 1, characterized in that, The gene expression prediction characteristics are obtained by predicting the expression of a target gene or its transcript as high expression or low expression based on the length and position information of reads around the transcription start site of the gene.
5. The use according to claim 1, characterized in that, The gene expression prediction feature is obtained by: for any gene or transcript, extracting position information and length information of reads at each position in a region between 1250 bp upstream of the TSS and 1250 bp downstream of the TSS, constructing a three-dimensional vector, wherein the first dimension X-axis is the position index of each position in the region between 1250 bp upstream of the TSS and 1250 bp downstream of the TSS relative to the TSS site, the second dimension Y-axis is the length information of the reads between 0 bp and 400 bp, and the third dimension Z-axis is a normalized value that can represent the severity of the aggregation of reads of any length at the position after Z value scaling with the second dimension as the index; dimension reduction is performed on the normalized value of the severity of the aggregation of reads of any length at the position to obtain a three-dimensional matrix of the gene or transcript; the three-dimensional matrices of the highly expressed genes and the lowly expressed genes with known gene expression amounts are input into a polynomial regression training as input values to obtain a model capable of predicting the expression score or expression of the gene or its transcript, and the three-dimensional matrix of the target gene or its transcript is input into the model to obtain the expression of the target gene or its transcript.
6. Use according to claim 5, characterized in that, The dimension reduction method includes a wavelet compression algorithm.
7. Use according to claim 1, characterized in that, The somatic copy number variation feature is obtained in the following manner: ; In the formula, i Indicates the first i CNV regions (1≤ i ≤ m ); j Indicates the first i The first CNV region j Genes (1≤ j ≤ n ); n This represents the total number of proto-oncogenes and tumor suppressor genes within the CNV region; m This indicates the total number of all CNV areas for the examinee; Indicates the first i The first CNV region j The weight of each tumor suppressor gene in the tumor suppressor gene database; Indicates the first i The first CNV region j The actual observed copy number of each tumor suppressor gene; Indicates the first i The first CNV region j The weight of each proto-oncogene in the proto-oncogene database; Indicates the first i The first CNV region j The actual observed copy number of each proto-oncogene.
8. The use according to claim 1, characterized in that, The fragment size distribution feature is obtained in the following manner: the number of reads with a length within a limited range in a limited window or the proportion of the number of reads in the total number of reads in the limited window is counted.
9. The use according to claim 1, characterized in that, Each marker in the terminal sequence distribution class marker is obtained in the following manner: the number of reads with a sequence length of 17 bp or less and aligned to a sequence 4 bp upstream of the position on the reference genome as a target terminal sequence or the proportion of the number of reads in the total number of reads is obtained.
10. The use according to claim 1, characterized in that, The product is also used for tumor primary tissue tracing.
11. A method for training a predictive model for detecting or assisting in the detection of a tumor, characterized in that, It comprises: The detection results of each marker in the biomarker or combination thereof in the training sample and the corresponding annotation results are obtained; wherein the biomarker or combination thereof is the biomarker or combination thereof in any one of claims 1-9, and the annotation result is a label representing whether the sample has a tumor; The detection results of the biomarker or combination thereof are input into a pre-constructed prediction model to obtain a prediction result; the prediction model is a machine learning model capable of predicting whether the sample has the tumor according to the detection results of the biomarker or combination thereof; The prediction model is updated in parameters based on the annotation results and the prediction results; The tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer, and nasopharyngeal cancer.
12. The training method of claim 11, wherein, The machine learning model comprises one or more layers of classifiers.
13. The training method of claim 12, wherein, The multi-layer classifier comprises: a first layer of classifiers to an n-th layer of classifiers, and n is a positive integer greater than 1. The multi-layer classifier comprises: a first layer of classifiers to an n-th layer of classifiers, and n is a positive integer greater than 1. The input of the first layer classifier is the detection result of each type of marker in the biomarker or combination thereof except for somatic copy number, and the output is the classification result of whether each type of marker except for somatic copy number is related to the tumor of the sample; the input of the nth layer classifier is the detection result of the somatic copy number type marker in the biomarker or combination thereof and the output of the corresponding (n-1)th layer classifier of each type of marker except for somatic copy number, and the output of the nth layer classifier is the final prediction result.
14. The training method of claim 12, wherein, The multi-layer classifier comprises a first layer classifier, a second layer classifier, a third layer classifier and a fourth layer classifier.
15. The training method of claim 14, wherein, The first layer classifier and / or the second layer classifier comprise any one or a combination of multiple of Kneighbors, LightGBM, SVM, SGD, MLP, GaussianProcess.
16. The training method of claim 14, wherein, The first layer classifier comprises Kneighbors, LightGBM and SVM.
17. The training method of claim 14, wherein, The second layer classifier comprises SGD, MLP and GaussianProcess.
18. The training method of claim 14, wherein, The third layer classifier comprises a logistic regression model.
19. The training method of claim 14, wherein, The fourth layer classifier comprises a logistic regression model.
20. A prognostic device for detecting or aiding in the detection of a tumor, characterized by It comprises: An acquisition module is configured to acquire the detection result of each marker in the biomarker or combination thereof in the sample to be tested; wherein the biomarker or combination thereof is the biomarker or combination thereof in any one of claims 1-9; A prediction module is configured to input the detection result of the biomarker or combination thereof into the prediction model trained by the training method in any one of claims 11-19 to obtain a prediction result. The tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer and nasopharyngeal cancer.
21. Use of a reagent combination for detecting biomarkers or combinations thereof in the manufacture of a product for tumor primary tissue provenance, characterized in that, The biomarker or combination thereof is the biomarker or combination thereof in any one of claims 1-9, and the tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer and nasopharyngeal cancer.
22. Use of a reagent combination for detecting biomarkers or combinations thereof in the manufacture of a product for tumor primary tissue provenance, characterized in that, The biomarker or combination thereof is the nucleosome distribution type marker, fragment size distribution type marker, terminal sequence type marker or gene expression prediction type marker in any one of claims 1-9, and the tumor is any one or a combination of liver cancer and nasopharyngeal cancer.
23. A kit comprising, It comprises: The reagent combination for detecting the biomarker or combination thereof in any one of claims 1-9.
24. A training method for a predictive model used for tracing the origin of primary tumor tissue, characterized in that, It comprises: An acquisition module is configured to acquire the detection result of each marker in the biomarker or combination thereof in the sample to be tested; wherein the biomarker or combination thereof is the biomarker or combination thereof in any one of claims 1-9; A prediction module is configured to input the detection result of the biomarker or combination thereof into the prediction model trained by the training method in any one of claims 11-19 to obtain a prediction result. A parameter updating module is configured to update the parameters of the prediction model based on the labeled result and the prediction result. The tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer, and nasopharyngeal cancer.
25. A training method for a predictive model used for tracing the origin of primary tumor tissue, characterized in that, It comprises: Obtaining the detection results of each marker in the biomarker or combination thereof in the training sample and the corresponding annotation results; wherein the biomarker or combination thereof is the biomarker or combination thereof described in claim 22, and the annotation result is a label representing the tissue provenance result of the tumor of the sample; Inputting the detection results of the biomarker or combination thereof into a pre-constructed prediction model to obtain a prediction result; the prediction model is a machine learning model capable of performing tissue provenance of a tumor on a sample according to the detection results of the biomarker or combination thereof; Updating the parameters of the prediction model based on the annotation results and the prediction results; The tumor is any one or a combination of liver cancer and nasopharyngeal cancer.
26. The training method of claim 24 or 25, wherein, The machine learning model comprises one or more layers of classifiers.
27. The training method of claim 26, wherein, The multi-layer classifier comprises: a first layer of classifiers to an n-th layer of classifiers, n is a positive integer greater than 1; The input of the first layer of classifiers is the detection result of each type of marker in the biomarker or combination thereof, and the output is the tissue provenance result of each type of marker on the sample; the input of the n-th layer of classifiers is the output of the n-1-th layer of classifiers corresponding to all types of markers, and the output of the n-th layer of classifiers is the final prediction result.
28. The training method of claim 26, wherein, The multi-layer classifier comprises: a first layer of classifiers, a second layer of classifiers, a third layer of classifiers, and a fourth layer of classifiers.
29. The training method of claim 28, wherein, The first layer of classifiers and / or the second layer of classifiers comprise any one or a combination of Kneighbors, LightGBM, SVM, SGD, MLP, GaussianProcess.
30. The training method of claim 28, wherein, The first layer of classifiers comprises Kneighbors, LightGBM, and SVM.
31. The training method of claim 28, wherein, The second layer of classifiers comprises SGD, MLP, and GaussianProcess.
32. The training method of claim 28, wherein, The third layer of classifiers comprises a logistic regression model.
33. The training method of claim 28, wherein, The fourth layer of classifiers comprises a logistic regression model.
34. A prediction device for tumor primary tissue provenance, comprising: It comprises: An acquisition module is configured to acquire detection results of each marker in a biomarker or combination thereof in a to-be-tested sample; wherein the biomarker or combination thereof is the biomarker or combination thereof described in claim 21; A prediction module is configured to input the obtained detection results of the biomarker or combination thereof into a prediction model trained by the training method of any one of claims 24, 26-33 to obtain a prediction result. The tumor is any one or a combination of liver cancer, intestinal cancer, esophageal cancer, pancreatic cancer, lung cancer, ovarian cancer, and nasopharyngeal cancer.
35. A prediction device for tumor primary tissue provenance, comprising: It comprises: An acquisition module is configured to acquire detection results of each marker in a biomarker or combination thereof in a to-be-tested sample; wherein the biomarker or combination thereof is the biomarker or combination thereof described in claim 22; A prediction module is configured to input the obtained detection results of the biomarker or combination thereof into a prediction model trained by the training method of any one of claims 25-33 to obtain a prediction result. The tumor is any one or a combination of the following: liver cancer and nasopharyngeal carcinoma.
36. An electronic device, comprising: The computer readable medium stores a computer program, which, when executed by a processor, implements any one of the following: a method for detecting or assisting in detecting a tumor, a method for tracing a tumor primary tissue, a training method of the prediction model for detecting or assisting in detecting a tumor according to any one of claims 11-19, and a training method of the prediction model for tracing a tumor primary tissue according to any one of claims 24-33. The method for detecting or assisting in detecting a tumor comprises: obtaining a detection result of each biomarker or a combination thereof in a sample to be detected, wherein the biomarker or the combination thereof is any one of the biomarkers or the combinations thereof according to any one of claims 1-9; inputting the obtained detection result of the biomarker or the combination thereof into the prediction model trained by the training method according to any one of claims 11-19 to obtain a prediction result; The method for detecting or assisting in detecting a tumor is not directly aimed at diagnosis or treatment of a disease. The method for tracing a tumor primary tissue comprises: obtaining a detection result of each biomarker or a combination thereof in a sample to be detected, wherein the biomarker or the combination thereof is any one of the biomarkers or the combinations thereof according to claim 21; inputting the obtained detection result of the biomarker or the combination thereof into the prediction model trained by the training method according to any one of claims 24, 26-33 to obtain a prediction result; or, obtaining a detection result of each biomarker or a combination thereof in a sample to be detected, wherein the biomarker or the combination thereof is any one of the biomarkers or the combinations thereof according to claim 22; inputting the obtained detection result of the biomarker or the combination thereof into the prediction model trained by the training method according to any one of claims 25-33 to obtain a prediction result; The method for tracing a tumor primary tissue is not directly aimed at diagnosis or treatment of a disease.
37. A computer readable medium characterized by The computer readable medium stores a computer program, which, when executed by a processor, implements any one of the following: a method for detecting or assisting in detecting a tumor, a method for tracing a tumor primary tissue, a training method of the prediction model for detecting or assisting in detecting a tumor according to any one of claims 11-19, and a training method of the prediction model for tracing a tumor primary tissue according to any one of claims 24-33. The method for detecting or assisting in detecting a tumor is the method for detecting or assisting in detecting a tumor according to claim 36. The method for tracing a tumor primary tissue is the method for tracing a tumor primary tissue according to claim 36.
Citation Information
Patent Citations
Kits, devices and methods for detecting chromosome copy number of embryo or tumor
CN103080336B
Size-based analysis of fetal DNA fraction in maternal plasma
CN104254618B
Analysis of fragmentation patterns in cell-free DNA
CN108026572B
Method for predicting tumor risk value based on plasma multi-omics multi-dimensional characteristics and artificial intelligence
CN112397143A
Model and kit for identifying benign and malignant pulmonary nodules or predicting risk of postoperative recurrence of lung cancer and application
CN115798582A