A method of simulating high-depth sequencing tss signatures

By extracting features from low-depth sequencing results and constructing models, the problem of unstable TSS region coverage in cell-free DNA sequencing was solved, achieving highly stable and economical TSS feature analysis in the non-invasive field.

CN116805512BActive Publication Date: 2026-03-27SUZHOU BASECARE MEDICAL DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Current technologies for low-depth sequencing of cell-free DNA suffer from problems such as small genome coverage and unstable TSS region sequencing fragment coverage, making it difficult to promote in the field of non-invasive procedures.

Method used

By acquiring sequencing data from cell-free DNA samples captured from a reference genome, features can be extracted and models constructed using low-depth sequencing results, eliminating the need for high-depth sequencing, expanding the analyzable TSS region, and increasing the stability of TSS region sequencing fragment coverage.

Benefits of technology

It achieves TSS features comparable to high-depth sequencing on the basis of low-depth sequencing, expands the analyzable TSS region, reduces sequencing costs, and improves the stability of TSS region sequencing fragment coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805512B_ABST
    Figure CN116805512B_ABST
Patent Text Reader

Abstract

The application discloses a method for simulating high-depth sequencing TSS characteristics. The method comprises the following steps: (1) obtaining reference genome captured free DNA sample sequencing data and sequence coverage of the upstream and downstream regions of the TSS of each gene; (2) obtaining features for constructing TSS values of simulated high-depth sequencing results based on the low-depth sequencing results of free DNA; and (3) constructing a TSS value model of simulated high-depth sequencing results based on the features of the low-depth sequencing results of free DNA. According to the application, the features are extracted based on the low-depth sequencing results, a model is constructed, the TSS characteristics equivalent to high-depth sequencing can be obtained, the number of analyzable gene TSS regions can be effectively increased, the stability of the sequencing fragment coverage of the analyzable gene TSS regions is increased, the sequencing cost is greatly reduced while the calculation accuracy of the TSS values is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of biotechnology, and relates to a method for simulating TSS characteristics of high-depth sequencing. BACKGROUND

[0002] Cell-free DNA (cfDNA) is DNA fragments from cells in the blood circulation system, mainly from fragmented DNA in the process of apoptosis, DNA fragments of necrotic cells, and exosomes secreted by cells. The length of cfDNA fragments is about 166 bp, which corresponds to the length of nucleosome DNA (145-147 bp) plus the length of connecting DNA (about 20 bp). It is stable and not degraded due to the protection of nucleosomes. The analysis of cfDNA is a rapidly developing field of liquid biopsy. The binding of RNA polymerase II to the genome affects the distribution of nucleosomes at transcription start sites (TSSs). Generally, the transcription core region of the gene promoter is usually not distributed with nucleosomes, and the whole genome map also shows that the nucleosome distribution of active and silent genes is significantly different. These results suggest that the distribution of nucleosomes is related to gene expression in eukaryotes. Therefore, by high-throughput sequencing of cfDNA, the distribution of sequencing sequences in the TSS region can be used to predict gene expression. Generally, the coverage depth of cfDNA at the TSS is negatively correlated with gene expression, that is, the lower the cfDNA fragment coverage of the TSS region of a gene, the more abundant the expression of the gene.

[0003] At present, the application of cell-free DNA is mostly low-depth sequencing (especially non-invasive prenatal diagnosis). Low-depth sequencing leads to the lack of coverage of some TSS regions on the one hand, and the coverage of some TSS regions is random on the other hand. High-depth sequencing of cfDNA can expand the coverage of the genome and increase the stability of the coverage of TSS regions, but also increases the cost of sequencing, which is difficult to directly promote in the non-invasive field.

[0004] CN113160889A discloses a cancer non-invasive early screening method based on cfDNA omics characteristics, including a cfDNA omics feature model and a machine learning training model, which comprises establishing a cfDNA omics feature model; extracting cfDNA through blood collection; library construction and sequencing of the extracted cfDNA; and extracting cfDNA omics features for comparison. By combining the length distribution characteristics of cfDNA, the copy number variation density distribution characteristics, and the openness characteristics around the promoter of cfDNA, the characteristics of cfDNA in gastric cancer patients are comprehensively described through the method of low-depth whole genome sequencing of cfDNA, and early gastric cancer patients are accurately identified. However, this method is low-depth sequencing, which has problems such as small genome coverage and unstable coverage of TSS region sequencing fragments.

[0005] CN113838533A discloses a cancer detection model and a construction method and kit thereof, by whole genome sequencing of plasma free DNA, nucleosome distribution characteristics, terminal sequence characteristics and fragment size distribution characteristics applicable to cancer detection are mined, by constructing a classification model of the three indicators, the prediction score of each indicator for the sample is obtained, then using a logistic regression model, the scores are integrated and the copy number variation characteristic information is added, to obtain a final classification prediction model, the cancer detection model significantly improves the efficiency and accuracy of cancer detection.

[0006] In summary, at present, the application of free DNA is mostly low-depth sequencing, which has problems such as small genomic coverage range and unstable TSS region sequencing fragment coverage. How to provide a sequencing method that is economical and convenient and has high accuracy and high stability has become one of the problems to be solved in the field of biotechnology. SUMMARY

[0007] In view of the deficiencies of the prior art and actual needs, the present application provides a method for simulating high-depth sequencing TSS characteristics, which solves the problems of few analyzable TSS regions and unstable TSS region sequencing fragment coverage in current free DNA low-depth sequencing, and achieves the effects of increasing the number of analyzable TSS regions and increasing the stability of TSS region sequencing fragment coverage.

[0008] To achieve the purpose of the present application, the following technical solutions are adopted:

[0009] In a first aspect, the present application provides a method for simulating high-depth sequencing TSS characteristics, the method comprising the following steps:

[0010] (1) obtaining the sequencing data of the reference genome captured free DNA sample and the sequence coverage of the upstream and downstream regions of each gene TSS;

[0011] (2) obtaining the features for constructing the TSS value of the simulated high-depth sequencing result based on the low-depth sequencing result of the free DNA;

[0012] (3) constructing the TSS value model of the simulated high-depth sequencing result based on the features obtained from the low-depth sequencing result of the free DNA.

[0013] The present application extracts features and constructs a model based on the low-depth sequencing result, which can obtain TSS characteristics comparable to high-depth sequencing without high-depth sequencing of cfDNA, effectively expand the analyzable TSS region, increase the stability of TSS region sequencing fragment coverage, and reduce the sequencing cost.

[0014] Preferably, the step (1) specifically comprises the following steps:

[0015] (1-1) The high-throughput sequencing raw data of the sample is aligned with the reference genome, and the coverage depth of each site in the sequence of the upstream and downstream regions of each gene TSS is counted;

[0016] (1-2) The TSS of each gene is standardized according to formula (1) to obtain the characteristics of the upstream and downstream of each gene TSS, respectively;

[0017] TSS inormalized = TSS idepth / total TSS depth *10 6 Formula (1);

[0018] Wherein TSS inormalized is the normalized value of the coverage depth of the upstream and downstream regions of the transcription start site region of gene i, TSS idepth is the coverage depth of the upstream and downstream regions of the transcription start site region of gene i, and total TSS depth is the sum of the coverage depths of the upstream and downstream regions of the transcription start site regions of all genes.

[0019] Preferably, the region is selected from a combination of at least three sites 0.5-1.5 kb, 1.5-2.5 kb, 2.5-3.5 kb, 3.5-4.5 kb, 4.5-5.5 kb or 5.5-6.5 kb upstream and downstream of the gene transcription start site.

[0020] For example, the region length can be 1 kb, 2 kb and 3 kb upstream and downstream; it can be 1 kb, 2 kb and 4 kb upstream and downstream; it can be 1 kb, 3 kb and 6 kb upstream and downstream.

[0021] The specific point value in the above 0.5-1.5 can be selected as 0.5, 0.6, 0.7, 0.9, 1, 1.2, 1.3, 1.4, 1.5, etc.

[0022] The specific point value in the above 1.5-2.5 can be selected as 1.5, 1.6, 1.7, 1.9, 2, 2.2, 2.3, 2.4, 2.5, etc.

[0023] The specific point value in the above 2.5-3.5 can be selected as 2.5, 2.6, 2.7, 2.9, 3, 3.2, 3.3, 3.4, 3.5, etc.

[0024] The specific point value in the above 3.5-4.5 can be selected as 3.5, 3.6, 3.7, 3.9, 4, 4.2, 4.3, 4.4, 4.5, etc.

[0025] The specific point value in the above 4.5-5.5 can be selected as 4.5, 4.6, 4.7, 4.9, 5, 5.2, 5.3, 5.4, 5.5, etc.

[0026] The specific point values in the above 5.5-6.5 can be selected from 5.5, 5.6, 5.7, 5.9, 6, 6.2, 6.3, 6.4, 6.5, etc.

[0027] Preferably, the step (2) specifically comprises the following steps:

[0028] (2-1) Extracting the chromosome information, start site, end site, positive and negative strand, gene length, intergenic distance to the previous adjacent gene, the previous adjacent gene and the next adjacent gene of each gene;

[0029] (2-2) According to the information of each gene obtained from (1-2) and (2-1), constructing feature engineering of each gene for predicting high-depth TSS results as input feature values for model construction.

[0030] Preferably, the feature values include:

[0031] The chromosome information, start site, end site, positive and negative strand, gene length, intergenic distance to the previous adjacent gene, three TSS values of each gene, the rank of three TSS values of each gene, the three TSS values of the previous adjacent gene, the rank of three TSS values of the previous adjacent gene, the three TSS values of the next adjacent gene, and the rank of three TSS values of the next adjacent gene of each gene. normalized normalized normalized normalized normalized normalized normalized 0.5-1.5kb normalized 1.5-2.5kb normalized 2.5-3.5kb normalized 3.5-4.5kb normalized 4.5-5.5kb normalized 5.5-6.5kb normalized , wherein the rank represents the ranking of TSS depth of all genes in the sample. depth

[0032] Preferably, the step (3) specifically comprises the following steps:

[0033] (3-1) Using a machine learning method to construct a mapping model based on the features confirmed in (2-2) in the training set;

[0034] (3-2) Using a 10-fold cross-validation method to optimize the model parameters and confirm the final prediction model.

[0035] ​​​​​​​​​​​​​Preferably, the machine learning method in step (3-1) comprises any one or a combination of at least two of logistic regression, support vector machine, random forest, decision tree, linear regression or naive Bayes

[0036] Preferably, the cell-free DNA sample comprises any one or a combination of at least two of plasma cell-free DNA, cell culture fluid cell-free DNA or seminal plasma cell-free DNA

[0037] In a second aspect, the present application provides a model simulating high-depth sequencing TSS characteristics, which is constructed by the method for simulating high-depth sequencing TSS characteristics according to the first aspect.

[0038] In a third aspect, the present application provides a device for simulating high-depth sequencing TSS characteristics, which comprises:

[0039] The cell-free DNA TSS analysis module, the feature construction module and the model construction module.

[0040] The cell-free DNA TSS analysis module is configured to perform the following steps:

[0041] Obtaining the sequence coverage of the cell-free DNA sample sequencing data captured by the reference genome and the length of the upstream and downstream regions of each gene TSS;

[0042] The feature construction module is configured to perform the following steps:

[0043] Constructing the features of the TSS values of the simulated high-depth sequencing results;

[0044] The model construction module is configured to perform the following steps:

[0045] Constructing the TSS value model of the simulated high-depth sequencing results based on the low-depth sequencing results of the cell-free DNA.

[0046] Preferably, the cell-free DNA TSS analysis module is configured to perform the following steps:

[0047] (1-1) Aligning the high-throughput sequencing raw data of the sample with the reference genome, and counting the coverage depth of each site in the upstream and downstream regions of each gene TSS;

[0048] (1-2) Standardizing each gene TSS according to formula (1) to obtain the features of the upstream and downstream of each gene TSS, respectively;

[0049] The feature construction module is configured to perform the following steps:

[0050] (2-1) extract the chromosome information, start site, end site, positive and negative strand, gene length, intergenic distance with the previous adjacent gene, the previous adjacent gene and the next adjacent gene of each gene;

[0051] (2-2) according to the information of each gene obtained from (1-2) and (2-1), construct the feature value of each gene for predicting the high-depth TSS result as the input value of model construction;

[0052] The model construction module is used to perform specifically includes:

[0053] (3-1) using machine learning method to construct mapping model based on the features confirmed in (2-2) in the training set;

[0054] (3-2) using 10-fold cross-validation method to optimize the model parameters, and confirming the final prediction model.

[0055] Compared with the prior art, the present application has the following beneficial effects:

[0056] The present application does not need to perform high-depth sequencing on cfDNA, only needs to construct a model on the basis of low-depth sequencing results to transform data, so that the TSS features equivalent to high-depth sequencing can be obtained, the TSS region that can be analyzed can be effectively expanded, the stability of the sequencing fragment coverage of the TSS region is increased, the number of regions with TSS value of 0 and the variability of TSS value in the sample are reduced, the sequencing cost is greatly reduced while the accuracy of TSS calculation value is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 is the analysis flowchart of the present application;

[0058] Figure 2 is the correlation statistical diagram of TSS values before and after model simulation of low-depth sequencing and high-depth sequencing TSS values respectively;

[0059] Figure 3 is the number diagram of TSS values and high-depth sequencing TSS values of 0 before and after model simulation of low-depth sequencing;

[0060] Figure 4 is the coefficient of variation diagram of TSS values and high-depth sequencing TSS values before and after model simulation of low-depth sequencing. DETAILED DESCRIPTION

[0061] In order to further illustrate the technical means adopted by the present application and its effects, the present application will be further described in conjunction with the embodiments and drawings. It can be understood that the specific embodiments described herein are only used to explain the present application, and not to limit the present application.

[0062] Unless otherwise specified in the examples, the techniques or conditions described in the literature were used, or the product manual was followed. Unless otherwise specified, the reagents or instruments used were commercially available conventional products.

[0063] Example 1

[0064] This example constructs a model simulating the TSS characteristics of high-depth sequencing.

[0065] (1) Acquisition of cell-free DNA samples

[0066] Thirty-four plasma cell-free DNA samples were obtained, and high-depth and low-depth sequencing was performed. The average sequencing depth of high-depth sequencing was more than 4X, and the average sequencing depth of low-depth sequencing was about 0.2X.

[0067] (2) Analysis of cell-free DNA

[0068] (2-1) Low-depth sequencing results: After high-throughput sequencing of cell-free DNA, the sequences were aligned with the human genome reference sequence hg19 to determine the position of each sequence on the human genome, and the coverage depth of each site of the sequence 1 kb, 2 kb, and 4 kb upstream and downstream of each gene TSS was counted. The coverage of 1 kb, 2 kb, and 4 kb upstream and downstream of each gene TSS was added to obtain the TSS of 1 kb, 2 kb, and 4 kb upstream and downstream of each gene TSS. depth The TSS of each gene was standardized according to the following formula to obtain the characteristics of 1 kb, 2 kb, and 4 kb upstream and downstream of each gene TSS,

[0069] TSS inomalized = TSS idepth / total TSS depth * 10 6 Equation (1);

[0070] (2-2) High-depth sequencing results: The TSS characteristics of high-depth sequencing results were extracted in the same way as described above for low-depth sequencing, but only the standardized feature values 1 kb upstream and downstream of the TSS were retained as the output values for model construction.

[0071] (3) Input feature construction

[0072] (3-1) Acquisition of gene information: Extract the chromosome information, start site, end site, positive and negative strands, gene length, intergenic distance with the previous adjacent gene, the previous adjacent gene (left gene), and the next adjacent gene (right gene) of each gene;

[0073] (3-2) Feature information extraction: According to the obtained information of each gene, the feature value of each gene for simulating high-depth TSS results is constructed, which specifically includes the following: chromosome information of each gene, start site, end site, positive and negative chain, gene length, intergenic distance with the previous adjacent gene, rank of TSS 1kb normalized of each gene, rank of TSS 2kb normalized of each gene, rank of TSS 4kb normalized of each gene, rank of TSS 1kb normalized of each gene, rank of TSS 2kb normalized of each gene, rank of TSS 4kb normalized of each gene, rank of TSS 1kb normalized of the previous adjacent gene, rank of TSS 2kb normalized of the previous adjacent gene, rank of TSS 4kb normalized of the previous adjacent gene, rank of TSS 1kb normalized of the previous adjacent gene, rank of TSS 2kb normalized of the previous adjacent gene, rank of TSS 4kb normalized of the next adjacent gene, rank of TSS 1kb normalized of the next adjacent gene, rank of TSS 2kb normalized of the next adjacent gene, rank of TSS 4kb normalized of the next adjacent gene, rank of TSS 1kb normalized of the next adjacent gene, rank of TSS 2kb normalized of the next adjacent gene, and rank of TSS 4kb normalized of the next adjacent gene.

[0074] wherein the rank represents the rank of the TSS depth of the gene in all TSS depth of the genes in the sample.

[0075] (4) Constructing a model simulating high-depth sequencing TSS features

[0076] (4-1) Based on the input features obtained in (3-2) and the output features obtained in (2-2), the TSS input features and output features of 28,000 genes are obtained for each sample. In the present application, 588,000 data of 28,000 genes of 21 plasma free DNA samples are used for model construction, and 364,000 data of 28,000 genes of 13 plasma free DNA samples are used as a test set to verify the model effect;

[0077] (4-2) In the data for model construction, the data is divided into a training set and a validation set by 4:1, the method of generalized linear model (glm.nb) in machine learning is used, and the model parameters are optimized by 10-fold cross-validation method to confirm the final model. The model construction code is as follows:

[0078]

[0079]

[0080] wherein trainx is the input feature obtained from (3-2), and trainy is the output feature obtained from (2-2).

[0081] (5) The effect of the calculation model in simulating high-depth sequencing TSS features in the model construction and sequencing model samples was calculated, and the results are shown in Table 1.

[0082] Table 1

[0083]

[0084] Results: In the model construction sample, the R of the model prediction value and the true value was 0.7472±0.0570, and the MAE (mean absolute error) was 8.5536±0.9723; in the test model sample, the R of the model prediction value and the true value was 0.7359±0.0589, and the MAE was 8.8391±1.0636; there was no significant difference in R and MAE between the model construction and test samples, indicating that the model effect was relatively stable, and there was no overfitting phenomenon in the model construction sample data set.

[0085] Example 2

[0086] This example verifies the effect of the model constructed in Example 1.

[0087] (1) The correlation of the TSS values of low-depth sequencing before and after model simulation with the TSS values of high-depth sequencing was compared. The results are shown in Table 2, and the correlation of the TSS values of low-depth sequencing after model simulation with the TSS values of high-depth sequencing was significantly higher than that of the original low-depth sequencing TSS values with the TSS values of high-depth sequencing; Figure 2

[0088] (2) The number of TSS values of 0 indicates that the region has no sequencing reads coverage. The number of TSS values of 0 of low-depth sequencing before and after model simulation and high-depth sequencing TSS values was compared. The results are shown in Table 3, and the median value of high-depth sequencing TSS values of 0 was 94.5; the number of low-depth sequencing original TSS values of 0 was the most, with a median value as high as 106; the number of TSS values of 0 of low-depth sequencing after model simulation was significantly reduced, with a median value of 62. This result suggests that through the simulation of low-depth sequencing data by the model, the number of TSS values of 0 can be effectively reduced; Figure 3

[0089] (3) The coefficient of variation (CV) of gene TSS value reflects the stability of sequencing reads coverage in the TSS region of different samples. The coefficient of variation (CV) of TSS values of low-depth sequencing before and after model simulation and high-depth sequencing TSS values was compared. The results are shown in Table 4. Figure 4 ​​As shown, the median value of CV of high-depth sequencing TSS value is 16.5%; the CV of low-depth sequencing TSS value is the highest, and the median value is as high as 34.7%; the CV of the TSS feature predicted by the model is the highest and is significantly reduced, and the median value is 13.9%. The results show that by simulating the low-depth sequencing data by the model, the variability of the gene TSS value in the sample can be effectively reduced.

[0090] In summary, the present application extracts feature engineering and constructs a model based on low-depth sequencing results, and can obtain TSS features comparable to high-depth sequencing without high-depth sequencing of cfDNA, can effectively expand the number of analyzable TSS regions, increase the stability of the sequencing fragment coverage of the TSS region, and greatly reduce the sequencing cost while ensuring the accuracy of the TSS value calculation.

[0091] The applicant declares that the present application illustrates the detailed method of the present application by the above-mentioned embodiments, but the present application is not limited to the above-mentioned detailed method, that is, it does not mean that the present application must rely on the above-mentioned detailed method to be implemented. It should be understood by those skilled in the art that any improvement of the present application, equivalent replacement of each raw material of the product of the present application, addition of auxiliary ingredients, selection of specific methods, etc. fall within the protection scope and disclosure scope of the present application.

Claims

1. A method for simulating the characteristics of high-depth sequencing TSS, characterized in that, The method includes the following steps: (1) Obtain sequencing data of cell-free DNA samples captured from the reference genome and the sequence coverage of upstream and downstream regions of each gene's TSS; (2) Based on the low-depth sequencing results of cell-free DNA, obtain the characteristics of TSS values ​​for constructing simulated high-depth sequencing results; (3) Construct a TSS value model to simulate high-depth sequencing results based on features obtained from low-depth sequencing results of free DNA; Step (1) specifically includes the following steps: (1-1) Obtain the raw high-throughput sequencing data of the sample and compare it with the reference genome to calculate the coverage depth of each site in the upstream and downstream regions of the TSS of each gene; (1-2) The TSS of each gene is standardized according to formula (1) to obtain the characteristics of the upstream and downstream of the TSS of each gene; Equation (1); TSS i normalized TSS is the normalized value of the coverage depth of the upstream and downstream regions of the transcription start site region of gene i. i depth The total TSS represents the coverage depth of the upstream and downstream regions of the transcription start site region of gene i. depth The sum of the coverage depths of the upstream and downstream regions of the transcription start site region for all genes; The region is selected from a combination of at least three of the following sites located 0.5-1.5 kb, 1.5-2.5 kb, 2.5-3.5 kb, 3.5-4.5 kb, 4.5-5.5 kb, or 5.5-6.5 kb upstream or downstream of the gene transcription start site; Step (2) specifically includes the following steps: (2-1) Extract the chromosome information, start site, end site, positive and negative strands, gene length, gene interval with the previous neighboring gene, the previous neighboring gene and the next neighboring gene for each gene; (2-2) Based on the information of each gene obtained in (1-2) and (2-1), construct feature engineering for each gene to predict high-depth TSS results, and use it as input feature value for model construction; The feature values ​​include: Chromosomal information for each gene: start site, end site, positive and negative strands, gene length, gene interval to the previous neighboring gene, and three TSSs for each gene. normalized Values, three TSS values ​​for each gene normalized rank, the three TSS of the previous neighboring gene normalized Value, the three TSS of the previous neighboring gene normalized The rank of the value, the three TSS of the next neighboring genes normalized The value and the three TSS of the next neighboring gene normalized The rank of the value; The three TSS normalized Values ​​include: TSS 0.5-1.5kb normalized TSS 1.5-2.5kb normalized TSS 2.5-3.5kb normalized TSS 3.5-4.5kb normalized TSS 4.5-5.5kb normalized or TSS 5.5-6.5kb normalized Combinations of at least three; where rank represents the TSS of the gene. depth TSS of all genes in the sample depth The sorting.

2. The method for simulating high-depth sequencing TSS characteristics according to claim 1, characterized in that, Step (3) specifically includes the following steps: (3-1) Use machine learning methods to construct a mapping model in the training set based on the features confirmed in (2-2); (3-2) The model parameters were optimized using a 10-time cross-validation method to confirm the final prediction model.

3. The method for simulating high-depth sequencing TSS characteristics according to claim 2, characterized in that, The machine learning methods described in step (3-1) include any one or a combination of at least two of the following: logistic regression, support vector machine, random forest, decision tree, linear regression, or Naive Bayes.

4. The method for simulating high-depth sequencing TSS characteristics according to claim 1, characterized in that, The cell-free DNA sample includes any one or a combination of at least two of the following: cell-free DNA from plasma, cell-culture medium, or seminal plasma.

5. A device for simulating the characteristics of high-depth sequencing TSS, characterized in that, The device includes: The free DNA TSS analysis module, feature construction module, and model construction module; The analysis module of the free DNA TSS is used to perform the following: Obtain sequencing data of cell-free DNA samples captured from the reference genome and sequence coverage of upstream and downstream regions of each gene's TSS; The feature construction module is used to perform the following: Constructing features for TSS values ​​that simulate high-depth sequencing results; The model building module is used to perform the following: A TSS value model simulating high-depth sequencing results was constructed based on low-depth sequencing results of free DNA. The analysis module of the free DNA TSS is used to perform specific tasks including: (1-1) Obtain the raw high-throughput sequencing data of the sample and compare it with the reference genome to calculate the coverage depth of each site in the upstream and downstream regions of the TSS of each gene; (1-2) The TSS of each gene is standardized according to formula (1) to obtain the characteristics of the upstream and downstream of the TSS of each gene; Equation (1); TSS i normalized TSS is the normalized value of the coverage depth of the upstream and downstream regions of the transcription start site region of gene i. i depth The total TSS represents the coverage depth of the upstream and downstream regions of the transcription start site region of gene i. depth The sum of the coverage depths of the upstream and downstream regions of the transcription start site region for all genes; The region is selected from a combination of at least three of the following sites located 0.5-1.5 kb, 1.5-2.5 kb, 2.5-3.5 kb, 3.5-4.5 kb, 4.5-5.5 kb, or 5.5-6.5 kb upstream or downstream of the gene transcription start site; The feature construction module is used to perform specific tasks including: (2-1) Extract the chromosome information, start site, end site, positive and negative strands, gene length, gene interval with the previous neighboring gene, the previous neighboring gene and the next neighboring gene for each gene; (2-2) Based on the information of each gene obtained in (1-2) and (2-1), construct feature values ​​for each gene to predict high-depth TSS results, and use them as input feature values ​​for model construction; The feature values ​​include: Chromosomal information for each gene: start site, end site, positive and negative strands, gene length, gene interval to the previous neighboring gene, and three TSSs for each gene. normalized Values, three TSS values ​​for each gene normalized rank, the three TSS of the previous neighboring gene normalized Value, the three TSS of the previous neighboring gene normalized The rank of the value, the three TSS of the next neighboring genes normalized The value and the three TSS of the next neighboring gene normalized The rank of the value; The three TSS normalized Values ​​include: TSS 0.5-1.5kb normalized TSS 1.5-2.5kb normalized TSS 2.5-3.5kb normalized TSS 3.5-4.5kb normalized TSS 4.5-5.5kb normalized or TSS 5.5-6.5kb normalized Combinations of at least three; where rank represents the TSS of the gene. depth TSS of all genes in the sample depth The sorting; The model building module is used to perform specific tasks including: (3-1) Use machine learning methods to construct a mapping model in the training set based on the features confirmed in (2-2); (3-2) The model parameters were optimized using a 10-time cross-validation method to confirm the final prediction model.

Citation Information

Patent Citations

  • Non-invasive cancer early screening system based on cfDNA omics characteristics

    CN113160889A

  • Cancer detection model and construction method and kit thereof

    CN113838533A

  • Gene detection method, feature extraction method, device, equipment and system

    CN113517022A

  • Screening system of tissue-specific gene marker based on peripheral blood free DNA high-throughput sequencing and application of screening system

    CN115019888A