A method and apparatus for ultra-high sensitivity sample component evaluation and provenance

By constructing a SNP combination database and a variant abundance background model, the deficiencies in sample component evaluation and traceability in existing technologies are resolved, and accurate qualitative evaluation of multi-source components within the sample and precise quantitative calculation of cross-contamination ratios are achieved, which has better applicability in ctDNA detection, in particular.

CN119517162BActive Publication Date: 2025-10-21ZHENYUE BIOTECHNOLOGY JIANGSU CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510088279.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-10-21
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Existing wet experimental techniques cannot accurately quantify the proportion of components present in the sample, nor can they trace the non-self sources of the sample. Bioinformatics methods have biases in the evaluation of single-sample scenarios, and the contamination ratio estimation lacks sensitivity, making it difficult to determine low-frequency contamination signals.

Method used

By acquiring sample data, identifying SNP sites and their genotypes, building a SNP combination database, using the double allele genotypes of SNP sites for quantitative statistics, and constructing a variation abundance background distribution model, we can achieve qualitative evaluation of multi-source components in the sample and quantitative calculation of cross-contamination ratios, and conduct traceability analysis of non-self sources.

Benefits of technology

It realizes the precise qualitative evaluation of multi-source components in the sample and the accurate quantitative calculation of the cross-contamination ratio, can accurately trace non-self sources, and constructs sample component evaluation indicators and calculation frameworks that do not rely on control samples, which has better applicability in ctDNA detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517162B_ABST
    Figure CN119517162B_ABST
Patent Text Reader

Abstract

The application discloses a kind of ultra-high sensitivity sample component evaluation and traceability method and device, the method comprises: obtaining sample, carries out nucleic acid extraction, library construction, capture enrichment and sequencing, generates sequencing data;Sequencing data is preprocessed, and high-quality sequencing data is obtained;Identify the SNP site and its genotype that exist in high-quality sequencing data, construct SNP combination database;Quantitative statistics is carried out to the double allele genotype of SNP site, and the variation abundance of suballele is calculated;According to the genotype information and allele variation abundance of SNP site, construct SNP site variation abundance background distribution model;Using the SNP combination database and variation abundance background distribution model constructed, cross contamination evaluation is carried out to sample.The method of the application realizes the qualitative evaluation of multiple source components in sample, the quantitative calculation of cross contamination proportion and the traceability analysis of non-self source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sample component evaluation and traceability, and in particular to a method and device for ultra-high sensitivity sample component evaluation and traceability. Background Art

[0002] Currently, dual indexing technology is used during the wet lab phase, giving each sample a unique dual-index combination. This significantly reduces the probability of crosstalk occurring simultaneously on two independent indices, thereby reducing contamination caused by label crosstalk and mismatches. A non-template control (NTC) is a commonly used control sample in molecular diagnostic experiments. It is a reaction performed without the addition of template DNA to test for false positive results during the reaction and is used to assess contamination in reagents, procedures, or the experimental environment.

[0003] Bioinformatics methods are used to identify genetic variants within samples that have the ability to distinguish between different populations. When a genotype that is inconsistent with or does not conform to expected characteristics is detected in a sample, it can be used as a basis for evaluating the sample composition. There are generally two types of molecular signatures: short tandem repeats (STRs), each of which has multiple combinations of repeats, depending on the number of tandem repeats; and single nucleotide polymorphisms (SNPs), which typically have three genotypes (AA, AB, and BB) depending on the allele A / B combination. Due to the molecular genetics of tumor samples, the lack of genomic mismatch repair function can lead to high levels of biological background noise at STR loci, and chromosomal microdeletions can also render STR information ineffective. SNPs, however, are ubiquitous and easily screened, making them more suitable for ctDNA testing.

[0004] Existing wet lab techniques are limited to specific steps and scenarios. They cannot specifically identify samples of non-self-origin, cannot accurately quantify the proportions of components present in a sample, and cannot trace the source. Existing bioinformatics methods and software, such as GATK CalculateContamination, provide some SNP-based contamination assessment methods. However, the GATK method requires paired control sample information, which can lead to biased assessments in single-sample scenarios. Furthermore, its contamination ratio estimation method still lacks sufficient sensitivity, making it difficult to determine low-frequency (<1%) contamination signals.

[0005] Therefore, it is an urgent problem for those skilled in the art to propose an ultra-sensitive method and device for sample component evaluation and traceability to solve the difficulties existing in the prior art. Summary of the Invention

[0006] The purpose of the present invention is to provide an ultra-sensitive method and device for sample component assessment and traceability, which can realize the qualitative assessment of multi-source components in the sample, the quantitative calculation of the cross-contamination ratio, and the traceability analysis of non-self-source components, and construct a set of sample component assessment indicators and calculation framework that do not rely on control samples.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A method for ultra-sensitive sample component assessment and traceability, comprising the following steps:

[0009] S1. Obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, then perform nucleic acid extraction, library construction, capture enrichment, and sequencing to generate sequencing data;

[0010] S2. Preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then map the sequencing reads back to the reference genome;

[0011] S3. Identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build a SNP combination database, and record the SNP combination genotype information of each test sample;

[0012] S4. Quantitatively count the double alleles of the SNP loci, record the number of supporting A alleles and the number of supporting B alleles respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of the minor alleles;

[0013] S5. Construct a SNP site variation abundance background distribution model based on the genotype information and allele variation abundance of the SNP site;

[0014] S6. Utilize the constructed SNP combination database and SNP site variation abundance background distribution model to conduct cross-contamination assessment on samples, achieve qualitative assessment of multi-source components within the sample, quantitative calculation of cross-contamination ratio, and traceability analysis of non-self sources.

[0015] Preferably, in S2, preprocessing the sequencing data to obtain high-quality sequencing data and posting the sequencing reads back to the reference genome specifically includes:

[0016] The sequencing data is split into connectors. The split sequencing data is usually in fastq format, containing the read segment name, base sequence, identification symbol and sequencing quality value corresponding to the base sequence; the split sequencing data is removed from the connector sequence, low-quality sequence and single base repeat sequence to obtain high-quality sequencing data; the processed sequencing reads are aligned back to the reference genome through sequence alignment to form a bam format file.

[0017] Preferably, in S3, selecting appropriate SNP sites to construct a SNP combination database specifically includes:

[0018] Identify SNPs and their genotypes within the coverage of high-quality sequencing data. Genotype information is divided into pure wild-type AA, heterozygous variant AB, and pure variant BB. Combined with population data, calculate the population frequency of the minor allele of the SNP site. The calculation formula is as follows:

[0019] ;

[0020] in, is the number of pure and variant genotypes in the population, is the number of heterozygous variant genotypes in the population, is the total number of people;

[0021] Select appropriate SNP sites whose population frequencies have the ability to identify individuals. Candidate sites should meet the following requirements: site depth is not less than 100; the site is not located in the HLA or other repeat unit intervals; the site allele frequency should be between 0.2-0.8; and there should not be more than two genotypes with a population frequency greater than 1% at the site;

[0022] For n candidate SNP sites, calculate their combined heterozygosity to obtain a set of variant combinations with individual identification capabilities. The combined heterozygosity should be between 0.4 and 0.6. The formula for calculating the combined heterozygosity is as follows:

[0023]

[0024] in, is the population frequency of the allele at the i-th SNP site;

[0025] A database is constructed to record the SNP combination genotype information of each test sample, which is called the germline library; the genotype information of the SNP site is encoded as 0, 1, and 2, representing AA, AB, and BB, respectively, and the missing genotype is encoded as -1. The test samples are labeled by PID and SampleID.

[0026] Preferably, in S4-S5, constructing the SNP site variation abundance background distribution model based on the genotype information and allele variation abundance of the SNP site specifically includes:

[0027] Quantitative statistics are performed on the double allele genotypes of SNP sites, and the support numbers are Count A and Count B , the allele types are divided into major alleles and minor alleles according to the support number, and the abundance of minor allele variation is calculated as follows:

[0028]

[0029] In the absence of control samples, statistical tests are performed based on the abundance of allele variation to determine the corresponding genotype; genotypes are divided into pure and heterozygous types, represented by 0 and 1 respectively; the minor allele genotype is inferred as follows:

[0030] ;

[0031] The SNP set was further filtered using population data and minor allele information to remove sites with excessively high NA minor allele proportions. Afterwards, the population samples were statistically analyzed to remove samples with excessively high NA proportions. The remaining samples were used to construct the background distribution model. Four sets of baseline data were generated for each site using the corresponding allele and genotype combination:

[0032] ,

[0033] ,

[0034] ,

[0035] ,

[0036] in, is the sub-allelic variation abundance set of n SNP sites, represents the sample with minor allele A and genotype 0, represents a sample with minor allele A and genotype 1, represents the sample with minor allele B and genotype 0, The sample with the minor allele B and genotype 1 is represented. The data distribution is fitted on this basis, and the distribution with the best fit is selected as the baseline model.

[0037] Preferably, in S6, performing cross contamination assessment on the sample specifically includes:

[0038] By calculating the proportion of SNP sites with a minor allele type of 0 in the sample, that is, the pure sum coefficient, the sample is judged whether there is cross contamination; if the pure sum coefficient is less than 0.2, the sample is a highly contaminated sample; if the pure sum coefficient is greater than or equal to 0.2, and the proportion of SNP sites with a minor allele type of NA and a minor allele frequency of less than 5% is greater than or equal to 0.05, the sample is a low-contamination sample;

[0039] The sites with genotype NA are the unique sites of the pollution source; according to the pollution patterns 0 / 1, 0 / 2, and 1 / 0, the expected frequency of pollution is calculated. 、 、 , calculated by The difference between the expected frequency of pollution and the pollution ratio is calculated using the least squares method ;

[0040] Pollution ratio The prior population probability of the SNP site is used to obtain the posterior probability of the contamination mode 0 / 1, 0 / 2, and 1 / 0, which is then converted into the germline list of the individual and the germline list of the contamination source, and the sample germline library is matched respectively, and traceability verification is performed based on the matching consistency score.

[0041] An ultra-high-sensitivity sample component assessment and traceability device, used to perform any of the above ultra-high-sensitivity sample component assessment and traceability methods, comprising:

[0042] The sample preparation and sequencing module is used to obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, and then perform nucleic acid extraction, library construction, capture enrichment and sequencing to generate sequencing data;

[0043] The off-machine data preprocessing module is used to preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then posting the sequencing reads back to the reference genome;

[0044] The SNP combination database construction module is used to identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build the SNP combination database, and record the SNP combination genotype information of each test sample;

[0045] The module for constructing a variation abundance background distribution model is used to perform quantitative statistics on the double allele genotypes of SNP sites, record the number of supporting A allele genotypes and the number of supporting B allele genotypes respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of minor alleles. Based on the genotype information and allele variation abundance of the SNP site, a variation abundance background distribution model of the SNP site is constructed;

[0046] The sample cross-contamination assessment module is used to conduct cross-contamination assessment on samples using the constructed SNP combination database and SNP site variation abundance background distribution model, to achieve qualitative assessment of multi-source components in the sample, quantitative calculation of cross-contamination ratio, and traceability analysis of non-self sources.

[0047] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned methods for ultra-high sensitivity sample component evaluation and traceability.

[0048] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0049] (1) Based on NGS sequencing data, this method selects SNP sites with good coverage and alleles, screens suitable SNP combinations as molecular tags for individual identification, and constructs an individual genotype database for sample matching and traceability. On this basis, the double allelic genotype of the SNP site is used to calculate the variation frequency of the minor allele. The sample-level pure sum coefficient is used to make a qualitative judgment of cross-contamination. The population frequency of each SNP site is determined by population data. Based on the prior population frequency and the background error rate of the variation direction, the contamination ratio is estimated and the genotype of the contamination source is inferred.

[0050] (2) This invention is used for sample component assessment of NGS sequencing data. It can achieve qualitative assessment of multi-source components within a sample, quantitative calculation of cross-contamination ratios, and traceability analysis of non-self-source components. It also constructs a set of sample component assessment indicators and calculation frameworks that are independent of control samples. For ctDNA detection scenarios, it can achieve accurate cross-contamination assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1A schematic flow chart of a method for ultra-high sensitivity sample component assessment and traceability provided by the present invention;

[0053] Figure 2 This is a diagram showing the simulation results of the pure sum coefficient of the present invention;

[0054] Figure 3 This is a background distribution diagram of the abundance of SNP alleles in Example 1 of the present invention;

[0055] Figure 4 This is a distribution diagram of consistency scores between samples with consistent sources (paired) and inconsistent sources (unpaired) in the present invention;

[0056] Figure 5 This is a background distribution diagram of the abundance of SNP alleles in Example 3 of the present invention;

[0057] Figure 6 A comparison chart of the contamination ratio results of the method of the present invention (marked as new) and the GATK-Contamination method (marked as pre);

[0058] Among them, (a) is the comparison chart of the contamination ratio results of all samples under logarithmic coordinates, and (b) is the comparison chart of the results above 1% in any method. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] like Figure 1 As shown, the present invention provides an ultra-high sensitivity method for sample component assessment and traceability, comprising the following steps:

[0062] S1. Obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, then perform nucleic acid extraction, library construction, capture enrichment, and sequencing to generate sequencing data;

[0063] S2. Preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then map the sequencing reads back to the reference genome;

[0064] S3. Identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build a SNP combination database, and record the SNP combination genotype information of each test sample;

[0065] S4. Quantitatively count the double alleles of the SNP loci, record the number of supporting A alleles and the number of supporting B alleles respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of the minor alleles;

[0066] S5. Construct a SNP site variation abundance background distribution model based on the genotype information and allele variation abundance of the SNP site;

[0067] S6. Utilize the constructed SNP combination database and SNP site variation abundance background distribution model to conduct cross-contamination assessment on samples, achieve qualitative assessment of multi-source components within the sample, quantitative calculation of cross-contamination ratio, and traceability analysis of non-self sources.

[0068] The method of the present invention specifically comprises:

[0069] 1. Sample Preparation and Sequencing

[0070] After receiving the sample, the patient information (PID) is entered into the laboratory management system (LIMS) and a sample number (SampleID) is assigned. Different samples from the same patient share the same PID, and the sample number is unique within the laboratory. Nucleic acid extraction is performed on body fluid and tissue samples in different operating rooms, and products from different operating batches are stored in 96-well plates. After final trimming, the nucleic acid product is connected to the sample-specific adapter sequence (barcode index) for library construction. Library products are distributed according to product reagents (assays), and samples processed with the same assay are mixed and captured and enriched. Based on the expected data volume of different samples, sequencing weights are assigned and sequencing is performed on the machine. Based on the above operations, a detailed sample information table is generated.

[0071] 2. Off-line data preprocessing

[0072] Adapter splitting generates sample data, typically in fastq format, containing read name, base sequence, identifier, and sequencing quality value corresponding to the base sequence. Adapter sequences, low-quality sequences, and single-base repeats are removed from the raw data to generate high-quality sequencing data. The processed sequencing reads are then aligned back to the reference genome, generating a bam file.

[0073] 3. Screening SNPs and building a SNP combination database

[0074] Identify the SNPs and their genotypes within the assay (product reagent) coverage. Genotype information is divided into pure and wild type AA, heterozygous variant AB, and pure and variant BB. Combined with population data, calculate the population frequency of the minor allele at the SNP site. The calculation formula is as follows:

[0075] ;

[0076] in, is the number of pure and variant genotypes in the population, is the number of heterozygous variant genotypes in the population, is the total number of people;

[0077] To select appropriate SNP sites, their population frequency must have the ability to identify individuals. Candidate sites should meet the following requirements: the site depth is not less than 100; the site is not located in the HLA or other repeat unit intervals; the site allele frequency should be between 0.2-0.8; and there should not be more than two genotypes with a population frequency greater than 1% at the site.

[0078] For n candidate SNP sites, calculate their combined heterozygosity (H) to obtain a set of variant combinations that require individual recognition capabilities. The combined heterozygosity should be between 0.4 and 0.6. The calculation formula is as follows:

[0079]

[0080] in, is the population frequency of the allele at the i-th SNP site; a database is constructed to record the SNP combination genotype information of each test sample, called the germline library. The genotype information of the SNP site is encoded as 0, 1, and 2, representing AA, AB, and BB respectively. The missing genotype is encoded as -1. The test samples are annotated by PID and SampleID. An example is shown in Table 1 below:

[0081]

[0082] A matching score calculation program is constructed to evaluate the consistency between samples. The matching score can be used to confirm the pairing of multiple samples from the same patient, deny the pairing of samples from different patients, and trace the source of contaminated samples. There are two ways to calculate the consistency assessment, as follows:

[0083] Calculation method 1 is to include all SNP sites for pairwise comparison. The same coding is the consistent site, the different coding is the differential site, the proportion of consistent sites is the consistency score, and the highest score is the best match.

[0084] Calculation method 2 is to include only heterozygous sites for pairwise comparison, 0-1, 1-0, 1-2, and 2-1 are heterozygous inconsistent sites, and the heterozygous inconsistent ratio is the LOH score.

[0085] 4. Constructing a SNP site variation abundance background distribution model

[0086] The biallelic genotypes of SNP sites are quantitatively counted, and their support numbers are CountA and CountB respectively. The alleles are divided into major alleles and minor alleles according to the support numbers. The variation abundance of the minor allele is calculated as follows:

[0087]

[0088] In the absence of control samples, statistical tests are performed based on the abundance of allele variation to determine the corresponding genotype. Genotypes can be divided into pure and heterozygous types, represented by 0 and 1 respectively. The minor allele genotype is inferred as follows:

[0089] ;

[0090] Using population data and this allele information, the SNP set is further filtered to remove sites where the minor allele is too high for NA. Afterwards, the population samples are statistically analyzed, and samples with too high a NA ratio are removed. The remaining samples are used to construct the background distribution model. For each site, four sets of baseline data are generated based on the corresponding allele and genotype combination:

[0091] ,

[0092] ,

[0093] ,

[0094] ,

[0095] in, is the sub-allelic variation abundance set of n SNP sites, represents the sample with minor allele A and genotype 0, represents a sample with minor allele A and genotype 1, represents the sample with minor allele B and genotype 0, Represents a sample with the minor allele B and genotype 1.

[0096] On this basis, the data distribution is fitted, the parameters of different distributions such as normal distribution, t distribution, Weibull distribution, etc. are estimated, and the distribution with the best fit is selected as the baseline model.

[0097] 5. Sample cross-contamination assessment

[0098] Based on the assumptions of a 0.5 heterozygosity rate and a 20% contamination rate for a single SNP combination, a simulation was performed to demonstrate the changes in minor allele and allele frequencies within the sample. There were 4*4=16 contaminating genotype combinations: 0 / 0, 0 / 1, 0 / 1, 0 / 2, 1 / 0, 1 / 0, 1 / 1, 1 / 1, 1 / 1, 1 / 2, 1 / 2, 2 / 0, 2 / 1, 2 / 1, and 2 / 2. After contamination, the proportion of minor allele 0 decreased from 50% to 12.5%, and the proportion of minor allele 1 decreased from 50% to 25%. Because genomic copy number variation and loss of heterozygosity are common in tumor samples, these changes may affect heterozygous SNPs, resulting in a decrease in the proportion of minor allele 1 while maintaining the proportion of 0. Therefore, the proportion of homozygous sites under this SNP combination was selected as a qualitative assessment of cross-contamination within the sample and is referred to as the homozygous coefficient.

[0099] The above allele type background distribution model and genotype inference model are used for contaminated samples to determine the potential contaminated sites and record their allele types and allele variation abundance. Based on the possible contamination patterns 0 / 1, 0 / 2, 1 / 0, the expected frequency of contamination is calculated. 、 、 , by calculating the observed The difference between the expected frequency of pollution and the pollution proportion is estimated using the least squares method The specific calculation is as follows:

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] in, is the actual value of the minor allele variation abundance at the site, is the expected value of the variable abundance of components from other sources, is the variation abundance when other source components are heterozygous, The variation abundance when other source components are pure and type; according to the pollution ratio The prior population probability of the SNP site is used to obtain the posterior probability of the contamination mode 0 / 1, 0 / 2, and 1 / 0, which is then converted into the germline list of the individual and the germline list of the contamination source, and the sample germline library is matched respectively, and traceability verification is performed based on the matching consistency score.

[0106] An ultra-high-sensitivity sample component assessment and traceability device, applying any of the above-mentioned ultra-high-sensitivity sample component assessment and traceability methods, comprising:

[0107] The sample preparation and sequencing module is used to obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, and then perform nucleic acid extraction, library construction, capture enrichment and sequencing to generate sequencing data;

[0108] The off-machine data preprocessing module is used to preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then posting the sequencing reads back to the reference genome;

[0109] The SNP combination database construction module is used to identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build the SNP combination database, and record the SNP combination genotype information of each test sample;

[0110] The module for constructing a variation abundance background distribution model is used to perform quantitative statistics on the double allele genotypes of SNP sites, record the number of supporting A allele genotypes and the number of supporting B allele genotypes respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of minor alleles. Based on the genotype information and allele variation abundance of the SNP site, a variation abundance background distribution model of the SNP site is constructed;

[0111] The sample cross-contamination assessment module is used to conduct cross-contamination assessment on samples using the constructed SNP combination database and SNP site variation abundance background distribution model, to achieve qualitative assessment of multi-source components in the sample, quantitative calculation of cross-contamination ratio, and traceability analysis of non-self sources.

[0112] Specifically, the data preprocessing module uses fastp software to remove adapter sequences and low-quality sequences from the sample raw data, and uses bwa software to align the sequencing data to the human reference genome (hg19) to obtain a bam file.

[0113] The SNP combination database construction module uses varscan software to perform variant identification on the bam file and generate a vcf file. Candidate SNPs are screened from the vcf file according to the following criteria:

[0114] The site depth is not less than 100;

[0115] The site is not located in the HLA or other repeat unit intervals;

[0116] The allele frequency of the locus should be between 0.2 and 0.8;

[0117] There are no more than two genotypes with a population frequency greater than 1% at the locus.

[0118] The variant abundance background distribution model building module uses the pysam pileup method to identify the SNP alleles and corresponding support numbers in the test samples. The SNP site coverage reads must meet the following conditions:

[0119] The base quality value of site sequencing is not less than 30;

[0120] The site coverage reads are not unmapped, secondary mapped, or soft-clip mapped;

[0121] The site is located in the overlap interval of pair-end reads, and the base sequence of the overlap is consistent.

[0122] Thus, the number of genotypes supporting A allele Count is obtained. A and the number of supporting B alleles Count B , and calculate the genotype information in two ways:

[0123] One is the ref-alt allele method, which is divided into pure wild-type AA, heterozygous variant AA, and pure variant BB based on the consistency with the reference genome. The three genotypes are coded as 0, 1, and 2. This information is used for recording and matching the germline library. Each SNP site is fixedly recorded as chrom-pos-ref-alt as an identifier;

[0124] One is the minor allele method, in which genotypes are divided into homozygous type 0 and heterozygous type 1. This information is used to assess contamination ratio and genotype inference, and each SNP site is recorded as chrom-pos-minor_allele.

[0125] The sample cross-contamination assessment module uses the above method to identify each site in the form of ref-alt alleles of the screened SNP combination, and randomly generates genotype information 0, 1, and 2 according to the heterozygosity rate of 0.5 to initialize the germline database.

[0126] The germline database has writing and searching functions. When a new sample enters the germline database, the sample number information and genotype information are recorded. At the same time, the database is traversed, and the consistency is compared with the recorded historical samples one by one. The optimal matching score is calculated and the matching sample number information is obtained.

[0127] New samples are identified using a pileup approach to identify variant alleles at each SNP locus within the SNP set. Genotype sequences are then determined based on the SNP locus's abundance. A locus with a variation abundance less than 20% is labeled 0, a variation abundance between [20% and 80%] is labeled 1, and a variation abundance greater than 80% is labeled 2.

[0128] The consistency is determined based on the best matching score and the matching sample number information. The specific rules for consistency determination are as follows:

[0129] If it is the first record, the consistency score is less than 0.75, the sample consistency level is B, and if the score is greater than or equal to 0.75, the consistency level is C;

[0130] If there are different samples with the same PID and the consistency score is greater than or equal to 0.75, the consistency grade is A. If the score is less than 0.75, the consistency grade is D.

[0131] Grades A and B indicate a PASS consistency assessment result, while grades C and D indicate a FAIL consistency assessment result. Specifically, for samples with a high tumor proportion, if the LOH score is greater than 0.2, the result is LOH.

[0132] Component evaluation is performed by simulating SNP sites. The data simulation method is as follows:

[0133] (1) Randomly distribute the sequences of [0, 1, 1, 2] to form a SNP set sequence of 100 samples;

[0134] (2) The variation abundance of sites with genotype 0 is randomly generated according to the normal distribution with mean 0 and variance 0.01. The variation abundance of sites with genotype 1 is randomly generated according to the normal distribution with mean 0.5 and variance 0.01. The variation abundance of sites with genotype 2 is randomly generated according to the normal distribution with mean 1 and variance 0.01.

[0135] (3) Generate simulated data of mixed components in different proportions. The simulated data are given in the form of minor alleles, providing information on minor allele genotypes and minor allele variation abundance;

[0136] (4) Calculate the sample pure sum coefficient.

[0137] like Figure 2The simulation results of the pure sum coefficient are shown, and 0.2 is preliminarily determined as the threshold to distinguish between low-proportion pollution and high-proportion pollution:

[0138] The samples are piled up to identify the variant allele type information of each site in the SNP combination. Based on the variation abundance of the minor allele, the minor allele type is inferred as follows:

[0139] ;

[0140] in is a statistical test under the SNP background distribution model. The genotype background distribution and The genotype background distribution of The genotypes in the CI95 interval of the distribution are determined to be the corresponding genotypes, otherwise they are NA.

[0141] The presence of cross-contamination in a sample is determined by calculating the proportion of SNP sites with a minor allele of 0, also known as the homozygous coefficient. If the homozygous coefficient is less than 0.2, the sample is presumed to be highly contaminated. If the homozygous coefficient is greater than or equal to 0.2 and the proportion of SNP sites with a minor allele of NA and a minor allele frequency less than 5% is greater than or equal to 0.05, the sample is presumed to be lowly contaminated. Otherwise, the sample is presumed to be from a single source.

[0142] The loci with NA genotype are the loci specific to the pollution source. Based on the possible pollution patterns 0 / 1, 0 / 2, and 1 / 0, calculate the expected frequency of pollution. 、 、 , by calculating the observed The difference between the expected frequency of pollution and the pollution proportion is estimated using the least squares method .

[0143] According to the pollution ratio The prior population probability of the SNP site is used to obtain the posterior probability of the contamination mode 0 / 1, 0 / 2, and 1 / 0, which is then converted into the germline list of the individual and the germline list of the contamination source, and the sample germline library is matched respectively, and traceability verification is performed based on the matching consistency score.

[0144] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned methods for ultra-high sensitivity sample component evaluation and traceability.

[0145] In specific example 1, SNP combinations were screened based on whole-exome tumor population data from 90 cases. Samples were obtained, and the sample patient information was entered into the laboratory management system and a sample number was assigned. Nucleic acid extraction, library construction, capture enrichment, and sequencing were then performed to generate sequencing data. The sequencing data was preprocessed, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and the sequencing reads were linked back to the reference genome. SNP sites and their genotypes in the high-quality sequencing data were identified, and the population frequency of the minor allele of the SNP site was calculated in combination with the population data. Appropriate SNP sites were selected to construct a SNP combination database, and the SNP combination genotype information of each test sample was recorded.

[0146] The double allele genotypes of SNP sites were quantitatively counted, and the number of supporting A allele type and the number of supporting B allele type were recorded respectively. The allele types were divided into major allele and minor allele according to the number of support, and the variation abundance of minor allele was calculated; based on the genotype information and allele variation abundance of SNP sites, the variation abundance background distribution model of SNP sites was constructed; using the constructed SNP combination database and SNP site variation abundance background distribution model, the samples were evaluated for cross-contamination, and qualitative evaluation of multi-source components in the samples, quantitative calculation of cross-contamination ratio and traceability analysis of non-self sources were achieved. A total of 100 SNP sites were screened for sample component evaluation. Figure 3 Shown is the background distribution of SNP allele abundance: Figure 4 The figure shows the distribution of consistency scores between samples with consistent sources (paired) and inconsistent sources (unpaired). The x-axis shows the results of SNP sets with different numbers of sites:

[0147] In Example 2, we describe the provenance assessment results from a customized panel. This assessment was performed based on whole-exome and customized data. Data analysis followed the same approach as in Example 1. Samples from 200 patients were collected from various patient groups. Tumor tissue samples from the day of surgery were captured and sequenced using whole-exome sequencing to generate a somatic variant profile, which was then used to design a customized panel. The SNP combination generated in Example 1 was incorporated into the customized panel. Cell-free plasma DNA samples from multiple postoperative time points were captured and sequenced using a customized panel with the same PID recorded during the experimental process. White blood cell control samples served as validation data and were captured and sequenced using both whole-exome and customized panels. In the provenance results from the customized panel data, two samples initially tested matched to non-autologous germline database records and were subsequently excluded from clinical analysis. Further record verification revealed that these two discrepant matches were due to errors in tissue sample recording. All other plasma samples tested multiple times matched to their corresponding germline database records.

[0148] In the specific example 3, the traceability evaluation results in the customized panel were introduced. Based on the data of more than 1,000 cases of 769-gene tumor population, the components were evaluated and the contamination ratio was calculated. The data analysis was carried out using the analysis method in the above example 1. A total of 229 SNP sites were screened for sample component evaluation. Figure 5 Shown is the background distribution of SNP allele abundance; Figure 6 (a)-(b) show the comparison results of the contamination ratio between the method of the present invention (marked as new) and the GATK-Contamination method (marked as pre);

[0149] The specific classification statistics are shown in Table 2:

[0150]

[0151] Among them, samples with a contamination ratio of more than 0.8 calculated by GATK were found to be mismatched with paired normal samples due to sample labeling errors after traceability analysis, and no cross-contamination actually occurred in the samples.

[0152] In the specific example 4, the quantitative evaluation results of the gradient dilution samples are introduced. The mixed sample of the P3 cell line is spiked into the P5 cell line, and the spike gradient is 1%, 0.2%, 0.1%, 0.05%, 0.025%, and 0.0125%. Each gradient is repeated 3 times, and the original P5 sample without spiked in is sequenced 2 times, for a total of 20 samples. The data analysis adopts the analysis method in the above example 1. The 229 SNP site set is used to evaluate the contamination ratio. The results are consistent with expectations and can accurately evaluate 0.01% of non-self-source components. The evaluation calculation results are shown in Table 3 below:

[0153] Table 3

[0154] sample concentration contamination suspected_snp_num sample1 0.00% 0.000% 1 sample2 0.00% 0.000% 0 sample3 0.01% 0.004% 5 sample4 0.01% 0.008% 8 sample5 0.01% 0.008% 9 sample6 0.03% 0.012% 12 sample7 0.03% 0.023% 19 sample8 0.03% 0.018% 12 sample9 0.05% 0.044% 37 sample10 0.05% 0.066% 42 sample11 0.05% 0.047% 39 sample12 0.10% 0.076% 39 sample13 0.10% 0.086% 47 sample14 0.10% 0.102% 50 sample15 0.20% 0.220% 64 sample16 0.20% 0.202% 62 sample17 0.20% 0.192% 63 sample18 1.00% 0.966% 71 sample19 1.00% 0.997% 72 sample20 1.00% 0.976% 71

[0155] This paper constructs a complete computational framework and evaluation metrics for sample composition assessment during NGS sequencing. Specifically, the combined heterozygosity rate is used to assess the validity of SNP combinations, SNP genotypes are used to construct germline databases, background distribution models are used to assess SNP loci typing, minor allele categories are used to estimate genotypes in the absence of control samples, the homozygosity coefficient is used for qualitative assessment of non-self-derived components, the contamination ratio is a quantitative result of non-self-derived components, and the consistency score and LOH score are used to match and identify sample sources.

[0156] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0157] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A method for ultra-high sensitivity sample component assessment and traceability, characterized in that: The following steps are involved: S1. Obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, then perform nucleic acid extraction, library construction, capture enrichment, and sequencing to generate sequencing data; S2. Preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then map the sequencing reads back to the reference genome; S3. Identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build a SNP combination database, and record the SNP combination genotype information of each test sample; S4. Quantitatively count the double alleles of the SNP loci, record the number of supporting A alleles and the number of supporting B alleles respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of the minor alleles; S5. Construct a SNP site variation abundance background distribution model based on the genotype information and allele variation abundance of the SNP site; S6. Using the constructed SNP combination database and SNP site variation abundance background distribution model, perform cross-contamination assessment on the sample to achieve qualitative assessment of multi-source components in the sample, quantitative calculation of cross-contamination ratio, and traceability analysis of non-self sources; In S2, the sequencing data is preprocessed to obtain high-quality sequencing data, and the sequencing reads are mapped back to the reference genome, specifically including: The sequencing data is split into connectors. The split sequencing data is usually in fastq format, containing the read segment name, base sequence, identifier, and sequencing quality value corresponding to the base sequence; Remove adapter sequences, low-quality sequences, and single-base repeat sequences from the split sequencing data to obtain high-quality sequencing data; The processed sequencing reads are aligned back to the reference genome to form a bam format file; In S3, selecting appropriate SNP sites to construct a SNP combination database specifically includes: Identify the SNPs and their genotypes within the coverage of the high-quality sequencing data. The genotype information is divided into pure wild-type AA, heterozygous variant AB, and pure variant BB. Combined with the population data, calculate the population frequency of the minor allele of the SNP site. The calculation formula is as follows: Among them, N hom is the number of pure and variant genotypes in the population, N het is the number of heterozygous variant genotypes in the population, and N is the total number of people; Select appropriate SNP sites whose population frequencies have the ability to identify individuals. Candidate sites should meet the following requirements: site depth is not less than 100; the site is not located in the HLA or other repeat unit intervals; the site allele frequency should be between 0.2-0.8; and there should not be more than two genotypes with a population frequency greater than 1% at the site; For n candidate SNP sites, calculate their combined heterozygosity to obtain a set of variant combinations with individual identification capabilities. The combined heterozygosity should be between 0.4 and 0.

6. The formula for calculating the combined heterozygosity is as follows: Among them, p i is the population frequency of the allele at the i-th SNP site; A database is constructed to record the SNP combination genotype information of each test sample, called the germline library; the genotype information of the SNP site is encoded as 0, 1, and 2, representing AA, AB, and BB respectively, and the missing genotype is encoded as -1. The test samples are annotated by PID and SampleID; In S4-S5, based on the genotype information and allele variation abundance of the SNP site, a SNP site variation abundance background distribution model is constructed, which specifically includes: Quantitative statistics are performed on the double allele genotypes of SNP sites, and the support numbers are Count A and Count B , the allele types are divided into major alleles and minor alleles according to the support number, and the abundance of minor allele variation is calculated as follows: In the absence of control samples, statistical tests are performed based on the abundance of allele variation to determine the corresponding genotype; genotypes are divided into pure and heterozygous types, represented by 0 and 1 respectively; the minor allele genotype is inferred as follows: The SNP set was further filtered using population data and minor allele information to remove sites with excessively high NA minor allele proportions. Afterwards, the population samples were statistically analyzed to remove samples with excessively high NA proportions. The remaining samples were used to construct the background distribution model. Four sets of baseline data were generated for each site using the corresponding allele and genotype combination: Baseline1={freq m,1 ,freq m,2 ,...,freq m,n },when G A =0 Baseline2={freq m,1 ,freq m,2 ,...,freq m,n },when G A =1 Baseline3={freq m,1 ,freq m,2 ,...,freq m,n },when G B =0 Baseline4={freq m,1 ,freq m,2 ,...,freq m,n },when G B =1 Among them, {freq m,1 ,freq m,2 ,...,freq m,n } is the sub-allelic variation abundance set of n SNP sites, G A =0 represents the sample with minor allele A and genotype 0, G A =1 represents the sample with minor allele A and genotype 1, G B =0 represents the sample with minor allele B and genotype 0, GB=1 represents the sample with minor allele B and genotype 1; on this basis, the data distribution is fitted, and the distribution with the best fit is selected as the baseline model; In S6, cross contamination assessment is performed on the sample, specifically including: By calculating the proportion of SNP sites with the minor allele type of 0 in the sample, that is, the pure sum coefficient, the sample is judged whether there is cross contamination; if the pure sum coefficient is less than 0.2, the sample is a highly contaminated sample; if the pure sum coefficient is greater than or equal to 0.2, and the proportion of SNP sites with the minor allele type of NA and the minor allele frequency less than 5% is greater than or equal to 0.05, the sample is a low-contamination sample; The sites with genotype NA are the unique sites of the pollution source; according to the pollution patterns 0 / 1, 0 / 2, and 1 / 0, the expected frequency of pollution freq is calculated. 0 / 1 、freq 0 / 2 、freq 1 / 0 , by calculating the freq m The difference between the expected frequency of pollution and the pollution ratio θ is calculated using the least squares method; The contamination ratio θ and the prior population probability of the SNP site are used to obtain the posterior probabilities of the contamination patterns 0 / 1, 0 / 2, and 1 / 0, which are then converted into the germline list of the individual and the germline list of the contamination source, and matched against the sample germline library respectively, and traceability verification is performed based on the matching consistency score.

2. An ultra-sensitive sample component assessment and traceability device, characterized in that: The method for performing the ultra-high sensitivity sample component assessment and traceability according to claim 1 is characterized by comprising: The sample preparation and sequencing module is used to obtain samples, enter sample patient information and assign sample numbers through the laboratory management system, and then perform nucleic acid extraction, library construction, capture enrichment and sequencing to generate sequencing data; The off-machine data preprocessing module is used to preprocess the sequencing data, including removing adapter sequences, low-quality sequences, and single-base repeat sequences to obtain high-quality sequencing data, and then posting the sequencing reads back to the reference genome; The SNP combination database construction module is used to identify SNP sites and their genotypes in high-quality sequencing data, calculate the population frequency of the minor allele of the SNP site based on population data, select appropriate SNP sites to build the SNP combination database, and record the SNP combination genotype information of each test sample; The module for constructing a variation abundance background distribution model is used to perform quantitative statistics on the double allele genotypes of SNP sites, record the number of supporting A allele genotypes and the number of supporting B allele genotypes respectively, divide the alleles into major alleles and minor alleles according to the number of support, and calculate the variation abundance of minor alleles. Based on the genotype information and allele variation abundance of the SNP site, a variation abundance background distribution model of the SNP site is constructed; The sample cross-contamination assessment module is used to use the constructed SNP combination database and SNP site variation abundance background distribution model to perform cross-contamination assessment on samples, realize qualitative assessment of multi-source components in the sample, quantitative calculation of cross-contamination ratio and traceability analysis of non-self sources.

3. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for ultra-high sensitivity sample component evaluation and tracing as claimed in claim 1 is implemented.

Citation Information

Patent Citations

  • Sample quality evaluation method and device

    CN111477277A

  • Genotype detection method, sample pollution detection method, device, equipment and medium

    CN115035950A

  • Method for detecting sample cross contamination based on SNP (Single Nucleotide Polymorphism) sites and application of method

    CN117059164A