Method and system for detecting chimeric rate of human cells based on micro haploid

By adopting microhaploid-based NGS technology and probe capture technology in chimeric rate detection, the problem of insufficient detection sensitivity of the existing technology in microchime is solved, and higher detection accuracy and reliability are achieved, meeting the needs of clinical and scientific research fields.

CN120026098APending Publication Date: 2025-05-23WUHAN HAIXI BIOTECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411940613.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing chimera rate detection technology has insufficient sensitivity, especially when detecting microchimera (<1%), which makes it difficult to meet the needs of early recurrence prediction.

Method used

The NGS technology based on microhaploids was used to detect human cells chimeric rates. By screening 101 microhaploids related to Asian populations, and combining probe capture technology, a high-quality sequencing library was established to achieve higher detection sensitivity.

Benefits of technology

It improves the accuracy and reliability of chimeric rate detection, with a sensitivity of at least 0.1%, which can meet the needs of high-precision chimeric rate detection in clinical and scientific research fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120026098A_ABST
    Figure CN120026098A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for detecting the chimeric rate of human cells based on a micro haploid, and belongs to the technical field of chimera analysis. According to the invention, the micro haploid is innovatively adopted as a basic unit of the genetic marker, and different from the prior art in which DNA fragments containing the genetic marker are generally enriched through a multiplex PCR amplification technology, a probe targeted capture technology is introduced for enrichment; the limitation of the multiple PCR amplification technology on the number of markers and the potential influence of the PCR amplification efficiency difference on accurate quantification are effectively avoided; meanwhile, the invention further develops a special data analysis algorithm which is used for processing sequencing data and accurately calculating the chimeric rate. A verification result shows that the scheme of the invention has excellent reliability and accuracy, and the detection sensitivity reaches the highest level of similar products on the current market. Therefore, the method has a wide application prospect in the fields needing micro chimera detection, such as hematopoietic stem cell transplantation, organ transplantation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chimera analysis, and in particular to a group of microhaploids and a method and system for detecting the chimera rate of human cells based on the group of microhaploids. Background Art

[0002] Chimerism analysis is a methodology for detecting and quantifying the difference between donor cells and recipient cells and calculating the corresponding proportions. The chimerism rate is usually used to measure the degree of chimerism. For example, in the context of hematopoietic stem cell transplantation, the proportion of cells in the recipient's blood that come from the recipient itself is called the recipient chimerism rate, and the proportion of cells that come from the donor is called the donor chimerism rate. Monitoring the chimerism rate is of great significance for the success of hematopoietic stem cell transplantation surgery and the assessment of the risk of recurrence. It is also of great reference significance in similar contexts such as organ transplantation and forensic identification.

[0003] In recent years, with the advancement and innovation of technology, the detection methods of chimerism have been significantly developed. Early non-molecular methods include red blood cell phenotype-based and cytogenetics-based methods, such as fluorescence in situ hybridization (FISH) of sex chromosomes; while recent molecular methods rely on the detection of genetic markers such as restriction fragment length polymorphism (RFLP), microsatellites or short tandem repeats (STRs), and short insertion / deletion polymorphisms (SIDP). STRs-based methods have widely replaced less sensitive methods in the past 20 years, especially the method of using fluorescent-labeled multiplex PCR to amplify STR loci combined with fully automated capillary electrophoresis. Because of its high sensitivity, good reproducibility, simplicity and efficiency, it has become the gold standard for detecting chimerism recommended by the International Bone Marrow Transplant Registry (CIBMTR). Despite this, with the increase in clinical monitoring needs, the above-mentioned STR-PCR methods are still relatively lacking in sensitivity, and their detection limits are usually 1%-5%. However, more and more studies have shown that more sensitive detection methods have more advantages in predicting early recurrence of the disease, such as the detection of microchimerism (<1%). Therefore, with the emergence and maturity of technologies such as real-time quantitative PCR (qPCR), digital droplet PCR (ddPCR) and next-generation sequencing (NGS), a variety of more sensitive chimerism detection technology routes have been formed, such as SNP-qPCR or SIDP-qPCR methods. They use more advanced fluorescence quantitative methods and carefully selected polymorphic sites such as biallelic SNPs and short insertions or deletions (SIDPs) as substitutes for STRs, which can push the detection limit to 0.1% or even lower. In addition, the ddPCR-based method is considered to be the most sensitive method at present because it can achieve absolute quantification. Its sensitivity has been shown to reach 0.01% in some studies. The chimerism detection method based on the NGS technology route has a sensitivity between qPCR and ddPCR. Compared with PCR technology, the NGS technology route is no longer limited in number for genetic marker candidates, which can avoid the use of STR genetic markers that are easily interfered by PCR errors (such as stutter peaks), and allow hundreds or thousands of SNP or Indel sites to be used as genetic markers at the same time. In addition, the acceptable range of sample DNA input is also wider, thereby achieving highly sensitive and more stable (such as ensuring that the number of markers containing valid information is sufficient) chimerism rate detection.

[0004] The SNP-NGS technology route has been proven to be applicable to the detection of microchimerism and has guiding significance for clinical prognosis. At present, there are relevant mature kits on the foreign market, such as AlloSeq HCT (www.caredx.com / alloseq-hct) developed by CareDx and NGStrack (https: / / www.gendx.com / product_line / ngstrack / ) developed by GenDX. According to the information on the company's official website, AlloSeqHCT uses 202 non-linkage disequilibrium SNPs distributed on all autosomes as genetic markers, and its detection limit can reach about 0.22%; while NGStrack uses 34 Indels distributed on 19 autosomes, and its sensitivity can reach at least 0.5%. In addition, Agena is also promoting a chimeric rate detection kit called ChimericID (https: / / china.agenabio.com / products / panel / chimeric-id-panel / ), which contains 92 independent SNPs. However, there are few NGS-based chimeric rate detection kits on the domestic market. In addition, it should be noted that products such as AlloSeq HCT and NGStrack rely on multiplex PCR amplification technology when enriching DNA fragments containing genetic markers, which limits the upper limit of the number of genetic markers that can be used and may affect the accuracy of the quantitative chimerism rate. Summary of the invention

[0005] In view of this, the present invention aims to develop a method and system for detecting human cell chimerism based on microhaploidy, so as to improve the accuracy and reliability of NGS technology in chimerism detection, thereby meeting the urgent needs of clinical and scientific research fields for high-precision chimerism detection technology.

[0006] In a first aspect, the present invention provides a group of microhaploids for use in detecting chimerism in human cells, which at least includes the 101 microhaploids shown in Table 1.

[0007] Table 1 Microhaploids used for human cell mosaicism detection

[0008] The core of mosaicism analysis lies in accurately identifying and quantifying the differences in specific genetic markers (hereinafter referred to as markers) between donor and recipient cells, so a suitable marker is one of the keys to mosaicism analysis. In current products on the market and related research, STR, SNP and Indel are widely used as genetic marker units, while the present invention innovatively uses micro-haploids as genetic marker units for mosaicism detection. Micro-haploids are DNA sequence fragments containing multiple adjacent SNPs, and their genotype combinations are more diverse than those of single SNPs. Therefore, under the premise of keeping the number of markers consistent, the use of micro-haploid markers is expected to more efficiently achieve effective differentiation of genetic material between individuals.

[0009] In the technical solution of the present invention, after a large number of experimental verifications, 101 microhaploids related to the Asian population as shown in Table 1 were finally screened out from the microhaploid MicroHapDB-0.11 database. Experimental data show that the use of these microhaploids can effectively achieve accurate analysis of the chimerism rate.

[0010] In a second aspect, the present invention provides a method for detecting human cell chimerism rate based on microhaploids, wherein the microhaploids used are selected from a plurality of those in Table 1, and the method comprises the following steps: S1. Collect samples from donors and recipients before and after transplantation; S2, extracting sample DNA, and using capture probes specifically targeting markers to establish a sequencing library, and sequencing; S3, aligning the sequencing data with the reference genome to obtain an alignment result; S4. According to the comparison results, the genotype of each marker and the number of reads of each genotype are obtained, and then the marker validity is judged according to the difference of marker genotypes between different individuals, and then the marker chimerism rate of each donor or recipient is calculated one by one based on the validity marker; S5. The chimerism rates of all markers in the donor or recipient are averaged to obtain the final chimerism rate.

[0011] In step S1 of the above method, the sample type includes but is not limited to peripheral blood, tissue, etc. The method of the present invention can obtain the difference in marker genotypes between the donor before transplantation and the recipient before transplantation and the quantitative difference of different marker genotypes in the post-transplantation sample by collecting samples, extracting DNA, constructing libraries, sequencing, and analyzing data, thereby providing a data source for calculating the chimerism rate. In order to facilitate the understanding of the types of samples required for the above method, hematopoietic stem cell transplantation is used as an example. In the monitoring of chimerism rate after hematopoietic stem cell transplantation, the following three types of samples need to be collected: ① Pre-transplant recipient (ie, patient) blood samples or tissue samples not affected by transplantation, which are used to determine the marker genotype in the patient's own genetic material, and this sample only needs to be collected once; ② Pre-transplant donor blood or tissue samples, used to determine the donor's marker genotype, also only need to be collected once; ③ Post-transplant recipient peripheral blood samples, which can be used to dynamically monitor changes in chimerism rate.

[0012] It is understood that in the method of the present invention, the donor can be one or more; if there are multiple donors, each donor needs to collect samples and obtain their marker genotype. In the method of the present invention, the genotype of the marker specifically refers to: the base sequence composed of the specific bases corresponding to the SNP site contained in the microhaploid.

[0013] Preferably, in the above method, step S2 includes the following operations: S21, extracting sample DNA; S22, fragmenting the DNA molecules and then performing PCR amplification; S23, performing hybridization reaction between the capture probe specifically targeting the marker and the library obtained in step S22 to capture the target DNA fragment containing the marker, and further construct a sequencing library after enrichment; S24, performing high-throughput sequencing on the sequencing library obtained in step S23.

[0014] It is understood that the DNA extraction method, nucleic acid fragmentation treatment method, design method of capture probes for specific target markers, enrichment method, etc. involved in step S2 of the above method are not limited and can be directly performed using existing technologies or commercially available products. For example, in some embodiments of the present invention, the nucleic acid fragmentation treatment method is ultrasonic treatment.

[0015] More preferably, in step S22 of the above method, the length of the DNA molecule after fragmentation is 300±50 bp.

[0016] More preferably, in step S24 of the above method, the sequencing depth of the samples of the recipient after transplantation is not less than 2000x, and the sequencing depth of the samples of the recipient and the donor before transplantation is not less than 700x.

[0017] Preferably, in the above method, step S3 includes the following operations: S31, preprocessing sequencing data to improve data quality, wherein the preprocessing includes removing sequencing adapters and low-quality reads; S32, aligning the preprocessed data with the reference genome (hg38 version) to determine the sequence position and variation information; S33. Remove duplicate sequences and obtain a sequencing read segment alignment result file.

[0018] Preferably, in the above method, a marker with a sequencing depth lower than a preset threshold of 1 is defined as an invalid marker, that is, the validity judgment and chimerism rate analysis of the invalid marker are abandoned, and the invalid marker is not used for the final chimerism rate calculation. In some embodiments of the present invention, the preset threshold is 150.

[0019] Preferably, in the above method, the marker genotype is obtained and counted as follows: when a pair of reads (double-end sequencing, sequencing reads read1 and read2) covers all SNP sites of the marker, the corresponding bases of all SNP sites in the marker can be extracted from the corresponding reads, and then the genotype is obtained and the genotype count is increased by 1. For example, if the corresponding bases of the three SNPs contained in a marker in a pair of reads are A, T, and T, respectively, the count of the corresponding genotype "ATT" is increased by 1. It should be noted that when read1 and reads2 cover the same SNP site and give inconsistent base sequencing results, the one with the higher base sequencing quality shall prevail.

[0020] Since humans are diploid organisms, a marker in an individual can only have one or two genotypes, namely homozygous and heterozygous. However, sequencing errors can lead to the presence of pseudogenotypes in the statistical marker genotypes. Taking the microhaploid mh01SCUZJ-0533772 as an example, it contains 3 SNPs, and in the sequencing results of a sample, the three genotypes of TGG, TCG, and TTG were statistically shown, and the number of reads of the three genotypes were 121, 113, and 1, respectively, which clearly indicates the presence of pseudogenotypes.

[0021] The problem of false genotypes can be solved in the following way: according to the reading count result, the genotypes measured by the marker are sorted in reverse order according to their reading counts, and all genotypes whose genotype frequency (i.e., the ratio of the number of reads of the current genotype to the total number of reads of all genotypes) is less than the preset threshold value 2 are excluded. At this time, if only one genotype is left, the genotype of the marker is unique, that is, in the current individual, the marker presents a homozygous genotype, otherwise, the top two genotypes are taken as the final genotype results, that is, in the current individual, the marker presents a heterozygous genotype. In some embodiments of the present invention, the preset threshold value 2 is 15%. Taking the above-mentioned microhaploid mh01SCUZJ-0533772 as an example, the proportion of genotype TTG is very low (1 / (121+113+1)), which should be considered as a pseudogenotype produced by sequencing errors, while the proportions of genotypes TGG and TCG are very close. Therefore, it is determined that the microhaploid mh01SCUZJ-0533772 in this sample presents a heterozygous genotype of "TGG|TCG".

[0022] Preferably, in the above method, the marker validity judgment is specifically as follows: the genotype contained in a marker in a certain individual sample is compared with the genotype of the marker in all other individual samples. When the individual sample has a genotype that is not present in other individual samples, the marker is a valid marker for the sample. A valid marker means that the marker can provide valid information for the individual, and the unique genotype of the marker can be directly used to calculate the chimerism rate. For example, in a transplant case, suppose there is a marker containing 3 SNPs, and the genotype of the marker in the recipient is "TTT|AAT", while the genotype of the marker in the first donor is "AAT|ATC", and the genotype of the marker in the second donor is "ATT|ATC"; then for the recipient, "TTT" is its unique genotype, and the marker is a validity maker, and its unique genotype can be used directly to calculate the marker mosaicism rate of the recipient; for the second donor, "ATT" is its unique genotype, and its unique genotype can be used directly to calculate the marker mosaicism rate of the donor; and the first donor does not have a unique genotype, so the marker mosaicism rate of the first donor can only be obtained indirectly through the method of elimination.

[0023] Preferably, in the above scheme, the marker mosaicism rate is calculated as follows: Accumulate the number of reads of all genotypes under the marker to get the total number of reads; The number of genotype reads from the individual is obtained, and the statistical method is: ① If the genotype of the individual is homozygous, the number of genotype reads from it is equal to the read count of its unique genotype; ② If the genotype of the individual is heterozygous, and one of them is a unique genotype, the number of genotype reads from it is twice the unique genotype read count; ③ If the genotype of the individual is heterozygous, and both genotypes are unique, the number of genotype reads from it is the sum of the read counts of its two unique genotypes; The chimerism rate of the marker in an individual can be calculated by dividing the genotype reads from the individual by the total number of reads.

[0024] The following examples are given to help those skilled in the art understand the above chimerism rate calculation method: In the case of only one donor, taking the genetic markers containing 3 SNPs as an example, the donor has a heterozygous genotype of "ATC|ATT", while the recipient has a homozygous genotype of "ATC|ATC", and the sequencing results show that the counts of the genotypes ATC and ATT are 150 and 100, respectively. Comparing the genotypes of the donor and the recipient, it can be found that ATT is a unique genotype of the donor, while the recipient does not have a unique genotype, so the recipient mosaicism rate cannot be directly calculated. However, in the case of only one donor, the sum of the donor's mosaicism rate and the recipient's mosaicism rate can be considered to be 1; therefore, after directly calculating the donor's mosaicism rate, the recipient's mosaicism rate can be indirectly calculated. At the same time, the donor is a heterozygous genotype, so its unique genotype ATT only represents half of the donor's genotype count, and the total genotype count of the donor and the recipient is ATC+ATT=250; then the donor marker mosaicism rate can be calculated according to the formula (2×ATT) / (ATC+ATT) as (2×100) / 250=0.8; then indirectly, the recipient marker mosaicism rate can be calculated as 1-0.8=0.2. Obviously, if the recipient and donor genotype combinations in this example are swapped, the recipient mosaicism rate can be directly calculated, and no detailed examples will be given here.

[0025] When there are multiple donors (such as 2), if the recipient genotype is "TTT|AAT", and the genotypes of the first and second donors are "AAT|ATC" and "ATT|AAT" respectively, the recipient's chimerism rate can be calculated based on the count of the recipient's unique genotype TTT. Of course, if the recipient does not have a unique genotype, but each of the other donors has a unique genotype, then the donor's chimerism rate can be calculated first, and then the recipient's chimerism rate can be calculated indirectly. This will not be described in detail here.

[0026] Preferably, in the above method, the marker is split into multiple sub-markers, and then the genotype of the sub-marker and the number of reads of each genotype are obtained, and then the validity judgment and chimerism rate calculation are performed, and the marker chimerism rate is the average chimerism rate of its corresponding sub-marker.

[0027] Among the 101 microhaploids provided by the present invention, one marker contains 2-7 SNPs, and the longest marker length reaches 276bp; when building the library for sequencing, the captured DNA fragments are randomly interrupted and of varying lengths, so the sequencing reads may only contain part of the marker information. Splitting a complete marker into sub-markers can effectively utilize DNA fragments containing only part of the marker, which is conducive to improving the accuracy of chimerism calculation. Taking a marker containing 3 SNPs (numbered 1, 2, and 3) as an example, it can be split into 7 sub-markers: 1, 2, 3, 12, 13, 23, and 123.

[0028] It is worth noting that the shorter the length of the marker, the more likely it is that the corresponding number of sequencing reads will be, and the more reliable its statistical power will be. However, the risk of sequencing errors corresponding to shorter markers is also higher (e.g., the probability of mis-measuring one SNP is higher than the probability of mis-measuring two SNPs at the same time). To solve this problem, the present invention solves it by calculating the average value of the sub-marker mosaic rate as the marker mosaic rate. In some embodiments of the present invention, the average value is calculated as follows: all sub-markers under a marker are sorted in reverse order according to the total number of genome reads corresponding to them, and the top three sub-marker results (if the counts are the same, they are all included) are extracted as the representative of the quantitative results of the marker's mosaic rate for the current individual, and the average mosaic rate of the extracted sub-markers is calculated.

[0029] Preferably, in the above method, the chimerism rate obtained in step (5) is corrected, and the correction method is: when the chimerism rates of both the acceptor and the donor can be directly calculated, but the cumulative sum of the chimerism rates is not 1, the detection result of the donor or acceptor with the largest chimerism rate is discarded, and then the chimerism rate is replaced by 1 minus the cumulative sum of the chimerism rates of the remaining donors or acceptors. This is because in practice, the inventors found that the sequencing results of the donors or acceptors with a higher proportion have more sequencing deviations, and the above operation helps to obtain more accurate quantitative results.

[0030] In a third aspect, the present invention provides a system for detecting human cell mosaicism based on microhaploids, wherein the system uses the 101 microhaploids shown in Table 1 as markers for detection and comprises at least the following modules: The sequencing module constructs high-quality sequencing libraries corresponding to the donor and acceptor based on the capture probe targeting the marker and performs sequencing; The data processing module is used to analyze the sequencing data and determine the genotype of the marker, the number of reads of each genotype, and the validity based on the analysis results; The chimerism rate calculation module calculates the chimerism rate based on the marker genotype, validity and number of reads.

[0031] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention innovatively introduces microhaploids as the basic unit of genetic markers. This strategy significantly improves the effectiveness of genetic markers. At the same time, during the data analysis process, multiple SNPs are simultaneously examined using microhaploids as the analysis unit, thereby effectively reducing the interference of sequencing errors on the results and improving the accuracy and reliability of the data.

[0032] (2) Based on the micro-haploid strategy, the present invention successfully uses probe capture technology to replace multiplex PCR amplification technology to enrich target DNA fragments, thereby achieving accurate quantification of chimerism rate. This innovation fundamentally breaks the bottleneck of limited marker number in traditional methods, making it possible to detect hundreds or thousands of genetic markers at the same time. In addition, by simply mixing the probes that capture these markers into NGS products that are also based on probe capture, the existing products can gain the additional function of detecting chimerism rate. One product has multiple functional uses, which greatly reduces the detection cost for clinical testing or medical research.

[0033] (3) Based on the micro-haploid + probe capture technology, the 101 micro-haploids provided by the present invention contain a total of 312 SNPs, which is much higher than the number of markers used in all existing products on the market. This is expected to avoid quantitative deviations caused by sampling errors or PCR amplification preferences, and provide more accurate and stable detection efficiency for the chimerism rate.

[0034] (4) With regard to the strategy of combining micro-haploid and probe capture, which is adopted for the first time in the present invention, the present invention provides a new analysis scheme for the obtained sequencing data, that is, it is no longer limited to the traditional SNP-by-SNP analysis mode, but rather performs synchronous or group analysis on multiple SNPs contained in the micro-haploid. This micro-haploid-based analysis strategy can effectively reduce the impact of sequencing errors, which is particularly critical for the accurate detection of micro-chimerism rate.

[0035] (5) The present invention uses NGS technology to perform chimerism analysis with a sensitivity of at least 0.1%. At the same time, it uses micro-haploids as genetic marker units, which has better sensitivity and can meet the clinical needs for micro-chimerism detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings used in the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and for those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 Flow chart of detecting human cell chimerism rate based on microhaploid in an embodiment of the present invention Figure 2 It is the correlation analysis result between the expected chimerism rate value and the actual chimerism rate detection value in the embodiment of the present invention. DETAILED DESCRIPTION

[0038] The following embodiments of the technical solution of the present invention are described in detail in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and are therefore only used as examples, and cannot be used to limit the protection scope of the present invention.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention belongs; the terms "including" and "having" and any variations thereof in the specification and claims of the present application are intended to cover non-exclusive inclusions.

[0040] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0041] With the development of technology, the traditional STR-PCR technology route faces significant limitations in the field of chimerism detection, and its detection limit is usually limited to the range of 1% to 5%. This limitation means that in key application scenarios such as early disease recurrence prediction, traditional technology may not provide sufficiently sensitive detection methods, especially when it comes to the detection of microchimerism (chimerism rate less than 1%). In order to overcome this challenge, researchers have developed new detection methods such as SNP-qPCR or SIDP-qPCR, which successfully reduce the detection limit to 0.1% or even lower by screening new biallelic SNPs and polymorphic sites such as short insertions or deletions. The ddPCR-based method has become one of the most sensitive detection methods at present, with its advantage of being able to achieve absolute quantification, and its sensitivity can even reach 0.01%. However, despite the significant progress in sensitivity of these methods, they may still be restricted by other factors in practical applications, so there are few related mature products on the market.

[0042] At the same time, the emergence of NGS technology routes has provided new possibilities for chimerism detection. Although its sensitivity is between qPCR and ddPCR, NGS technology is highly favored for its high resolution (accurate to the base level) and wide selection of genetic markers, and there are more and more related research reports and market products. Although NGS-based chimerism detection products have appeared on the foreign market, such as AlloSeq HCT and NGStrack, these products still rely on multiplex PCR amplification technology when enriching DNA fragments containing genetic markers. The technical difficulty of this technical route in multiple primer design increases sharply with the increase in the number of markers, and thus still potentially limits the upper limit of the number of markers that can be used. Existing studies (https: / / doi.org / 10.1016 / j.cca.2022.05.026) point out that for biallelic markers (commonly used are SNP / Indel), generally at least 40 markers are required to meet 99% of clinical application scenarios. Based on the results of literature retrieval, it can be found that the number of markers contained in most products or studies, especially those based on PCR technology, is usually less than 40. When the number of markers is insufficient, the stability and reliability of the product's detection capabilities will be affected. Especially in transplantation cases between direct relatives or relatives, due to the high similarity of DNA, the number of markers that can distinguish between donors and recipients will be significantly reduced, which in turn affects the accuracy of the quantitative chimerism rate.

[0043] In order to make up for the defects of products used for mosaicism detection in the prior art and provide a richer and more reliable choice for mosaicism detection, the present invention has developed a scheme for human cell mosaicism detection based on micro-haploids. Different from the AlloSeq HCT and NGStrack methods that use independent SNPs or Indel as genetic markers, the scheme of the present invention uses micro-haploids as genetic markers (the term "micro-haploid" was proposed in 2013 and is now widely used). The term "micro-haploid" specifically refers to a fragment length of less than 300bp, containing two or more allele combinations with linkage disequilibrium SNPs. The micro-haplotype locus contains multiple SNP sites, so it belongs to a multi-allelic genetic marker, which contains richer genetic information than a single SNP site, but its mutation rate is 5 to 6 orders of magnitude lower than STR. Micro-haplotypes currently have many studies and applications in the field of forensic medicine, and can be used for DNA mixture differentiation and paternity testing. More than 3,000 publicly published micro-haplotype type markers have been included in the current MicroHapDB database.

[0044] On the one hand, the embodiment of the present invention screens out a group of microhaploids from the microhaploid database by optimizing the marker selection strategy, specifically including 101 microhaploids, each of which contains 2 to 7 SNPs, totaling 312 SNPs, which is much higher than the number of markers used in existing products, and can provide higher individual differentiation. Especially in the case of multi-donor transplantation, the number of markers of the present invention is more sufficient, and it also helps to avoid quantitative deviations caused by sampling errors or PCR amplification preferences, thereby improving the reliability of the detection results.

[0045] In a second aspect, an embodiment of the present invention provides a method for detecting the chimerism rate of human cells based on microhaploidy, such as Figure 1 As shown, it at least includes the following steps: S1. Collect samples from donors and recipients before and after transplantation; S2, extracting sample DNA, and using capture probes specifically targeting markers to establish a sequencing library, and sequencing; S3, aligning the sequencing data with the reference genome to obtain an alignment result; S4. According to the comparison results, the genotype of each marker and the number of reads of each genotype are obtained, and then the marker validity is judged according to the difference of marker genotypes between different individuals, and then the marker chimerism rate of each donor or recipient is calculated one by one based on the validity marker; S5. The chimerism rates of all markers in the donor or recipient are averaged to obtain the final chimerism rate.

[0046] Existing chimerism detection schemes based on NGS technology usually use multiplex amplification PCR technology to enrich genetic marker fragments. In this scheme, the product has a single purpose, and the difficulty of primer design increases exponentially with the increase in the number of markers used, resulting in the number of markers commonly reported to be less than 50, and very few can reach 202 SNPs. Compared with the prior art, the present invention uniquely uses probe targeted capture technology to enrich genetic marker fragments. The conversion of this technology breaks through the upper limit of the number of markers that can be used. At the same time, the inventor's research shows that these probes for chimerism detection can be mixed with the probes of the existing NGS panel to achieve the purpose of simultaneously realizing chimerism detection and the panel itself in the same panel, which is difficult to achieve based on the multiplex PCR amplification technology scheme.

[0047] However, compared with multiplex PCR amplification technology, probe-targeted capture technology is very likely to contain incomplete marker sequences in the enrichment process due to the fragmentation of nucleic acids, which affects the effective use of data and the reliability of analysis results. To address this problem, the present invention effectively solves the problem by splitting markers into sub-markers and coordinating intra-group calculations.

[0048] Some specific examples are listed below. It should be noted that the examples described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. If no specific techniques or conditions are specified in the examples, the techniques or conditions described in the literature in this field or the product instructions are used. If the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be obtained commercially.

[0049] Example 1 This example provides a set of microhaploid markers that can be used for mosaicism analysis. The process of obtaining this set of haploids includes the following: First, following the designed screening criteria, microhaplotypes were screened from the microhaplotype database (GitHub-bioforensics / MicroHapDB: Portable database of microhaplotype marker and allele frequencydata) as candidate genetic markers. Specifically, the screening criteria include: ① The marker must be located on the autosome; ② The marker length (the span from the first SNP to the last SNP) ranges from 25 to 300 bp; ③ Based on the genotype frequency in the East Asian population, the effective number of alleles (AE value) is calculated and screened. ④ The number of SNPs in the microhaploid is limited to between 2 and 7; ⑤ The SNP in the marker must not be located in the low-complexity region of the genome or the single-base tandem (repeat unit length greater than 4) region; ⑥ Among the known genotypes of the marker, if there is only one base inconsistency between more than two genotypes and the difference is "A / T" or "C / G", then the marker will be discarded. The purpose of this step of filtering is to minimize the quantitative interference caused by common sequencing errors. The microhaploids in the database that meet the above conditions are determined as candidate genetic markers. Secondly, we attempted to design NGS capture probes for the candidate genetic markers obtained, and through analysis of pre-experimental data and the sequence environment of the SNP, we eliminated markers with consistently low sequencing depth or poor mosaicism prediction results. Ultimately, we retained 101 high-quality markers that could be used for subsequent mosaicism analysis, as shown in Table 2.

[0050] Table 2 Information of microhaploid markers

[0051] Note: The starting coordinates, ending coordinates and SNP location coordinates all correspond to the GRCH38 version.

[0052] Example 2 Based on the microhaploid provided in Example 1, this example provides a method for detecting the chimerism rate of human cells, comprising the following steps: (1) Collect pre-transplant donor and recipient samples and post-transplant recipient samples.

[0053] (2) The samples collected in step (1) are processed as follows, including the following operations: ① Extract DNA molecules from the sample; ② Fragment the DNA molecules (e.g., ultrasonic treatment) into about 300±50bp; ③ The obtained DNA fragments are further processed to construct a library suitable for sequencing, including fragment end repair, ligation of adapter sequences, etc.; ④ Amplify the library by PCR. The number of amplification cycles can be adjusted according to the initial DNA input to ensure the adequacy of the library and the feasibility of sequencing; ⑤ Perform hybridization reaction between the capture probe that specifically targets the marker and the library obtained in step ④ to accurately capture the gene sequence containing the marker, and then use the magnetic bead method to efficiently enrich the target area captured by the probe, thereby obtaining a high-quality sequencing library; ⑥Perform a comprehensive quality test on the library constructed in step ⑤ to ensure that the concentration, purity, fragment size distribution and other indicators of the library meet the requirements of the sequencing platform; ⑦ Use the Illumina sequencing platform and the PE150 sequencing strategy to perform high-throughput sequencing on the library; among them, the sequencing depth of the library corresponding to the recipient sample after transplantation should reach about 2000x, and the sequencing depth of the library corresponding to the donor and recipient samples before transplantation should be no less than 700x.

[0054] In the above step ⑤, since the microhaploid sequence is known and designing specific capture probes for specific sequences is an existing technology, this example will not elaborate on the capture probes. Those skilled in the art can design them by themselves or directly entrust a biosynthesis company to prepare them.

[0055] (3) Processing of sequencing data.

[0056] First, Fastp (https: / / github.com / OpenGene / fastp) software was used to remove sequencing adapters and low-quality reads (removed according to the software default parameters) to improve data quality; then, Bwa mem (https: / / github.com / lh3 / bwa) software was used to align the obtained data with the hg38 reference genome to determine the sequence position and variation information; finally, Gencore (https: / / github.com / OpenGene / gencore) software was used to remove the repeated sequences generated by PCR according to the alignment position of the sequencing reads and the software default parameters to obtain the final sequencing read alignment result file.

[0057] (4) Chimerism analysis

[0058] The analysis process is mainly implemented using the Python programming language. The data analysis logic mainly uses pysam (https: / / github.com / pysam-developers / pysam) to parse the result file obtained in the comparison step (3), thereby determining the genotypes and counts of each marker in the recipient and donor, and then calculating the chimerism rate. The specific operations include the following: ① Filter out markers with a sequencing depth of less than 150 among the 101 markers and define them as invalid markers, that is, they will not participate in the subsequent chimerism rate analysis; ② Split the retained marker (i.e., the marker with complete sequence, hereinafter referred to as the main marker) into multiple sub-markers, and count the reads of each genotype containing the sub-marker. The counting method is: when a pair of reads covers all SNP sites of the sub-marker, it is a valid count.

[0059] ③ By comparing the sub-marker genotypes of the donor and recipient, determine whether the sub-marker is effective; ④ After counting the genotype read counts of each sub-marker, the chimerism rate analysis is performed for each donor or recipient one by one, and calculated using the following formula: Mosaicism rate = number of genotype reads from a certain individual / total number of reads; Among them, the total number of reads is the sum of the number of reads of all genotypes of the sub-marker. The number of genotype reads from an individual is specifically as follows: if the sub-marker genotype of the individual is homozygous, the number of genotype reads from it is equal to the read count of its unique genotype; if the sub-marker genotype of the individual is heterozygous, and one of them is a unique genotype, the number of genotype reads from it is twice the unique genotype read count; if the sub-marker genotype of the individual is heterozygous, and both genotypes are unique, the number of genotype reads from it is the sum of the read counts of its two unique genotypes.

[0060] ⑤ Summarize the chimerism analysis details of all sub-markers. After the summary, conduct a detailed analysis for each individual (donor or recipient).

[0061] For the mosaicism rate of individuals, they were grouped according to the main marker and sorted in reverse order according to the total number of genotype reads corresponding to the sub-marker within the group; then, the top three sub-marker results (if the counts were the same, they were all included) were extracted as the representative of the quantitative results of the main marker for the mosaicism rate of the current individual; after processing all markers, the Grubbs method was used to identify and eliminate outliers in the mosaicism rate to ensure the robustness of the results; when identifying outliers, the alpha parameter was set to 0.05.

[0062] Next, the results of the sub-markers were grouped again according to the main marker to which they belonged, and the average value within the group was calculated as a representative of the chimerism rate indicated by the main marker.

[0063] Finally, the average chimerism rates indicated by all main markers are further averaged to obtain the final chimerism rate of the current individual.

[0064] After calculating the average chimerism of all donors and recipients, check whether the sum of the chimerism of the recipient and the donor is 1. If it is not 1, find the maximum chimerism in the donor or recipient and subtract the sum of the remaining chimerism except the maximum chimerism from 1 to get the corrected maximum value. For example, if the average chimerism of the donor and recipient is 0.1 and 0.92 respectively (the sum is not 1), the adjusted chimerism of the donor and recipient will be 0.1 and 0.9 respectively (that is, 1 minus 0.1).

[0065] In this method, splitting the marker into sub-markers can maximize the use of all sequencing reads. At the same time, during the data analysis process, the accuracy and robustness of the chimerism rate calculation are effectively improved by integrating the information of multiple sub-markers and the main marker.

[0066] Example 3 Based on the method in Example 2, this example used pre-transplantation samples and donor samples from real patients to construct simulated samples covering 15 chimerism gradients, and used the simulated samples to verify the accuracy and reliability of the method of the present invention, including the following operations: (1) Sample preparation.

[0067] Sufficient amount of DNA was extracted from the patient samples and donor samples before transplantation, and the DNA was mixed according to the 15 preset recipient and donor ratios (i.e., chimerism gradient: 0.001, 0.003, 0.005, 0.01, 0.02, 0.05, 0.1, 0.25, 0.5, 0.75, 0.9, 0.99, 0.995, 0.997, 0.999) to prepare 15 simulated samples.

[0068] (2) Detection.

[0069] Each simulated sample was tested three times; the real patient pre-transplantation sample and donor sample were tested once each to determine the genotype information of the donor and recipient. Therefore, the sequencing data of 47 samples were tested and analyzed in this case, and the test results are shown in Table 3.

[0070] Table 3 Statistics of chimerism detection results of different simulated samples

[0071] (3) Data analysis.

[0072] Linear regression analysis was used to compare the correlation between the expected chimerism rate values ​​and the true chimerism rate detection values.

[0073] The results are as follows Figure 2 As shown, the linear correlation coefficient between the two is as high as 0.9977, the R-square is close to 1, and the coefficient of variation of the repeated sample detection results is less than 10%, indicating the stability and accuracy of the method of the present invention.

[0074] Example 4 Based on the method in Example 2, clinical samples were collected in this example for chimerism detection, as follows: (1) Sample collection.

[0075] In this case, peripheral blood samples of 8 clinical patients and their corresponding donor samples were collected. Among them, 3 patients received two transplants from different donors (one bone marrow transplant and one umbilical cord blood transplant), and the rest received only bone marrow transplants. In addition, since the peripheral blood sample before transplantation of one patient could not be obtained, the nail sample was used as a substitute to obtain the genotype information before transplantation. In addition, 2 patients underwent chimerism monitoring 2 and 3 times at different time points after transplantation.

[0076] (2) The chimerism rate of each patient's post-transplantation sample was tested (repeated 3 times), and the detection results are shown in Table 4.

[0077] Table 4 Statistics of chimerism detection results of different clinical samples

[0078] As a control, this case also used the existing STR-PCR technology to test each patient's post-transplantation sample once. Comparing the two methods, the following conclusions can be drawn: 1. When the chimerism rate is higher than 1%, from a qualitative perspective, whether the two methods are completely consistent in detecting the chimerism rate; that is, if one technology detects chimerism, the other technology also detects chimerism; if one technology does not detect chimerism, the other technology also does not detect it.

[0079] 2. In the range of chimerism rate between 1% and 10%, although there are certain differences in the quantitative results between the method of the present invention and the STR-PCR technology, there is a high linear correlation between the two. This shows that although the specific values ​​may be different, the chimerism trends detected by the two technologies in this chimerism rate range are consistent.

[0080] 3. When the chimerism rate is higher than 20%, the quantitative results of the method of the present invention and the STR-PCR technology are relatively consistent, indicating that the two technologies have similar detection accuracy within this high chimerism rate range.

[0081] 4. For some post-transplant samples, the method of the present invention indicated the presence of microchimerism. This finding suggests that the method of the present invention may have a higher sensitivity and be able to detect low-level chimerism.

[0082] In summary, the microhaploid and the scheme for detecting human cell chimerism rate based on microhaploid provided by the present invention can successfully realize chimerism rate detection with high reliability, high accuracy and high sensitivity, and can meet the urgent needs of clinical and scientific research fields for high-precision chimerism rate detection technology.

[0083] It should be noted that the present invention is not limited to the above-mentioned embodiments. The above-mentioned embodiments are only examples, and the embodiments having the same structure as the technical idea and exerting the same effect within the scope of the technical solution of the present invention are all included in the technical scope of the present invention. In addition, within the scope of the main purpose of the present invention, various modifications that can be thought of by those skilled in the art are applied to the embodiments, and other methods of combining some of the constituent elements in the embodiments are also included in the scope of this application.

Claims

1. Application of a group of microhaploids in detecting chimerism in human cells, characterized in that: At least including the following microhaplotypes: mh07SCUZJ-0428708, mh10WL-004, mh06WL-028.v1, mh13SCUZJ-0304208, mh19WL-028, mh11WL-039, mh10SCUZJ-0206566, mh06WL-064, mh16PK-83544, mh06USC-6pA, mh01ZHA-012.v1, mh18LS-18qA, mh07LS-7pF, mh18WL-026, mh13LS-13qE, mh02LS-2qG, mh13LS-13qC, mh16HYP-36, mh03LS-3pC, mh07WL-007.v1, mh01HYP-02, mh07LS-7qE, mh03LS-3pA, mh18WL-027.v2, mh02LS-2qB, mh03LS-3pD, mh06LS-6pE, mh03LS-3pE, mh14USC-14qB, mh02LS-2qC, mh04WL-091, mh09USC-9qA, mh08HYP-23, mh06KK-031, mh10ZHA-002, mh04WL-074.v2, mh02KK-215, mh03HYP-09, mh11KK-091, mh02LS-2qH, mh10LS-10pA, mh02LS-2qF, mh04USC-4qA, mh10LS-10qA, mh11KK-039, mh01LS-1qB, mh10LS-10qG, mh06LS-6qA, mh02USC-2pA, mh05KK-122, mh11KK-040, mh03KK-006, mh15CP-003, mh15LS-15qD, mh13KK-213.v4, mh02WL-020, mh06HYP-18, mh08LS-8qD, mh05LS-5qF, mh20NH-26, mh08LS-8qA, mh22KK-069, mh10LS-10qE, mh10KK-086, mh10LS-10qD, mh17LS-17qD, mh03NH-07, mh05LS-5qC, mh04LS-4pA, mh05KK-120.v1, mh02USC-2pB, mh18ZHA-005, mh04LS-4qF, mh18LW-46, mh10WL-060, mh01WL-115, mh05WL-049, mh06WL-080, mh03WL-087, mh07WL-079, mh06WL-063.v2, mh04WL-092, mh01WL-038.v3、mh08SCUZJ-0476324、mh10WL-055、mh01WL-044、mh13WL-014、mh20WL-023.v3、mh20HYP-41、mh04KK-015、mh04LS-4pE、mh05LS-5qD、mh01USC-1qC.v1、mh12LS-12qB、mh10KK-101、mh07LS-7pB、mh09KK-035.v1、mh01SCUZJ-0533772、mh03LS-3qA、mh04LS-4pD、mh16LS-16qA。.

2. A method for detecting human cell chimerism based on microhaploidy, characterized in that: A plurality of microhaploids are used as markers to detect the chimerism rate of human cells, and the microhaploids are selected from mh07SCUZJ-0428708, mh10WL-004, mh06WL-028.v1, mh13SCUZJ-0304208, mh19WL-028, mh11WL-039, mh10SCUZJ-0206566, mh06WL-064, mh16PK-83544, mh06USC-6pA, mh01ZHA-012.v1, mh18LS-18qA, mh07LS-7pF, mh18WL-026, mh13LS-13qE, mh02LS-2qG, mh 13LS-13qC, mh16HYP-36, mh03LS-3pC, mh07WL-007.v1, mh01HYP-02, mh07 LS-7qE, mh03LS-3pA, mh18WL-027.v2, mh02LS-2qB, mh03LS-3pD, mh06LS-6 pE, mh03LS-3pE, mh14USC-14qB, mh02LS-2qC, mh04WL-091, mh09USC-9qA, m h08HYP-23, mh06KK-031, mh10ZHA-002, mh04WL-074.v2, mh02KK-215, mh03 HYP-09, mh11KK-091, mh02LS-2qH, mh10LS-10pA, mh02LS-2qF, mh04USC-4qA, mh10LS-10qA, mh11KK-039, mh01LS-1qB, mh10LS-10qG, mh06LS-6qA, mh 02USC-2pA, mh05KK-122, mh11KK-040, mh03KK-006, mh15CP-003, mh15LS-15qD, mh13KK-213.v4, mh02WL-020, mh06HYP-18, mh08LS-8qD, mh05LS-5qF, mh20NH-26, mh08LS-8qA, mh22KK-069, mh10LS-10qE, mh10KK-086, mh10LS-10qD, mh17LS-17qD, mh03NH-07, mh05LS-5qC, mh04LS-4pA, mh05KK-120.v 1. mh02USC-2pB, mh18ZHA-005, mh04LS-4qF, mh18LW-46, mh10WL-060, mh01 WL-115, mh05WL-049, mh06WL-080, mh03WL-087, mh07WL-079, mh06WL-063.v2, mh04WL-092, mh01WL-038.v3, mh08SCUZJ-0476324, mh10WL-055, mh01WL-044, mh13WL-014, mh20WL-023.v3, mh20HYP-41, mh04KK-015, mh04LS-4pE, mh05LS-5qD, mh01USC-1qC.v1, mh12LS-12qB, mh10KK-101, mh07LS-7pB, mh09KK-035.v1, mh01SCUZJ-0533772, mh03LS-3qA, mh04LS-4pD, and mh16LS-16qA.

3. The method according to claim 2, characterized in that The following steps are involved: S1. Collect samples from donors and recipients before and after transplantation; S2, extracting sample DNA, and using capture probes specifically targeting markers to establish a sequencing library, and sequencing; S3, aligning the sequencing data with the reference genome to obtain an alignment result; S4. According to the comparison results, the genotype of the marker and the number of reads of each genotype are obtained, the marker validity is judged according to the difference of marker genotypes between different individuals, and the chimerism rate of each donor or recipient is calculated one by one based on the validity marker; S5. The chimerism rates of all markers in the donor or recipient are averaged to obtain the final chimerism rate.

4. The method according to claim 3, characterized in that Step S2 includes the following: S21, extracting sample DNA; S22, fragmenting the DNA molecules and then amplifying them; S23, performing hybridization reaction between the capture probe specifically targeting the marker and the library obtained in step S22 to capture the target fragment containing the marker, and further construct a sequencing library after enrichment; S24, sequencing the sequencing library obtained in step S23.

5. The method according to claim 4, characterized in that In step S24, the sequencing depth of the samples of the recipient after transplantation is not less than 2000x, and the sequencing depth of the samples of the recipient and the donor before transplantation is not less than 700x.

6. The method according to claim 3, characterized in that Step S3 includes the following: S31, preprocessing sequencing data to improve data quality, wherein the preprocessing includes removing sequencing adapters and low-quality reads; S32, comparing the preprocessed data with the reference genome to determine the sequence position and variation information; S33. Remove duplicate sequences and obtain a sequencing sequence comparison result file.

7. The method according to claim 3, characterized in that The marker validity judgment is specifically as follows: the genotype contained in a marker in a certain individual sample is compared with the genotype of the marker in all other individual samples. When the individual sample has a genotype that is not present in other individual samples, the marker is a validity marker for the sample.

8. The method according to claim 3, characterized in that The calculation formula of the mosaic rate in step S4 is: mosaic rate = number of genotype reads from a certain individual / total number of reads; wherein the total number of reads is the sum of the number of reads of all marker genotypes, and the counting method of the number of genotype reads from a certain individual is: if the marker genotype of the individual is homozygous, the number of genotype reads from the individual is equal to the read count of its unique genotype; if the marker genotype of the individual is heterozygous, and one of them is a unique genotype, the number of genotype reads from the individual is twice the unique genotype read count; if the marker genotype of the individual is heterozygous and both genotypes are unique, the number of genotype reads from the individual is the sum of the two unique genotype read counts.

9. The method according to claim 3, characterized in that: In step S4, a marker is split into multiple sub-markers, and then the genotype of the sub-marker and the number of reads of each genotype are obtained, and then the validity judgment and mosaicism rate calculation are performed, and the marker mosaicism rate is the average mosaicism rate of its corresponding sub-marker.

10. A system for detecting human cell chimerism based on microhaploid, characterized in that: The system uses the 101 microhaploids described in claim 1 as markers for detection and comprises the following modules: The sequencing module constructs high-quality sequencing libraries corresponding to the donor and acceptor based on the capture probe targeting the marker and performs sequencing; A data processing module is used to analyze the sequencing data and determine the genotype of the marker and the number of reads of each genotype based on the analysis results; The chimerism rate calculation module calculates the chimerism rate based on the marker genotype, validity and number of reads.

Citation Information

Cited By

  • Method and system for detecting chimeric rate based on next-generation sequencing data

    CN120877853A

  • A method and system for detecting chimerism rate based on next-generation sequencing data

    CN120877853B