A method and system for detecting chimerism rate based on next-generation sequencing data
By using microhaplotype and probe capture technology in second-generation sequencing data analysis, the problem of existing chimerism detection methods in the absence of samples or multiple donors has been solved. This method achieves high sensitivity and high resolution chimerism assessment, is suitable for complex multiple donor scenarios, and reduces detection costs and complexity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN HAIXI BIOTECHNOLOGY CO LTD
- Filing Date
- 2025-07-03
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for detecting chimerism cannot be effectively analyzed in the absence of pre-transplant samples or genetic information, and are difficult to detect in multi-donor scenarios, with limited resolution and sensitivity, thus failing to meet the needs of precision medicine.
Using microhaplotypes as genetic markers, combined with probe capture technology to enrich target DNA fragments, and through next-generation sequencing data analysis, a method and system for detecting chimerism based on microhaplotypes were developed. This method supports blind testing and multi-donor chimerism detection. Taking advantage of the high polymorphism of microhaplotypes and the high resolution of next-generation sequencing, the chimerism calculation is optimized by combining the SLSQP algorithm.
It enables chimerism assessment in the absence of donor and recipient genetic information, improves the sensitivity and resolution of detection, is suitable for multi-donor scenarios, reduces detection cost and complexity, and achieves a sensitivity of up to 0.03%.
Smart Images

Figure CN120877853B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chimerism detection technology, and specifically to a method and system for detecting chimerism rate based on second-generation sequencing data. Background Technology
[0002] Allogeneic hematopoietic stem cell transplantation (allo-HSCT) is an effective treatment for various hematologic malignancies and genetic diseases. However, monitoring chimerism after transplantation is crucial for assessing transplant efficacy, early detection of relapse, and adjustment of treatment regimens. While traditional short tandem repeat (STR) analysis methods are relatively mature in clinical applications for chimerism monitoring, their resolution and sensitivity are limited, making it difficult to meet the growing demand for precision medicine.
[0003] In recent years, with the development of next-generation sequencing (NGS) technology, NGS-based chimerism detection methods have gradually become a research hotspot. NGS, with its high throughput, high resolution, and high sensitivity, has shown great potential in chimerism detection. According to the latest literature, at least two commercially available NGS kits (AlloSeq HCT™ by CareDx® and NGStrack™ by GenDX®) can detect chimerism rates after stem cell transplantation, and their results show high consistency with traditional STR-based techniques and the latest digital PCR (dPCR) techniques. Typically, regardless of whether PCR or NGS is used, pre-transplant patient and hematopoietic stem cell donor samples need to be analyzed to obtain specific genotype information of their respective genetic markers. The difference lies in the genetic markers used in NGS (Next-Generation Sequencing-Based Chimerism Analysis). First, genetic markers distinguishing between donor and recipient must be identified experimentally beforehand. Then, selective measurements are performed on all valid markers. Especially in qPCR, to ensure relative quantitative stability, pre-transplant samples may be required as a reference for each measurement. In contrast, with NGS, the identification of valid markers and determination of chimerism can be performed simultaneously. Furthermore, due to the high resolution of NGS (accurate to each base), even if pre-transplant genetic information or samples from the recipient or donor are missing or insufficient, chimerism can theoretically still be inferred after excluding sequencing errors. A recent paper (Utility of Next-Generation Sequencing-Based Chimerism Analysis for Early Relapse Prediction following Allogenic Hematopoietic Cell Transplantation) mentions that the commercial kit AlloSeq HCT™ supports this "blind" analysis mode, enabling effective chimerism assessment even in the absence of pre-transplant samples, including cases where recipient, donor, or both information are missing. However, no related patent information or published algorithms have been found.
[0004] Furthermore, current commercial NGS kits use several independent single nucleotide polymorphisms (SNPs) as genetic markers and rely on multiplex PCR amplification technology, thus limiting the upper limit of the number of genetic markers that can be used, which may affect the accuracy of chimerism quantification. The inventors' previous research (patent application CN202411940613.3) used microhaplotypes (MHs) as genetic markers and utilized probe-targeted capture technology to enrich target genetic markers, further improving the accuracy and reliability of chimerism detection. Microhaplotypes refer to multiple closely linked SNPs located on the same chromosome, which are genetically transmitted as a whole; compared with traditional STR markers, microhaplotypes have higher polymorphism, providing more refined genetic information and are more suitable for complex multi-donor chimerism scenarios. However, the data analysis scheme provided in the inventors' previous research cannot support "blind" analysis, therefore, effective analysis cannot be performed when donor and / or pre-transplant recipient genetic information or samples are lacking. Summary of the Invention
[0005] In view of the technical problems existing in the background art, the present invention provides a method and system for detecting chimerism rate based on second-generation sequencing data, aiming to solve the defect that existing algorithms cannot cope with blind testing mode.
[0006] The specific technical solution of the present invention is as follows:
[0007] In a first aspect, the present invention provides a method for detecting chimerism rate based on next-generation sequencing data, comprising the following steps:
[0008] S1. Analyze the second-generation sequencing data of chimeric samples to obtain the candidate genotypes and the corresponding support read counts for each target genetic marker (hereinafter referred to as marker); wherein, the marker type is microhaplotype, each marker contains 2-7 SNPs and the length is ≤300bp;
[0009] S2. For each marker, perform the following analysis step by step:
[0010] S21. Based on the candidate genotypes and the number of contributors for this marker, arrange and combine the candidate genotypes to obtain all possible genotype pairings for this marker.
[0011] S22. Filter the above genotype pairings based on read counts;
[0012] S23. For the retained genotype pairings, calculate the predicted chimerism rate of the contributors;
[0013] S3. Based on the predicted chimerism rates of all markers obtained in step S2, calculate the final chimerism rate of the contributors.
[0014] Preferably, the microhaploids are multiple of the 101 microhaploids shown in Table 1.
[0015] Table 1
[0016]
[0017] Preferably, the method for obtaining second-generation sequencing data is as follows: using probe capture technology to enrich target DNA fragments containing markers to construct sequencing libraries, and then performing high-depth (≥2000x) sequencing.
[0018] Preferably, step S1 includes the following steps:
[0019] S11. Preprocessing sequencing data to improve data quality, including removing sequencing adapters and low-quality reads;
[0020] S12. Align the preprocessed data with the hg38 reference genome to determine sequence positions and variation information;
[0021] S13. Remove the repetitive sequences generated by PCR to obtain the sequencing read alignment result file;
[0022] S14. Based on the comparison results, obtain the candidate genotypes under each marker and the read count of each candidate genotype.
[0023] In the method of this invention, contributors specifically refer to the recipient and the donor; it is understood that there is only one recipient and at least one donor. Taking a marker containing two SNPs as an example, step S21 is as follows: if it is determined from the second-generation sequencing data that there are three candidate genotypes, AA, GC and GA, in the chimeric sample, and the number of contributors is two (i.e., one recipient + one donor), and given that humans are diploid organisms, the genotypes of the two contributors obtained by permutation and combination may be one of [('AA', 'AA'), ('AA', 'GC'), ('AA', 'GA'), ('GC', 'GC'), ('GC', 'GA'), ('GA', 'GA')]. The candidate list of genotype pairing combinations of the two contributors can be obtained by permutation and combination again, as shown in Table 2.
[0024] Table 2
[0025]
[0026] Preferably, the filtering method in step S22 is to exclude the following genotype pairings:
[0027] a) A candidate genotype has a read count higher than a preset threshold but does not contain a genotype pairing;
[0028] b) Genotype pairings where two genotypes of a contributor are unique, but the read count ratios of the two genotypes differ by more than 5 times.
[0029] Step S22 can effectively exclude genotype pairings with low probability. Specifically, for type a) combinations, the preset threshold can be set according to actual needs. For example, in some embodiments of the present invention, the preset threshold is 50. For type b) combinations, the exclusion principle is based on diploid characteristics.
[0030] Preferably, in the above method, step S23 includes the following steps:
[0031] i) For any genotype pairing combination, first determine whether the marker is effective in this case;
[0032] ii) If not valid, no further analysis will be performed; if valid, the chirp rate will be calculated in any of the following ways:
[0033] Method 1: The chimerism rate of a contributor is equal to the number of contributor reads divided by the total number of reads, where the total number of reads is the sum of the read counts of all genotypes under this marker, and the number of contributor reads is calculated based on the read counts of their unique genotypes;
[0034] Method 2: Calculate the theoretical frequency and detection frequency of all genotypes under this marker, and optimize the chimerism rate of contributors by minimizing the difference between the theoretical frequency and the detection frequency.
[0035] In the method of this invention, the criterion for judging the validity of the marker is as follows: the genotype of the current contributor is compared with the genotypes of all other contributors. If the contributor has a unique genotype, then the marker can provide effective distinguishing information for that contributor, i.e., it is valid. The number of reads by the contributor is calculated as follows: when the contributor has a homozygous genotype, the number of reads is equal to the count of reads of its unique genotype; when the contributor has a heterozygous genotype and one of its genotypes is unique, the number of reads is equal to twice the count of reads of its unique genotype; when the contributor has a heterozygous genotype and both of its genotypes are unique, the number of reads is equal to the sum of the counts of reads of its two unique genotypes. For example, for a marker containing 3 SNPs, the genotype of this marker among the current contributors is "TTT|AAT", while the genotypes of the other two contributors are "AAT|ATC" and "ATT|ATC". Then, for the current contributor, "TTT" is its unique genotype, and the chimerism rate of the receptor can be calculated based on this feature. Similarly, "ATT" is the unique genotype of another contributor, and its chimerism rate can also be calculated directly based on this feature. However, the remaining third contributor does not have a unique genotype, and its chimerism rate can only be obtained indirectly through the elimination method.
[0036] In the method of this invention, the second method for calculating the embedding rate can use the SLSQP (Sequential Least Squares Programming) algorithm to solve this nonlinear optimization problem. In this optimization problem, the embedding rate can be explicitly restricted to a range of 0 to 1, and the sum of all embedding rates is 1; that is, first, an objective function is constructed, representing the sum of squares of the differences between the theoretical frequency and the detection frequency, then boundary conditions for the embedding rate and equality constraints with a sum of 1 are set, and finally, the SLSQP algorithm is used to solve the constrained nonlinear optimization problem to find the optimal embedding rate.
[0037] It is understandable that, for the two chimerism rate calculation methods provided by this invention, method one should be preferred for analysis. When method one cannot provide the chimerism rate of all contributors or the chimerism rate of a certain contributor is not in the range of 0 to 1 (this may be due to sequencing errors or deviations in capturing the target DNA fragment), method two can be used for calculation to obtain a relatively more reasonable chimerism rate. Of course, method two can also be used for calculation only.
[0038] Preferably, step S3 includes the following steps:
[0039] S31. Preprocess all the chimerism prediction results obtained in step S2 to filter out low-quality prediction results.
[0040] S32. Based on the filtered chimerism rate prediction results, perform chimerism rate analysis on the contributors to obtain the estimated chimerism rate.
[0041] S33. By comparing the predicted chimerism rate with the filtered chimerism rate prediction results, the chimerism rate prediction results are further filtered.
[0042] S34. Based on the predicted chimerism rate after filtering in step S33, perform chimerism rate analysis on the contributors again (analysis method is the same as in step S32) to obtain the final chimerism rate.
[0043] More preferably, the marker is split into multiple sub-markers before analysis, and each sub-marker contains 1 to 3 SNPs.
[0044] In some embodiments of the present invention, the filtering in step S31 includes the following operations:
[0045] Based on the total number of read segments for each sub-marker, the splicing prediction results of sub-markers with a total number of read segments < 300 are filtered out;
[0046] For a chimerism prediction result, if the minimum absolute deviation between the genotype detection frequency and the theoretical frequency in all SNP combinations is >0.1, it will be filtered out.
[0047] For a single sub-marker, if it corresponds to multiple genotype combinations and their chirality prediction results are obtained, then the results are filtered based on the number of duplicates after removing the mean square deviation of the corresponding genotype frequencies. Specifically, when the number of contributors is 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 2, or when the number of contributors is greater than 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 4, the chirality prediction results of that sub-marker are excluded.
[0048] For a single sub-marker, if it corresponds to multiple genotype combinations with chimerism prediction results, the prediction result with the smallest sum of genotype base editing distances among the contributors is retained.
[0049] For a single sub-marker, if it corresponds to multiple genotype combinations and the predicted chimerism rate is obtained, only the prediction result with the smallest mean square deviation of genotype frequency is retained.
[0050] All chimerism predictions were sorted according to the mean squared deviation of genotype frequencies, and chimerism predictions between the 5th and 90th percentiles were retained.
[0051] In some embodiments of the present invention, step S32 includes the following operations:
[0052] The sub-markers are grouped according to the markers, and within each group, they are sorted in reverse order based on the total number of reads corresponding to the genotypes of the sub-markers. The chimerism prediction results of the top three sub-markers are extracted as the representative of the main marker's quantitative chimerism result for the current contributor.
[0053] The Grubbs method was used to identify outliers in the mean squared deviation of genotype frequencies and outliers in the chimerism rate, and the chimerism rate prediction results of the corresponding sub-markers were filtered out.
[0054] The sub-markers are then grouped again according to their respective markers, and the average value within each group is calculated to obtain the chimerism rate of the respective markers.
[0055] The Grubbs method is used to identify outliers in marker chirp rate and filter out the corresponding markers.
[0056] Based on the remaining markers, calculate the average value, which is the estimated splice rate of the current contributor;
[0057] The estimated splicing rate of all contributors was calculated one by one.
[0058] In some embodiments of the present invention, the filtering in step S33 includes the following operations:
[0059] For a single chirp prediction result, calculate the ratio of the current predicted chirp to the estimated chirp. If the ratio of a certain contributor is greater than 4 or less than 1 / 3, then filter it out.
[0060] For a single sub-marker, if it corresponds to multiple genotype combinations and the predicted chimerism rate is obtained, only the prediction result with the smallest mean square deviation of genotype frequency is retained.
[0061] For a single splice rate prediction result, calculate the absolute deviation between the estimated splice rate and the current predicted splice rate, sort all splice rate prediction results according to the absolute deviation, and retain the splice rate prediction results that are below or equal to the 85th percentile.
[0062] For a single sub-marker, if it corresponds to multiple genotype combinations and the chimerism prediction results are obtained, the prediction result with the smallest sum of base editing distances between the genotypes of the contributors is retained.
[0063] For a single sub-marker, if it corresponds to multiple genotype combinations and the chimerism prediction results are obtained, the result with the smallest mean square deviation of genotype frequency is retained.
[0064] All chimerism predictions were sorted according to the mean squared deviation of genotype frequencies, and predictions below and equal to the 85th percentile were retained.
[0065] All splicing rate predictions are sorted again based on the absolute deviation between the estimated splicing rate and the current predicted splicing rate, and predictions below and equal to the 85th percentile are retained.
[0066] Secondly, the present invention provides an analysis system for performing the above-described method for detecting the chimerism rate, specifically comprising the following modules:
[0067] Data input module: The input data must include at least the human reference genome, marker locus information, and next-generation sequencing data of the chimeric sample to be tested;
[0068] Data processing module: Used to process sequencing data and analyze each aligned sequencing read for each marker to determine the specific genotype information of the target genetic marker in the sample and its corresponding read count;
[0069] The chimerism analysis module is used to calculate all possible chimerism predictions for each marker.
[0070] Contributor chimerism calculation module: Based on the chimerism prediction results of markers, the final chimerism rate of all contributors is calculated.
[0071] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0072] (1) The present invention allows for application in multi-donor scenarios. Due to limitations in the number and type of markers, almost no chimerism detection products currently on the market mention whether they can detect multi-donor chimerism. Theoretically, the difficulty of multi-donor chimerism detection is significantly higher than that of single-donor chimerism detection, because it requires more markers or markers with more diverse genotypes to distinguish different donors, especially when there is a blood relationship. Based on a certain number of microhaplotypes as markers, the algorithm provided by the present invention has been proven to detect multi-donor chimerism; for example, in hematopoietic stem cell transplantation cases where the donor and recipient have a direct or collateral blood relationship, the algorithm proposed in this invention can accurately detect the chimerism of all donors and recipients in cases with one or two donors, and has been validated by digital PCR.
[0073] (2) Allows for blind testing. Most products on the market cannot detect chimerism when the genetic marker genotype is unknown; that is, most solutions require prior access to pre-transplant samples or genotype information from both the recipient and donor. In contrast, this invention, when the number of markers is sufficient, only requires NGS data from any non-chimeric sample from either the recipient or donor, along with NGS data from the post-transplant sample to be tested, to calculate the chimerism rate between the recipient and donor. This blind testing mode is very useful in clinical applications where donor information cannot be collected (e.g., due to the donor's unexpected death or insufficient umbilical cord blood samples), and it also reduces the cost of continuous testing. In more specific cases, even without providing pre-transplant samples or genotype information from both the donor and recipient, the solution of this invention can still calculate the chimerism rate, but it cannot distinguish between donor and recipient, and when the number of donors is large, the reliability of the test results will decrease.
[0074] (3) High sensitivity. The highest sensitivity of similar NGS chimerism detection products on the market is currently 0.1%. Based on micro-haplotypes as markers, the proposed method has been proven to achieve the same level of sensitivity through batch sample testing. In conventional analysis mode, the current algorithm can conservatively estimate the limit of quantitation to be 0.1%, while the limit of detection can be 0.03%. In blind testing mode, the limit of quantitation can reach 0.3%, and the limit of detection can reach 0.1%. Attached Figure Description
[0075] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0076] Picture 1 This is the correlation analysis result between the expected chimerism rate value and the actual chimerism rate detection value in Embodiment 3 of the present invention. Detailed Implementation
[0077] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.
[0078] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs; the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the invention; the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion.
[0079] The term "chimeric sample" specifically refers to a blood or tissue sample that is chimeric after transplantation, such as a peripheral blood sample from a recipient after bone marrow transplantation.
[0080] The term "genotype" specifically refers to the base sequence consisting of all the SNP sites corresponding to the specific bases in a marker or sub-marker. Since humans are diploid organisms, in a normal individual, one marker or sub-marker contains one or two genotypes, namely a homozygous genotype and a heterozygous genotype.
[0081] The counting rule for the term "read count" is as follows: when a pair of reads (paired-end sequencing, read1 and read2) covers all SNP sites of a marker, the corresponding bases for all SNP sites in that marker can be extracted from the corresponding reads, thus obtaining its genotype and incrementing the count by 1. For example, if a marker contains 3 SNPs whose corresponding bases in a pair of reads are A, T, and T, then the count for the corresponding genotype "ATT" is incremented by 1. It is important to note that when read1 and read2 cover the same SNP site and give inconsistent base sequencing results, the one with the higher base sequencing quality takes precedence. For example, taking a sub-marker containing 3 SNPs as an example, the following genotypes and their corresponding counts may be obtained: ATC (310 times), CTT (302 times), ACC (3 times), and CTC (2 times). To obtain more accurate chimerism quantification, the counts can be corrected based on the similarity between the pre-determined or pre-defined contributor genotypes and the actual statistically obtained genotypes. Specifically, if any of the statistically obtained genotypes differ from the expected donor genotypes, they are considered pseudogenotypes resulting from sequencing errors. To reassign the counts corresponding to these pseudogenotypes to the expected genotypes, the base editing distance between the pseudogenotypes and the expected genotypes can be calculated first. Then, the counts corresponding to the pseudogenotypes are added to the counts of the expected genotypes with the closest base editing distance. If there are genotypes with the same distance, they are distributed equally. In the example above, if ATC and CTT are the expected genotypes of the contributors, while ACC and CTT are pseudogenotypes resulting from sequencing errors, ACC is most similar to ATC, and CTC is equally similar to both ATC and CTT. Therefore, the count of ATC can be corrected to 314, and the count of CTT can be corrected to 303.
[0082] The term "sub-marker" refers to a sub-marker derived from a complete marker (i.e., the main marker) (discontinuous sub-marking is allowed). For example, a marker containing 4 SNPs (SNPs numbered 0, 1, 2, and 3) can be subdivided into 14 sub-markers: 0, 1, 2, 3, 4, 01, 02, 03, 12, 13, 23, 012, 013, 023, and 123. It is worth noting that shorter sub-markers may correspond to a higher number of sequencing reads and more reliable statistical power, but shorter sub-markers also have a higher risk of sequencing errors. Furthermore, longer sub-markers contain greater genotypic diversity, which is more advantageous for differentiating multiple donors. Therefore, in this invention, the number of SNPs in the sub-marker is controlled to be between 1 and 3. Splitting the master marker is more beneficial for utilizing sequencing reads to reduce the impact of sampling errors on chimerism quantification and to minimize the impact of sequencing errors on chimerism quantification. This is because when obtaining the sequenced DNA fragments using probe capture technology, a pair of sequencing reads may only contain partial SNP information of the marker. However, when the number of contributors is small (e.g., equal to 2), a difference of only one base is enough to distinguish between the donor and recipient. This allows DNA fragments containing only partial markers to also transmit effective information.
[0083] The "theoretical frequency" of a genotype refers to the sequencing frequency (or proportion) of a genotype calculated based on known chimerism and genotype combinations. In other words, under ideal sequencing conditions (such as no sequencing errors, sufficient sequencing depth, and no PCR amplification bias), the genotype frequency calculated from the sequenced read count should equal this theoretical value. Assuming a chimeric sample contains... N Each of the contributors represents a percentage of [number]. C 1, C 2,…, CN ,and C 1+ C 2+...+ CN =1; for a given marker, this N The diploid genotypes of the contributors are as follows: g 1, G 1),( g 2, G 2),…,( gN , GN Each contributor's genotype corresponds to a percentage coefficient, which depends on whether the contributor is homozygous: if the contributor is homozygous ( gi = GiIf the genotype is 1, then the proportion coefficient is 1; if it is a heterozygous genotype (gi≠Gi), then the proportion coefficient is 0.5. Therefore, the proportion coefficient for each genotype among the contributors can be set as ( k 1,1− k 1),( k 2,1− k 2),…,( kN ,1− kN Therefore, in the chimeric sample, the theoretical frequency of genotype gi or Gi originating from the i-th contributor is: T × Ki or T ×(1− Ki Ultimately, the theoretical frequency of each genotype is the sum of the theoretical frequencies of that genotype from all contributor sources. Taking a microhaploid containing 3 SNPs as an example, assuming the genotype of the first contributor is "ATC|ATC" and the genotype of the second contributor is "ACC|ATC", and the chimerism rates of the first and second contributors in the chimeric sample are 0.1 and 0.9 respectively, then the theoretical frequencies of different genotypes in the chimeric sample can be directly calculated: the frequency of the first SNP being A in the marker is equal to 1, because the genotype of both contributors at this SNP locus is A. Similarly, the frequency of the third SNP being C is also 1, while the frequency of the second SNP being T is 0.1 + 0.9 × 0.5 = 0.55, and the frequency of the second SNP being C is 0 + 0.9 × 0.5 = 0.45. Assuming the numbers of these three SNPs are 0, 1, and 2, then the genotype frequencies of any combination of these 3 SNPs can be deduced as: {(0,): {'A': 1.0}, (1,): {'T': 0.55,'C': 0.45}, (2,): {'C': 1.0}, (0, 1): {'AT': 0.55, 'AC': 0.45}, (0, 2): {'AC': 1.0}, (1, 2): {'TC': 0.55, 'CC': 0.45},(0, 1, 2): {'ATC': 0.55, 'ACC':0.45}}.
[0084] The "detection frequency" of a genotype is equal to the read count of that genotype divided by the total number of reads. Taking a marker containing 3 SNPs as an example, assuming the genotype of the first contributor is "ATC|ATC" and the genotype of the second contributor is "ACC|ATC", based on sequencing alignment results, the read count for the ATC genotype is 60, while the read count for the ACC genotype is 40. Therefore, the detection frequency for the ATC genotype is 0.6, and the detection frequency for the ACC genotype is 0.4.
[0085] The term "genotype frequency mean squared bias" refers to the sum of squares of the genotype frequency biases divided by the total number of SNP combinations. The genotype frequency bias equals the theoretical frequency minus the detection frequency. That is, for a marker containing multiple SNPs, the genotype frequency bias for all its combinations needs to be calculated before further calculation. For example, if a marker with two SNPs has the genotype AT, the calculation will consider the genotype frequency biases for the combinations A, T, and AT. If a marker contains three SNPs with the genotype ATC, then its SNP combinations include the seven combinations A, T, C, AT, AC, TC, and ATC.
[0086] The term "base edit distance" refers to the minimum number of editing operations required to transform one DNA sequence into another, including inserting, deleting, and replacing bases in the sequence.
[0087] With the development of allo-HSCT technology, higher requirements have been placed on chimerism monitoring. However, existing chimerism detection methods have the following problems: ① Limited resolution and sensitivity: Although traditional STR analysis methods are relatively mature in chimerism monitoring, their resolution and sensitivity are limited, making it difficult to detect low-level chimerism (e.g., <1%), which is particularly important in early disease recurrence prediction; ② Reliance on pre-transplant samples: Most current PCR-based techniques (e.g., qPCR) require the prior identification of genetic markers that can distinguish between donors and recipients, and pre-transplant samples are needed as a reference for each measurement. However, in clinical practice, pre-transplant samples may be missing or insufficient, which limits the application scope of these methods to some extent; ③ Insufficient capacity to process complex samples: In cases involving multiple donors or multi-source chimerism, especially in kinship transplantation cases, qPCR products based on dimorphic markers (e.g., Indel sites) may face the problem of insufficient effective markers, mainly because markers with only two genotypes often cannot distinguish three or more contributors; ④ High cost and technical complexity: Although NGS technology has the advantages of high throughput and high resolution, existing NGS kits (e.g., AlloSeq) are limited in their capacity and complexity. HCT™ and NGStrack™ rely on multiplex PCR amplification to enrich target regions in the early stages. When there are many targets, the requirements for primer design are high, and it is easier to face amplification bias and primer dimer problems, which increases the complexity of the experiment.
[0088] To address the technical challenges of chimerism detection, this invention provides an NGS-based chimerism detection method. This method utilizes micro-haplotypes as genetic markers, which exhibit higher polymorphism compared to single SNPs, enabling more sensitive detection of low-level chimerism. This improves resolution and sensitivity, making it more suitable for complex multi-donor applications. Furthermore, this method employs probe capture technology instead of multiple primer amplification to enrich genetic markers, reducing experimental costs and complexity. It is also more applicable to scenarios with a larger number of genetic markers, laying a more solid foundation for stable chimerism monitoring. Moreover, this invention creatively develops a blind testing mode, enabling chimerism assessment even in the absence of donor or recipient genetic information. This is particularly suitable for situations where pre-transplant samples are difficult to obtain (such as emergency transplants or cases of improper sample preservation), thus expanding the application scope of chimerism monitoring. It also reduces continuous monitoring costs without significantly reducing detection sensitivity.
[0089] The following are some specific embodiments. It should be noted that the embodiments described below are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they shall be performed in accordance with the techniques or conditions described in the literature in this field or according to the product instructions. Reagents or instruments used, unless otherwise specified, are all conventional products that can be obtained commercially.
[0090] Example 1
[0091] This example provides a method for detecting chimerism based on next-generation sequencing data, suitable for "blind" analysis. The specific steps are as follows:
[0092] (1) Analyze the sequencing data of the chimeric samples.
[0093] Based on all markers in Table 1, target DNA fragments containing markers are enriched using capture probes to construct sequencing libraries, and then high-depth sequencing is performed to obtain sequencing data of chimeric samples (this is prior art, for details please refer to patent application CN 202411940613.3).
[0094] Preprocessing of sequencing data: First, Fastp (https: / / github.com / OpenGene / fastp) software was used to remove sequencing adapters and low-quality reads (using the software's default parameters) to improve data quality. Then, Bwa mem (https: / / github.com / lh3 / bwa) software was used to align the obtained data with the hg38 reference genome to determine sequence positions and variant information. Finally, Gencore (https: / / github.com / OpenGene / gencore) software was used to remove repetitive sequences generated by PCR based on the alignment positions of the sequencing reads and using the software's default parameters to obtain the final sequencing read alignment result file.
[0095] The marker is split into several sub-markers, and the alignment result file obtained is parsed using pysam (https: / / github.com / pysam-developers / pysam) to determine the candidate genotype and corresponding read count of each sub-marker under each marker in the chimeric sample.
[0096] (2) For each sub-marker, perform the following analysis step by step:
[0097] (2.1) Based on the candidate genotypes of the sub-marker and the number of contributors, the candidate genotypes are arranged and combined to obtain all possible diploid genotype pairings.
[0098] (2.2) Filter the above genotype pairings based on read counts, specifically filtering out the following two combinations:
[0099] A genotype pairing combination where the read count of a candidate genotype is higher than 50 but does not contain that candidate genotype;
[0100] A genotype pairing where two contributors have unique genotypes but their read count ratios differ by more than 5 times.
[0101] (2.3) For the remaining genotype pairings under this sub-marker, calculate the chimerism prediction results for all genotype pairings, and the calculation method includes the following operations:
[0102] (2.3.1) For any genotype pairing, first determine whether the sub-marker is effective in this case;
[0103] (2.3.2) If valid, the splicing rate prediction result of the sub-marker is calculated using the following two methods:
[0104] Method 1: Calculate the splicing rate directly based on the segment count, which includes the following steps:
[0105] ① Sum the read counts of all genotypes under this sub-marker to obtain the total number of reads;
[0106] ② For each contributor, calculate the number of reads they sourced, specifically: when the contributor is homozygous, the number of reads they sourced is equal to the read count of their unique genotype; when the contributor is heterozygous and one of their genotypes is unique, the number of reads they sourced is twice the read count of their unique genotype; when the contributor is heterozygous and both genotypes are unique, the number of reads they sourced is equal to the sum of the read counts of their two unique genotypes.
[0107] ③ The chimerism rate of a contributor can be obtained by dividing the number of reads by the total number of reads. For contributors without a unique genotype, their chimerism rate can be obtained indirectly through the exclusion method, as detailed in patent application CN202411940613.3, which will not be elaborated here.
[0108] Method 2: Calculate the theoretical and detection frequencies of all genotypes under this marker, and optimize the chimerism rate of each contributor by minimizing the difference between the theoretical and detection frequencies. Specifically, the SLSQP algorithm is used to solve this nonlinear optimization problem. In this optimization problem, first construct the objective function, representing the sum of squares of the differences between the theoretical and detection frequencies. Then, set boundary conditions for the chimerism rate and equality constraints with a sum of 1. Finally, use the SLSQP algorithm to solve the constrained nonlinear optimization problem to find the optimal chimerism rate.
[0109] In this example, we first use method one to calculate the chimerism rate. If method one cannot give the chimerism rate of all contributors or the chimerism rate of a certain contributor is not in the range of 0 to 1, then we use method two to calculate it.
[0110] (3) Summary calculation of chimerism rate under blind test mode.
[0111] After analyzing each sub-marker of each marker one by one using the method in step (2), a summary table is obtained. It is worth noting that since each sub-marker may correspond to multiple genotype pairings, there are also multiple chimerism prediction results.
[0112] Based on the summarized master table, perform the following calculations step by step:
[0113] (3.1) Based on the total number of read segments for each sub-marker, filter out the splicing prediction results of sub-markers with a total number of read segments < 300;
[0114] (3.2) The estimation of the chimerism rate includes the following steps:
[0115] (3.2.1) For each chimerism prediction result, if the minimum absolute deviation between the detection frequency and the theoretical frequency of the genotype in all SNP combinations is greater than 0.1, then the result is considered unreliable and is excluded.
[0116] (3.2.2) For a sub-marker, if there are multiple genotype combinations with predicted chimerism, the results are filtered based on the number of duplicates after removing the mean square deviation of the corresponding genotype frequencies. Specifically, when the number of contributors is 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 2, or when the number of contributors is greater than 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 4, the predicted chimerism of the sub-marker is excluded.
[0117] (3.2.3) For a sub-marker, if there are multiple chimerism prediction results for its corresponding genotype combinations, the prediction result with the smallest sum of base editing distances between the genotypes of the contributors shall be retained (if multiple minimum values are encountered, all shall be retained).
[0118] (3.2.4) For a sub-marker, if there are multiple genotype combinations with predicted chimerism, only the prediction result with the smallest mean square deviation of genotype frequency is retained (if there are multiple minimum values, all are retained).
[0119] (3.2.5) Sort all prediction results of all sub-markers in the total table according to the mean square deviation of genotype frequency, and retain only the prediction results between the 5th percentile and the 90th percentile.
[0120] (3.2.6) Based on the filtered master table from the above steps, perform a chimerism analysis for each contributor.
[0121] ① Determine whether the chimerism rate is 0.
[0122] The number of sub-markers with a chirality rate below 0.000001 is counted as zero-chirality; the number of sub-markers with a chirality rate above 0.05 and a minimum genotype read count greater than or equal to 99 is counted as high-chirality; and the number of sub-markers with a chirality rate less than or equal to 0.05 and a maximum genotype read count greater than 10 times the minimum genotype read count is counted as low-chirality. The proportion of each of these three types of chirality to the total number of sub-markers is then calculated. When the zero-chirality rate exceeds 85%, the current contributor's chirality rate is directly determined to be 0. When the zero-chirality rate exceeds 45%, the low-chirality rate is less than 5%, and the high-chirality rate is less than 50%, the current contributor's chirality rate is also directly determined to be 0. Otherwise, the current contributor's chirality rate is determined to be non-zero.
[0123] ② Continue the analysis for cases where the chimerism rate was not determined to be 0.
[0124] The sub-markers were grouped according to the main marker, and then sorted in reverse order according to the total number of reads for the genotype corresponding to the sub-marker within each group. The chimerism prediction results of the top three sub-markers (if the counts are the same, all of them are included) were extracted as representative of the main marker's quantitative results on the chimerism of the current individual.
[0125] ③ Filtering based on the Grubbs method.
[0126] The Grubbs method was used to identify outliers (alpha = 0.01) in the mean squared deviation of genotype frequencies, and the corresponding sub-markers were filtered out. The Grubbs method was also used to identify outliers (alpha = 0.01) in the chimerism rate, and the corresponding sub-markers were filtered out.
[0127] ④ Calculation of the chimerism rate of the master marker.
[0128] Sub-markers are grouped again according to their respective primary markers, and the average value within each group is calculated. This average value is used as a representative of the chirp rate indicated by the primary marker, resulting in a summary table of the chirp rates of the primary markers.
[0129] ⑤ Calculation of the estimated chimerism rate of contributors.
[0130] Based on the summary table of chirp rates of the main markers, the Grubbs method is used to identify outliers in the chirp rate (alpha = 0.05) and filter out the corresponding main markers. Finally, based on the remaining main markers, the average chirp rate is calculated as the estimated chirp rate of the current contributor.
[0131] The estimated splicing rate of all contributors was calculated one by one.
[0132] (3.3) Final determination of chimerism rate.
[0133] Starting from the initial summary table, perform the following analysis steps in sequence:
[0134] (3.3.1) For each chirp prediction result, calculate the ratio of the current predicted chirp to the estimated chirp. If the ratio of a certain contributor is greater than 4 or less than 1 / 3, then ignore the current chirp prediction result.
[0135] (3.3.2) Similar to the prediction stage, only the results with smaller mean square bias of genotype frequencies are retained;
[0136] (3.3.3) For each chirp prediction result, calculate the absolute deviation between the estimated chirp and the current predicted chirp. Then sort the total table according to the deviation. Finally, if the 85th percentile of the deviation is greater than 0, only the prediction results below or equal to the 85th percentile are retained. If the above conditions are not met, no filtering is performed.
[0137] (3.3.4) For each sub-marker, after the above filtering steps, if there are multiple chimerism prediction results for its corresponding genotype combinations, retain the prediction result with the smallest sum of genotype base editing distances between contributors (if the sum of genotype base editing distances of two combinations is the same, then both combinations are retained).
[0138] (3.3.5) For each sub-marker, after the above filtering steps, if there are multiple genotype combinations with predicted chimerism rates, the result with the smallest mean square deviation of genotype frequency is retained.
[0139] (3.3.6) For the total table, sort according to the mean square deviation of genotype frequency, and finally retain the prediction results below and equal to the 85th percentile.
[0140] (3.3.7) For the summary table, sort again according to the absolute deviation between the estimated splice rate and the current predicted splice rate. Finally, if the 85th percentile of the deviation is greater than 0, only the prediction results below and equal to the 85th percentile are retained.
[0141] (3.3.8) Based on the above filtered master table, the final chimerism rate analysis is performed for each contributor, and the analysis process is the same as step (3.2.6).
[0142] Additionally, if pre-transplant sequencing data from some contributors is provided, it is necessary to determine which of the predicted unknown contributors, after blind testing, is a match. The matching logic is as follows: check the degree of genotypic similarity (identical genotypes) between the known contributors and the predicted unknown contributors; the unknown contributor with the highest similarity ratio will be identified as a known contributor.
[0143] Example 2
[0144] Unlike Example 1, this example provides a method for detecting chimerism based on next-generation sequencing data suitable for routine analysis (i.e., obtaining genetic information from all contributors or pre-transplant samples), specifically including the following steps:
[0145] (1) Analyze the sequencing data of the chimeric samples.
[0146] The same as step (1) in Example 1.
[0147] (2) Calculation of the chimerism rate of sub-markers.
[0148] In the conventional analysis mode, the genotypes and read counts of all sub-markers for each contributor can be obtained through sequencing data, thus directly determining the contributor's genotype combination. Then, following the method in step (2.3) of Example 1, the corresponding chimerism rate is calculated; unlike the blind test mode, in the conventional analysis mode, one sub-marker corresponds to only one specific chimerism rate prediction result.
[0149] (3) Calculation of chimerism rate under conventional analysis mode.
[0150] After analyzing each sub-marker of the main marker one by one using the method in step (2), a master table is obtained. Based on this master table, the following operations are performed step by step:
[0151] (3.1) Filtering.
[0152] First, exclude the chimerism prediction results of sub-markers with a total number of reads < 700. Then, use the Grubbs method to identify and remove outliers with mean squared deviation of genotype frequencies (set the alpha parameter to 0.05), that is, exclude the chimerism prediction results of the corresponding sub-markers with outliers in the mean squared deviation of genotype frequencies.
[0153] (3.2) Based on the filtered master table, the splicing rate of each contributor is analyzed.
[0154] ① For current contributors, exclude chimerism predictions for sub-markers that do not have their own unique genotype.
[0155] ② Group according to the main marker, and sort the sub-markers in reverse order according to the total number of reads within the group. Extract the chirp prediction results of the top three sub-markers (if the counts are the same, all of them are included) as the representative of the main marker's quantitative result on the chirp of the current individual.
[0156] ③ Determine whether the current contributor's chirality rate is 0, and the determination rules are as follows: count the number of sub-markers with a chirality rate lower than 0.000001 as the number of zero chimerisms; count the number of sub-markers with a chirality rate greater than 0.05 and a minimum genotype read count greater than or equal to 99 as the number of high chimerisms; count the number of sub-markers with a chirality rate less than or equal to 0.05 and a maximum genotype read count greater than 10 times the minimum genotype read count as the number of low chimerisms. Then calculate the proportion of these three types of chimerisms to the total number of sub-markers. When the proportion of zero chimerism exceeds 72%, the current contributor's chirality rate is directly determined to be 0. When the proportion of zero chimerism exceeds 45%, the proportion of low chimerism is less than 5%, and the proportion of high chimerism is less than 50%, the current contributor's chirality rate is also directly determined to be 0. Otherwise, the current contributor's chirality rate is determined to be non-zero.
[0157] ④ If the current contributor's chimerism rate is not determined to be 0, then continue to calculate the chimerism rate.
[0158] The Grubbs method is used to identify and remove outliers in the chirp rate (alpha = 0.05), filtering out the chirp rate prediction results of sub-markers with outliers. Next, the sub-markers are grouped again according to their respective main markers, and the average value within each group is calculated as a representative of the chirp rate indicated by the main marker. Then, the same outlier removal method is used to remove outliers from the chirp rates obtained from all main markers. Finally, the average chirp rate obtained from all remaining main markers is further averaged to obtain the final chirp rate of the current contributor.
[0159] After calculating the chimerism rate of all contributors using the method described above, the chimerism rates of all contributors are sorted. Finally, the maximum chimerism rate is corrected by subtracting the sum of the remaining chimerism rates excluding the maximum chimerism rate from 1. For example, if the average chimerism rates of the donor and recipient are 0.1 and 0.92 respectively (the sum is not necessarily not equal to 1), then the adjusted chimerism rates of the donor and recipient will be 0.1 and 0.9 respectively (i.e., 1 minus 0.1).
[0160] Example 3
[0161] This example uses pre-transplant samples from real patients and donor samples to construct a simulated sample covering 15 chimerism rate gradients. The accuracy and reliability of the method of this invention are verified using this simulated sample, including the following operations:
[0162] (1) Sample preparation.
[0163] Sufficient DNA was extracted from both the pre-transplant patient and donor samples. The DNA was then mixed according to 15 preset recipient-donor ratios (i.e., chimerism gradients: 0.001, 0.003, 0.005, 0.01, 0.02, 0.05, 0.1, 0.25, 0.5, 0.75, 0.9, 0.99, 0.995, 0.997, 0.999) to prepare 15 simulated samples.
[0164] (2) Detection.
[0165] Each simulated sample was tested three times; real pre-transplant samples from patients and donor samples were tested once each to determine the genotype information of the donor and recipient. Therefore, this case involved the testing and analysis of sequencing data from a total of 47 samples.
[0166] The methods in Example 1 were used to perform part-blind analysis (without inputting donor sample information) and full-blind analysis (without providing pre-transplant samples from either the donor or recipient). The methods in Example 2 were used to perform normal analysis (with complete pre-transplant samples from both the donor and recipient). The results of the recipient chimerism detection under different modes are shown in Table 3.
[0167] Table 3
[0168]
[0169] (3) Data analysis.
[0170] The correlation between the expected chimerism rate and the actual chimerism rate was compared using linear regression analysis.
[0171] The analysis results of all analytical models show that the linear correlation coefficient between the two exceeds 0.99, and the R-squared value approaches 1. Picture 1 The coefficients of variation for repeated sample testing results were all less than 15% (Table 4). The results of blind test analysis were close to those of conventional analysis, but the results of conventional analysis were closer to the expected values.
[0172] Table 4
[0173]
[0174] Example 4
[0175] This example collected real clinical samples for chimerism detection and analysis, as detailed below:
[0176] (1) Sample collection.
[0177] Peripheral blood samples were collected from eight clinical patients and their corresponding donors. Three patients received two transplants (one bone marrow transplant and one umbilical cord blood transplant) from different donors, while the rest received only a bone marrow transplant. Since a pre-transplant peripheral blood sample was unavailable for one patient (patient 4), a nail sample was used as a substitute to obtain their pre-transplant genotype information. In addition, two patients (patients 5 and 4) underwent two and three chimerism monitoring sessions, respectively, at different time points after transplantation.
[0178] (2) Detection.
[0179] Chimerism was tested on post-transplant samples from each patient (repeated 3 times). Three analytical modes were used: routine analysis (with complete pre-transplant samples from both donor and recipient), semi-blind analysis (with only one sample provided, either the recipient's or the donor's), and fully blind analysis (without either donor or recipient's pre-transplant samples). The testing methods were the same as in Example 3. Additionally, existing STR and dPCR technologies were used to test post-transplant samples from each patient once. The results of recipient chimerism under different testing methods are shown in Table 5.
[0180] Table 5
[0181]
[0182] Note: None indicates that there are no corresponding analysis results.
[0183] (3) Results analysis.
[0184] Based on the analysis results of several clinical samples, it can be confirmed that STR-based capillary electrophoresis cannot detect microchimerism, while the method in this invention can detect microchimerism, and this has been validated by digital PCR. Furthermore, at low chimerism levels (1%–10%), the NGS detection results and digital PCR results are highly consistent, while the STR-based quantitative results are nearly twice as high. In addition, the results of the blinded analysis mode and the conventional analysis mode of this invention are highly consistent, as are the results of the semi-blind analysis mode and the fully blind analysis mode. Therefore, the method of this invention is significantly superior to STR-based methods, can detect microchimerism, and the accuracy of the quantitative results has been validated by digital PCR. However, for multi-donor chimeric samples, the results in the fully blind analysis mode are unstable.
[0185] In summary, the proposed solution not only enables chimerism analysis in conventional modes but also supports blind testing. In non-blind testing, this solution achieves a lower limit of detection (LOD) (0.03%) than similar products, surpassing the 0.22% LOD reported by AlloSeqHCT™, and a 0.1% LOD, superior to AlloSeq HCT's 0.36%. Furthermore, 0.1% is the lowest LOD reported in the literature for similar NGS products. Particularly in blind testing, the proposed solution achieves a 0.1% LOD and a 0.3% LOD, providing a reliable basis for clinical decision-making.
[0186] It should be noted that the present invention is not limited to the above-described embodiments. The above embodiments are merely examples, and any embodiments that have the same structure and perform the same effects as the technical concept within the scope of the present invention are included within the scope of the present invention. Furthermore, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways of constructing by combining some of the constituent elements of the embodiments, without departing from the spirit of the present invention, are also included within the scope of the present invention.
Claims
1. A method for detecting chimerism rate based on next-generation sequencing data, characterized in that, Including the following: S1. Analyze the sequencing data of the chimeric sample to obtain the candidate genotypes of all markers and their corresponding read counts, wherein the markers are microhaplotypes; S2. For each marker, it is divided into several sub-markers, and each sub-marker contains 1-3 SNP sites. Then, the following analysis is performed step by step for each sub-marker: Based on the candidate genotypes and the number of contributors of the sub-marker, all possible genotype pairings under that sub-marker are obtained; The genotype pairings are filtered based on read counts. The filtering excludes the following genotype pairings: genotype pairings where the read count of a candidate genotype is higher than a preset threshold but the candidate genotype is not present, and genotype pairings where two genotypes of a contributor are unique but the read count ratio is greater than 5. For any retained genotype pairing, first determine the validity of the sub-marker in that case, and then calculate the chimerism prediction result of the contributors for the valid sub-markers using any of the following methods: Method 1: The chimerism rate of a contributor is equal to the number of contributor reads divided by the total number of reads, where the total number of reads is the sum of the read counts of all genotypes under this sub-marker, and the number of contributor reads is calculated based on the read counts of their unique genotypes; Method 2: Calculate the theoretical frequency and detection frequency of all genotypes under this sub-marker, and optimize the chimerism rate of contributors by minimizing the difference between the theoretical frequency and the detection frequency; S3. Based on the predicted splice rates of all sub-markers, calculate the final splice rate of contributors using the following operations: S31. Preprocess all the chimerism prediction results obtained in step S2 to filter out low-quality prediction results. S32. Based on the filtered chimerism rate prediction results, perform chimerism rate analysis on the contributors to obtain the estimated chimerism rate. S33. By comparing the predicted chimerism rate with the filtered chimerism rate prediction results, the chimerism rate prediction results are further filtered. S34. Based on the predicted chimerism rate after filtering in step S33, chimerism rate analysis is performed on the contributors again to obtain the final chimerism rate. The filtering in step S31 includes the following operations in sequence: Based on the total number of read segments for each sub-marker, the splicing prediction results of sub-markers with a total number of read segments < 300 are filtered out; For a chimerism prediction result, if the minimum absolute deviation between the genotype detection frequency and the theoretical frequency in all SNP combinations is >0.1, it will be filtered out. For a single sub-marker, if it corresponds to multiple genotype combinations and their chirality prediction results are obtained, then the results are filtered based on the number of duplicates after removing the mean square deviation of the corresponding genotype frequencies. Specifically, when the number of contributors is 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 2, or when the number of contributors is greater than 2 and the number of values of the mean square deviation of the corresponding genotype frequencies is greater than 4, the chirality prediction results of that sub-marker are excluded. For a single sub-marker, if it corresponds to multiple genotype combinations with chimerism prediction results, the prediction result with the smallest sum of genotype base editing distances among the contributors is retained. For a single sub-marker, if it corresponds to multiple genotype combinations and the predicted chimerism rate is obtained, only the prediction result with the smallest mean square deviation of genotype frequency is retained. All chimerism prediction results were sorted according to the mean squared deviation of genotype frequencies, and chimerism prediction results between the 5th percentile and the 90th percentile were retained. Step S32 includes the following operations: The sub-markers are grouped according to the markers, and within each group, they are sorted in reverse order based on the total number of reads corresponding to the genotypes of the sub-markers. The chimerism prediction results of the top three sub-markers are extracted as the representative of the main marker's quantitative chimerism result for the current contributor. The Grubbs method was used to identify outliers in the mean squared deviation of genotype frequencies and outliers in the chimerism rate, and the chimerism rate prediction results of the corresponding sub-markers were filtered out. The sub-markers are then grouped again according to their respective markers, and the average value within each group is calculated to obtain the chimerism rate of the respective markers. The Grubbs method is used to identify outliers in marker chirp rate and filter out the corresponding markers. Based on the remaining markers, calculate the average value, which is the estimated splice rate of the current contributor; The estimated splicing rate of all contributors was calculated one by one; The filtering in step S33 includes the following operations: For a single chirp prediction result, calculate the ratio of the current predicted chirp to the estimated chirp. If the ratio of a certain contributor is greater than 4 or less than 1 / 3, then filter it out. For a single sub-marker, if it corresponds to multiple genotype combinations and the predicted chimerism rate is obtained, only the prediction result with the smallest mean square deviation of genotype frequency is retained. For a single splice rate prediction result, calculate the absolute deviation between the estimated splice rate and the current predicted splice rate, sort all splice rate prediction results according to the absolute deviation, and retain the splice rate prediction results that are below or equal to the 85th percentile. For a single sub-marker, if it corresponds to multiple genotype combinations and the chimerism prediction results are obtained, the prediction result with the smallest sum of base editing distances between the genotypes of the contributors is retained. For a single sub-marker, if it corresponds to multiple genotype combinations and the chimerism prediction results are obtained, the result with the smallest mean square deviation of genotype frequency is retained. All chimerism predictions were sorted according to the mean squared deviation of genotype frequencies, and predictions below and equal to the 85th percentile were retained. All splicing rate predictions are sorted again based on the absolute deviation between the estimated splicing rate and the current predicted splicing rate, and predictions below and equal to the 85th percentile are retained.
2. The method according to claim 1, characterized in that, The marker is selected from at least two of the micro-haplotypes shown in Table 1.
3. An analysis system for detecting chimerism rate based on next-generation sequencing data, characterized in that, The analysis system is used to perform the method of claim 1 or 2, and comprises the following modules: Data input module: The input data includes the human reference genome, marker site information, and next-generation sequencing data of the chimeric sample to be tested; Data processing module: Used to process sequencing data and analyze each marker individually to determine the specific genotype information of the marker in the sample and its corresponding read count; The chimerism analysis module is used to calculate all possible chimerism predictions for each marker. The chimerism rate calculation module: Based on the chimerism rate prediction results of the markers, it summarizes and calculates the final chimerism rate of the contributors.