Detection method for dominant clone caused by exogenous DNA insertion mutation
Through the detection method based on the PCR primer similarity ranking algorithm, the problem of difficult to evaluate the potential cancer risk caused by integrated vectors in gene therapy is solved, accurate detection of dominant clones and credible assessment of tumor risk is achieved, and the accuracy and clinical practicality of the detection are improved.
Patent Information
- Application Number
- PCT/CN2024/133762
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-19
AI Technical Summary
During gene therapy, host cell genomic instability and rearrangement caused by integrated vectors may lead to cancer, and existing detection methods are difficult to accurately evaluate the tumorigenicity and tumorigenicity of insertion mutations.
The detection method based on PCR primer similarity ranking algorithm is used to achieve the detection of dominant clones through positive and negative chain correction, redundant integration preprocessing, identification and removal of false positive results caused by non-specific amplification, cluster analysis and cloning plane evaluation.
It significantly reduces false positive results for IS identification, improves detection accuracy, can credibly evaluate tumor risk caused by insertion mutations, and increases the practicality of the detection at the clinical level.
Smart Images

Figure CN2024133762_19062025_PF_FP_ABST
Abstract
Description
Detection method based on dominant clones caused by exogenous DNA insertion mutation Technical Field
[0001] The present invention relates to the technical field related to second-generation sequencing data analysis in the field of bioinformatics, and specifically to a detection method for dominant clones caused by exogenous DNA insertion mutations. Background Art
[0002] Gene therapy refers to a treatment method that uses molecular biological methods to introduce exogenous DNA into the genome of genetically defective cells to restore normal cellular function. However, safety has always been a major issue plaguing the development of gene therapy. Integration vectors are a type of vector DNA sequence commonly used in gene therapy that are used to carry exogenous DNA fragments and integrate them into the host genome by insertion. Among them, lentiviral vectors (LVs) and adenovirus-like vectors (AAV) have become ideal tool vectors for gene therapy due to their efficient gene introduction efficiency or stable expression ability in target cell genes.
[0003] Once integrated into the host genome, an integrating vector can cause genomic instability and rearrangements in the host cell, potentially leading to disrupted gene expression, evading host immune recognition, and maintaining long-term survival, ultimately leading to cancer. Carcinogenicity due to genomic integration has been observed in preclinical studies and some clinical trials using viral vectors. In a gene therapy trial using AAV in dogs, researchers found that some of the therapeutic gene fragments carried by the AAV virus integrated into the dog's chromosomes near genes that control growth. Furthermore, some liver cells in these dogs divided more rapidly than other cells, forming clumps of daughter cells, potentially inducing cancer. Therefore, testing for the carcinogenic risk of vector insertion is a crucial component of gene therapy after the marketing of gene therapy products. However, during ongoing cell therapy, the onset of this potential tumorigenicity often has a significant lag and is difficult to predict. Currently, there is no reliable method to quantitatively assess the tumorigenicity and tumorigenicity of insertional mutations. Most current IS detection methods are based on PCR amplification of the ends of inserted fragments to identify IS. In practical applications, when the proportion of target fragments is low, nonspecific amplification will occur, leading to false positive results in IS identification. Summary of the Invention
[0004] To address the problem in existing technologies that the onset of potential tumorigenicity during continuous cell therapy often has a significant lag and is difficult to predict, there is currently no good method to quantitatively evaluate the tumorigenicity and tumorigenicity of insertional mutations; current IS detection methods are mostly based on PCR amplification of the ends of inserted fragments to identify IS. In practical applications, when the proportion of target fragments is low, nonspecific amplification will occur, resulting in false positive results in IS identification. The present invention provides a detection method based on dominant clones caused by exogenous DNA insertion mutations.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] The present invention is based on a method for detecting dominant clones caused by exogenous DNA insertion mutations, comprising the following steps:
[0007] Step 1: Pre-process the virus IS information data by correcting positive and negative chains and integrating redundancy;
[0008] Step 2: After pretreatment, IS caused by nonspecific amplification is identified to remove false positive IS test results caused by nonspecific PCR amplification;
[0009] Step 3: Merge and cluster the ISs at adjacent locations according to the site clustering algorithm;
[0010] Step 4: Analyze the clone plane based on the clustering results and describe the dominant clones.
[0011] As a preferred technical solution of the present invention, the virus IS information data in step 1 is a collection of virus ISs, wherein each IS indicates the integration position on the host genome.
[0012] As a preferred technical solution of the present invention, the specific operation of performing positive and negative chain correction on the viral IS information data in step 1 is that if the gene where the IS is located is a negative chain and the IS read segment is a negative chain, the IS is converted into a positive chain.
[0013] As a preferred technical solution of the present invention, the method of redundantly integrating viral IS information data is as follows: if the IS site and positive and negative chain information are completely consistent, they are merged into one record; the site and chromosome information corresponding to the IS are encoded and simplified into one number as the basis of the site clustering algorithm.
[0014] As a preferred technical solution of the present invention, the method for identifying IS caused by non-specific amplification in step 3 comprises three steps:
[0015] Step A: Extract all PCR primer sequences and calculate the reverse complementary sequence based on the primer sequence to form the primer sequence library S1; extract several base pairs before and after each IS according to its position to form the sequence library to be matched S2; on the target genome, randomly select a sufficient number of random sites to form the background sequence library S3;
[0016] Step B, matching and scoring, for each sequence in S2 and S3, match each primer sequence in the primer sequence library and its reverse complementary sequence respectively; in the matching results of the sequence and its reverse complementary sequence, select the highest matching score among all scores as the score of the similarity between the IS and the primer; after the matching is completed, each IS in S2 and S3 obtains a similarity score vector, and the vector length is equal to the number of primers; for each primer, the similarity scores corresponding to the primer in S2 and S3 are ranked, and the ranked scores are converted to percentiles, such as scores ranked in the top 1%, all converted to 1, and so on; after the conversion is completed, each IS corresponds to a percentile score vector, which indicates the ranking of the similarity between the sequence and the primer sequence among all sequences;
[0017] Step C: Determination of nonspecific amplification. If nonspecific amplification does not occur, the sequences in S2 and S3 are completely random, and the corresponding similarity ranking scores are theoretically evenly distributed. However, in the event of nonspecific amplification, the ranking distribution of the scores in S2 will be biased toward the top-ranked sites. By calculating the ranking vector of each IS in step 2, the probability that each IS in step 2 is nonspecifically amplifiable is obtained. The calculation method is to multiply the percentile score vectors of each IS. If the probability is less than a certain threshold T, the IS is marked as nonspecifically amplified. IS marked as specifically amplified will be removed from subsequent analyses.
[0018] As a preferred technical solution of the present invention, in step 3, the ISs at adjacent positions are merged according to the site clustering algorithm, and the clustering method is to group the ISs according to chromosomes, and the ISs in each chromosome; construct an IS distance matrix according to the distance between each IS genomic site, and then cluster the distance matrix using a hierarchical clustering algorithm; after the hierarchical clustering is completed, the result clusters of the hierarchical clustering are extracted and identified according to a certain distance threshold. The IS cluster is called UIS, and the representative site of the UIS is determined by the site with the highest read support number among all the sites that constitute the IS cluster.
[0019] As a preferred technical solution of the present invention, the method of analyzing the clone plane according to the clustering results in step 4 and describing the dominant clones is to use the proportion of the read segment support number of the top 10 IS in the sample to calculate and evaluate the dominant clones of the sample from two dimensions, namely the diversity of cell types at independent integration sites of the cells and the uniformity of the proportion of cell types at independent integration sites, that is, clone diversity and clone uniformity.
[0020] The beneficial effects of the present invention are:
[0021] This method for detecting dominant clones caused by exogenous DNA insertional mutations utilizes a PCR primer similarity ranking algorithm to identify and eliminate nonspecific amplifications potentially caused by PCR primer similarity. This significantly reduces false positives in IS identification and improves accuracy in practical applications. By assessing the IS diversity and IS uniformity of a sample, the present invention enables the detection of dominant clones, thereby providing a more reliable assessment of tumors caused by insertional mutations. By calculating the dominant clone plane, this method quantitatively assesses the state of disordered cell replication caused by IS insertions, predicts the tumorigenicity of insertional mutations, and enhances the clinical practicality of IS detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0023] FIG1 is a schematic flow chart of a method for detecting dominant clones based on exogenous DNA insertion mutations according to the present invention;
[0024] FIG2 is a sample schematic diagram of the selection of dominant clones based on exogenous DNA insertion mutations in Example 2 of the present invention;
[0025] FIG3 is a schematic diagram showing the calculation results of diversity and uniformity of 9 samples based on the detection method of dominant clones caused by exogenous DNA insertion mutations in Example 2 of the present invention;
[0026] FIG4 is a schematic diagram showing the prediction results of the cloning plane for nine samples based on the detection method of dominant clones caused by exogenous DNA insertion mutations in Example 2 of the present invention;
[0027] 5 is a schematic diagram comparing the similarity of primer sequences and the ranking distribution of random sites in the method for detecting dominant clones caused by exogenous DNA insertion mutations in Example 2 of the present invention;
[0028] FIG6 is a graph showing the IS ranking distribution curve of the samples after filtering according to the method for detecting dominant clones caused by exogenous DNA insertion mutations in Example 2 of the present invention;
[0029] 7 is a schematic diagram of the cloning plane calculation results of the detection method of the dominant clones caused by exogenous DNA insertion mutation in Example 2 of the present invention;
[0030] FIG8 is a schematic diagram of the top 10 IS ratios of 9 samples in the method for detecting dominant clones caused by exogenous DNA insertion mutations in Example 2 of the present invention. DETAILED DESCRIPTION
[0031] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0032] Example: As shown in FIG1 , the present invention is based on a method for detecting dominant clones caused by exogenous DNA insertion mutations, comprising the following steps:
[0033] Step 1: Pre-process the virus IS information data by correcting positive and negative chains and integrating redundancy;
[0034] Step 2: After pretreatment, IS caused by nonspecific amplification is identified to remove false positive IS test results caused by nonspecific PCR amplification;
[0035] Step 3: Merge and cluster the ISs at adjacent locations according to the site clustering algorithm;
[0036] Step 4: Analyze the clone plane based on the clustering results and describe the dominant clones.
[0037] The virus IS information data in step 1 is a collection of virus ISs, wherein each IS indicates the integration position on the host genome.
[0038] The specific operation of performing positive and negative strand correction on the virus IS information data in step 1 is that if the gene where the IS is located is a negative strand and the IS read segment is a negative strand, the IS is converted into a positive strand.
[0039] Among them, the method of redundant integration of virus IS information data is: if the site and positive and negative chain information of IS are completely consistent, they are merged into one record; the site and chromosome information corresponding to IS are encoded and simplified into one number as the basis of the site clustering algorithm.
[0040] The method for identifying IS caused by nonspecific amplification in step 3 comprises three steps:
[0041] Step A: Extract all PCR primer sequences and calculate the reverse complementary sequence based on the primer sequence to form the primer sequence library S1; extract several base pairs before and after each IS according to its position to form the sequence library to be matched S2; on the target genome, randomly select a sufficient number of random sites to form the background sequence library S3;
[0042] Step B, matching and scoring, for each sequence in S2 and S3, match each primer sequence in the primer sequence library and its reverse complementary sequence respectively; in the matching results of the sequence and its reverse complementary sequence, select the highest matching score among all scores as the score of the similarity between the IS and the primer; after the matching is completed, each IS in S2 and S3 obtains a similarity score vector, and the vector length is equal to the number of primers; for each primer, the similarity scores corresponding to the primer in S2 and S3 are ranked, and the ranked scores are converted to percentiles, such as scores ranked in the top 1%, all converted to 1, and so on; after the conversion is completed, each IS corresponds to a percentile score vector, which indicates the ranking of the similarity between the sequence and the primer sequence among all sequences;
[0043] Step C: Determination of nonspecific amplification. If nonspecific amplification does not occur, the sequences in S2 and S3 are completely random, and the corresponding similarity ranking scores are theoretically evenly distributed. However, in the event of nonspecific amplification, the ranking distribution of the scores in S2 will be biased toward the top-ranked sites. By calculating the ranking vector of each IS in step 2, the probability that each IS in step 2 is nonspecifically amplifiable is obtained. The calculation method is to multiply the percentile score vectors of each IS. If the probability is less than a certain threshold T, the IS is marked as nonspecifically amplified. IS marked as specifically amplified will be removed from subsequent analyses.
[0044] Among them, in the step 3, the ISs at adjacent positions are merged according to the site clustering algorithm, and the clustering method is to group the ISs according to chromosomes, and the ISs in each chromosome; construct an IS distance matrix according to the distance between each IS genomic site, and then cluster the distance matrix using a hierarchical clustering algorithm; after the hierarchical clustering is completed, the IS cluster after extraction and identification of the hierarchical clustering result clusters according to a certain distance threshold is called UIS, and the representative site of UIS is determined by the site with the highest read support number among all the sites that make up the IS cluster.
[0045] Among them, the method of analyzing the cloning plane according to the clustering results in step 4 and describing the dominant clones is to use the proportion of the read segment support number of the top 10 IS in the sample to calculate and evaluate the dominant clones of the sample from two dimensions, namely the diversity of cell types at independent integration sites of the cells and the uniformity of the proportion of cell types at independent integration sites, that is, clonal diversity and clonal uniformity.
[0046] The diversity index is calculated from Simpson's Diversity Index. The calculation method is D = 1-Σ(ni / N) 2
[0047] Where D is the Simpson index, n is the number of reads supporting each detected UIS, and N is the total number of UIS.
[0048] The calculation method of the uniformity index is based on the Shannon index. The specific calculation method is as follows: H = -Σ(pi*log2(pi))
[0049] Where H is the Shannon index, p is the proportion of each IS in the total IS, log2 represents the binary logarithm, ln represents the base of the binary logarithm, and S represents the total number of IS.
[0050] In a coordinate system, diversity and evenness are used as the x-axis and y-axis metrics, respectively, and a sample can be projected onto this axis as a point. This method uses over 210 samples from a local database to construct a machine learning discriminant model based on a support vector machine. The diversity and evenness indicators of the samples are used as independent variables, and clonality is used as the dependent variable. The support vector machine is trained to find a dividing hyperplane in the sample space that separates samples with dominant clones from those without dominant clones (as shown in Figure 4). After obtaining the dividing hyperplane, this method uses it to assess the dominant clonality of new samples. This plane can determine whether a cell has a dominant clone. The specific judgment rule is: if the sample is below the cloning plane, it is a dominant clone; otherwise, it is not. If the results of the cloning plane determine that the cell has not undergone a dominant clone, the sample is polyclonal; if the results of the cloning plane determine that the cell has undergone a dominant clone, the sample is oligoclonal.
[0051] This method utilizes a PCR primer similarity ranking algorithm to identify and eliminate nonspecific amplification caused by possible PCR primer similarity, significantly reducing false positives in IS identification and improving accuracy in practical applications. By assessing the IS diversity and IS uniformity of a sample, the present invention enables the detection of dominant clones, thereby providing a more reliable assessment of tumors caused by insertional mutations. By calculating the dominant clone plane, this method quantitatively assesses the state of disordered cell replication caused by IS insertions, predicts the tumorigenicity of insertional mutations, and enhances the clinical practicality of IS detection.
[0052] Example 2, as shown in Figures 2 to 8, 4 samples with dominant clones and 5 samples without dominant clones were selected respectively, and then the method was used to detect dominant clones on the 9 samples (the top 10 IS percentages of the 9 samples, of which samples 2, 4, 5, and 6 had obvious dominant clones); the detection of the example samples illustrates the impact of false positive IS produced by nonspecific amplification in the samples on the true results; and after treatment by this method, the false positives were significantly reduced, indicating that this method has a good removal effect on false positive IS caused by nonspecific amplification. First, this method extracts all PCR primer sequences used in the experiment, calculates its reverse complementary sequence based on the primer sequence, and becomes the primer sequence library S1. Then, this method extracts the flanking sequences of 100bp before and after each IS in the two samples to form a sequence library S2 to be matched, which contains the IS flanking sequences from the two samples. At the same time, 1000 random sites are randomly selected on the target genome to form a random site background sequence library S3.
[0053] After obtaining the three sequence libraries S1, S2, and S3, this method matches each sequence in the sample sequence library (S2) and the random site sequence library (S3) to the three primer sequences and their respective reverse complements in S1. Within the similarity scores a and b between each primer sequence and its reverse complement, M = max(a, b) is selected as the similarity score between the sequence and the primer. Finally, the similarity scores M1, M2, and M3 for each primer are used as the final score vector, as shown in Table 2. After scoring, each sequence in S2 and S3 is assigned a set of similarity score vectors. The similarity score vectors corresponding to different primers in S2 and S3 are ranked, and the ranking scores are converted to percentiles. For example, scores in the top 1% are converted to 1, scores in the top 1%-2% are converted to 2, and so on. After the conversion, each sequence in S2 and S3 is assigned a ranking vector of length 3, indicating the rank of the best similarity score between the sequence and the three primer sequences among all sequences.
[0054] Comparison of the density distribution of IS scores in samples and scores from random sequences. The density distribution shows that, compared to random sequences, the ranking distribution of the two example samples is significantly biased towards the top-ranked sites, indicating that the sequences near the IS in these two samples are more similar to the PCR primers. Figure 5 shows a comparison of the similarity of primer sequences to the ranking distribution of random sites in two samples (Sample 1 and Sample 2) with significant nonspecific amplification leading to false positives.
[0055] A distinct characteristic of nonspecific amplification is its multiple occurrence in different samples. ISs randomly integrated into the human genome are unlikely to appear in both samples. We identified multiple ISs that were detected simultaneously in both samples. Comparing the scores of these ISs with those of other ISs revealed that the score density distribution for ISs that appeared simultaneously in both samples was more severely biased toward the top-ranked sites. As shown in Figure 6, the ISs detected in both samples 1 and 2 showed the highest similarity with the primers. However, after filtering using this method, the similarity ranking distribution between the ISs and primers shifted more towards that of random sites. This indicates that this method is effective in reducing false positives in IS identification and improving accuracy, which also indirectly illustrates the correlation between the top scores and nonspecific amplification.
[0056] Figure 7 shows the results of the cloning plane calculation. The x-axis represents sample diversity, and the y-axis represents sample uniformity. As can be seen, samples with dominant clones are below the cloning plane, while samples without dominant clones are above the cloning plane. This demonstrates that the cloning plane defined by this method can accurately identify samples with dominant clones.
[0057] Finally, this method multiplies the ranking scores in S2 to convert them into the probability that the IS is not nonspecifically amplified with the three primers. ISs with a probability higher than T are then labeled as specifically amplified sequences and removed. After removing ISs marked with nonspecific amplification, the skewness of the overall score curve distribution decreases significantly, resulting in a nearly uniform distribution, demonstrating that this method is effective in identifying nonspecific amplification of ISs.
[0058] After obtaining the diversity and uniformity indicators for each sample, this method establishes a coordinate system with diversity as the x-axis and uniformity as the y-axis, respectively. The calculated diversity and uniformity results for each sample are used as the x-axis and y-axis coordinates of a point, respectively, projected onto these coordinate systems. The cloning plane, determined through machine learning, is then plotted on this coordinate system. The relative position of the sample with respect to the cloning plane is observed (as shown in Figure 7). If a sample is below the cloning plane, it is a dominant clone; otherwise, it is not. As can be seen in Figure 7, samples 2, 4, 5, and 6, which exhibit dominant cloning, are all below the cloning plane, represented by dots; while samples 1, 3, 7, 8, and 9, which do not exhibit dominant cloning, are all above the cloning plane, represented by triangles. The judgment results are shown in Table 5. These results demonstrate that this method can accurately determine the clonality of all samples.
[0059] In summary, this method utilizes a PCR primer similarity ranking algorithm to identify and eliminate nonspecific amplification caused by potential PCR primer similarity. This significantly reduces false positives in IS identification and improves the accuracy of IS identification in practical applications. Furthermore, by assessing IS diversity and IS uniformity across samples, this method enables the detection of dominant clones, providing a more reliable assessment of tumors caused by insertional mutations and enhancing the clinical utility of IS detection.
[0060] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for detecting dominant clones caused by exogenous DNA insertion mutation, characterized in that: The following steps are involved: Step 1: Preprocess the virus IS information data by positive and negative chain correction and redundant integration; Step 2: After the pretreatment is completed, the IS caused by nonspecific amplification is identified to remove the false positive IS test results caused by nonspecific PCR amplification; Step 3: According to the site clustering algorithm, the ISs at adjacent locations are merged and clustered; Step 4: Analyze the clone plane according to the clustering results and describe the dominant clones.
2. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 1, characterized in that: The virus IS information data in step 1 is a collection of virus ISs, wherein each IS indicates the integration position on the host genome.
3. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 1, characterized in that: The specific operation of performing positive and negative strand correction on the virus IS information data in step 1 is that if the gene where the IS is located is a negative strand and the IS read segment is a negative strand, the IS is converted into a positive strand.
4. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 3, characterized in that: The method for redundantly integrating virus IS information data is as follows: if the IS site and positive and negative chain information are completely consistent, they are merged into one record; the site and chromosome information corresponding to the IS are encoded and simplified into one number as the basis of the site clustering algorithm.
5. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 1, characterized in that: The method for identifying IS caused by non-specific amplification in step 3 includes three steps: Step A: extract all PCR primer sequences, calculate the reverse complementary sequence according to the primer sequence, and form the primer sequence library S1; extract several base pairs before and after each IS according to its position, and form the sequence library S2 to be matched; on the target genome, randomly select a sufficient number of random sites to form the background sequence library S3; Step B, matching and scoring, for each sequence in S2 and S3, each primer sequence in the primer sequence library and its reverse complementary sequence are matched respectively; in the matching results of the sequence and its reverse complementary sequence, the highest matching score among all the scores is selected as the score of the similarity between the IS and the primer; after the matching is completed, each IS in S2 and S3 obtains a similarity score vector, and the vector length is equal to the number of primers; for each primer, the similarity scores corresponding to the primer in S2 and S3 are sorted, and the sorted scores are converted by percentiles, such as the scores ranked in the top 1%, all converted to 1, and so on; after the conversion is completed, each IS corresponds to a percentile score vector, which indicates the ranking of the similarity between the sequence and the primer sequence among all sequences; Step C, judgment of non-specific amplification; if there is no non-specific amplification, the sequences in S2 and S3 are completely random, and the distribution of the corresponding similarity ranking scores is theoretically uniform; in the case of non-specific amplification, the ranking distribution of the scores in S2 will be biased towards the top-ranked sites; by calculating the ranking vector of each IS in step 2, the probability of each IS in step 2 being non-specifically amplifiable is obtained; The calculation method is to multiply the percentile score vectors of each IS. If the probability is less than a certain threshold T, the IS is marked as non-specific amplification; the IS marked as specific amplification will be removed in subsequent analysis.
6. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 1, characterized in that: In the step 3, the ISs at adjacent positions are merged according to the site clustering algorithm, and the clustering method is to group the ISs according to chromosomes, and the ISs in each chromosome; construct an IS distance matrix according to the distance between each IS genomic site, and then cluster the distance matrix using a hierarchical clustering algorithm; after the hierarchical clustering is completed, the IS cluster after extracting and identifying the result clusters of the hierarchical clustering according to a certain distance threshold is called UIS, and the representative site of UIS is determined by the site with the highest read support number among all the sites that make up the IS cluster.
7. The method for detecting dominant clones based on exogenous DNA insertion mutation according to claim 1, characterized in that: The method for analyzing the cloning plane according to the clustering results in step 4 and describing the dominant clones is to use the proportion of the read support number of the top 10 IS in the sample to calculate and evaluate the dominant clones of the sample from two dimensions, namely, the diversity of cell types at independent integration sites of the cells and the uniformity of the proportion of cell types at independent integration sites, that is, clonal diversity and clonal uniformity.
Citation Information
Patent Citations
Automatic sequencing analysis method and device for tiny residual lesions based on NGS
CN111261226A
Method for analyzing lymphocytes or plasma cells in sample by applying immune repertoire sequencing method, application of method and kit of method
CN112852936A
IGH hypermutation detection method and system based on NGS amplicon sequencing technology
CN115433768A
Detection method based on dominant clone caused by exogenous DNA insertion mutation
CN117746982A
Identification of splicing disrupting mutations and use thereof
WO2023223330A1