Primer design method and device, disease subtyping diagnosis method and device, composition, kit, and use thereof
By designing primers and standard quality grains, the problems of false positives, false negatives, and PCR amplification bias in TRD gene rearrangement detection were solved, enabling efficient and accurate lymphoma diagnosis and monitoring of small lesions.
Patent Information
- Application Number
- PCT/CN2024/115631
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Existing methods for detecting TRD gene rearrangements suffer from false positives, false negatives, low detection sensitivity, and inability to detect minute lesions. In addition, high-throughput sequencing methods have problems such as PCR amplification bias and incomplete coverage of clone types.
A primer design method was designed to select conservative regions by calculating conservation scores, generate primer combinations for efficient amplification, and combine them with standard quality plasmid design and sequencing library construction to improve detection accuracy and sensitivity.
It improves the detection rate and accuracy of lymphoma, and can effectively detect TRD gene rearrangements, enabling accurate diagnosis of lymphoma and monitoring of small lesions.
Smart Images

Figure CN2024115631_05032026_PF_FP_ABST
Abstract
Description
Primer design, disease typing and diagnostic methods and devices, compositions, kits and their uses Technical Field
[0001] This disclosure relates to, but is not limited to, the field of biotechnology, and particularly to primer design, disease typing and diagnostic methods and apparatus, compositions, kits and their uses. Background Technology
[0002] The T cell receptor (TCR) is a membrane receptor on the surface of T lymphocytes, composed of two polypeptide chains: α(TRA) and β(TRB) or γ(TRG) and δ(TRD). α(TRA) and γ(TRG) are light chains, while β(TRB) and δ(TRD) are heavy chains, with α(TRA) and β(TRB) forming the vast majority of TCRs. The TRD gene consists of a variable region (V), a diversity region (D), a joining region (J), and a constant region (C). Specifically, the TRD V / D / J gene clusters each contain multiple V, D, or J gene segments. During lymphocyte development, a gene segment is randomly selected from each of the V, D, or J gene clusters, and under the action of recombinase, it is cleaved and linked together to form a complete gene encoding the TRD heavy chain function—a process known as gene rearrangement. Because the V, D, or J segments that make up the TRD gene are diverse, and varying numbers of bases are randomly inserted or deleted between DJ or V-DJ, the TRD protein exhibits diversity, i.e., polyclonal TRD gene rearrangements. Lymphocytes carry specific TRD rearrangement sequences. During lymphoma development, a particular lymphocyte undergoes malignant proliferation accompanied by the proliferation of its specific TRD rearrangement sequence, i.e., lymphoma TRD gene rearrangement monoclonal. Both polyclonal and monoclonal TRD gene rearrangements provide important auxiliary means for lymphoma diagnosis.
[0003] Traditional detection methods for TRD gene rearrangements involve capillary electrophoresis combined with fluorescent fragment analysis based on first-generation sequencing platforms. This involves designing specific PCR primers and labeling them with fluorescence at the 5' end, obtaining the target fragment through PCR amplification, and then separating the amplification products by capillary electrophoresis to form a fragment size distribution peak map, thereby determining whether the TRD is monoclonal or polyclonal and aiding in the diagnosis of lymphoma. However, this method has the following drawbacks:
[0004] (a) False positive results exist: This analytical method is based on the size of PCR product fragments, which leads to fragments with different sequences but the same length being mixed together to form false positive peaks;
[0005] (b) False negative results exist: the fragment distribution peak diagram is limited to a certain range, causing positive peaks outside the range to be ignored;
[0006] (c) Limited clinical application: This method is mainly used to determine tumors or hyperplasia in lymphatic system diseases. Because it is impossible to sequence the specific sequence of each clone, it cannot be used to monitor small residual lesions, etc.
[0007] (d) Low detection sensitivity: The inability to accurately assess and correct PCR amplification bias leads to low detection sensitivity.
[0008] In recent years, high-throughput sequencing technology has been widely used in the detection of TRD gene rearrangements. This involves specific amplification of TRD fragments using multiplex PCR primers, construction of sequencing libraries using adapter ligation or PCR amplification, and identification of TRD gene rearrangement monoclonal or polyclonal sequences through TRD clone sequence analysis and frequency statistics. However, currently available products suffer from problems such as significant PCR amplification bias and incomplete coverage of clone types, leading to low detection accuracy.
[0009] Summary of the Invention
[0010] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0011] This disclosure provides a primer design method, including:
[0012] Obtain one or more reference data, each of which includes multiple reference gene sequences;
[0013] For each type of reference data, the following operations are performed: align the multiple reference gene sequences by site, determine the conservation score list for each site, obtain the conservation scores of multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generate K primer combinations for the K conservation intervals; screen the primers in the K primer combinations and evaluate the screened K primer combinations, and obtain the final primer combination based on the evaluation results.
[0014] This disclosure also provides a primer design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the primer design method according to any embodiment of this disclosure based on the instructions stored in the memory.
[0015] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the primer design method described in any embodiment of this disclosure.
[0016] This disclosure also provides a program product including instructions that, when executed by a computer, perform the primer design method as described in any embodiment of this disclosure.
[0017] The primer design method and apparatus of this disclosure calculate conservation scores to obtain multiple conservation intervals, and then design, screen and evaluate primers for the conservation intervals to finally obtain a set of primers for efficient amplification. By using this primer set for high-throughput sequencing, the problem of low sensitivity of traditional capillary electrophoresis + fragment analysis is solved, as well as the problems of high bias and low coverage of PCR amplification in previous high-throughput sequencing methods, which effectively improves the detection rate and accuracy of lymphoma.
[0018] This disclosure also provides a standard quality grain design method, including:
[0019] Obtain a first gene cluster, a second gene cluster, and a third gene cluster. The first gene cluster includes multiple first gene sequence fragments, the second gene cluster includes multiple second gene sequence fragments, and the third gene cluster includes multiple third gene sequence fragments. Each of the first, second, and third gene clusters contains multiple functional fragments.
[0020] Multiple first gene sequence fragments, multiple second gene sequence fragments, and multiple third gene sequence fragments are combined to obtain multiple fragment groups, each fragment group containing one first gene sequence fragment, one second gene sequence fragment, and one third gene sequence fragment;
[0021] A non-human sequence was inserted between the second and third gene sequence fragments in each fragment group to obtain multiple plasmid sequences.
[0022] The ratio of each plasmid sequence was determined to obtain the designed standard plasmid.
[0023] This disclosure also provides a composition comprising, obtained by the methods described herein:
[0024] TRD upstream decoy oligonucleotides, wherein the TRD upstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO: 1-9; and
[0025] TRD downstream decoy oligonucleotides, wherein the TRD downstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO:10-15.
[0026] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:1-9; and the downstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:10-15.
[0027] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:1-9; and the downstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:10-15.
[0028] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRD is selected from all sequences in SEQ ID NO:1-9; the downstream decoy oligonucleotide of the TRD is selected from all sequences in SEQ ID NO:10-15.
[0029] In some exemplary embodiments, the bait oligonucleotide is one or more of the primer and probe.
[0030] In some exemplary embodiments, the bait oligonucleotides may be primers for amplifying the TRD gene (e.g., a combination of TRD-specific primers, including primers for the TRD V region, TRDδ2, TRD J region, and TRDδ3), and are divided into upstream TRD bait oligonucleotides and downstream TRD bait oligonucleotides. The upstream TRD bait oligonucleotides include specific primer sequences that are complementary to the upstream primers of the TRD V region and TRDδ2, and the downstream TRD bait oligonucleotides include specific primer sequences that are complementary to the downstream primers of the TRD J region and TRDδ3.
[0031] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRD further comprises a forward adapter primer sequence, and the downstream decoy oligonucleotide of the TRD further comprises a reverse adapter primer sequence; the forward adapter primer sequence and the reverse adapter primer sequence are complementary to the adapter primer sequence used for sequencing.
[0032] In some exemplary embodiments, the adapter primers used for sequencing are selected from one or more of the Nextera series adapters and the TruSeq series adapters.
[0033] In some exemplary embodiments, the forward adapter primer sequence and the reverse adapter primer sequence are located at both ends of each pair of bait oligonucleotides in this application, for subsequent addition of primers to both ends of the PCR product.
[0034] In some exemplary embodiments, the forward adapter primer sequence is shown in SEQ ID NO:16, and the reverse adapter primer sequence is shown in SEQ ID NO:17.
[0035] In some exemplary embodiments, the upstream decoy oligonucleotide of the TRD is selected from one or more sequences shown in SEQ ID NO:18-26; and
[0036] The downstream decoy oligonucleotide of the TRD is selected from one or more sequences shown in SEQ ID NO:27-32.
[0037] This disclosure also provides the use of the compositions described herein in amplifying the TRD gene and / or detecting TRD gene rearrangements.
[0038] In some exemplary embodiments, the decoy oligonucleotides described herein (e.g., upstream and downstream TRD decoy oligonucleotides) can be used to amplify the TRD gene to obtain rearranged PCR products, including VJ rearrangement products, DJ rearrangement products, DD rearrangement products, and VD rearrangement products.
[0039] In some exemplary embodiments, adapter primers can be added to both ends of the bait oligonucleotides described herein to obtain PCR-amplified bait oligonucleotides. PCR is then performed using these PCR-amplified bait oligonucleotides, and the resulting PCR products can be sequenced to obtain the sequence of each rearrangement product. The rearrangement of the TRD gene can thus be determined more accurately and efficiently.
[0040] This disclosure also provides a kit comprising the compositions described herein.
[0041] In some exemplary embodiments, the kit further comprises:
[0042] Twelve standard quality particles, each containing a UMI sequence, and each standard quality particle containing a different UMI sequence, so that each standard quality particle can be uniquely identified by the UMI sequence;
[0043] The UMI sequence is 12 bp in length, with the first 8 bp being a non-human random sequence and the last 4 bp being the first 4 bases of the TRD J region sequence.
[0044] Each of the standard quality grains further comprises a TRD V region sequence, a TRD D region sequence, and a TRD J region sequence, wherein the TRD V region sequence, the TRD D region sequence, the UMI sequence, and the TRD J region sequence are sequentially linked end-to-end in each of the standard quality grains.
[0045] In some exemplary embodiments, the last 4 bp segment of the UMI sequence is selected from ACAC, CTTT, CTCC, CCAG, or CACA.
[0046] In some exemplary embodiments, the UMI sequence is as shown in SEQ ID NO:33-44.
[0047] In some exemplary embodiments, the 12 standard quality grains are standard quality grains that respectively contain the following sequences: TRDV1*01, TRDD1&TRDD2 and TRDJ1*01; TRDV2*01, TRDD1&TRDD3 and TRDJ2*01; TRDV3*01, TRDD2&TRDD3 and TRDJ3*01; TRAV14 / DV4*01, TRDD1&TRDD2 and TRDJ4*01; TRAV23 / DV6*01, TRDD1&TRDD3 and TRDJ1*01; TRAV29 / DV5*01, TRDD2&TRDD3 and TRDJ2*01; TRAV36 / DV7*01, TRDD1&TRDD2 and TRDJ3*01; TRAV38-2 / DV8*01, TRDD2&TRDD3 and TRDJ4*01; TRD-δ2-F-intron, TRDD2&TRDD3 and TRDJ3*01; TRD-δ2-F -intron, TRDD2&TRDD3 and TRD-δ3-R-intron; TRDV2*01, TRDD1&TRDD3 and TRD-δ3-R-intron; or TRAV23 / DV6*01, TRDD2&TRDD3 and TRD-δ3-R-intron.
[0048] In some exemplary embodiments, the 12 standard quality particles are mixed in an equimolar ratio to obtain a uniformity standard.
[0049] By mixing one or more of the 12 types of standard quality grains in a high proportion and the other standard quality grains in a low proportion, an experimental standard for simulating monoclonal rearrangements is obtained. In some exemplary embodiments, the high proportion of standard quality grains in the experimental standard can be set as needed, and can be one or more types (two, three, four, five or more). The concentration of the high proportion of standard quality grains can be much higher than that of the low proportion of standard quality grains, for example, 100 times, 1000 times, 10000 times, 100,000 times, 1 million times or more of the concentration of the low proportion of standard quality grains. In some exemplary embodiments, the concentration of the high proportion of standard quality grains in the experimental standard can be slightly higher than that of the low proportion of standard quality grains, for example, 1.1 times, 1.25 times, 1.5 times, 2 times, 4 times, 8 times, 10 times or more of the concentration of the low proportion of standard quality grains.
[0050] This disclosure also provides the use of the kit described herein in evaluating the amplification efficiency of multiple primers used to amplify the TRD gene.
[0051] This disclosure also provides embodiments of the kit described herein for use in detecting TRD gene rearrangements.
[0052] This disclosure also provides a disease typing diagnostic method for detecting TRD gene rearrangements, including the following steps:
[0053] 1) Obtain the genomic DNA of the sample to be tested;
[0054] 2) Perform PCR on the genomic DNA obtained in step 1) using the composition described herein to obtain PCR products;
[0055] 3) Sequencing the PCR products obtained in step 2) and analyzing the sequencing results to determine whether the TRD gene rearrangement in the sample is monoclonal or polyclonal.
[0056] In some exemplary embodiments, step 1) further includes incorporating the homogeneity standard described herein into the genomic DNA. In some exemplary embodiments, different standards (e.g., homogeneity standards or experimental standards) may be incorporated in step 1) for different experimental purposes.
[0057] This disclosure also provides a standard plasmid design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the standard plasmid design method according to any embodiment of this disclosure based on the instructions stored in the memory.
[0058] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the standard quality grain design method described in any embodiment of this disclosure.
[0059] This disclosure also provides a program product including instructions that, when executed by a computer, perform a standard quality grain design method as described in any embodiment of this disclosure.
[0060] The standard plasmid design method and apparatus of this disclosure obtain multiple plasmid sequences by inserting a non-human sequence between the first gene sequence fragment and the second gene sequence fragment in each fragment group. The plasmid sequences can be identified and effectively separated from the mixed sample, thereby enabling the quantification of the mixed sample.
[0061] This disclosure also provides a method for quantitative analysis of standard quality grains, including:
[0062] Obtain paired-end sequencing data corresponding to the standard quality plasmid, wherein the paired-end sequencing data includes Reads1 and Reads2 sequences;
[0063] Determine whether the Reads1 and Reads2 sequences include a UMI sequence identifier;
[0064] When both the Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequences corresponding to the Reads1 and Reads2 sequences have been identified.
[0065] When only one of the Reads1 and Reads2 sequences includes a UMI sequence identifier, the Hamming distance between the Reads sequence that does not include the UMI sequence identifier and the UMI sequence identifier is determined; when the determined Hamming distance is less than or equal to a, the plasmid sequence corresponding to the Reads1 and Reads2 sequences is identified, where a is a natural number less than or equal to 2.
[0066] This disclosure also provides a standard plasmid quantification analysis apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the standard plasmid quantification analysis method according to any embodiment of this disclosure based on the instructions stored in the memory.
[0067] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the standard plasmid quantitative analysis method described in any embodiment of this disclosure.
[0068] This disclosure also provides a program product including instructions that, when executed by a computer, perform a standard plasmid quantitative analysis method as described in any embodiment of this disclosure.
[0069] The standard plasmid quantitative analysis method and apparatus of this disclosure can quantify mixed samples by identifying the plasmid sequence corresponding to the read sequence based on the UMI sequence identifier and Hamming distance. For example, it can identify whether the experimental mixing ratio meets expectations (such as whether the plasmid addition ratio is consistent with the sequencing detection ratio).
[0070] This disclosure also provides a primer characterization method, including:
[0071] Obtain sequencing data corresponding to the primer to be tested, and determine the 5' end position of the primer to be tested based on the sequencing data;
[0072] Multiple sequences are designed based on the 5' end position of the primer to be detected. Each of the multiple sequences starts at the 5' end position of the primer to be detected, and the 3' end cutoff positions of two adjacent sequences differ by a single base.
[0073] Design multiple nucleic acid structures based on multiple designed sequences;
[0074] Identify the nucleic acid structures among the multiple nucleic acid structures that can complementarily pair with the primers to be detected;
[0075] The sequence of the primer to be detected is determined based on the nucleic acid structure that can complementarily pair with the primer to be detected.
[0076] The primer characterization method of this disclosure determines the 5' end position of the primer to be detected based on sequencing data, designs multiple sequences and multiple corresponding nucleic acid structures based on the 5' end position of the primer to be detected, then determines the nucleic acid structure that can complementarily pair with the primer to be detected among the multiple nucleic acid structures, and determines the sequence of the primer to be detected based on the nucleic acid structure that can complementarily pair with the primer to be detected, thus enabling the sequencing of any unknown primer.
[0077] This disclosure also provides a method for constructing a sequencing library, including:
[0078] Extract DNA from the genome to be tested;
[0079] Design specific primers for the genome to be detected, mix the specific primers with the extracted DNA, perform the first round of PCR amplification to obtain the first round of PCR product, and purify the first round of PCR product;
[0080] The purified first-round PCR product was mixed with universal adapter primers and subjected to a second-round PCR amplification to obtain a second-round PCR product. The second-round PCR product was then purified to obtain the constructed sequencing library.
[0081] The sequencing library construction method of this disclosure uses a two-step PCR method to ligate sequencing adapters to the 5' and 3' ends of the target fragment, respectively, to convert the target DNA to be sequenced into a form that can be read by the sequencer.
[0082] This disclosure also provides a method for disease classification and diagnosis, including:
[0083] Obtain sequencing data of the sample to be tested, and perform data preprocessing on the obtained sequencing data;
[0084] The preprocessed data is compared with the reference gene sequences of multiple pre-defined clone species to obtain the sequence proportion corresponding to each clone species.
[0085] Sort the proportions of multiple sequences from largest to smallest, and label the proportions of multiple sequences in the sorting order as the proportion of the first sequence to the proportion of the M1th sequence, where M1 is the number of clone species;
[0086] The method detects whether the differences between the proportions of the first and third sequences, and between the proportions of the second and third sequences, exceed preset difference thresholds. When both the differences between the proportions of the first and third sequences and between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample. When the differences between the proportions of the first and third sequences exceed preset difference thresholds but the differences between the proportions of the second and third sequences do not exceed preset difference thresholds, the sample to be tested is determined to be a monoclonal sample. When neither the differences between the proportions of the first and third sequences nor the differences between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.
[0087] This disclosure also provides a disease typing diagnostic apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of the disease typing diagnostic method described in any embodiment of this disclosure based on the instructions stored in the memory.
[0088] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the disease classification and diagnosis method described in any embodiment of this disclosure.
[0089] This disclosure also provides a program product including instructions that, when executed by a computer, perform a disease subtyping diagnosis method as described in any embodiment of this disclosure.
[0090] The disease typing diagnosis method and apparatus of this disclosure identify each TRD clone sequence and determine whether the sample to be tested is a monoclonal sample, oligoclonal sample or polyclonal sample based on whether the difference between the proportions of the top three sequences with the largest proportion exceeds a preset difference threshold. It can determine the specific sequence of each clone, thereby accurately typing and diagnosing the sample to be tested, and thus realizing the tracking and monitoring of residual small lesions of lymphoma.
[0091] After reading and understanding the accompanying diagrams and detailed descriptions, other aspects can be understood.
[0092] Overview of the attached figures
[0093] The accompanying drawings are provided to further illustrate the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure. The shapes and sizes of the components in the drawings do not reflect actual proportions and are only intended to illustrate the content of this disclosure.
[0094] Figure 1 is a flowchart illustrating a primer design method provided by an exemplary embodiment of this disclosure;
[0095] Figure 2 is a schematic diagram of the TRD sequence structure;
[0096] Figure 3 is a schematic diagram of the TRD detection result of a PBMC negative sample provided by an exemplary embodiment of this disclosure;
[0097] Figure 4 is a schematic diagram of the TRD detection results of a lymphoma-positive sample provided by an exemplary embodiment of this disclosure;
[0098] Figure 5 is a flowchart illustrating a standard quality grain design method provided by an exemplary embodiment of this disclosure;
[0099] Figure 6 is a graph showing the detection results of the proportion of TRD plasmids obtained by using uniformity standards and TRD multiple primers in an exemplary embodiment of this disclosure.
[0100] Figure 7 is a schematic flowchart of a standard quality grain quantitative analysis method provided by an exemplary embodiment of the present disclosure;
[0101] Figure 8 is a flowchart illustrating a primer characterization method provided in an exemplary embodiment of this disclosure;
[0102] Figure 9 is a schematic diagram of a set (10) nucleic acid structures provided in an exemplary embodiment of this disclosure;
[0103] Figure 10 is a schematic diagram of the ligation products of the nucleic acid structure and primers shown in Figure 9;
[0104] Figure 11 is a schematic diagram of the process of performing Sanger fragment analysis on the ligation products shown in Figure 10;
[0105] Figure 12 is a flowchart illustrating a sequencing library construction method provided by an exemplary embodiment of this disclosure;
[0106] Figure 13 is a schematic diagram of the lymphoma TRD gene rearrangement detection library construction process provided by an exemplary embodiment of the present disclosure;
[0107] Figure 14 is a flowchart illustrating a disease classification and diagnosis method provided by an exemplary embodiment of this disclosure;
[0108] Figures 15A to 15C are schematic diagrams showing the clone types and corresponding sequencing sequence numbers of three TRD samples provided in the exemplary embodiments of this disclosure;
[0109] Figure 16 is a schematic diagram of a primer design device provided in an exemplary embodiment of the present disclosure;
[0110] Figure 17 is a schematic diagram of a standard quality grain design device provided by an exemplary embodiment of the present disclosure;
[0111] Figure 18 is a schematic diagram of a standard quality grain quantitative analysis device provided by an exemplary embodiment of the present disclosure;
[0112] Figure 19 is a schematic diagram of the structure of a disease typing diagnostic device provided by an exemplary embodiment of the present disclosure.
[0113] Detailed Explanation
[0114] To make the objectives, technical solutions, and advantages of this disclosure clearer, the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be arbitrarily combined with each other.
[0115] Unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, but do not exclude other elements or objects.
[0116] As shown in Figure 1, this disclosure provides a primer design method, including:
[0117] Step 101: Obtain one or more reference data sets, each of which includes multiple reference gene sequences;
[0118] Step 102: For each type of reference data, perform the following operations: Align multiple reference gene sequences by site, determine the conservation score list for each site, obtain the conservation scores of multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generate K primer combinations for the K conservation intervals; screen the primers in the K primer combinations and evaluate the screened K primer combinations, and obtain the final primer combination based on the evaluation results.
[0119] The primer design method of this disclosure involves aligning multiple reference gene sequences by site to determine a conservation score list for each site. Based on the conservation score list for each site, multiple conservation intervals are obtained with conservation scores. K conservation intervals with high conservation scores are selected, where K is a natural number greater than or equal to 1. K primer combinations are generated for these K conservation intervals. The primers in the K primer combinations are screened and evaluated. Based on the evaluation results, the final primer combination is obtained. This method can design PCR primers with high specificity and good uniformity at each target site, thereby effectively solving the problems of large PCR amplification bias and incomplete coverage of clone types in current products, which leads to low detection accuracy.
[0120] In this embodiment of the disclosure, a conservative interval refers to the interval in which the sequence similarity between multiple sequences exceeds a preset similarity score threshold. Within a conservative interval, different sequences exhibit a high degree of similarity, which is typically quantified by the percentage of alignment scores. For example, if multiple sequences have a similarity score percentage of over 90% in a region, then this region can be considered very conservative, i.e., this region is a conservative interval.
[0121] In this embodiment of the disclosure, the conservation score list for each site represents the proportion of different base types at each site in multiple sequences, and the conservation score for each conservation interval represents the overall similarity between different sites in each conservation interval and different sequences.
[0122] In some exemplary embodiments, the reference data may include at least one of the following: TRDV reference gene sequence, TRDD reference gene sequence, and TRDJ reference gene sequence.
[0123] TRD gene rearrangement testing is an important medical diagnostic tool, primarily used to examine clonal lymphoma, which is particularly significant in the diagnosis of hematological diseases. Using TRD multiplex amplification technology to identify molecular subtypes of lymphoma offers the following advantages compared to other techniques:
[0124] (1) High specificity: TRD multiplex amplification can detect clonal immunoglobulin genes in lymphoma cells, thereby determining the molecular subtype of lymphoma, which has high specificity;
[0125] (2) High sensitivity: TRD multiplex amplification can detect very small amounts of lymphoma cells, even in low concentrations of mixed cell samples;
[0126] (3) Fast speed: Multiplex amplification using TRD can be performed quickly, usually yielding results within a few hours, which helps to determine the molecular subtype of lymphoma as early as possible;
[0127] (4) High reliability: The developed TRD multiplex amplification has high accuracy and reliability, and can provide reliable diagnostic and treatment guidance for clinicians.
[0128] For example, when the reference data includes the TRDV reference gene sequence and the TRDJ reference gene sequence, the final primer combination is a TRDV-J region multi-specific primer, which can be used to detect the VJ gene rearrangement of the TRD gene; when the reference data includes the TRDD reference gene sequence and the TRDJ reference gene sequence, the final primer combination is a TRD DJ region multi-specific primer, which can be used to detect the DJ gene rearrangement of the TRD gene; when the reference data includes the TRDV reference gene sequence and the TRDD reference gene sequence, the final primer combination is a TRD VD region multi-specific primer, which can be used to detect the VD gene rearrangement of the TRD gene; and when the reference data includes the TRDD reference gene sequence, the final primer combination is a TRD D region multi-specific primer, which can be used to detect the DD gene rearrangement of the TRD gene.
[0129] For example, when the reference data includes the TRDV reference gene sequence, the TRDD reference gene sequence, and the TRDJ reference gene sequence, the final primer combination designed is a TRD VJ region, VD region, DD region, and DJ multispecific primer, which can be used to detect VJ gene rearrangement, VD gene rearrangement, DD gene rearrangement, and DJ gene rearrangement of the IGK gene.
[0130] Specifically, the primer design method of this disclosure can design a set of highly efficient TRD primer combinations, designing multiplex PCR primers targeting the TRD V region, TRDD region, and TRD J region, covering all TRD gene rearrangement clonal types, and simultaneously identifying VJ, DJ, DD, and VD rearrangement types. By using the primer combinations designed in this disclosure, TRD sequences (including amplification of TRD VJ, DJ, DD, and VD gene rearrangements) can be amplified multiple times, which can be used for lymphoma diagnosis and typing.
[0131] In this embodiment of the disclosure, the TRDV reference gene sequence, TRDD reference gene sequence, and TRDJ reference gene sequence can be obtained by downloading the TRD V, D, and J sequences from Gemerline in the IGMT database.
[0132] In some exemplary embodiments, the conservation score list for each site includes five base types A, T, G, C, and N, as well as the percentage of each base type in multiple reference gene sequences.
[0133] Taking reference data including TRDV and TRDJ reference gene sequences as an example, as shown in Figure 2, we first need to determine the conservation score list for each site across the entire TRDV and TRDJ intervals. First, for the TRDV interval, we align multiple TRDV reference gene sequences by site, assuming P... i For the position i, there are five base types in multiple TRDV reference gene sequences: A, T, G, C, and N, where N represents an unknown base type. The five base types at each position are sorted from highest to lowest percentage. Therefore, the conservation score list can be represented as a set, where each element contains a base type and its percentage in the sequence. For example, the conservation score list for a certain position {'A': 0.25, 'T': 0.20, 'G': 0.18, 'C': 0.15, 'N': 0.12} represents the percentages of the five base types A, T, G, C, and N at the corresponding positions in multiple reference gene sequences, which are 0.25, 0.20, 0.18, 0.15, and 0.12, respectively. After this calculation, the conservation score list for each position in the TRDV interval is obtained.
[0134] Similarly, a list of conservation scores for each site in the TRDJ interval can be calculated.
[0135] In some exemplary embodiments, the conservation scores of multiple conservation intervals are obtained based on a list of conservation scores for each site, including:
[0136] Based on the pre-set initial conservative interval [start, end] and the sliding window step size W, multiple conservative intervals are obtained through the sliding window method;
[0137] The conservatism score for each conservatism interval is calculated using the following formula: W i =log1 / R i C i =W i ×R i , Among them, W i R represents the weight of the i-th site in each conservatism interval. i C represents the percentage of the most prevalent base type at the i-th site in each conserved interval. i Let S be the conservatism score of the i-th position in each conservatism interval, where S is the conservatism score of each conservatism interval, and n is the length of each conservatism interval, and n = end - start + 1.
[0138] In this embodiment of the disclosure, the conservative range can be obtained by sliding window method or not, and this disclosure does not limit it.
[0139] In this embodiment of the disclosure, the length n of each conservatism interval can be the primer length defined experimentally.
[0140] In this embodiment of the disclosure, the weight W of each site is first calculated based on the proportion of the highest-proportion base type at each site in each conservative interval. Then, the conservative score of each site is calculated based on the weight of each site and the proportion of the highest-proportion base type at each site. Finally, the sum of the conservative scores of each site in the entire interval is taken as the conservative score of the entire conservative interval.
[0141] After obtaining the conservatism scores of all conservatism intervals, all conservatism intervals can be sorted from high to low according to their conservatism scores. The top K conservatism intervals with the highest conservatism scores are selected, where K is a natural number greater than or equal to 1. A primer combination is generated for each of these K conservatism intervals, that is, K primer combinations are generated.
[0142] In some exemplary embodiments, generating K primer combinations for K conserved regions includes:
[0143] For each of the K conservative intervals, perform the following operation:
[0144] Determine the possible base types at each site in the conservatism interval, wherein the possible base types at each site are base types whose proportion is greater than or equal to a preset proportion threshold;
[0145] A primer set is generated based on the possible base types at each site within the conserved region. The number of primers in the primer set is m, where m = ∏m. i m i denoted as the number of possible base types at the i-th site in the conservative interval, ∏ as the quadrature symbol, i being between 1 and n, and n being the length of the conservative interval.
[0146] In this embodiment of the disclosure, for each site in each conservatism interval, an indicator function f can be defined. i (j), f i (j) indicates whether the proportion of the j-th base type at the i-th site is greater than or equal to a preset proportion threshold, where j is between 1 and 4. Since N bases generally have a low proportion, they are not considered here. If the proportion of the j-th base type at the i-th site is greater than the preset proportion threshold θ, then f i (j) = 1; otherwise f i (j) = 0. When generating primer combinations, it is necessary to consider all f values at each site. i For base types where (j) = 1, all f at each site... i By arranging and combining the base types (j) = 1, we can obtain all possible primer sequences for each conserved region, that is, generate a set of primer combinations for each conserved region.
[0147] For example, suppose that the first position of a certain conserved region contains three base types, such as ['A', 'T', 'G'], the second position contains two base types, such as ['G', 'C'], the third position contains only one base type ['T'], and so on. Then the primer combinations generated for this conserved region are:
[0148] [['ACT…'],
[0149] ['AGT…']
[0150] ['TCT…']
[0151] ['TGT…'],
[0152] ['GCT…']
[0153] ['GGT…']
[0154] ...
[0155] ]
[0156] In some exemplary embodiments, primers in the K primer combinations are screened based on at least one of the following: dimer, hairpin structure, annealing temperature, and GC content.
[0157] In this embodiment, primers in the K-group primer combination can be screened based on factors such as dimer composition, hairpin structure, Tm temperature (annealing temperature), and GC content. However, this disclosure does not limit this, and users can also screen primers in the K-group primer combination based on other conditions.
[0158] The selection criteria for primers based on dimer formation are as follows: at the experimental temperature Tt, primers should avoid dimer formation as much as possible, where T is the preset first experimental temperature and t is the preset first temperature difference. For example, assuming the preset first experimental temperature T is 60℃ and the preset first temperature difference t is 15℃, if the primer does not form a dimer at an experimental temperature of 45℃, then the primer is retained; if the primer forms a dimer at an experimental temperature of 45℃, then the primer is deleted. Dimers are polymers formed by the combination of complementary bases on two primers during a PCR reaction. The presence of dimers is equivalent to a reduction in the amount of raw material chains that could be used for amplification, thus reducing amplification efficiency. Therefore, it is best to avoid the formation of such substances.
[0159] The selection criteria for primers based on hairpin structure are as follows: at the experimental temperature Tt, the primers should avoid forming hairpin structures as much as possible, where T is the preset first experimental temperature and t is the preset first temperature difference. For example, assuming the preset first experimental temperature T is 60℃ and the preset first temperature difference t is 15℃, if the primer does not form a hairpin structure at the experimental temperature of 45℃, then the primer is retained; if the primer forms a hairpin structure at the experimental temperature of 45℃, then the primer is deleted.
[0160] The selection criteria for primers based on Tm temperature are as follows: the annealing temperature of the primers should be within the preset experimental temperature range. For example, suppose the preset experimental temperature range is [Tm]. low T high ], where T low The lowest temperature, T high The highest temperature is [T]. If the primer annealing temperature is [T], low T high If the primer is within the range of [T], then retain the primer; if the primer's annealing temperature is not within [T], then retain the primer. low T high If the primer is within the specified range, then delete it. For example, [T] low T high The temperature can be [50℃, 60℃], however, this disclosure does not limit it.
[0161] The selection criteria for primers based on GC content are as follows: the GC content of the primers should be within a preset GC content range. For example, assuming the preset GC content range is [G... low G high ], where G low For the lowest GC content, Ghigh The highest GC content is achieved if the primer's GC content is within [G]. low G high If the GC content of the primer is within the range of [G], then retain the primer; if the GC content of the primer is not within the range of [G], then retain the primer. low G high If the primer is within the specified range, then delete it. For example, [G] low G high The percentage can be [40%, 60%], however, this disclosure does not limit it.
[0162] In some exemplary embodiments, the K primer combinations are evaluated based on at least one of the following: dimer, hairpin structure, amplification coverage, nonspecific amplification rate (or specificity).
[0163] In this embodiment of the disclosure, when screening primers in a set of primer combinations, one or more redundant primer sequences in the set of primer combinations will be deleted; and when evaluating K sets of primer combinations, one or more sets of primer combinations with poor evaluation results will be deleted, or in other words, the best or better set of primer combinations will be selected from multiple sets of primer combinations.
[0164] Complementarity, dimers, or hairpin structures at the 3' ends of primers can all lead to PCR reaction failure. Therefore, when evaluating a primer combination, if a dimer or hairpin structure is formed in the combination, the combination should be deleted.
[0165] Amplification coverage and specificity are the two most important metrics for evaluating primer effectiveness. Amplification coverage refers to the proportion of target sequences captured by the target primers in an existing database. Specificity refers to the proportion of amplified sequences targeted by a primer combination, i.e., the ratio of specifically amplified sequences to the total sequence.
[0166] For example, for TRD gene rearrangement detection, the final TRD-specific primer combinations are shown in Table 1 (SEQ ID NO: 1-15) by screening primers in the K-group primer combination and evaluating the screened K-group primer combinations.
[0167] Table 1
[0168] Secondary structure and hairpin structure analysis showed that this TRD-specific primer combination did not generate secondary structures or hairpin structures under experimental temperatures greater than or equal to 45 degrees Celsius. Amplification simulation analysis confirmed that this TRD-specific primer combination could amplify all functional TRDV, TRDD, and TRDJ genes, with an amplification coverage of 100%.
[0169] The following analysis uses the TRD rearrangement detection of the REH cell line as an example to examine the specificity of primers and / or kits. REH cells are precursor B cells isolated from the peripheral blood of patients with acute lymphoblastic leukemia (ALL). They have a lymphoblast-like morphology and do not belong to either the B cell or T cell type.
[0170] The PCR reaction system was configured as follows: Thermo AmpliTaq Gold was selected as the PCR amplification reagent. TM 360 DNA polymerase or equivalent amplification reagents are used. The primer components and ratios in the primer combination are shown in Table 2. This primer combination and ratio can reduce PCR amplification bias and thus improve detection accuracy. PCR primers are added to DNase-Free & RNase-Free water according to the synthesis report in Table 2 (SEQ ID NO:18-32) to a primer concentration of 100 μM. The TRD primer pool is formed by mixing TRD V, TRDD (5' intron region), TRD J, and TRDD (3' intron region) primers.
[0171] Table 2
[0172] The italicized, underlined regions are the linker sequences, while the regular, ununderlined regions are the primer sequences in Table 1. Specifically, TRD V1 to TRD V8 correspond to primers V1 to V8 in Table 1, TRDD 5' intron region 1 corresponds to primer δ2-1 in Table 1, TRD J1 to TRD J5 correspond to primers J1 to J5 in Table 1, and TRDD 3' intron region 1 corresponds to primer δ3-1 in Table 1.
[0173] REH uses the Meiji Bio Universal DNA Extraction Pre-packed Kit or equivalent kit to extract genomic DNA; the PCR amplification system configuration is shown in Table 3 (First Round PCR Amplification System Table) and Table 4 (Second Round PCR Amplification System Table), and the PCR amplification conditions are shown in Table 5 (First Round PCR Amplification Conditions Table) and Table 6 (Second Round PCR Amplification Conditions Table).
[0174] Table 3
[0175] In Table 3, AmpliTaq Gold 360 buffer is a buffer for PCR amplification, dNTPs mixture is a mixture, AmpliTaq Gold 360 DNA polymerase is a polymerase, DNase and RNase-free water is nuclease-free water (DNase-free and RNase-free water), X is the volume calculated from 100 ng of template, and T represents Total.
[0176] Table 4
[0177] In Table 4, VAHTS HiFi amplification mixture is a single mixture, P5 adapter primer is a P5 adapter primer, P7 adapter primer is a P7 adapter primer, the first round PCR product is the product purified after the first PCR, and T represents Total.
[0178] Table 5
[0179] Table 6
[0180] The testing steps are as follows:
[0181] 1) AmpliTaq Gold TM 360 DNA polymerase, buffer, MgCl2, dNTPs, polymerase, GC enhancer, REH DNA, TRD primer pool, DNase-Free & RNase-Free water were taken out and thawed on ice;
[0182] 2) Take a PCR reaction tube, add each component according to Table 3, and mix well by pipetting.
[0183] 3) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 5, and run the PCR instrument;
[0184] 4) After PCR amplification, the PCR product is purified using magnetic beads, with the ratio of magnetic beads being 0.8 times the amplification volume;
[0185] 5) Take a new PCR reaction tube, add the components according to Table 4, and mix well by pipetting.
[0186] 6) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 6, and run the PCR instrument;
[0187] 7) After PCR amplification, the PCR product is purified using magnetic beads at a ratio of 0.8 times the amplification volume. The purified product is the sequencing library.
[0188] 8) The library was subjected to high-throughput sequencing using an Illumina NovaSeq sequencer with a read length of PE150.
[0189] The TRD rearrangement sequence of the REH cell line was accurately identified using high-throughput sequencing libraries. The results of various indicators are shown in Table 7.
[0190] Table 7
[0191] The total number of reads refers to all sequences obtained from sequencing. The number of filtered reads is the number of sequences obtained after data preprocessing (filtering out low-quality sequences, adapter sequences, etc.). The effective read ratio represents the ratio of the number of filtered reads to the total number of reads. The target read number represents the number of target sequences in the filtered reads. The target read ratio is the ratio of the target reads to the number of filtered reads. The higher the target read ratio, the better the specificity. As shown in Table 7, the target read ratio of the REH cell line reached 88.76%, indicating that the primers and / or kits of this embodiment have good specificity.
[0192] The accuracy of primer and / or kit rearrangement detection will be analyzed using the TRD rearrangement detection of a peripheral blood mononuclear cell (PBMC) negative sample and a lymphoma positive sample as examples.
[0193] The PCR reaction system was configured as follows: Thermo AmpliTaq Gold was selected as the PCR amplification reagent. TM 360 DNA polymerase or equivalent amplification reagents are used. The primer components and ratios in the primer combination are shown in Table 2. This primer combination and ratio can reduce PCR amplification bias and thus improve detection accuracy. PCR primers are added to DNase-Free & RNase-Free water according to the synthesis report in Table 2, with a primer concentration of 100 μM. The TRD primer pool is formed by mixing TRD V, TRDD (5' intron region), TRD J, and TRDD (3' intron region) primers.
[0194] Lymphoma-positive and PBMC-negative samples were selected as PCR templates. Genomic DNA was extracted using the Meiji Bio Universal DNA Extraction Pre-packed Kit or an equivalent kit. The PCR amplification system configuration is shown in Table 3 (First Round PCR Amplification System Table) and Table 4 (Second Round PCR Amplification System Table). The PCR amplification conditions are shown in Table 5 (First Round PCR Amplification Conditions Table) and Table 6 (Second Round PCR Amplification Conditions Table).
[0195] The testing steps are as follows:
[0196] 1) AmpliTaq Gold TM 360 DNA polymerase, buffer, MgCl2, dNTPs, polymerase, GC enhancer, gDNA, TRD primer pool, DNase-Free & RNase-Free water were taken out and thawed on ice;
[0197] 2) Take a PCR reaction tube, add each component according to Table 3, and mix well by pipetting.
[0198] 3) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 5, and run the PCR instrument;
[0199] 4) After PCR amplification, the PCR product is purified using magnetic beads, with the ratio of magnetic beads being 0.8 times the amplification volume;
[0200] 5) Take a new PCR reaction tube, add the components according to Table 4, and mix well by pipetting.
[0201] 6) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 6, and run the PCR instrument;
[0202] 7) After PCR amplification, the PCR product is purified using magnetic beads at a ratio of 0.8 times the amplification volume. The purified product is the sequencing library.
[0203] 8) The library was subjected to high-throughput sequencing using an Illumina NovaSeq sequencer with a read length of PE150.
[0204] Figure 3 shows the TRD detection results for PBMC-negative samples, and Figure 4 shows the TRD detection results for lymphoma-positive samples. In Figures 3 and 4, the horizontal axis represents the CDR3 sequence length (i.e., the rearrangement position), and the vertical axis represents the proportion of detected plasmid sequences. The same vertical bar includes the proportion of plasmid sequences with the same CDR3 sequence length but different CDR3 sequences (multiple plasmid sequences with different CDR3 sequences on the same vertical bar are separated by horizontal lines). The proportion of plasmid sequences with a certain CDR3 sequence length is the ratio of the number of plasmid sequences with that CDR3 sequence length detected to the total number of sequences detected. As can be seen from Figures 3 and 4, the proportion of sequences detected in PBMC-negative samples with different CDR3 sequence lengths is relatively close, showing polyclonal TRD gene rearrangement. In contrast, lymphoma-positive samples show specific TRD rearrangement sequence proliferation at CDR3 sequence length 51, showing monoclonal TRD. This indicates that the rearrangement detection accuracy of this TRD multiple primer combination and / or kit is high.
[0205] As shown in Figure 5, this disclosure also provides a standard quality grain design method, including:
[0206] Step 501: Obtain the first gene cluster, the second gene cluster, and the third gene cluster. The first gene cluster includes multiple first gene sequence fragments, the second gene cluster includes multiple second gene sequence fragments, and the third gene cluster includes multiple third gene sequence fragments. The first gene cluster, the second gene cluster, and the third gene cluster all contain multiple functional fragments.
[0207] Step 502: Combine multiple first gene sequence fragments, multiple second gene sequence fragments, and multiple third gene sequence fragments to obtain multiple fragment groups. Each fragment group contains one first gene sequence fragment, one second gene sequence fragment, and one third gene sequence fragment.
[0208] Step 503: Insert a non-human sequence between the second and third gene sequence fragments in each fragment group to obtain multiple plasmid sequences;
[0209] Step 504: Determine the ratio of each plasmid sequence to obtain the designed standard plasmid.
[0210] In conventional analytical methods, plasmid sequences do not carry UMI sequence identifiers. Plasmid sequences are identified through sequence alignment. However, this analytical method has high identification rate in samples containing only plasmid sequences. But when peripheral blood samples are mixed, only most plasmid sequences can be identified. It is difficult to identify whether some sequences are from peripheral blood samples or plasmid sequences.
[0211] The standard plasmid design method of this disclosure obtains multiple plasmid sequences by inserting a non-human sequence between the second and third gene sequence fragments of each fragment group. All plasmid sequences can be distinguished from rearranged sequences derived from peripheral blood samples by the non-human sequence, which means that plasmid sequences can be accurately identified and quantified.
[0212] In this embodiment of the disclosure, the length of the inserted non-human sequence is n3 bp, where n3 is between 6 and 10. For example, n3 = 8.
[0213] In this embodiment of the disclosure, the n2 bp of the inserted non-human sequence and the third gene sequence fragment form a (n2+n3)bp unique molecular marker (UMI) sequence identifier, and the UMI sequence identifiers in different plasmid sequences have at least one site with different base types.
[0214] UMI sequence identifiers, as unique identifiers for sequences, are used for sequence identification and classification in subsequent analyses. Specifically, different UMI sequence identifiers distinguish DNA templates from different sources, differentiating between false-positive mutations caused by random errors during PCR amplification and sequencing, and mutations truly carried by the patient, thereby improving the sensitivity and specificity of the detection.
[0215] In this embodiment of the disclosure, n2 is between 2 and 6. For example, n2 = 4.
[0216] In some exemplary embodiments, the first gene cluster may be the TRDV gene cluster, the second gene cluster may be the TRDD gene cluster, and the third gene cluster may be the TRDJ gene cluster.
[0217] For example, the first gene cluster can be a TRDV gene cluster containing all functional fragments, the second gene cluster can be a TRDD gene cluster containing all functional fragments, and the third gene cluster can be a TRDJ gene cluster containing all functional fragments. The TRD gene sequence can be sourced from the IMGT database. All TRDV, TRDD, and TRDJ fragments are randomly combined, and an 8 bp non-human sequence is added to each combination to obtain multiple plasmid sequences. In each plasmid sequence, the 8 bp non-human sequence and the 4 bp TRDJ front end sequence together form a 12 bp UMI sequence. The generated UMI sequence combinations are shown in Table 8 (SEQ ID NO: 33-44).
[0218] Table 8
[0219] Table 8 shows that the non-human random sequence inserted into each UMI sequence is just an example. Users can redesign the plasmid sequence and the inserted non-human random sequence according to their needs, as long as the base types of at least one site are different in different UMI sequence identifiers.
[0220] In some exemplary embodiments, when the standard plasmid is a uniformity standard, the ratio of each plasmid sequence is a uniform ratio of equal concentration.
[0221] In some exemplary embodiments, when the standard plasmid is an experimental standard, the proportion of one or more plasmid sequences is greater than the proportion of the remaining plasmid sequences (plasmid sequences other than one or more plasmid sequences).
[0222] In this embodiment of the disclosure, the uniformity standard is a standard prepared by mixing each plasmid in the same proportion. The experimental standard is a standard prepared by mixing one or more plasmids in a certain high proportion. In some exemplary embodiments, a control standard may also be provided, which is a standard in which no plasmid is added to the sample and pure water is used instead.
[0223] Homogeneity standards can be used to verify the amplification efficiency of primer combinations in a single experiment; control standards can be used to verify whether there is contamination in a single experiment; experimental standards are used to simulate polyclonal or monoclonal experiments; a certain proportion of homogeneity standards can be added to quantify unknown experimental samples to detect the rearrangement type and quantification of the sample itself; homogeneity standards or experimental standards can be used to verify the influence of different experimental reagents and conditions on experimental results.
[0224] The following example uses a PCR amplification uniformity experiment using uniformity standards to verify the amplification efficiency of TRD multiplex primer combinations and / or kits.
[0225] The PCR reaction system was configured as follows: Thermo AmpliTaq Gold was selected as the PCR amplification reagent. TM 360 DNA polymerase or equivalent amplification reagents; PCR primers were prepared by adding DNase-Free and RNase-Free water according to the synthesis report in Table 2, with a primer concentration of 100 μM. The TRD primer pool was formed by mixing TRD V, TRDD (5' intron region), TRD J, and TRDD (3' intron region) primers; 12 TRDV-DJ standard plasmid sequences were designed, each plasmid sequence including a fragment of TRD V / D / J sequences randomly composed of a UMI sequence identifier composed of 8 bases, as shown in Table 8 above.
[0226] The test sample was a homogeneity standard, with a total volume of approximately 10,000 copies. The plasmid concentration was quantified and its molar concentration calculated using Qubit 4.0. The plasmids were mixed equimolarly, amplified using M13F and M13R primers (universal primers), and a sequencing library was constructed. The number and proportion of each plasmid were counted using UMI (Uniqueness Index), and the molar ratio of each plasmid was adjusted to 0.95-1.05, which constituted the homogeneity standard.
[0227] The configuration of the PCR amplification system is shown in Table 3 (first round PCR amplification system) and Table 4 (second round PCR amplification system) above, and the PCR amplification conditions are shown in Table 5 (first round PCR amplification conditions) and Table 6 (second round PCR amplification conditions) above.
[0228] The testing steps are as follows:
[0229] 1) AmpliTaq Gold TM 360 DNA polymerase, buffer, MgCl2, dNTPs, polymerase, GC enhancer, homogeneity standard, TRD primer pool, DNase-Free & RNase-Free water were taken out and dissolved on ice;
[0230] 2) Take a PCR reaction tube, add each component according to Table 3, and mix well by pipetting.
[0231] 3) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 5, and run the PCR instrument;
[0232] 4) After PCR amplification, the PCR product is purified using magnetic beads, with the ratio of magnetic beads being 0.8 times the amplification volume;
[0233] 5) Take a new PCR reaction tube, add the components according to Table 4, and mix well by pipetting.
[0234] 6) Place the PCR reaction tubes in the PCR instrument, set the PCR amplification program according to Table 6, and run the PCR instrument;
[0235] 7) After PCR amplification, the PCR product is purified using magnetic beads at a ratio of 0.8 times the amplification volume. The purified product is the sequencing library.
[0236] 8) Perform high-throughput sequencing on the sequencing library using an Illumina NovaSeq sequencer with a read length of PE150.
[0237] Figure 6 is a detection result diagram of the proportion of TRD plasmids obtained by using uniformity standards and TRD multiplex primers according to an exemplary embodiment of the present disclosure. In Figure 6, the horizontal axis represents the plasmid sequence number, and the vertical axis represents the proportion of detected plasmid sequences. The proportion of plasmid sequences with a certain number is the ratio of the number of detected plasmid sequences with that number to the total number of detected sequences. As can be seen from Figure 6, the PCR amplification uniformity of the TRD multiplex primer combination and / or kit is good.
[0238] As shown in Figure 7, this disclosure also provides a method for quantitative analysis of standard quality grains, including:
[0239] Step 701: Obtain the paired-end sequencing data corresponding to the standard quality plasmid, which includes the Reads1 sequence and the Reads2 sequence;
[0240] Step 702: Determine whether the Reads1 and Reads2 sequences include a UMI sequence identifier;
[0241] Step 703: When both Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequences corresponding to Reads1 and Reads2 sequences have been identified.
[0242] Step 704: When only one of the Reads1 and Reads2 sequences contains a UMI sequence identifier, determine the Hamming distance between the Reads sequence that does not contain a UMI sequence identifier and the UMI sequence identifier; when the determined Hamming distance is less than or equal to a, determine that the plasmid sequence corresponding to the Reads1 and Reads2 sequences has been identified, where a is a natural number less than or equal to 2.
[0243] The standard plasmid quantitative analysis method disclosed herein determines whether Reads1 and Reads2 sequences include a UMI sequence identifier. When both Reads1 and Reads2 sequences include the same UMI sequence identifier, the plasmid sequence corresponding to Reads1 and Reads2 sequences is identified. When only one of Reads1 and Reads2 sequences includes the UMI sequence identifier, the Hamming distance between the Reads sequence without the UMI sequence identifier and the UMI sequence identifier is determined. When the determined Hamming distance is less than or equal to 'a', the plasmid sequence corresponding to Reads1 and Reads2 sequences is identified, where 'a' is a natural number less than or equal to 2. This method can accurately perform quantitative analysis of standard plasmids, and further, the amplification results of samples can be accurately inferred based on the quantitative analysis results of standard plasmids.
[0244] In this embodiment of the disclosure, when neither the Reads1 sequence nor the Reads2 sequence contains a UMI sequence identifier, or when only one of the Reads1 sequence and the Reads2 sequence contains a UMI sequence identifier and the Hamming distance between the Reads sequence containing a UMI sequence identifier and the UMI sequence identifier is greater than a, it is determined that no plasmid sequence corresponding to the Reads1 sequence and the Reads2 sequence has been identified.
[0245] In this embodiment of the disclosure, a can be equal to 1. When a is equal to 1, the identification method is more rigorous, which allows for more accurate quantitative analysis of the standard quality particles.
[0246] Paired-end sequencing performs sequencing from both ends of the insert fragment. The ATCG sequence read from each end is called a read. Each insert fragment will generate two reads, namely reads1 and reads2. The reads1 and reads2 data corresponding to a sample are stored in two compressed packages.
[0247] In this embodiment of the disclosure, when the sequencing data corresponding to the standard plasmid is single-end sequencing data, it is determined whether the Reads sequence includes a UMI sequence identifier; when the Reads sequence includes a UMI sequence identifier, it is determined that the plasmid sequence corresponding to the Reads sequence has been identified; when the Reads sequence does not include a UMI sequence identifier, it is determined that the plasmid sequence corresponding to the Reads sequence has not been identified.
[0248] As shown in Figure 8, this disclosure also provides a primer characterization method, including:
[0249] Step 801: Obtain the sequencing data corresponding to the primer to be tested, and determine the 5' end position of the primer to be tested based on the sequencing data;
[0250] Step 802: Design multiple sequences based on the 5' end position of the primer to be detected. Each of the multiple sequences starts at the 5' end position of the primer to be detected, and the 3' end cutoff positions of two adjacent sequences differ by a single base.
[0251] Step 803: Design multiple nucleic acid structures based on the designed sequences;
[0252] Step 804: Identify the nucleic acid structures among multiple nucleic acid structures that can complementary pair with the primers to be detected;
[0253] Step 805: Determine the sequence of the primer to be tested based on the nucleic acid structure that can complement the primer to be tested.
[0254] The primer characterization method of this disclosure determines the 5' end position of the primer to be detected based on sequencing data, designs multiple sequences and multiple corresponding nucleic acid structures based on the 5' end position of the primer to be detected, then determines the nucleic acid structure that can complementarily pair with the primer to be detected among the multiple nucleic acid structures, and determines the sequence of the primer to be detected based on the nucleic acid structure that can complementarily pair with the primer to be detected, thus enabling the sequencing of any unknown primer.
[0255] In some exemplary embodiments, the method further includes, prior to: performing PCR amplification on the primers to be detected, and obtaining sequencing data and the corresponding 5' end position of the primers to be detected by high-throughput sequencing.
[0256] In some exemplary implementations, the length of each sequence is between 18 bp and 27 bp.
[0257] In some exemplary embodiments, the number of sequences is 10.
[0258] Since primers are typically between 18 and 27 bases in length, a set of 10 sequences is designed based on the 5' end position of the primer to be detected, with each sequence differing by a single base at the 3' end, and each sequence being between 18 and 27 bp in length.
[0259] In some exemplary embodiments, each nucleic acid structure includes a hairpin structure, each nucleic acid structure has an inverse complementary sequence fused to its end, and each nucleic acid structure has a fluorescently labeled end base.
[0260] As shown in Figure 9, a set (10) of nucleic acid structures for detection primers were designed. Each nucleic acid structure includes an artificially designed hairpin structure, with reverse complementary sequences of different lengths fused to the ends, and fluorescent labels modified with bases at the ends.
[0261] In some exemplary embodiments, identifying nucleic acid structures among multiple nucleic acid structures that can complementaryly pair with the primer to be detected includes:
[0262] For each nucleic acid structure, the following steps were performed: the primer to be tested was mixed with the nucleic acid structure, and denaturation, annealing, and ligation were performed to obtain the ligation product; the length of the ligation product was then detected.
[0263] Select the nucleic acid structure corresponding to the ligation product with a length greater than the preset length threshold as a nucleic acid structure that can complementarily pair with the primer to be detected.
[0264] For example, mix the primers to be detected with nucleic acid structures at equimolar concentrations according to Table 9:
[0265] Table 9
[0266] After mixing, place the mixture in a boiling water bath for 5 minutes, turn off the heating switch, and let it stand to room temperature. This step usually takes 8 to 12 hours.
[0267] Configure the connection system according to Table 10:
[0268] Table 10
[0269] Mix all components in the connection system thoroughly, centrifuge the liquid to the bottom of the tube, and react at 25°C for 30 minutes.
[0270] As shown in Figure 10, the primers to be tested anneal to the paired nucleic acid structures to form double-stranded structures, which are then ligated by T4 DNA ligase. Structures that cannot be paired cannot be ligated, resulting in nucleic acid structures and ligation products of different lengths.
[0271] In this embodiment of the disclosure, the ligation product can be detected using Sanger fragment analysis. As shown in Figure 11, the T4 DNA ligase ligates the primer to be tested to an artificially designed nucleic acid structure, which is then denatured into a single strand by formamide. The corresponding ligation product is the longest, while the unligated nucleic acid structure and the free primer are shorter. Sanger fragment analysis can determine the 3' cutoff position of the primer to be tested, thereby obtaining the full-length sequence of the primer to be tested.
[0272] As shown in Figure 12, this disclosure also provides a method for constructing a sequencing library, including:
[0273] Step 1201: Extract DNA from the genome to be tested;
[0274] Step 1202: Design specific primers for the genome to be detected, mix the specific primers with the extracted DNA, perform the first round of PCR amplification to obtain the first round of PCR products, and purify the first round of PCR products;
[0275] Step 1203: Mix the purified first-round PCR product with universal adapter primers, perform a second-round PCR amplification to obtain the second-round PCR product, purify the second-round PCR product, and obtain the constructed sequencing library.
[0276] The sequencing library construction method of this disclosure uses a two-step PCR method to ligate sequencing adapters to the 5' and 3' ends of the target fragment, respectively, to convert the target DNA to be sequenced into a form that can be read by the sequencer.
[0277] As shown in Figure 13, the universal adapter primers include the P5 adapter primer and the P7 adapter primer.
[0278] For example, the genome to be detected is the TRD genome, and the specific primers include a forward primer targeting the TRD V / D 5' intron region and a reverse primer targeting the TRD J / D 3' intron region.
[0279] The design methods for specific primers can refer to the primer design methods described above, and will not be repeated here.
[0280] As shown in Figure 14, this embodiment of the present disclosure also provides a disease classification and diagnosis method, including:
[0281] Step 1401: Obtain the sequencing data of the sample to be tested, and perform data preprocessing on the obtained sequencing data;
[0282] Step 1402: Compare the preprocessed data with the reference gene sequences of multiple preset clone species to obtain the sequence proportion corresponding to each clone species;
[0283] Step 1403: Sort the proportions of multiple sequences from largest to smallest, and mark the proportions of multiple sequences in the sorting order as the proportion of the first sequence to the proportion of the M1th sequence, where M1 is the number of clone species;
[0284] Step 1404: Detect whether the differences between the proportions of the first and third sequences, and between the proportions of the second and third sequences, exceed preset difference thresholds; when both the differences between the proportions of the first and third sequences and between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample; when the differences between the proportions of the first and third sequences exceed preset difference thresholds but the differences between the proportions of the second and third sequences do not exceed preset difference thresholds, the sample to be tested is determined to be a monoclonal sample; when neither the differences between the proportions of the first and third sequences nor the differences between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.
[0285] In this embodiment of the disclosure, the proportions of multiple sequences can be regarded as multiple clonal distribution peaks. When the differences between the proportions of the first sequence and the third sequence, as well as the differences between the proportions of the second sequence and the third sequence, both exceed a preset difference threshold, it can be considered that there are two main peaks among the multiple clonal distribution peaks, and the sample to be detected is determined to be an oligoclonal sample. When the difference between the proportions of the first sequence and the third sequence exceeds a preset difference threshold, but the difference between the proportions of the second sequence and the third sequence does not exceed a preset difference threshold, it can be considered that there is only one main peak among the multiple clonal distribution peaks, and the sample to be detected is determined to be a monoclonal sample. When the differences between the proportions of the first sequence and the third sequence, as well as the differences between the proportions of the second sequence and the third sequence, do not exceed a preset difference threshold, it can be considered that there is no main peak among the multiple clonal distribution peaks, and the sample to be detected is determined to be a polyclonal sample.
[0286] In this embodiment of the disclosure, when the sample to be tested is detected as an oligoclonal sample or a monoclonal sample, the diagnostic result of the sample to be tested can be considered as positive, and the corresponding disease subtype can be determined based on the number of clonal distribution peaks and the type of clones; when the sample to be tested is determined to be a polyclonal sample, the diagnostic result of the sample to be tested can be considered as negative.
[0287] The disease typing and diagnosis method of this disclosure identifies each TRD clone sequence and determines whether the sample to be tested is a monoclonal sample, oligoclonal sample, or polyclonal sample based on whether the difference between the proportions of the top three sequences with the largest proportion exceeds a preset difference threshold. It can determine the specific sequence of each clone, thereby accurately typing and diagnosing the sample to be tested, and thus realizing the tracking and monitoring of residual small lesions in lymphoma.
[0288] In some exemplary embodiments, the acquired sequencing data undergoes data preprocessing, including:
[0289] Filter out primer sequences, low-quality sequences, sequences with a high proportion of N bases, and sequences whose length is lower than a preset length threshold from the sequencing data;
[0290] The sequencing data were subjected to adapter sequence detection and adapter sequence removal.
[0291] Remove the low-quality bases at both ends of the sequence.
[0292] In this embodiment of the disclosure, during data preprocessing, the sequencing data is first statistically analyzed, including the amount of sequencing data and its quality value. Then, the sequencing data is filtered, including filtering out sequences containing adapter sequences and with a 5' end of polyN, sequences with an average quality value lower than a preset quality threshold Q, sequences with an excessively high proportion of N bases, and sequences with a length lower than a preset length threshold. For example, the preset quality threshold Q can be 25. When the quality value of a base is greater than or equal to 25, it can be considered reliable sequencing, with a corresponding error rate of approximately 0.3%. When the proportion of N bases in a sequence is higher than a preset proportion threshold M2 (for example, M2 can be set to 10%), it is considered to have a high proportion of N bases. When the average quality value of a sequence is lower than the preset quality threshold Q or the proportion of N bases in a sequence is higher than the preset proportion threshold M, the sequence is filtered out. Furthermore, when the sequence length is lower than a preset length threshold L, the sequence is filtered out.
[0293] Data preprocessing also includes the removal of adapter sequences and low-quality bases: adapter sequences are detected and removed from the sequence, using general-purpose software such as trim_galore or cutadapter. When the base quality values at both ends of the sequence are lower than a preset quality threshold Q, those bases are removed.
[0294] By preprocessing data, we can obtain high-quality data for subsequent analysis.
[0295] In some exemplary embodiments, the reference gene sequences for multiple clone species are: TRDV reference gene sequence, TRDD reference gene sequence, and TRDJ reference gene sequence.
[0296] Assume the proportion of the i-th sequence is A i In some exemplary embodiments, detecting whether the difference between the proportion of the first sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio A1 / A3 of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X; when A1 / A3≥X, it is determined that the difference between A1 and A3 exceeds the preset difference threshold; when the ratio A1 / A3 of the proportion of the first sequence to the proportion of the third sequence is <X, it is determined that the difference between A1 and A3 does not exceed the preset difference threshold, where X is the preset difference threshold and X is greater than or equal to 2.
[0297] Similarly, detecting whether the difference between the proportion of the second sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the second sequence to the proportion of the third sequence, A2 / A3, is greater than or equal to X. When A2 / A3≥X, it is determined that the difference between A2 and A3 exceeds the preset difference threshold; when the ratio of the proportion of the second sequence to the proportion of the third sequence, A2 / A3<X, it is determined that the difference between A2 and A3 does not exceed the preset difference threshold, where X is the preset difference threshold, and X is greater than or equal to 2.
[0298] In this embodiment of the disclosure, the preprocessed sequence is subjected to TRD rearrangement identification. To determine whether the sequence is a VDJ rearrangement, the sequence is aligned to the IMGT database to identify the V, D, and J fragments contained within the sequence. When the sequence contains the V and J sequences of TRD, the sequencing data is considered to contain the VJ gene rearrangement of TRD; when the sequence contains the D 5' intron region sequence and the J sequence of TRD, the sequencing data is considered to contain the DJ gene rearrangement of TRD; when the sequence contains the V and D 3' intron region sequences of TRD, the sequencing data is considered to contain the VD gene rearrangement of TRD; when the sequence contains both the D 5' and D 3' intron region sequences of TRD, the sequencing data is considered to contain the DD gene rearrangement of TRD. If no TRD sequence is identified in the sequence, it is considered a non-specific amplification sequence. The number of detected clone sequences corresponding to each rearrangement type is determined.
[0299] The obtained clone sequences are used to identify the clonal distribution of the sample to be tested, mainly including polyclonal, oligoclonal and monoclonal sequences.
[0300] Monoclonal sample: A single main peak is detected in the amplification product (taking X=2 as an example, the ratio of the height of the first highest clonal distribution peak to the height of the third highest clonal distribution peak is greater than or equal to twice, and the ratio of the height of the second highest clonal distribution peak to the height of the third highest clonal distribution peak is less than twice).
[0301] Oligoclonal sample: Two main peaks were detected in the amplification product (taking X=2 as an example, the ratio of the height of the first highest clonal distribution peak to the height of the third highest clonal distribution peak and the ratio of the height of the second highest clonal distribution peak to the height of the third highest clonal distribution peak are both greater than or equal to two).
[0302] Multiple clone samples: Multiple gene rearrangement distribution peaks were detected in the amplification products, with no main peak (taking X=2 as an example, the ratio of the height of the first highest clone distribution peak to the height of the third highest clone distribution peak is less than two times).
[0303] Assuming the sequencing data contains M1 clone species, let A represent the sequence proportion of the i-th clone species. i M1 represents the clone species index, where i is the index of the clone species, and 1 ≤ i ≤ M1. For example, assuming the sequencing data contains VJ gene rearrangements of 4 TRDs, VD gene rearrangements of 5 TRDs, DJ gene rearrangements of 2 TRDs, and DD gene rearrangements of 1 TRD, then M1 = 12.
[0304] Sort all clone species in descending order of their sequence proportions to obtain an ordered sequence A1, A2, ..., A M1 Where A1≥A2...≥A M1 .
[0305] The method detects whether the differences between the proportions of the first sequence A1 and the third sequence A3, and between the proportions of the second sequence A2 and the third sequence A3, exceed preset difference thresholds. When both the differences between the proportions of the first sequence A1 and the third sequence A3, and between the proportions of the second sequence A2 and the third sequence A3, exceed the preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample. When the differences between the proportions of the first sequence A1 and the third sequence A3 exceed the preset difference thresholds, but the differences between the proportions of the second sequence A2 and the third sequence A3 do not exceed the preset difference thresholds, the sample to be tested is determined to be a monoclonal sample. When neither the differences between the proportions of the first sequence A1 and the third sequence A3, nor the differences between the proportions of the second sequence A2 and the third sequence A3, exceed the preset difference thresholds, the sample to be tested is determined to be a polyclonal sample.
[0306] In this embodiment of the disclosure, when the sample to be tested is a polyclonal sample, the sample to be tested is determined to be a negative sample; when the sample to be tested is an oligoclonal sample or a monoclonal sample, the sample to be tested is determined to be a positive sample.
[0307] When performing multiplex primer amplification experiments, adjustments to reagents, primer combinations, and amplification temperatures may be necessary multiple times, requiring evaluation of the results of each experiment. The following indicators can be used to assess the amplification effectiveness of each experiment.
[0308] Amplification homogeneity: This assesses whether the sequence coverage is uniform due to amplification bias. Quantitative analysis is performed using known standard plasmids. Relative standard deviation (RSD) is used to quantify amplification homogeneity. Known homogeneity standards involve proportional input of all plasmids, but amplification results will still exhibit bias. RSD = (Standard Deviation / Mean) * 100. A smaller RSD value indicates higher amplification homogeneity, meaning the amplification results of multiple plasmids are more similar and the amplification efficiency is more uniform; conversely, a larger RSD value indicates less uniform amplification efficiency.
[0309] Amplification specificity: This is used to assess whether the amplification products specifically match the target sequence. It is evaluated by calculating the ratio of the number of amplification products in non-target regions to the total number of amplification products. The higher the specificity, the lower the proportion of non-target amplification.
[0310] Amplification coverage: assess the extent to which the standard plasmid is covered by the sequencing sequence to identify whether any plasmids have not been amplified or have been over-amplified.
[0311] The following experimental analysis example illustrates the disease classification and diagnostic method of this disclosure. In this example, TRD sample 1 is a negative standard, TRD sample 2 and TRD sample 3 are both positive standards, and TRD sample 4 is pure water (blank control experiment).
[0312] Specifically, the data analysis process includes the following steps:
[0313] 1) Perform data statistics on the sequencing data of the samples to be tested, including the amount of sequencing data and the Q30 ratio, as shown in Table 11.
[0314] Table 11
[0315] 2) Subsequent data preprocessing included sequencing data statistics and low-quality data filtering. Low-quality data included sequences containing only primer sequences with polyN at the end, sequences with low average base quality, sequences containing structural sequences, and sequences with excessively high N content. In this example, the filtering conditions were: Q value below 20 was considered low-quality bases; sequences with adapters longer than 3 bp at the end were removed; and sequences with a length less than 50 bp after removal were discarded. The filtering results are shown in Table 12.
[0316] Table 12
[0317] 3) Clone identification and authentication.
[0318] The clone rearrangements in the samples were identified and statistically analyzed, including the total number of clone sequences, clone types, and the percentage of clone sequences in the sequence (Ratio). The identification results are shown in Table 13, which show that the sequencing data of most samples contain a high proportion of clone sequences.
[0319] Table 13
[0320] Figures 15A to 15C are schematic diagrams illustrating the clone types and corresponding sequencing sequence numbers of the three samples provided in the exemplary embodiments of this disclosure. As shown in Figures 15A to 15C, Sample 1 is a negative standard, and its amplification efficiency was evaluated. Specific amplification sequences accounted for 72.34%, indicating good amplification specificity. Amplification covered all standard plasmids, achieving 100% coverage. Samples 2 and 3 are positive standards, with distinct main peaks and specific amplification rates of 85.73% and 82.28%, respectively, demonstrating high amplification specificity. The amplification primers performed as expected. Amplification covered all standard plasmids, achieving 100% coverage. Sample 4 did not contain any plasmid sequences, and a small amount of sequences could be detected in the sequencing results, indicating a certain level of contamination, which was within a controllable range.
[0321] This disclosure utilizes TRD multiplex amplification to achieve highly accurate and sensitive diagnosis and typing of lymphoma. It employs algorithmic design of highly specific multiplex amplification primers, standard quality granules for uniform amplification and quantification of experimental samples, and the aforementioned data analysis process for typing and diagnosis. This approach solves the problems of low sensitivity associated with traditional capillary electrophoresis and fragment analysis, as well as the issues of high bias and low coverage in previous high-throughput sequencing PCR amplification methods, effectively improving the detection rate and accuracy of lymphoma.
[0322] This disclosure also provides a primer design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the primer design method as described in any embodiment of this disclosure based on the instructions stored in the memory.
[0323] As shown in Figure 16, in one example, the primer design device may include: a first processor 1610, a first memory 1620, a first bus system 1630, and a first transceiver 1640, wherein the first processor 1610, the first memory 1620, and the first transceiver 1640 are connected through the first bus system 1630, the first memory 1620 is used to store instructions, and the first processor 1610 is used to execute the instructions stored in the first memory 1620 to control the first transceiver 1640 to transmit and receive signals. Specifically, the first transceiver 1640, under the control of the first processor 1610, can acquire one or more reference data. Each type of reference data includes multiple reference gene sequences. For each type of reference data, the first processor 1610 performs the following operations: aligning the multiple reference gene sequences by site, determining a conservation score list for each site, obtaining conservation scores for multiple conservation intervals based on the conservation score list for each site, selecting K conservation intervals with higher conservation scores, where K is a natural number greater than or equal to 1, generating K primer combinations for the K conservation intervals; screening the primers in the K primer combinations and evaluating the screened K primer combinations, obtaining the final primer combination based on the evaluation results.
[0324] It should be understood that the first processor 1610 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0325] The first memory 1620 may include read-only memory and random access memory, and provides instructions and data to the first processor 1610. A portion of the first memory 1620 may also include non-volatile random access memory. For example, the first memory 1620 may also store device type information.
[0326] In addition to the data bus, the first bus system 1630 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the first bus system 1630 in Figure 16.
[0327] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the first processor 1610 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the first memory 1620. The first processor 1610 reads information from the first memory 1620 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.
[0328] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the primer design method as described in any embodiment of this disclosure. The primer design method driven by executing executable instructions is essentially the same as the primer design method provided in the above embodiments of this disclosure, and will not be described in detail here.
[0329] In some possible implementations, various aspects of the primer design methods provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the primer design methods according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the primer design methods described in the embodiments of this disclosure.
[0330] This disclosure also provides a standard plasmid design apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the standard plasmid design method as described in any embodiment of this disclosure based on the instructions stored in the memory.
[0331] As shown in Figure 17, in one example, the standard quality grain design device may include: a second processor 1710, a second memory 1720, a second bus system 1730, and a second transceiver 1740, wherein the second processor 1710, the second memory 1720, and the second transceiver 1740 are connected through the second bus system 1730, the second memory 1720 is used to store instructions, and the second processor 1710 is used to execute the instructions stored in the second memory 1720 to control the second transceiver 1740 to transmit and receive signals. Specifically, the second transceiver 1740, under the control of the second processor 1710, can acquire a first gene cluster, a second gene cluster, and a third gene cluster. The first gene cluster includes multiple first gene sequence fragments, the second gene cluster includes multiple second gene sequence fragments, and the third gene cluster includes multiple third gene sequence fragments. Each of the first, second, and third gene clusters contains multiple functional fragments. The second processor 1710 combines the multiple first gene sequence fragments, the multiple second gene sequence fragments, and the multiple third gene sequence fragments to obtain multiple fragment groups. Each fragment group contains one first gene sequence fragment, one second gene sequence fragment, and one third gene sequence fragment. A non-human sequence is inserted between the second gene sequence fragment and the third gene sequence fragment in each fragment group to obtain multiple plasmid sequences. The ratio of each plasmid sequence is determined to obtain the designed standard plasmid.
[0332] It should be understood that the second processor 1710 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0333] The second memory 1720 may include read-only memory and random access memory, and provides instructions and data to the second processor 1710. A portion of the second memory 1720 may also include non-volatile random access memory. For example, the second memory 1720 may also store device type information.
[0334] In addition to the data bus, the second bus system 1730 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the second bus system 1730 in Figure 17.
[0335] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the second processor 1710 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the second memory 1720. The second processor 1710 reads information from the second memory 1720 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.
[0336] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the standard protogranule design method as described in any embodiment of this disclosure. The standard protogranule design method driven by executing executable instructions is essentially the same as the standard protogranule design method provided in the above embodiments of this disclosure, and will not be described in detail here.
[0337] In some possible implementations, various aspects of the standard protogranule design method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the standard protogranule design method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the standard protogranule design method described in the embodiments of this disclosure.
[0338] This disclosure also provides a standard plasmid quantification analysis apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the standard plasmid quantification analysis method as described in any embodiment of this disclosure based on the instructions stored in the memory.
[0339] As shown in Figure 18, in one example, the standard quality grain quantitative analysis device may include: a third processor 1810, a third memory 1820, a third bus system 1830, and a third transceiver 1840. The third processor 1810, the third memory 1820, and the third transceiver 1840 are connected through the third bus system 1830. The third memory 1820 is used to store instructions, and the third processor 1810 is used to execute the instructions stored in the third memory 1820 to control the third transceiver 1840 to transmit and receive signals. Specifically, the third transceiver 1840, under the control of the third processor 1810, can acquire paired-end sequencing data corresponding to the standard plasmid. The paired-end sequencing data includes Reads1 and Reads2 sequences. The third processor 1810 determines whether Reads1 and Reads2 sequences include UMI sequence identifiers. When both Reads1 and Reads2 sequences include UMI sequence identifiers and the included UMI sequence identifiers are the same, it is determined that the plasmid sequence corresponding to Reads1 and Reads2 sequences has been identified. When only one of Reads1 and Reads2 sequences includes a UMI sequence identifier, the Hamming distance between the Reads sequence that does not include a UMI sequence identifier and the UMI sequence identifier is determined. When the determined Hamming distance is less than or equal to a, it is determined that the plasmid sequence corresponding to Reads1 and Reads2 sequences has been identified, where a is a natural number less than or equal to 2.
[0340] It should be understood that the third processor 1810 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0341] The third memory 1820 may include read-only memory and random access memory, and provides instructions and data to the third processor 1810. A portion of the third memory 1820 may also include non-volatile random access memory. For example, the third memory 1820 may also store device type information.
[0342] In addition to the data bus, the third bus system 1830 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the third bus system 1830 in Figure 18.
[0343] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the third processor 1810 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the third memory 1820. The third processor 1810 reads information from the third memory 1820 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.
[0344] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the standard plasmid quantitative analysis method as described in any embodiment of this disclosure. The standard plasmid quantitative analysis method driven by executing executable instructions is essentially the same as the standard plasmid quantitative analysis method provided in the above embodiments of this disclosure, and will not be described in detail here.
[0345] In some possible implementations, various aspects of the standard plasmid quantification method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the standard plasmid quantification method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the standard plasmid quantification method described in the embodiments of this disclosure.
[0346] This disclosure also provides a disease subtyping diagnostic apparatus, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the disease subtyping diagnostic method as described in any embodiment of this disclosure based on the instructions stored in the memory.
[0347] As shown in Figure 19, in one example, the disease typing diagnostic device may include: a fourth processor 1910, a fourth memory 1920, a fourth bus system 1930, and a fourth transceiver 1940. The fourth processor 1910, the fourth memory 1920, and the fourth transceiver 1940 are connected through the fourth bus system 1930. The fourth memory 1920 is used to store instructions, and the fourth processor 1910 is used to execute the instructions stored in the fourth memory 1920 to control the fourth transceiver 1940 to transmit and receive signals. Specifically, the fourth transceiver 1940, under the control of the fourth processor 1910, acquires sequencing data of the sample to be tested. The fourth processor 1910 performs data preprocessing on the acquired sequencing data; compares the preprocessed data with the reference gene sequences of multiple preset clone species to obtain the sequence proportion corresponding to each clone species; sorts the multiple sequence proportions from largest to smallest, and labels the multiple sequence proportions in sorting order as the first sequence proportion to the M1th sequence proportion, where M1 is the number of clone species; detects whether the difference between the first sequence proportion and the third sequence proportion, and the difference between the second sequence proportion and the third sequence proportion, exceeds a preset difference threshold; when the differences between the first sequence proportion and the third sequence proportion, and the differences between the second sequence proportion and the third sequence proportion, both exceed the preset difference threshold, the sample to be tested is determined to be an oligoclonal sample; when the difference between the first sequence proportion and the third sequence proportion exceeds the preset difference threshold but the difference between the second sequence proportion and the third sequence proportion does not exceed the preset difference threshold, the sample to be tested is determined to be a monoclonal sample; when the differences between the first sequence proportion and the third sequence proportion, and the differences between the second sequence proportion and the third sequence proportion, both do not exceed the preset difference threshold, the sample to be tested is determined to be a polyclonal sample.
[0348] It should be understood that the fourth processor 1910 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0349] The fourth memory 1920 may include read-only memory and random access memory, and provides instructions and data to the fourth processor 1910. A portion of the fourth memory 1920 may also include non-volatile random access memory. For example, the fourth memory 1920 may also store device type information.
[0350] In addition to the data bus, the fourth bus system 1930 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as the fourth bus system 1930 in Figure 19.
[0351] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the fourth processor 1910 or through software instructions. That is, the method steps of this embodiment can be executed by the hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the fourth memory 1920. The fourth processor 1910 reads information from the fourth memory 1920 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.
[0352] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the disease subtyping and diagnosis method as described in any embodiment of this disclosure. The disease subtyping and diagnosis method driven by executing executable instructions is essentially the same as the disease subtyping and diagnosis method provided in the above embodiments of this disclosure, and will not be described in detail here.
[0353] In some possible implementations, various aspects of the disease subtyping diagnosis method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the disease subtyping diagnosis method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the disease subtyping diagnosis method described in the embodiments of this disclosure.
[0354] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0355] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0356] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.
Claims
A primer design method, comprising: Acquire multiple reference data, including TRDV reference gene sequence, TRDD reference gene sequence and TRDJ reference gene sequence; For each type of reference data, the following operations are performed: align multiple reference gene sequences by site, determine the conservation score list for each site, obtain conservation scores for multiple conservation intervals based on the conservation score list for each site, select K conservation intervals with higher conservation scores (K is a natural number greater than or equal to 1), generate K primer combinations for the K conservation intervals, screen the primers in the K primer combinations, evaluate the screened K primer combinations, and obtain the final primer combinations based on the evaluation results. According to the method of claim 1, wherein, The process of obtaining conservation scores for multiple conservation intervals based on the conservation score list for each site includes: Based on the pre-set initial conservative interval [start, end] and the sliding window step size W, multiple conservative intervals are obtained through the sliding window method; The conservatism score for each conservatism interval is calculated using the following formula: W i =log 1 / R i C i =W i ×R i , Among them, W i R represents the weight of the i-th site in each conservatism interval. i C represents the percentage of the most prevalent base type at the i-th site in each conserved interval. i Let S be the conservatism score of the i-th position in each conservatism interval, where S is the conservatism score of each conservatism interval, and n is the length of each conservatism interval, and n = end - start + 1. The method according to claim 2, wherein, The generation of K primer combinations for the K conserved regions includes: For each of the K conservative intervals, perform the following operation: Determine the possible base types at each site in the conservatism interval, wherein the possible base types at each site are the base types whose proportion at each site is greater than or equal to a preset proportion threshold; A primer combination is generated based on the possible base types present at each site in the conserved region, wherein the number of primers in the primer combination is m, and m = ∏m i m i denoted as the number of possible base types at the i-th site in the conservative interval, ∏ as the product symbol, i being between 1 and n, and n being the length of the conservative interval. A primer design apparatus includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the primer design method as described in any one of claims 1 to 3 based on the instructions stored in the memory. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the primer design method as described in any one of claims 1 to 3. A computer program product includes instructions that, when executed by a computer, perform the primer design method as described in any one of claims 1 to 3. A composition comprising: obtained by the method according to any one of claims 1 to 3: TRD upstream decoy oligonucleotides, wherein the TRD upstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96% of any one or more sequences in SEQ ID NO:1-9. 97%, 98%, 97% or 100%; and TRD downstream decoy oligonucleotides, wherein the TRD downstream decoy oligonucleotides are derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of any one or more sequences in SEQ ID NO:10-15. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:1-9; the downstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of several sequences in SEQ ID NO:10-15. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:1-9; the downstream decoy oligonucleotide of the TRD is derived from at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 97%, or 100% of all sequences in SEQ ID NO:10-15. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRD is selected from all sequences in SEQ ID NO:1-9; the downstream decoy oligonucleotide of the TRD is selected from all sequences in SEQ ID NO:10-15. The composition according to claim 7, wherein, The upstream decoy oligonucleotide of the TRD further includes a forward adapter primer sequence, and the downstream decoy oligonucleotide of the TRD further includes a reverse adapter primer sequence; the forward adapter primer sequence and the reverse adapter primer sequence are complementary to the adapter primer sequence used for sequencing. The composition according to claim 11, wherein, The adapter primers used for sequencing are selected from one or more of the Nextera series adapters and the TruSeq series adapters. The composition according to claim 11, wherein, The forward adapter primer sequence is shown in SEQ ID NO:16, and the reverse adapter primer sequence is shown in SEQ ID NO:
17. The composition according to claim 11, wherein, The upstream decoy oligonucleotide of the TRD is selected from one or more sequences shown in SEQ ID NO:18-26; and The downstream decoy oligonucleotide of the TRD is selected from one or more sequences shown in SEQ ID NO:27-32. Use of the composition of any one of claims 7 to 14 in amplifying the TRD gene and / or detecting TRD gene rearrangements. A kit comprising the composition of any one of claims 7 to 14. The kit according to claim 16, wherein, The kit also contains: Twelve standard regiofluids are provided, each containing a unique UMI sequence, ensuring that each regiofluid is uniquely identified by its UMI sequence. Identification in one location; The UMI sequence is 12 bp in length, with the first 8 bp being a non-human random sequence and the last 4 bp being the first 4 bases of the TRD J region sequence. Each of the standard quality grains further comprises a TRD V region sequence, a TRD D region sequence, and a TRD J region sequence, wherein the TRD V region sequence, the TRD D region sequence, the UMI sequence, and the TRD J region sequence are sequentially linked end-to-end in each of the standard quality grains. The kit according to claim 17, wherein, The last 4 bp segment of the UMI sequence is selected from ACAC, CTTT, CTCC, CCAG, or CACA. The kit according to claim 17, wherein, The UMI sequence is shown in SEQ ID NO:33-44. The kit according to claim 17, wherein, The 12 standard quality grains are standard quality grains containing the following sequences respectively: TRDV1*01, TRDD1&TRDD2 and TRDJ1*01; TRDV2*01, TRDD1&TRDD3 and TRDJ2*01; TRDV3*01, TRDD2&TRDD3 and TRDJ3*01; TRAV14 / DV4*01, TRDD1&TRDD2 and TRDJ4*01; TRAV23 / DV6*01, TRDD1&TRDD3 and TRDJ1*01; TRAV29 / DV5*01, TRDD2&TRDD3 and TRDJ2*01; TRAV36 / DV7*01, TRDD1&TRDD2 and TRDJ3*01; TRAV38-2 / DV8*01, TRDD2&TRDD3 and TRDJ4*01; TRD-δ2-F-intron, TRDD2&TRDD3 and TRDJ3*01; TRD-δ2-F-in tron, TRDD2&TRDD3 and TRD-δ3-R-intron; TRDV2*01, TRDD1&TRDD3 and TRD-δ3-R-intron; or TRAV23 / DV6*01, TRDD2&TRDD3 and TRD-δ3-R-intron. The kit according to claim 20, wherein, The uniformity standard is obtained by mixing the 12 standard quality particles in an equimolar ratio. By mixing one or more of the 12 standard quality particles in a high proportion and the other standard quality particles in a low proportion, an experimental standard for simulating monoclonal rearrangement is obtained. Use of the kit according to any one of claims 16 to 21 in evaluating the amplification efficiency of multiple primers used to amplify the TRD gene. Use of the kit according to any one of claims 16 to 21 in the detection of TRD gene rearrangements. A disease typing diagnostic method for detecting TRD gene rearrangements includes the following steps: 1) Obtain the genomic DNA of the sample to be tested; 2) Perform PCR on the genomic DNA obtained in step 1) using the composition of any one of claims 7 to 14 to obtain PCR products; 3) Sequencing the PCR products obtained in step 2) and analyzing the sequencing results to determine whether the TRD gene rearrangement in the sample is monoclonal or polyclonal. The method according to claim 24, wherein, It also includes, in step 1), incorporating the uniformity standard as defined in claim 21 into the genomic DNA. The method according to claim 24, wherein, Analyze the sequencing results using the following steps: Perform data preprocessing on the sequencing data; The preprocessed data is compared with reference gene sequences of multiple predefined cloning species to obtain the results for each cloning species. The proportion of sequences corresponding to the class, wherein the reference gene sequences include: TRDV reference gene sequence, TRDD reference gene sequence and TRDJ reference gene sequence; Sort the multiple sequence proportions from largest to smallest, and label the proportions of the multiple sequences in the sorting order as the proportion of the first sequence to the proportion of the M1th sequence, where M1 is the number of clone species: The method detects whether the differences between the proportions of the first and third sequences, and between the proportions of the second and third sequences, exceed preset difference thresholds. When both the differences between the proportions of the first and third sequences and between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be an oligoclonal sample. When the differences between the proportions of the first and third sequences exceed preset difference thresholds but the differences between the proportions of the second and third sequences do not exceed preset difference thresholds, the sample to be tested is determined to be a monoclonal sample. When neither the differences between the proportions of the first and third sequences nor the differences between the proportions of the second and third sequences exceed preset difference thresholds, the sample to be tested is determined to be a polyclonal sample. The method according to claim 26, wherein, Detecting whether the difference between the proportion of the first sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X; when the ratio of the proportion of the first sequence to the proportion of the third sequence is greater than or equal to X, determining that the difference between the proportion of the first sequence and the proportion of the third sequence exceeds the preset difference threshold; when the ratio of the proportion of the first sequence to the proportion of the third sequence is less than X, determining that the difference between the proportion of the first sequence and the proportion of the third sequence does not exceed the preset difference threshold, where X is the preset difference threshold and X is greater than or equal to 2; Detecting whether the difference between the proportion of the second sequence and the proportion of the third sequence exceeds a preset difference threshold includes: detecting whether the ratio of the proportion of the second sequence to the proportion of the third sequence is greater than or equal to X; when the ratio of the proportion of the second sequence to the proportion of the third sequence is greater than or equal to X, determining that the difference between the proportion of the second sequence and the proportion of the third sequence exceeds the preset difference threshold; when the ratio of the proportion of the second sequence to the proportion of the third sequence is less than X, determining that the difference between the proportion of the second sequence and the proportion of the third sequence does not exceed the preset difference threshold. A disease typing diagnostic apparatus includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the disease typing diagnostic method as described in any one of claims 24 to 27 based on the instructions stored in the memory. A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the disease subtyping diagnostic method as described in any one of claims 24 to 27. A computer program product includes instructions that, when executed by a computer, perform the disease subtyping diagnostic method as described in any one of claims 24 to 27.
Citation Information
Patent Citations
Method for carrying out high-throughput sequencing on TCR (T cell receptor) or BCR (B cell receptor) and method for correcting multiplex PCR (polymerase chain reaction) primer deviation by utilizing tag sequences
CN103710454A
Method and device for determining whether variable-region amplification primers have deviation or not and application of method and device
CN105331680A
Modified cells and methods of therapy
CN108472314A
Methods and compositions for detecting single t cell receptor affinity and sequence
CN109312402A
Multiplex primer group and method for constructing human T cell immune group library based on high-throughput sequencing using primer group
CN109554440A